Forskningsradar

Science Journals

Peer-reviewade publikationer — 53080 artiklar

Statistical Limits and Efficient Algorithms for Differentially Private Federated Learning
arXiv:2605.18656v1 Announce Type: cross Abstract: Federated Learning is a leading framework for training ML and AI models collaboratively across numerous user devices or databases. We study the trade-offs among estimation accuracy, privacy constraints, and communication cost for differentially private (DP) federated M estimation. The two standard methods in the literature are FedAvg, which may suffer from high federation bias, and FedSGD, which can incur high communication cost. Aimed at improving accuracy at a reduced communication cost, we propose FedHybrid, which uses FedSGD starting with an improved initialization by the FedAvg estimator. We propose FedNewton, which averages local Newton iterations to reduce bias in FedAvg, achieving an estimation accuracy comparable to FedSGD with much fewer communication rounds when the number of clients grows sufficiently slowly. We establish finite sample upper bounds on the mean-squared error rates of the DP versions of these estimators as functions of the number of clients, local sample sizes, privacy budget, and number of iterations. We further derive a minimax lower bound on the MSE of any iterative private federated procedure that provides a benchmark to assess the optimality gap of these methods. We numerically evaluate our methods for training a logistic regression and a neural network on the computer vision datasets MNIST and CIFAR-10.
UPSim: UxNB Propagation Simulator for 3D Map-Driven FR3 Deployments
arXiv:2605.17378v1 Announce Type: cross Abstract: We introduce UPSim (UxNB Propagation Simulator), a ray tracing-calibrated, semi-deterministic solution for spatially consistent FR3 air-to-ground propagation modeling in uncrewed aerial vehicle (UAV) networks. Instead of launching rays for every receiver position, UPSim derives deterministic visibility regions from 3D building geometry via shadow projection. It then augments these regions with line-of-sight (LOS) state-specific and altitude-aware path loss, correlated large-scale fading, and small-scale fading. Calibration and validation against FR3 ray tracing data using the global 3D-GloBFP building dataset demonstrate that UPSim accurately reproduces empirical channel distributions. Furthermore, the resulting maps support route-based analysis of channel evolution over complex urban layouts, exposing critical trajectory-level statistics such as outage distances. Consequently, UPSim offers a highly scalable, practical middle ground between computationally expensive full ray tracing and purely stochastic channel generation for mobility-aware planning and radio-map construction in aerial access scenarios.
HYVINT: Intensity-Driven Hypergraph Generation with Variational Representations
arXiv:2605.16836v1 Announce Type: cross Abstract: Hypergraphs provide a principled framework for modeling polyadic interactions, with applications in recommendation systems, social networks, and molecular modeling. Hypergraph generation remains challenging because incidence structures are discrete, sparse, and governed by heterogeneous higher-order interactions. Existing generators often rely on implicit latent spaces or continuous incidence decoders, which provide limited mechanistic interpretation of how node-hyperedge incidences arise. To address these limitations, we propose HYVINT, an intensity-driven hypergraph generative framework. Our key innovations are twofold: (i) we develop an intensity-driven incidence formation mechanism for hypergraphs that links latent interaction strength to binary incidence, and (ii) we derive a tractable lower-bound variational estimator for learning latent representations. We provide generation error bounds with asymptotic convergence rates and empirically show that HYVINT achieves strong fidelity while maintaining substantial novelty and diversity on synthetic and real-world hypergraphs.
From order to chaos in a chip-scale Kerr parametric oscillator
arXiv:2605.18690v1 Announce Type: new Abstract: Integrated photonics has enabled a wide class of chip-scale light sources and quantum technologies. Within this field, microresonator-based degenerate optical parametric oscillators (DOPOs) have gained prominence. Above a critical power threshold, these systems undergo spontaneous symmetry breaking to settle into one of two stable, {\pi}-phase-shifted states -- a mechanism successfully used for quantum random number generation and photonic Ising machines. Here, we show that DOPOs based on the Kerr nonlinearity host a significantly broader range of nonlinear dynamics than previously explored. Using a silicon nitride microring resonator, we experimentally identify Hopf bifurcations that trigger a transition from stationary operation to self-sustained oscillations at MHz frequencies. By adjusting pump detunings and powers, we achieve turnkey control over these oscillatory regimes, navigating the system between stable binary states and periodic limit cycles. Furthermore, we report the experimental observation of period-doubling bifurcations, which numerical simulations reveal as the precursor to a cascading instability culminating in chaos at elevated pump powers. Our results establish a framework for controlling nonlinear instabilities in chip-scale parametric oscillators, with applications in programmable photonic hardware and dynamical optical computing.
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation
arXiv:2512.23994v4 Announce Type: replace Abstract: Text-to-audio-video (T2AV) generation is central to applications such as filmmaking and world modeling. However, current models often fail to produce physically plausible sounds. Previous benchmarks primarily focus on audio-video temporal synchronization, while largely overlooking explicit evaluation of audio-physics grounding, thereby limiting the study of physically plausible audio-visual generation. To address this issue, we present PhyAVBench, the first benchmark that systematically evaluates the audio-physics grounding capabilities of T2AV, image-to-audio-video (I2AV), and video-to-audio (V2A) models. PhyAVBench offers PhyAV-Sound-11K, a new dataset of 25.5 hours of 11,605 audible videos collected from 184 participants to ensure diversity and avoid data leakage. It contains 337 paired-prompt groups with controlled physical variations that drive sound differences, each grounded with an average of 17 videos and spanning 6 audio-physics dimensions and 41 fine-grained test points. Each prompt pair is annotated with the physical factors underlying their acoustic differences. Importantly, PhyAVBench leverages paired text prompts to evaluate this capability. We term this evaluation paradigm the Audio-Physics Sensitivity Test (APST) and introduce a novel metric, the Contrastive Physical Response Score (CPRS), which quantifies the acoustic consistency between generated videos and their real-world counterparts. We conduct a comprehensive evaluation of 17 state-of-the-art models. Our results reveal that even leading commercial models struggle with fundamental audio-physical phenomena, exposing a critical gap beyond audio-visual synchronization and pointing to future research directions. We hope PhyAVBench will serve as a foundation for advancing physically grounded audio-visual generation. Prompts, ground-truth, and generated video samples are available at https://github.com/imxtx/PhyAVBench.
Mapping the Turn: An Eulerian Binormal-Axis Diagnostic for Recirculating 3D Flows
arXiv:2605.18439v1 Announce Type: new Abstract: Three-dimensional (3D) recirculating flows are often interpreted qualitatively from selected streamline visualizations. In separated flows, such recirculating motion is central to the drag modulation, but the local orientation of recirculation remains difficult to quantify in a field-based form. This work introduces an Eulerian binormal-axis diagnostic that locally evaluates the orientation of streamline turning at each point in the velocity field, yielding a spatially resolved field of the recirculating direction. Motivated by the Frenet-Serret binormal direction of a curved streamline, the diagnostic uses the velocity vector and its convective acceleration to extract the local streamline-turning axis without requiring explicit streamline integration. The resulting direction is encoded with barycentric RGB weights to visualize streamwise, spanwise, and wall-normal turning axis contributions. The diagnostic is first applied to Hill's spherical vortex, which provides a controlled analytic example of 3D recirculating motion for interpreting the binormal-axis direction and the associated barycentric RGB encoding. It is then applied to the mean field of a pressure-gradient-induced 3D separation bubble. The resulting visualizations show that the diagnostic reveals orientation changes that are not apparent from streamline visualization. The proposed diagnostic therefore converts qualitative streamline impressions into a spatially resolved measure of local streamline-turning orientation, providing a quantitative complement to conventional 3D flow visualization.
Optimising CSRNet with parameter-free attention mechanisms for crowd counting in public transport
arXiv:2605.18349v1 Announce Type: new Abstract: Occupancy estimation and crowd counting are critical tasks in designing smart and efficient public transport vehicles. Given that public transport loading can vary from sparse to crowded, classical models for occupancy estimation must be adapted to suit this purpose. Attention mechanisms have shown remarkable capability in enhancing the representational power of deep neural networks for crowd counting in congested scenes with occlusion, complex backgrounds, and perspective distortion. However, conventional approaches, often implemented as parameterized sub-networks within convolutional layers, inevitably increase model size and computational cost, limiting deployment on resource-constrained edge devices. This paper investigates the effectiveness of state-of-the-art parameter-free attention mechanisms for crowd counting and density map estimation in highly congested scenes. We evaluate channel-wise (PFCA), spatial-wise (SA), and 3-D (SimAM) modules and compare their performance with parameterized attention modules constrained to introduce no more than 1% additional parameters. Furthermore, we present a novel combination of attention mechanisms that combines the strengths of PFCA and SA (PFCASA) customized for analyzing video streams onboard public transport systems. Using CSRNet as the backbone, experiments on the ShanghaiTech dataset demonstrate that parameter-free attention mechanisms achieve comparable or superior accuracy without introducing additional model parameters. A detailed performance analysis further reveals that PFCASA outperforms other attention modules in scenes with fewer than 40 individuals, while PFCA shows greater effectiveness as crowd density increases, underscoring their potential applicability for integration into smart public transport modalities.
Dynamical system analysis of quantum tunneling in an asymmetric double-well potential
arXiv:2510.24100v2 Announce Type: replace-cross Abstract: We study quantum tunneling in an asymmetric double-well potential using a dynamical systems--based approach rooted in the Ehrenfest formalism. In this framework, the time evolution of a Gaussian wave packet is governed by a hierarchy of coupled equations linking lower- and higher-order position moments. An approximate closure scheme, required to render the system tractable, yields a reduced dynamical system for the mean and variance, with skewness explicitly entering due to the potential's asymmetry. Stability analysis of this system identifies energy thresholds for detectable tunneling across the barrier and reveals regimes where tunneling, though theoretically allowed, remains practically undetectable. Comparison with full numerical solutions of the time-dependent Schr\"odinger equation shows that, beyond reproducing key tunneling features, the dynamical systems approach provides an interpretable description of quantum transport through tunneling in an effective asymmetric two-level system.
Leveraging Graph Structure in Seq2Seq Models for Knowledge Graph Link Prediction
arXiv:2605.18211v1 Announce Type: new Abstract: We introduce Graph-Augmented Sequence-to-Sequence (GA-S2S), a novel framework that integrates a T5-small encoder-decoder with a Relational Graph Attention Network (RGAT) to improve link prediction in knowledge graphs. While existing Seq2Seq models rely solely on surface-level textual descriptions of entities and relations and at best, flatten the neighborhoods of a query entity into a single linear sequence, thereby discarding the inherent graph structure, GA-S2S jointly encodes both textual features and the full $k$-hop subgraph topology surrounding the query entity. By integrating raw encoder outputs with RGAT's relation-aware embeddings, our model captures and leverages richer multi-hop relational patterns and textual information. Our preliminary experiments on the CoDEx dataset demonstrate that GA-S2S outperforms competitive Seq2Seq-based baseline models, achieving up to a 19\% relative gain in link prediction accuracy.
The MixCount Dataset: Bridging the Data Gap for Open-Vocabulary Object Counting
arXiv:2605.18063v1 Announce Type: new Abstract: Object counting is a foundational vision task with over a decade of dedicated research, yet state-of-the-art models still fail systematically in the mixed-object setting that dominates real-world applications such as industrial inspection and product sorting. We show that this gap is strongly driven by limitations in existing training and evaluation data: real counting datasets are prohibitively expensive to annotate and suffer from labeling noise, while existing synthetic alternatives lack diversity and realism. We address this with MixCount, a dataset and benchmark for mixed-object counting designed to target the failure modes of current counting models. To overcome the high cost of constructing and labeling such data, we develop an automatic generation pipeline that synthesizes images, fine-grained textual descriptions, and pixel-perfect counting annotations at scale, eliminating the labeling ambiguity that plagues prior datasets. Evaluating state-of-the-art counting models on MixCount exposes severe degradation in the mixed-object setting. More importantly, training these models on our synthesized data yields substantial gains on real-world benchmarks, reducing MAE by 20.14% on FSC-147 and by 18.3% on PairTally. These results establish MixCount as both a benchmark and a training dataset for fine-grained counting, and demonstrate that our pipeline, which produces effectively unlimited labeled data, helps address a long-standing bottleneck in counting models.
PROTEA: Offline Evaluation and Iterative Refinement for Multi-Agent LLM Workflows
arXiv:2605.18032v1 Announce Type: new Abstract: Multi-agent LLM workflows -- systems composed of multiple role-specific LLM calls -- often outperform single-prompt baselines, but they remain difficult to debug and refine. Failures can originate from subtle errors in intermediate outputs that propagate to downstream nodes, requiring developers to inspect long traces and infer which agent to modify. We present PROTEA, a unified interface for offline, test-driven improvement of multi-agent workflows. PROTEA executes a workflow, scores intermediate node outputs with configurable rubrics, and overlays per-node states and rationales on the workflow graph to localize likely bottlenecks. To support complex systems where final-answer references are the primary supervision, PROTEA performs backward node evaluation: it generates candidate node-level expectations from final-answer references and graph context, then compares them with observed node outputs. For selected nodes, PROTEA presents targeted prompt revisions as editable before/after comparisons, then automatically reruns and re-evaluates the workflow to show output changes and score trajectories within the same interface. In two production-adjacent workflows, PROTEA improved document-inspection accuracy from 64.3% to 83.9% and recommendation Hit@5 from 0.30 to 0.38. In a formative study with six experienced LLM developers, participants valued graph-level localization, per-node rationales, and editable before/after prompt revisions.
Scalable Decision-Focused Learning through Cost-Sensitive Regression
arXiv:2605.18005v1 Announce Type: new Abstract: Many real-world combinatorial problems involve uncertain parameters, which can be predicted given contextual features and historical data. These `predict-then-optimize' or `contextual optimization' problems have gained significant attention: end-to-end training methods can now minimize the downstream task cost rather than the predictive error. However, despite their effectiveness, these decision-focused learning (DFL) approaches often rely on repeated solving of the underlying combinatorial optimization problem during training, making them computationally expensive and difficult to scale. We reframe the learning problem as a cost-sensitive multi-output regression problem: multi-output due to the combinatorial problem having multiple uncertain parameters, and cost-sensitive due to the downstream task cost being the real target. Our technical contribution is the formalization of multiple loss function components that follow from this reframing: cost-insensitive normalization, decision-aware asymmetric penalization of over- and underpredictions, and instance-based costs that mimic the true downstream task-based loss locally. These components require zero or one solve per training data instance, while requiring no further solves during training. Experiments show that the combination of loss components achieves comparable downstream task quality to the state of the art, while being significantly more efficient, enabling scaling to problem sizes that have not been tackled before with DFL.
Training data attribution in diffusion models via mirrored unlearning and noise-consistent skew
arXiv:2605.17938v1 Announce Type: new Abstract: Training data attribution (TDA) should enable generative model interpretability and foster a variety of related downstream tasks. Nonetheless, current TDA approaches lack reliability and robustness, preventing their adoption in real-world setups. In this paper, we take a decisive step towards more reliable and robust TDA for diffusion models. We propose to perform TDA with mirrored unlearning and noise-consistent skew (MUCS). The idea is to fine-tune a second model with bounded mirrored gradient ascent, and to measure the normalized skew of this model with respect to the original one using consistent noise samples. We show that, while being conceptually simple and generic, MUCS systematically outperforms existing methods on three different datasets by a large margin. We additionally study the effect that core design choices have on final performance, and analyze novel aspects regarding the overlap of influential instances across generated items and the potential of ensembling TDA approaches. We believe that our findings may have broader implications for more general unlearning setups, as well as for tasks requiring the comparison of diffusion losses.
Elastic wave propagation governs impulse enhancement in pulsed jets through flexible nozzles
arXiv:2605.17319v1 Announce Type: new Abstract: Inspired by cephalopod jet propulsion through compliant funnels, this study investigates elastic wave propagation and energy exchange in passively deforming cylindrical nozzles through three-dimensional, two-way fluid-structure interaction simulations. Flexible nozzles with varying stiffness ($Eh = 75 - 500~\mathrm{N\,m^{-1}}$, where $E$ and $h$ are Young's modulus and nozzle thickness, respectively) are subjected to a pulsatile jet inflow at $Re \sim 4000$. Increasing nozzle flexibility reduces the deformation-wave speed in accordance with Moens-Korteweg scaling, thereby prolonging the nozzle expansion phase. This delayed expansion enhances jet entrainment and elastic energy storage while suppressing early shear-layer roll-up and vortex formation. During contraction, the stored elastic energy is released, thereby enhancing jet acceleration and vortex formation. For the most flexible nozzle, the primary vortex-ring circulation increases by 52.13%, the vortex convection distance by 9.00%, and the peak outlet kinetic energy flux by a factor of 4.62 compared with a rigid nozzle. These effects collectively yield a 61.92% increase in total hydrodynamic impulse. These findings identify passive wave-speed tuning via nozzle compliance as a mechanism to enhance pulsed-jet thrust for bio-inspired underwater propulsion.
GREEN GRID: A Web-Based E-Waste Recycling Platform
arXiv:2605.17924v1 Announce Type: new Abstract: Electronic waste (e-waste) is one of the fastest-growing waste streams worldwide due to rapid technological advancements and shorter device lifespans. Improper disposal releases hazardous substances that harm the environment and human health, while valuable materials such as gold, copper, and aluminum are lost if not recycled. In 2022, approximately 62 million metric tonnes of e-waste were generated globally, but only about 22% was formally recycled. India generated around 1.751 million metric tonnes in 2023-24, with only 43% processed through authorized channels. Green Grid is a full-stack web-based platform designed to simplify and encourage e-waste recycling through an E-Dumper Locator, Green Rewards System, Insights and Awareness Hub, Scheduled Pickup Service, Recycling Impact Calculator, Eco AI Assistant, and Eco-Marketplace. Developed using React.js, Node.js, Express.js, SQL, Google Maps API, and JWT authentication, the platform transforms e-waste recycling into a transparent, educational, and rewarding process. By combining technology, awareness, and incentives, Green Grid promotes responsible disposal and supports circular economy practices for a more sustainable future.
FLEX-MoE: Federated Mixture-of-Experts with Load-balanced Expert Assignment for Edge Computing
arXiv:2512.23070v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) models enable scalable neural networks through conditional computation, offering enhanced effectiveness and efficiency for next-generation wireless communications. However, deploying MoE with federated learning (FL) over wireless and IoT edge networks faces two critical challenges: 1) resource-constrained clients cannot store large AI models with full expert sets, and 2) non-IID data distributions cause severe expert load imbalance that degrades model performance. To this end, we propose FLEX-MoE, a federated MoE framework that jointly optimizes expert assignment and load balancing under limited client capacity. Specifically, our approach introduces client-expert fitness scores that quantify expert suitability for local datasets through training feedback, and employs an optimization-based algorithm to maximize client-expert specialization while enforcing balanced expert utilization system-wide. Unlike greedy methods that focus solely on personalization while ignoring load imbalance, FLEX-MoE addresses expert utilization skew, which is particularly severe in heterogeneous edge FL. Our experimental results demonstrate superior accuracy and consistently balanced expert utilization across diverse resource-constrained scenarios for edge computing.
Domain Transfer Becomes Identifiable via a Single Alignment
arXiv:2605.17918v1 Announce Type: new Abstract: Domain transfer (DT) maps source to target distributions and supports tasks such as unsupervised image-to-image translation, single-cell analysis, and cross-platform medical imaging. However, DT is fundamentally ill-posed: push-forward mappings are generally non-identifiable, as measure-preserving automorphisms (MPAs) preserve marginals while altering cross-domain correspondences, leading to content-misaligned translation. Recent work shows that MPAs can be eliminated by jointly transferring multiple corresponding source/target conditional distributions, but supervision signals labeling such conditionals are not always available in practice. We develop an alternative route to DT identifiability. Under a structural sparsity condition on the Jacobian support pattern, we show that distribution matching together with a single paired anchor sample suffices to identify the ground-truth transfer -- requiring substantially less supervision than prior approaches. To enable practical high-dimensional learning, we further propose an efficient Jacobian sparsity regularizer based on randomized masked finite differences, yielding a scalable surrogate without explicit Jacobian evaluation. Empirical results on synthetic and real-world DT tasks validate the theory.
The Bayesian Geometry of Transformer Attention
arXiv:2512.22471v5 Announce Type: replace Abstract: Transformers often appear to perform Bayesian reasoning in context, but verifying this rigorously has been impossible: natural data lack analytic posteriors, and large models conflate reasoning with memorization. We address this by constructing \emph{Bayesian wind tunnels} -- controlled environments where the true posterior is known in closed form and memorization is provably impossible. In these settings, small transformers reproduce Bayesian posteriors with $10^{-3}$-$10^{-4}$ bit accuracy, while capacity-matched MLPs fail by orders of magnitude, establishing a clear architectural separation. Across two tasks -- bijection elimination and Hidden Markov Model (HMM) state tracking -- we find that transformers implement Bayesian inference through a consistent geometric mechanism: residual streams serve as the belief substrate, feed-forward networks perform the posterior update, and attention provides content-addressable routing. Geometric diagnostics reveal orthogonal key bases, progressive query-key alignment, and a low-dimensional value manifold parameterized by posterior entropy. During training this manifold unfurls while attention patterns remain stable, a \emph{frame-precision dissociation} predicted by recent gradient analyses. Taken together, these results demonstrate that hierarchical attention realizes Bayesian inference by geometric design, explaining both the necessity of attention and the failure of flat architectures. Bayesian wind tunnels provide a foundation for mechanistically connecting small, verifiable systems to reasoning phenomena observed in large language models.
An Instrument for Physical Vapor Deposition onto Cryo-EM Samples for Microsecond Time-Resolved Cryo-EM
arXiv:2512.20522v2 Announce Type: replace Abstract: Laser flash melting and revitrification experiments have recently improved the time resolution of cryo-electron microscopy (cryo-EM) to the microsecond timescale, making it fast enough to observe many of the protein motions that are associated with function. The technique has also opened up a new dimension for cryo-EM sample preparation, making it possible to deposit compounds onto a cryo-EM sample while it is frozen, so that upon flash melting, the embedded particles experience an altered environment. For example, we have recently shown that depositing ultrathin silicon dioxide membranes onto a cryo-EM sample causes particles to detach from the interface upon flash melting, removing preferred particle orientation. These experiments also point towards a new strategy for initiating protein dynamics in time resolved experiments by depositing reagents, which will then mix with the sample upon flash melting. Here, we describe an apparatus for physical vapor deposition of compounds onto cryo-EM samples, detailing its design and operation. As a demonstration, we determine that the minimum thickness of silicon dioxide sealing membranes in a laser flash melting experiment is just over two monolayers. We propose that our design can form the basis for an integrated platform for microsecond time-resolved cryo-EM experiments.
WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform
arXiv:2605.17912v1 Announce Type: new Abstract: World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, existing embodied world model benchmarks are still largely confined to vision-only prediction, offline embodied applications, and simulator-based evaluation, making them insufficient for assessing increasingly comprehensive world models. In this work, we introduce WorldArena 2.0, an expanded benchmark that systematically broadens embodied world model evaluation along three dimensions: modality, functionality, and platform. Along the modality dimension, WorldArena 2.0 extends evaluation from vision-only to visuotactile modalities, enabling assessment of multimodal perception and prediction. Along the functionality dimension, it extends beyond policy evaluation and planning to assess world models as interactive RL environments for policy optimization. Along the platform dimension, it moves beyond simulator-only evaluation to a diverse suite of simulated and real-world robotic settings across multiple embodiments. Under a standardized protocol, WorldArena 2.0 comprehensively evaluates perceptual quality, interactive utility, and cross-platform performance, providing a comprehensive testbed for tracking progress toward embodied world models. The benchmark is available at: https://world-arena.ai.
QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation
arXiv:2512.19134v2 Announce Type: replace Abstract: Dynamic Retrieval-Augmented Generation adaptively determines when to retrieve during generation to mitigate hallucinations in large language models (LLMs). However, existing methods rely on model-internal signals (e.g., logits, entropy), which are fundamentally unreliable because LLMs are typically ill-calibrated and often exhibit high confidence in erroneous outputs. We propose QuCo-RAG, which shifts from subjective confidence to objective statistics computed from pre-training data. Our method quantifies uncertainty through two stages: (1) before generation, we identify low-frequency entities indicating long-tail knowledge gaps; (2) during generation, we verify entity co-occurrence in the pre-training corpus, where zero co-occurrence often signals hallucination risk. Both stages leverage Infini-gram for millisecond-latency queries over 4 trillion tokens, triggering retrieval when uncertainty is high. Experiments on multi-hop QA benchmarks show QuCo-RAG achieves EM gains of 5--12 points over state-of-the-art baselines with OLMo-2 models, and transfers effectively to models with undisclosed pre-training data (Llama-3, Qwen2.5, GPT-4.1/5-chat), improving EM by up to 14 points. Generalization to long-form generation and biomedical QA further validates the robustness of our paradigm. These results establish corpus-grounded verification as a principled, practically model-agnostic paradigm for dynamic RAG. Our code is publicly available at https://github.com/ZhishanQ/QuCo-RAG.
ShareChat: A Dataset of Chatbot Conversations in the Wild
arXiv:2512.17843v4 Announce Type: replace Abstract: By evaluating Large Language Models (LLMs) through uniform, text-only interfaces, current academic benchmarks obscure how the unique designs and affordances of distinct commercial platforms shape real-world user behavior and system performance. To bridge this gap, we present ShareChat, the first large-scale corpus of 142,808 conversations (660,293 turns) collected from publicly shared URLs on ChatGPT, Perplexity, Grok, Gemini, and Claude. ShareChat preserves native platform affordances, including citations, thinking traces, and code artifacts, across 95 languages and the period from April 2023 to October 2025, complementing existing corpora that homogenize these interactions. To demonstrate the dataset's evaluative utility, we present three case studies: a conversation completeness analysis assessing cross-platform differences in intent satisfaction, a source grounding analysis comparing citation strategies between search-augmented systems, and a temporal analysis revealing divergent response latency dynamics. Together, these analyses demonstrate research questions that are inaccessible to single-platform or stripped-affordance corpora. The dataset is publicly available.
Learning under Distributional Drift: Prequential Reproducibility as an Intrinsic Statistical Resource
arXiv:2512.13506v4 Announce Type: replace Abstract: Statistical learning under distributional drift remains poorly characterized, especially in closed-loop settings where learning alters the data-generating law. We introduce an intrinsic drift budget $C_T$ that quantifies cumulative information-geometric motion of the data distribution along the realized learner-environment trajectory, measured in Fisher-Rao distance. The budget separates exogenous environmental change from policy-sensitive feedback induced by the learner's actions. This gives a rate-based characterization of prequential reproducibility: when performance on the realized stream is used to predict one-step-ahead performance under the next distribution, the drift contribution enters through the average motion rate $C_T/T$, not through cumulative drift alone. We prove a drift-feedback bound of order $T^{-1/2}+C_T/T$, up to controlled second-order remainder terms, and establish a matching sharpness lower bound for the same prequential reproducibility gap on a canonical regular subclass. Thus the dependence on the average Fisher-Rao motion rate is tight up to constants: $C_T/T$ is sufficient for upper control and unavoidable on regular hard subclasses. We further prove an information-theoretic indistinguishability result showing that order-$C/T$ effects on the one-step-ahead target need not be identifiable from the realized performance stream alone. Finally, we show that fixed monitoring channels induce contracted observable Fisher motion, and experiments, including a misspecified real-data feedback setting, indicate that appropriately chosen channels can retain risk-relevant drift signal when the intrinsic data-generating law is unavailable. The resulting theory treats exogenous drift, adaptive data analysis, and performative feedback as different sources of Fisher-Rao motion along the same learner-environment trajectory.
On the Accuracy of Newton Step and Influence Function Data Attributions
arXiv:2512.12572v2 Announce Type: replace Abstract: Data attribution aims to explain model predictions by estimating how they would change if certain training points were removed, and is used in a wide range of applications, from interpretability and credit assignment to unlearning and privacy. Even in the relatively simple case of logistic regressions, existing mathematical analyses of leading data attribution methods such as Influence Functions (IF) and single Newton Step (NS) remain limited in two key ways. First, they rely on global strong convexity assumptions which are often not satisfied in practice. Second, the resulting bounds scale very poorly with the number of parameters ($d$) and the number of samples removed ($k$). As a result, these analyses are not tight enough to answer fundamental questions such as "what is the asymptotic scaling of the errors of each method?" or "which of these methods is more accurate for a given dataset?" In this paper, we introduce a new analysis of the NS and IF data attribution methods for convex learning problems. To the best of our knowledge, this is the first analysis of these questions that does not assume global strong convexity and also the first explanation of [KATL19] and [RH25a]'s observation that NS data attribution is often more accurate than IF. We prove that for sufficiently well-behaved logistic regressions, our bounds are asymptotically tight up to poly-logarithmic factors, yielding scaling laws for the errors in the average-case sample removals. \[ \mathbb{E}_{T \subseteq [n],\, |T| = k} \bigl[ \|\hat{\theta}_T - \hat{\theta}_T^{\mathrm{NS}}\|_2 \bigr] = \widetilde{\Theta}\!\left(\frac{k d}{n^2}\right), \qquad \mathbb{E}_{T \subseteq [n],\, |T| = k} \bigl[ \|\hat{\theta}_T^{\mathrm{NS}} - \hat{\theta}_T^{\mathrm{IF}}\|_2 \bigr] = \widetilde{\Theta}\!\left( \frac{(k + d)\sqrt{k d}}{n^2} \right). \]
A Conservative Discontinuous Galerkin Algorithm for Particle Kinetics on Smooth Manifolds
arXiv:2512.05298v2 Announce Type: replace Abstract: A novel, conservative discontinuous Galerkin algorithm is presented for particle kinetics on manifolds. The motion of particles on the manifold is represented using using both canonical and non-canonical Hamiltonian formulations. Our schemes apply to either formulations, but the canonical formulation results in a particularly efficient scheme that also conserves particle density and energy exactly. The collisionless update is coupled to a Bhatnagar-Gross-Krook (BGK) collision operator that provides a simplified model for relaxation to local thermodynamic equilibrium. An iterative scheme is constructed to ensure collisional invariants (density, momentum and energy) are preserved numerically. Rotation of the manifold is incorporated by modifying the Hamiltonian while ensuring a canonical formulation. Several test problems, including a kinetic version of the classical Sod-shock problem, Kelvin-Helmholtz instability on the surfaces of a sphere and a paraboloid, with and without rotations, is presented. A prospectus for further development of this approach to simulation of kinetic theory in general relativity is presented.