arXiv:2512.22930v3 Announce Type: replace
Abstract: In this paper, we show that the equational theory of relational Kleene algebra with the graph loop operator (a.k.a. fixset) is PSpace-complete. Here, the graph loop is the unary operator that restricts a binary relation to the identity relation. We further show that this PSpace-completeness still holds by extending the terms with top, tests, converse, and nominals, over relational models. Notably, for Kleene algebra with tests (KAT), while the equational theory of relational KAT with antidomain is ExpTime-complete, we show that the equational theory of relational KAT with domain is PSpace-complete, thereby resolving a problem left open in previous works.
To this end, we introduce a novel automaton model on relational structures (graphs), called loop-automata. Loop-automata extend nondeterministic finite automata with a transition type that tests whether the current vertex has a loop. Using this model, we can give a polynomial-time reduction from the equational theories above to the language inclusion problem for 2-way alternating automata.
Science Journals
arXiv:2512.24168v2 Announce Type: replace
Abstract: Turbulence is omnipresent in the atmosphere and a long-standing scientific conundrum that makes flight complex. This complexity is little understood; surprisingly, when turbulence arises, air vehicles struggle while birds seem to thrive. Birds often encounter intense turbulence during takeoff and landing, because of turbulent boundary layer effects. During landing, birds respond by fanning their tail over a wide range of spreads and angles of attack. How their tail functions aerodynamically under these conditions is little understood. Here, we use a bio-hybrid feathered robot model of a pigeon tail in a wind tunnel to compare its aerodynamics in laminar versus turbulent flow. We measured the lift and drag forces generated by the tail as a function of angle of attack, tail spread, and flow condition. We found tail spread scarcely changes tail aerodynamic lift and drag force coefficients, despite large aspect ratio variations. Consequently, tail spread primarily changes force via tail area modulation, simplifying flight control. The effect of laminar versus turbulent flow is pronounced; at the same tail spread and angle of attack, turbulence increases lift and drag by approximately a factor two. Quantitative flow measurement analysis with proper orthogonal decomposition shows force enhancement is linked to modifications in the spatial and temporal structure of the wake. The results suggest a wake instability that arises in laminar flow is suppressed in turbulent flow, which enhances tail efficiency, benefiting flight control. These insights may inspire engineers to design aerial vehicle tails with improved flight control in turbulence.
arXiv:2607.17462v1 Announce Type: new
Abstract: A platform-specific API is implemented for a particular platform (e.g., operating system), thus, it may not work on other platforms than the target one. Detecting the usage of such APIs is important for supporting software maintenance, as it allows maintainers to be alerted about APIs that could pose potential risks. This paper proposes PSASpotter, a tool to detect the usage of platform-specific APIs in Python systems. PSASpotter also identifies whether the platform-specific APIs are used within a defensive code, such as try/except blocks or if blocks that check the current platform. PSASpotter can support the development of novel empirical studies about the usage of platform-specific APIs in the Python ecosystem. Moreover, the defensive code detected by PSASpotter may contain alternative solutions for unavailable APIs, which can provide insights for software development and testing across multiple platforms. PSASpotter is available at: https://github.com/ricardojob/PSASpotter. Tool video: https://youtu.be/d3WyozTAKS8.
arXiv:2410.01021v4 Announce Type: replace
Abstract: Chains of co-B\"uchi automata (COCOA) have recently been introduced as a new canonical representation of omega-regular languages. The co-B\"uchi automata in a chain assign each omega-word its natural color, which depends only on the language itself and not on the chosen automaton representation. Automata in such a chain can be minimized in polynomial time and are good-for-games, making this representation attractive for verification and reactive synthesis. However, in these applications, specifications are usually given in linear temporal logic (LTL). To make COCOA useful, an LTL specification must first be translated into the chain of automata. The only translation currently known proceeds via deterministic parity automata (LTL$\,{\to}\,$DPA$\,{\to}\,$COCOA), where the first step ignores natural colors and requires involved constructions due to Safra or Esparza et al. This raises the question of whether, by exploiting the definition of the natural color of words, one can avoid such constructions and obtain a direct translation from LTL to COCOA.
In this paper, we present a simple yet optimal translation from LTL to COCOA, as well as a variant that translates LTL into DPA. The translation represents a new path from LTL to DPA and exploits the definition of natural colors. It relies on standard operations on weak alternating automata, the Miyano-Hayashi breakpoint construction, the subset construction, and simple graph algorithms. Starting from weak alternating automata, the procedure also applies to specifications in linear dynamic logic. The procedure runs in asymptotically optimal doubly exponential time and produces automata of asymptotically optimal size.
arXiv:2512.24679v2 Announce Type: replace
Abstract: Intelligent fault diagnosis has become an indispensable technique for ensuring machinery reliability. However, existing methods suffer significant performance decline in real-world scenarios where models are tested under unseen working conditions, while domain adaptation approaches are limited to their reliance on target domain samples. Moreover, most existing studies rely on single-modal sensing signals, overlooking the complementary nature of multi-modal information for improving model generalization. To address these limitations, this paper proposes a multi-modal cross-domain mixed fusion model with dual disentanglement for fault diagnosis. A dual disentanglement framework is developed to decouple modality-invariant and modality-specific features, as well as domain-invariant and domain-specific representations, enabling both comprehensive multi-modal representation learning and robust domain generalization. A cross-domain mixed fusion strategy is designed to randomly mix modality information across domains for modality and domain diversity augmentation. Furthermore, a triple-modal fusion mechanism is introduced to adaptively integrate multi-modal heterogeneous information. Extensive experiments are conducted on induction motor fault diagnosis under both unseen constant and time-varying working conditions. The results demonstrate that the proposed method consistently outperforms advanced methods and comprehensive ablation studies further verify the effectiveness of each proposed component and multi-modal fusion. The code is available at: https://github.com/xiapc1996/MMDG.
arXiv:2601.02451v2 Announce Type: replace
Abstract: Graph Neural Networks (GNNs) suffer from over-smoothing in deep architectures and expressiveness bounded by the 1-Weisfeiler-Leman (1-WL) test. We adapt Manifold-Constrained Hyper-Connections, recently proposed for Transformers, to graph neural networks. Our method, \mhcgnn{}, expands node representations across $n$ parallel streams and constrains stream-mixing matrices to the Birkhoff polytope of doubly stochastic matrices via Sinkhorn-Knopp normalization. We prove that \mhcgnn{} mitigates over-smoothing via a layer-wise residual lower bound showing that node-pair differences decay at rate $(1{-}\varepsilon)^L$ (where $\varepsilon$ measures deviation of the mixing matrix from identity), far slower than the standard $(1{-}\gamma)^L$ collapse rate driven by the spectral gap $\gamma$. This two-regime analysis, via the protected orthogonal subspace for $L < n$ and the layer-wise contraction for $L \geq n$, provides architecture-agnostic rate guarantees absent from prior methods. With independent random stream initialization, \mhcgnn{} can distinguish graphs beyond 1-WL by maintaining stream diversity across layers via doubly stochastic mixing. Depth experiments spanning 2 to 128 layers reveal that standard GNNs collapse to near-random performance beyond 16 layers, while \mhcgnn{} maintains over 74\% accuracy at 128 layers, with improvements exceeding 50 percentage points at extreme depths. Ablations confirm that manifold constraints are essential: removing them causes up to 82\% performance degradation. Experiments on heterophilic graphs (roman-empire, penn94, genius) and expressiveness benchmarks (EXP) further validate the contribution. Code is available at https://github.com/smishra-lab/mhc-gnn
arXiv:2601.04231v3 Announce Type: replace
Abstract: Wastewater surveillance, which regularly measures pathogen biomarkers in wastewater samples, is a valuable tool for monitoring infectious diseases circulating in communities. Yet, most wastewater-based epidemiology methods that use wastewater surveillance results to infer disease trends implicitly assume that individuals excrete only at their residential locations and that the populations contributing to wastewater samples are static. These simplifying assumptions ignore daily mobility, social interactions, and heterogeneous toilet-use patterns, which can bias the interpretation of wastewater results, especially at upstream sampling locations such as neighborhoods, institutions, or buildings. Here, we introduce an agent-based geospatial simulation framework. Building on an established Patterns of Life model, we simulate daily human activities within a realistic urban environment and extend the framework with a physiologically motivated defecation cycle and toilet-use patterns. We couple this behavioral model with an infectious disease model to simulate transmission through spatial and social interactions. When an infected agent defecates, a pathogen-shedding model determines the amount of pathogen released in the feces. By integrating population mobility, disease transmission, toilet-use behavior, and pathogen shedding, the framework can simulate the spatiotemporal dynamics of wastewater pathogen loads. Using a case study of 10,000 simulated agents in Fulton County, Georgia, we examine how varying infection rates alter epidemic trajectories, wastewater pathogen loads, and the spatial distribution of pathogen shedding over time. Our results show that mobility and toilet use can substantially decouple residential disease prevalence from wastewater pathogen loads and demonstrate how behaviorally grounded simulations can support interpretation, scenario analysis, and wastewater surveillance
arXiv:2601.07121v2 Announce Type: replace
Abstract: Large language models (LLMs) are increasingly used not only for problem solving but also for creative ideation; however, generating ideas that are both novel and coherent remains challenging. While high-temperature sampling can promote originality, it often compromises consistency and usefulness. Here, we propose ReMIND, a four-stage framework comprising wake, which establishes a stable semantic baseline through low-temperature generation; dream, which performs high-temperature exploratory generation; judge, which evaluates candidate outputs for consistency and extracts salient novel ideas; and rewake, which consolidates selected ideas into coherent final outputs. By assigning these functions to independent LLM modules, ReMIND explicitly separates exploration from stabilization. We systematically evaluated diverse model configurations using multiple creative ideation tasks. External evaluations showed that novelty enhancement could emerge through two separable stages: first from the wake to the dream phase through high-temperature exploration, and subsequently from the dream to the rewake phase through judge-mediated selection and consolidation. Notably, different LLM families exhibited distinct functional tendencies. These findings suggest that serendipitous ideation in LLMs is not solely a property of individual models but emerges from interactions among models assigned to specialized cognitive roles. ReMIND provides a general framework for investigating computational creativity and demonstrates how modular LLM orchestration can bridge exploratory generation and coherent idea formation.
arXiv:2607.17327v1 Announce Type: cross
Abstract: Quantum machine learning models define probabilistic input--output maps through coherent quantum evolution and measurement. While such models can exhibit computational advantages, their internal functioning and decision making generally resists interpretation in terms of stochastic trajectories through intermediate configurations. In contrast to classical (Markovian) stochastic processes, quantum dynamics generically violates the Chapman--Kolmogorov divisibility condition, preventing a decomposition into probabilistically meaningful intermediate transitions. We develop a probabilistic framework for representing quantum learning models as stochastic processes over configuration spaces where the dynamics are modeled as linear maps on probability distributions. Starting from a fixed POVM, arbitrary quantum channels induce transition kernels on the associated probability representation. For informationally complete POVMs, and in particular SIC-POVMs, these kernels are Markovian but generally quasi-stochastic, with non-classicality appearing as negativity. By contrast, projective spaces admit positive stochastic kernels but generally require non-Markovian dynamics due to the failure of Chapman--Kolmogorov divisibility. This yields a trade-off between negativity and dependence on past configurations, i.e. quantum dynamics can be represented either by Markovian quasi-stochastic maps or by positive stochastic processes with higher Markov order. We discuss how such representations of quantum dynamics can be interpreted as stochastic walks through a memory space in the spirit of Projective Simulation, a model of learning and agency in which decisions arise from random walks over an episodic memory network. We further outline how finite-order stochastic kernels can approximate such quantum deliberation processes and show in what regimes the classical machine learning model is recovered.
arXiv:2607.17365v1 Announce Type: cross
Abstract: High-contrast imaging demands extremely sensitive wavefront sensing to correct atmospheric effects and surface errors. While the limits of the sensitivity of a wavefront sensor are well known, a practical, robust design that saturates these limits remains elusive. This work further investigates the PIAA-ZWFS (Phase-Induced Amplitude Apodization-Zernike Wavefront Sensor). In previous work, we developed a framework to optimise its design, maximising Fisher information per frame in the presence of phase aberrations. In these proceedings, we study the effect of various manufacturing and alignment errors in the system on the overall performance. We employ both traditional Monte Carlo sampling and a 2nd order expansion using our auto-differentiable simulator. The performance of the PIAA-ZWFS is not significantly degraded by these errors at values typical for manufacturing.
arXiv:2507.05278v4 Announce Type: replace
Abstract: The TRIUMF Ultracold Advanced Neutron (TUCAN) collaboration has been developing a high-intensity ultracold neutron (UCN) source aimed at searching for the neutron electric dipole moment (EDM) with a sensitivity goal of $10^{-27}\ e{\rm cm}$. This article reports on recent progress in commissioning of the UCN source and in the development of the neutron EDM spectrometer. In its final configuration, the accelerator-driven super-thermal UCN source will enable a neutron EDM experiment with two orders of magnitude improved statistics compared to the current best experiment. Substantial progress in 2024 allowed the collaboration to operate the complete source system, with the exception of the liquid deuterium cold moderator, resulting in the first production of UCNs. The status of the EDM spectrometer is also presented, with emphasis on UCN handling components and magnetic subsystems relevant to field control, shielding, and magnetometry.
arXiv:2601.10379v2 Announce Type: replace
Abstract: Sparse regression provides a compact and interpretable route for nonlinear system modeling by selecting a small number of active terms from a candidate dictionary. Most sparse regressors, however, are constructed offline and then used as static predictors. In online operation, changing load, material properties, ambient conditions, or equipment states may alter both the coefficient values and the effective active support within the dictionary. Moreover, a direct recursive update over a rich dictionary may spread the adaptation over many weakly relevant terms, causing an initially sparse model to become increasingly dense. The key problem is therefore to maintain a sparse regressor online, so that it can absorb streaming data while keeping a compact but revisable active structure. This paper develops a Bayesian recursive sparse learning (BRSL) method for online sparse identification over candidate dictionary terms. The coefficient distribution is updated through a Bayesian posterior recursion, where sliding-window likelihood-ratio information recursion incorporates new samples, removes expired samples, and discounts historical information in a unified update. To preserve sparsity during recursion, posterior-guided shrinkage is introduced to suppress weakly supported dictionary terms and revise the active structure according to posterior evidence. The posterior update is performed in a candidate subspace with an adaptive information floor to keep the recursive solve well posed, and a bounded-error relation is given to clarify the influence of shrinkage, residual information, coefficient drift, and information conditioning. The proposed method is evaluated on sparse coefficient tracking and a power-plant-oriented multi-input multi-output (MIMO) nonlinear time-varying identification benchmark.
arXiv:2601.10922v2 Announce Type: replace
Abstract: We study data curation for multimodal reasoning in a fixed-protocol fine-tuning regime, where the base model, optimizer, training schedule, and evaluation pipeline are held constant and the main degree of freedom is the training data. Using the NeurIPS 2025 Data Curation for Vision--Language Reasoning (DCVLR) challenge as a controlled testbed, we analyze how source-dataset alignment, model-relative difficulty, dataset size, diversity heuristics, and rewritten synthetic mixtures affect downstream reasoning accuracy. Among the tested interventions, difficulty filtering on an aligned source corpus provides the strongest gains at matched scale. The effect is not explained only by LiveXivTQA weighting: a per-benchmark decomposition shows that much of the improvement over random sampling comes from OlympiadBench, the largest non-LiveXivTQA benchmark. Qwen-derived difficulty scores also transfer to some additional model families, though the benefit is architecture-dependent. In contrast, increasing dataset size beyond roughly 1k aligned examples mainly reduces run-to-run variance under the fixed recipe, while the diversity and rewritten CoSyn mixtures we tested do not improve over the difficulty-filtered baseline. These results provide a scoped empirical recipe for data-constrained multimodal reasoning fine-tuning, rather than a universal claim about data selection across all training regimes.
arXiv:2601.16991v3 Announce Type: replace
Abstract: Adapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense weight updates, which hinders their use in resource-constrained environments. Low-rank Adaptation (LoRA) reduces trainable parameters by factorizing weight updates, yet the underlying dense weights still impose high storage and computation costs. Magnitude-based pruning can yield sparse models but typically degrades LoRA's performance when applied naively. In this paper, we introduce SALR (Sparsity-Aware Low-Rank Representation), a novel fine-tuning paradigm that unifies low-rank adaptation with sparse pruning under a rigorous mean-squared-error framework. We prove that statically pruning only the frozen base weights minimizes the pruning error bound, and we recover the discarded residual information via a truncated-SVD low-rank adapter, which provably reduces per-entry MSE by a factor of $(1 - r/\min(d,k))$. To maximize hardware efficiency, we fuse multiple low-rank adapters into a single concatenated GEMM, and we adopt a bitmap-based encoding with a two-stage pipelined decoding + GEMM design to achieve true model compression and speedup. Empirically, SALR attains 50\% sparsity on various LLMs while matching the performance of LoRA on GSM8K and MMLU, reduces model size by $2\times$, and delivers up to a $1.7\times$ inference speedup.
arXiv:2601.17950v2 Announce Type: replace
Abstract: The space of task-agnostic feature upsampling has emerged as a promising area of research to efficiently create denser features from pre-trained visual backbones. These methods act as a shortcut to achieve dense features for a fraction of the cost by learning to map low-resolution features to high-resolution versions. While early works in this space used iterative upsampling approaches, more recent works have switched to cross-attention-based methods, which risk falling into the same efficiency scaling problems of the backbones they are upsampling. In this work, we demonstrate that iterative upsampling methods can still compete with cross-attention-based methods; moreover, they can achieve state-of-the-art performance with lower inference costs. We propose UPLiFT, an architecture for Universal Pixel-dense Lightweight Feature Transforms. We also propose an efficient Local Attender operator to overcome the limitations of prior iterative feature upsampling methods. This operator uses an alternative attentional pooling formulation defined fully locally. We show that our Local Attender allows UPLiFT to maintain stable features throughout upsampling, enabling state-of-the-art performance with lower inference costs than existing pixel-dense feature upsamplers. In addition, we apply UPLiFT to generative downstream tasks and show that it achieves competitive performance with state-of-the-art Coupled Flow Matching models for VAE feature upsampling. Altogether, UPLiFT offers a versatile and efficient approach to creating denser features.
arXiv:2601.18823v4 Announce Type: replace
Abstract: Variational autoencoders (VAE) encode data into lower-dimensional latent vectors before decoding those vectors back to data. Once trained, one can hope to detect out-of-distribution (abnormal) latent vectors, but several issues arise when the latent space is high dimensional. This includes an exponential growth of the hypervolume with the dimension, which severely affects the generative capacity of the VAE. In this paper, we draw insights from high dimensional statistics: in these regimes, the latent vectors of a standard VAE are distributed on the `equators' of a hypersphere, challenging the detection of anomalies. We propose to formulate the latent variables of a VAE using hyperspherical coordinates, which allows compressing the latent vectors towards a given direction on the hypersphere, thereby allowing for a more expressive approximate posterior. We show that this improves both the fully unconditional-OOD and conditional-OOD anomaly detection ability of the VAE, achieving the best performance on the datasets we considered, outperforming existing methods. For the unconditional-OOD and conditional-OOD modalities, respectively, these are: i) detecting unusual landscape from the Mars Rover camera and unusual Galaxies from ground based imagery (complex, real world datasets); ii) standard benchmarks like Cifar10 and subsets of ImageNet as the in-distribution (ID) class.
arXiv:2601.21372v3 Announce Type: replace
Abstract: We present NEMO, a system that translates Natural-language descriptions of decision problems into formal Executable Mathematical Optimization implementations using autonomous coding agents (ACAs). Existing approaches rely on specialized large language models (LLMs) or bespoke task-specific agents that are often brittle and frequently generate syntactically invalid or non-executable code. NEMO instead treats ACAs as a first-class abstraction analogous to API-based interaction with LLMs; their sandboxed execution guarantees code is executable by construction and supports automated validation and repair. We introduce novel coordination patterns including asymmetric validation loops between independently generated optimizer and simulator implementations, external memory for experience reuse, and robustness enhancements via minimum Bayes risk (MBR) decoding and self-consistency. Across nine established optimization benchmarks, NEMO achieves state-of-the-art performance on the majority of tasks with substantial margins on several datasets, demonstrating the power of execution-aware agentic architectures for automated optimization modeling.
arXiv:2510.05743v3 Announce Type: replace
Abstract: We review the historical development and current trends of artificially intelligent agents (agentic AI) in the social and behavioral sciences: from the first programmable computers, and social simulations soon thereafter, to today's experiments with large language models. This overview emphasizes the role of AI in the scientific process and the changes brought about, both through technological advancements and the broader evolution of science from around 1950 to the present. Some of the specific points we cover include: the challenges of presenting the first social simulation studies to a world unaware of computers, the rise of social systems science, intelligent game theoretic agents, the age of big data and the epistemic upheaval in its wake, and the current enthusiasm around applications of generative AI, and many other topics. A pervasive theme is how deeply entwined we are with the technologies we use to understand ourselves.
arXiv:2601.21428v2 Announce Type: replace
Abstract: In this work (Part I), we study three time-discretization schemes for the Dynamical Low-Rank Approximation (DLRA) of high-dimensional stochastic differential equations (SDEs). Specifically, we consider the Dynamically Orthogonal (DO) method for DLRA proposed and analyzed in arXiv:2308.11581v4, which approximates the true solution by a linear combination of few products between deterministic orthonormal modes and stochastic modes, both time-dependent. The first scheme considered consists in a forward discretization in time of both deterministic and stochastic components, in a Euler-Maruyama style. Its convergence is proven subject to a time-step restriction dependent on the smallest singular value of the Gram matrix associated to the stochastic modes, which, on its turn, is shown to be always positive, provided that the SDE under study is driven by a non-degenerate noise. The second and the third schemes, on the other hand, are staggered ones, alternating updates of the deterministic and the stochastic modes in half steps, and have a projector splitting nature. We show stability of the second scheme and prove convergence with constants independent of the smallest singular value. The third scheme works better in practice, although our theoretical convergence bounds are worse than those for the second one. Computational experiments support our theoretical results. In this work we do not consider the discretization in probability, which will be the topic of Part II.
arXiv:2601.22259v2 Announce Type: replace
Abstract: While tabular foundation models have achieved remarkable success in classification and regression, adapting them to model time-to-event outcomes for survival analysis is non-trivial due to right-censoring, where data observations may end before the event of interest occurs. We utilize a classification-based framework that reformulates both static and dynamic survival analysis as a series of binary classification problems by discretizing event times. Censored observations are naturally handled as examples with missing labels at certain time points. This classification formulation enables existing tabular foundation models (TFMs) to perform survival analysis through in-context learning without explicit training. In contrast to classical approaches that use binary classifiers to model discrete-time hazards, our approach directly models cumulative failure probabilities, which we find empirically to be more robust to the number of discretization bins by avoiding multiplicative accumulation of per-bin errors. We prove that under standard censoring assumptions, minimizing our binary classification loss recovers the true survival probabilities as the training set size increases. We demonstrate through evaluation across 48 real-world datasets (43 static and 5 dynamic) that off-the-shelf TFMs with this classification formulation outperform classical and deep learning baselines on average over multiple survival metrics.
arXiv:2602.02025v2 Announce Type: replace
Abstract: ML models critically depend on feature quality, yet in real-world settings, useful features are often distributed across multiple relational tables rather than a single dataset. Feature augmentation addresses this problem by automatically discovering and joining additional tables to enrich a base table with predictive features. However, scaling feature augmentation to complex schemas with many tables and multi-hop relationships is challenging. It requires exploring a large space of join paths, executing costly joins, and selecting useful features from noisy results. Existing approaches suffer from either limited effectiveness or efficiency. Restricting exploration to simple joins limits predictive performance, while more expressive methods rely on expensive training data, lack scalability, or fail to fully exploit schema-level semantics. We present Hippasus, a cost-aware, LLM-augmented feature discovery framework over relational schemas that addresses these challenges. Hippasus combines lightweight statistical signals with adaptive semantic reasoning, invoking stronger (LLM-based) analysis only when necessary. It further introduces efficient multi-way join execution with cross-path feature consolidation, and a hybrid feature selection strategy that integrates statistical relevance with semantic refinement. Experiments on real-world datasets show that Hippasus improves feature augmentation accuracy by up to 26.8% over state-of-the-art methods, while achieving a favorable effectiveness-cost tradeoff.
arXiv:2602.02138v3 Announce Type: replace
Abstract: Despite the remarkable success that Multi-Agent Code Generation Systems (MACGS) have achieved, the inherent complexity of multi-agent architectures produces substantial volumes of intermediate outputs. To date, the individual importance of these intermediate outputs to the system correctness remains opaque, which impedes targeted optimization of MACGS designs. To address this challenge, we propose CAM, the first \textbf{C}ausality-based \textbf{A}nalysis framework for \textbf{M}ACGS that systematically quantifies the contribution of different intermediate features for system correctness. By comprehensively categorizing intermediate outputs and systematically simulating realistic errors on intermediate features, we identify the important features for system correctness and aggregate their importance rankings.
We conduct extensive empirical analysis on the identified importance rankings. Our analysis reveals intriguing findings: first, we uncover context-dependent features\textemdash features whose importance emerges mainly through interactions with other features, revealing that quality assurance for MACGS should incorporate cross-feature consistency checks; second, we reveal that hybrid backend MACGS with different backend LLMs assigned according to their relative strength achieves up to 7.3\% Pass@1 improvement, underscoring hybrid architectures as a promising direction for future MACGS design. We further demonstrate CAM's practical utility through two applications: (1) failure repair which achieves a 73.6\% success rate by optimizing top-3 importance-ranked features and (2) feature pruning that reduces up to 33.6\% intermediate token consumption while maintaining generation performance. Our work provides actionable insights for MACGS design and deployment, establishing causality analysis as a powerful approach for understanding and improving MACGS.
arXiv:2602.03061v2 Announce Type: replace
Abstract: Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-variance accuracy estimates and unstable rankings across platforms. On difficult problems, an LLM may fail to produce a correct final answer, yet still provide reliable pairwise comparison signals indicating which of two candidate solutions is better. We leverage this observation to design a statistically efficient evaluation framework that combines standard labeled outcomes with pairwise comparison signals obtained by having models judge auxiliary reasoning chains. Treating these comparison signals as control variates, we develop a semiparametric estimator based on the efficient influence function (EIF) for the setting where auxiliary reasoning chains are observed. This yields a one-step estimator that achieves the semiparametric efficiency bound, guarantees strict variance reduction over naive sample averaging, and admits asymptotic normality for principled uncertainty quantification. Across simulations, our one-step estimator substantially improves ranking accuracy, with gains increasing as model output noise grows. Experiments on GPQA Diamond, AIME 2025, and GSM8K further demonstrate more precise performance estimation and more reliable model rankings, especially in small-sample regimes where conventional evaluation is pretty unstable.
arXiv:2602.04976v3 Announce Type: replace
Abstract: The TRISTAN detector is an upgrade to the KATRIN experiment to enable a differential measurement of the tritium $\beta$-decay spectrum to search for sterile neutrinos with keV masses. This entails performing precision electron spectroscopy with over one thousand silicon drift detector pixels, each responsible for recording incident electron rates of $10^5$ counts per second. A project specific data acquisition (DAQ) system is developed to meet the experimental challenges through a remote analog to digital conversion (RADC) design. In this work, the conceptual design of the RADC DAQ is presented along with the built system for operating the TRISTAN detector upgrade. The system includes flexible signal processing logic and data management that is optimized for the high-rate precision measurement.
arXiv:2602.05463v2 Announce Type: replace
Abstract: Modern AI systems achieve remarkable capabilities at the cost of substantial energy consumption. To connect intelligence to physical efficiency, we propose two complementary bits-per-joule metrics under explicit accounting conventions: (1) Thermodynamic Epiplexity per Joule, new bits of structure about a specified environment-instance variable encoded in an agent's state per unit energy, and (2) Empowerment per Joule, sensorimotor channel capacity per expected energetic cost over a fixed horizon. These give two axes of physical intelligence, recognition versus control, but the resulting numbers are benchmark-relative rather than universal. Drawing on stochastic thermodynamics, we formulate a Landauer-scale closed-cycle benchmark for epiplexity acquisition by combining a thermodynamic-learning inequality with data processing, and clarify why boundary closure is required; conversely, a decoupling construction shows that without such assumptions information gain and in-boundary dissipation need not be tightly linked. For empirical settings where the latent structure variable is unavailable, we recommend compute-bounded MDL epiplexity / compression-gain surrogates. Finally, we propose a unified efficiency framework with a minimal checklist of conventions for relative bits-per-joule comparisons, and give a compact language-model reporting example.