arXiv:2605.16208v1 Announce Type: cross Abstract: Flexible continuous-time survival modeling is critical for capturing complex time-varying hazard dynamics in high-dimensional data; however, training such models remains challenging due to the intractable integral required for likelihood estimation. We introduce QSurv, a scalable deep learning framework that enables nonparametric continuous-time modeling without relying on time discretization or restrictive distributional assumptions. We propose a training objective based on Gauss-Legendre numerical quadrature, which approximates the cumulative hazard with high-order accuracy while facilitating efficient end-to-end training via standard backpropagation. Furthermore, to effectively capture non-stationary hazard dynamics in complex architectures, we introduce time-conditioned low-rank adaptation, a mechanism that conditions general neural backbones on time by dynamically modulating weights via low-rank updates. We provide theoretical analysis establishing approximation error bounds for cumulative-hazard evaluation. Comprehensive experiments across synthetic benchmarks, large-scale real-world tabular datasets, and high-dimensional medical imaging tasks demonstrate that QSurv achieves competitive predictive performance with advantages in instantaneous hazard function estimation, enabling more interpretable characterization of time-varying risk patterns.
Science Journals
arXiv:2605.16230v1 Announce Type: cross Abstract: Magnetic order is a fundamental property of materials, governing collective behavior and enabling a broad range of functionalities. Yet magnetic structure remains difficult to determine: experiments are costly and specialized, while first-principles methods often struggle with the noncollinear and incommensurate orders found in real materials. Here we introduce magnetic structure network (MSN), an E(3) equivariant graph neural network that predicts both collinear and non-collinear magnetic structures directly from atomic crystal structures, trained directly on experimentally determined structures from MAGNDATA. By proposing the primitive modulated structure representation (PMSR), we are able to encode commensurate and incommensurate structures in a unified way without symmetry assumptions. The model achieves strong performance across all modulation components and reconstructs experimental magnetic structures with high fidelity. Our approach provides a scalable framework for rapid magnetic structure prediction and opens a route to data-driven discovery of magnetic materials.
arXiv:2605.16159v1 Announce Type: new Abstract: Real-time event detection in IoT mesh sensor networks must balance sensitivity against false-positive load on a constrained mesh radio. We present a Monte Carlo comparison of the Temporal Spectral Noise-Floor Adaptation (TSNFA) detector against four classical comparators drawn from the radar Constant False Alarm Rate (CFAR) family and from sequential change detection: the Lipski FFT energy detector, Cell-Averaging CFAR (CA-CFAR), Ordered-Statistic CFAR (OS-CFAR), and state-machine Cumulative Sum (CUSUM). All five detectors are implemented to fit a Cortex-M0+ class envelope, process a 1-D 100 Hz time series in 128-sample frames, and use temporal reference windows in place of the spatial reference cells of conventional radar CFAR. Across a factorial set of four configurations (10 and 50 nodes; 12 dB and 18 dB SNR), each replicated five times over 24 hours, TSNFA achieves 99.97 to 100% event detection rate with 100% event precision and zero false-positive clusters per node. The classical comparators each succeed on one quality dimension and fail on another. Lipski FFT (k = 3), CA-CFAR, and OS-CFAR all maintain near-perfect detection rate but with event precision below 3% and per-node bandwidth between 145 kB/h and 1.2 MB/h. CA-CFAR and OS-CFAR are indistinguishable in false-alarm performance, both saturating the same broadband-statistic failure mode. CUSUM shows an SNR-dependent detection-rate drop from about 70% at 18 dB to 51% at 12 dB. TSNFA is the only algorithm tested that simultaneously achieves high detection rate, high precision, and low per-node bandwidth.
arXiv:2605.16179v1 Announce Type: new Abstract: Agricultural landscape segmentation in the Global South is challenging as it is characterized by fragmented plots, high intra-class variance, and a scarcity of labeled training data. Recent advances in segmentation have been made by Multimodal Large Language Models (MLLMs). However, current approaches encounter critical context length bottlenecks and a domain alignment gap in understanding satellite features. We address these limitations through MAgSeg, a novel, decoder-free MLLM segmentation approach. MAgSeg is an architecturally efficient approach that enables standard MLLMs to perform segmentation of complex smallholder agricultural landscapes from high-resolution satellite imagery, without requiring auxiliary vision decoders. We introduce a novel instruction tuning data format designed to enable scalable fine-tuning and post-training on high resolution satellite imagery, which enables MAgSeg to learn from the global context of the image while generating text tokens for only a patch within the image. Extensive evaluations on datasets spanning three countries in the Global South demonstrate that MAgSeg significantly outperforms state-of-the-art MLLM baselines, offering a scalable solution to map smallholder agricultural environments.
arXiv:2406.18224v3 Announce Type: replace Abstract: We provide the first fully polynomial-time randomized approximation scheme for the following two counting problems: 1. Given a Context Free Grammar $G$ over alphabet $\Sigma$, count the number of words of length exactly $n$ generated by $G$. 2. Given a circuit $\varphi$ in Decomposable Negation Normal Form (DNNF) over the set of Boolean variables $X$, compute the number of assignments to $X$ such that $\varphi$ evaluates to 1. Finding polynomial time algorithms for the aforementioned problems has been a longstanding open problem. Prior work could either only obtain a quasi-polynomial runtime (SODA 1995) or a polynomial-time randomized approximation scheme for restricted fragments, such as non-deterministic finite automata (JACM 2021) or non-deterministic tree automata (STOC 2021).
arXiv:2504.00663v2 Announce Type: replace Abstract: Existing hardware-aware NAS (HW-NAS) methods typically assume access to precise information circa the target device, either via analytical approximations of the post-compilation latency model, or through learned latency predictors. Such approximate approaches risk introducing estimation errors that may prove detrimental in risk-sensitive applications. In this work, we propose a two-stage HW-NAS framework, in which we first learn an architecture controller on a distribution of synthetic devices, and then directly deploy the controller on a target device. At test-time, our network controller deploys directly to the target device without relying on any pre-collected information, and only exploits direct interactions. In particular, the pre-training phase on synthetic devices enables the controller to design an architecture for the target device by interacting with it through a small number of high-fidelity latency measurements. To guarantee accessibility of our method, we only train our controller with training-free accuracy proxies, allowing us to scale the meta-training phase without incurring the overhead of full network training. We benchmark on HW-NATS-Bench, demonstrating that our method generalizes to unseen devices and searches for latency-efficient architectures by in-context adaptation using only a few real-world latency evaluations at test-time.
arXiv:2503.16589v2 Announce Type: replace Abstract: A key trait of stochastic optimizers is that multiple runs of the same optimizer in attempting to solve the same problem can produce different results. As a result, their performance is evaluated over several repeats, or runs, on the problem. However, the accuracy of the estimated performance metrics depends on the number of runs and should be studied using statistical tools. We present a statistical analysis of the common metrics, and develop guidelines for experiment design to measure the optimizer's performance using these metrics to a high level of confidence and accuracy. To this end, we first discuss the confidence interval of the metrics and how they are related to the number of runs of an experiment. We then derive a lower bound on the number of repeats in order to guarantee achieving a given accuracy in the metrics. Using this bound, we propose an algorithm to adaptively adjust the number of repeats needed to ensure the accuracy of the evaluated metric. Our simulation results demonstrate the utility of our analysis and how it allows us to conduct reliable benchmarking as well as hyperparameter tuning and prevent us from drawing premature conclusions regarding the performance of stochastic optimizers.
arXiv:2212.13347v3 Announce Type: replace Abstract: We report a diffuse Maxwellian illumination scheme for wide-field retinal laser Doppler holography. Inserting an engineered diffuser in the illumination arm transforms a spatially concentrated near-infrared laser focus into an angularly diversified illumination pattern, thereby reducing local irradiance near the anterior segment while preserving coherent interferometric detection. This configuration allows the eyepiece to be positioned closer to the cornea, increasing the digitally reconstructed retinal field of view without producing a localized corneal hot spot. We compare three illumination geometries: focused non-diffuse illumination, diffuse illumination at the same cornea--eyepiece distance, and diffuse Maxwellian illumination. Diffuse Maxwellian illumination expands the retinal field of view while preserving Doppler contrast in broad and high-frequency fluctuation bands. Light-hazard assessment is limited to the current ophthalmic standards ISO 15004-2:2024 and ANSI Z80.36-2021. Based on measured beam profiles, the recommended operating power at 852 nm is set by the most restrictive relevant exposure condition among the assessed anterior-segment, iris, and retinal limits. These results support diffuse illumination as a practical route toward safer, non-mydriatic, wide-field Doppler holography of the human retina.
arXiv:2605.16236v1 Announce Type: cross Abstract: We theoretically investigate acoustic spin resonance in a spatially homogeneous spinor polariton condensate. A longitudinal acoustic wave generates a time-periodic strain-induced effective magnetic field acting on the condensate pseudospin. When this field is transverse to the static in-plane linear-polarization splitting, it resonantly drives polarization oscillations. We show that spin-dependent interactions shift the resonance and produce nonlinear line shapes, while gain, reservoir dynamics, and spin relaxation make the response dissipative and history-dependent, producing amplitude hysteresis. In the presence of lifetime anisotropy, the condensate can develop a bifurcated stationary state with finite circular polarization, and a resonant acoustic drive can switch between the corresponding out-of-plane branches. A Zeeman splitting provides an additional conservative knob for tuning the resonance frequency. Our results identify coherent acoustic driving as a route to resonant, nonlinear, and switchable control of polariton pseudospin dynamics.
arXiv:2605.15734v1 Announce Type: new Abstract: The use of large language models to assess user states in conversational and adaptive systems is based on the assumption that the metrics used for such assessment are stable and interpretable at the level of individual scores. This paper empirically tests this assumption, focusing on the psychometric reliability of artificial intelligence (AI) measures of user states. This study employed replication evaluation procedures to assess the repeatability of a broad set of metrics across three different bimodal large language models (GPT-4o audio, Gemini 2.0 Flash, Gemini 2.5 Flash). Analyses include both individual score reliability and aggregated reliability, allowing us to distinguish metrics potentially useful for real-time adaptation from those that retain their value only in aggregated analyses. The results demonstrate that metric reliability cannot be considered a default property in interpretive domains. The lack of stability at the level of individual scores precludes the interpretation of such scores as indicators of user state in real-time adaptive systems, even if these metrics demonstrate stability after aggregation. At the same time, the study indicates that individually unstable metrics can retain analytical utility in post-hoc studies, identifying rules governing interactions and their relationships with user experience parameters such as satisfaction, trust, and engagement. The main contribution of this work, besides quantifying the severity of the problem (only 31 of 213 metrics met the criteria), is the proposal of a replicable evaluation framework, enabling measurable evaluations of metric applicability. This approach supports more responsible AI design of adaptive systems, in which the interpretation of results requires explicit validation of reliability and monitoring for violations over time.
arXiv:2512.04457v2 Announce Type: replace Abstract: Removing specific data influence from large language models (LLMs) remains challenging, as retraining is costly and existing approximate unlearning methods are often unstable. The challenge is exacerbated when the forget set is small or imbalanced. We introduce RapidUn, an influence-driven and parameter-efficient unlearning framework. It first estimates per-sample influence through a fast estimation module, then maps these scores into adaptive update weights that guide selective parameter updates -- forgetting harmful behavior while retaining general knowledge. On Mistral-7B and Llama-3-8B across Dolly-15k and Alpaca-57k, RapidUn achieves up to 100 times higher efficiency than full retraining and consistently outperforms Fisher, GA, and LoReUn on both in-distribution and out-of-distribution forgetting. These results establish influence-guided parameter reweighting as a scalable and interpretable paradigm for LLM unlearning.
arXiv:2503.14852v2 Announce Type: replace Abstract: Machine learning (ML) has shown promise in vulnerability detection, but ML detectors may rely on irrelevant code features, causing them to highlight non-vulnerable lines as suspicious. Such misleading predictions increase developers' manual effort and may lead to incorrect patching strategies, motivating the need to identify untrustworthy predictions automatically. We present UntrustVul, an approach for detecting untrustworthy vulnerability predictions by identifying suspicious lines that are inherently unrelated to vulnerabilities. UntrustVul leverages patterns from historical vulnerable lines and flags predictions as untrustworthy when the highlighted lines neither match known vulnerability patterns nor influence lines that do. A line is considered vulnerability-irrelevant if it does not resemble historical vulnerabilities and all its successors in the data and control dependency graph are also vulnerability-irrelevant. The approach is designed conservatively to minimise misclassifying trustworthy predictions as untrustworthy. We evaluate UntrustVul on 115K predictions from four models across the BigVul, MegaVul, SARD, and PrimeVul datasets. Results show that UntrustVul achieves AUC scores of 70%-88% and F1-scores of 82%-94%, outperforming existing approaches by 6%-59% in AUC and 13%-92% in F1-score.
arXiv:2603.18527v4 Announce Type: replace Abstract: High-frequency Helmholtz problems in heterogeneous media remain challenging for both classical iterative methods and end-to-end neural PDE solvers. We propose Neural Preconditioned Born Series (NPBS), a learned iterative preconditioning framework that operates in preconditioned residual coordinates induced by the Convergent Born Series (CBS). Existing learned Born-series methods primarily use Born-style unrolling for forward wavefield prediction, while learned Helmholtz preconditioners are usually formulated in physical residual coordinates. NPBS fills this gap by recasting Born-series iteration as shifted-Laplacian left preconditioning, and replacing the CBS preconditioner with a learned residual-to-correction map in the Born-preconditioned coordinates. The left preconditioner further induces a residual metric, which yields a metric-matched training objective that aligns optimization with the preconditioned geometry used at inference. On heterogeneous Helmholtz benchmarks, metric-matched NPBS reduces iteration counts by up to $1.9\times$ over direct residual learning, with gains increasing from $1.2\times$ to $1.9\times$ as the wavenumber rises. Compared to classical CBS, learned NPBS reduces stationary iteration counts by over $20\times$; when used as a preconditioner for FGMRES, it further achieves the lowest wall-clock time among all evaluated methods. The same metric-matched formulation also improves convergence on convection--diffusion--reaction systems and Newton linear systems for nonlinear PDEs, indicating that residual-metric matching is a general design principle for neural preconditioners.
arXiv:2604.26139v2 Announce Type: replace Abstract: Diffusion large language models generate text through multi-step denoising, where hallucination signals may emerge throughout the trajectory rather than only in the final output. Existing detectors mainly rely on output uncertainty or coarse trace statistics, which often fail to capture the richer hidden dynamics of D-LLMs. We propose HIVE, a hidden-evidence verification framework that extracts compressed hidden evidence from denoising trajectories, selects informative step-layer evidence, and conditions a verifier language model on the selected evidence through prefix embeddings. HIVE produces both a continuous hallucination score from verifier decision logits and structured verification outputs, including hallucination types, evidence pairs, and short rationales. Across two D-LLMs and three QA benchmarks, HIVE consistently outperforms eight strong baselines and achieves up to 0.9236 AUROC and 0.9537 AUPRC. Ablation studies further confirm the importance of hidden-evidence conditioning, learned evidence selection, two-stream evidence representation, and step-layer embeddings. These results suggest that selected hidden evidence from denoising trajectories provides a stronger and more usable hallucination signal than output-only uncertainty or coarse trace statistics.
arXiv:2605.16118v1 Announce Type: new Abstract: The source distribution in conditional flow matching is a design parameter that can be calibrated to data, not a default isotropic prior. We exploit this in Multi-Fidelity Flow Matching (MFFM), a cascade refinement framework for parametric PDE solutions: the source is calibrated to the empirical low-to-high-fidelity residual scale with local Gaussian-blur correlation, and the velocity network is conditioned on the low-fidelity solution. Conditioning makes the residual refinement problem substantially easier than unconditional field generation, while residual-calibrated source noise improves the flow-matching training geometry. A multi-resolution cascade applies the same construction independently between adjacent fidelities. After level-wise flow-matching pretraining, we fine-tune the composed cascade end-to-end with a deterministic one-step rollout, which makes one velocity evaluation per cascade level the optimized operating point at inference. The result is a learned analog of multigrid refinement that reaches the finest grid in $L$ deterministic network evaluations per query. We validate MFFM on eight benchmarks: two super-resolution problems and six spatiotemporal forecasting tasks from PDEBench, The Well, and the FNO Navier--Stokes dataset.
arXiv:2302.09758v5 Announce Type: replace Abstract: A recently proposed superconducting linear collider with energy recovery (ERLC) and multiple beam reuse employs twin RF structures to eliminate parasitic collisions in the linacs. Such a collider can operate in either pulsed or continuous-wave (CW) mode, achieving a luminosity of ${\cal O}(10^{36})$ cm$^{-2}$s$^{-1}$ at $2E_0$ = 250--500 GeV. This paper demonstrates that in pulsed mode, the ERLC luminosity is independent of the accelerating gradient for a fixed total power, enabling operation at the highest available gradients. A similar independence holds for the CW mode when the available power significantly exceeds the operational threshold. The luminosity scales with the cavity quality factor as $L\propto Q_0^{1/2}$. We also present, for the first time, a study of a twin $e^-e^-$ ERLC and estimate its performance. This configuration is simpler than the $e^+e^-$ version as it eliminates the need for beam recirculation; electrons can be generated anew for each cycle. In this case, the luminosity scales as $L\propto Q_0^{1/4}$. Furthermore, the use of traveling-wave (TW) RF structures allows for higher gradients and reduced thermal loading. We show that an ERLC with $G$ = 40 MeV/m can operate in CW mode, reaching luminosities of $L_{e^+e^-}$= (1-2.5)$\times 10^{36}$ and $L_{e^-e^-}$= (3-7)$\times 10^{36}$ cm$^{-2}$s$^{-1}$ at $2E_0$ = 250 and 500 GeV, respectively, with a total power consumption of 150-300 MW. These results position the ERLC as a highly promising candidate for a future Higgs factory.
arXiv:2504.09544v3 Announce Type: replace Abstract: Recent advances in self-supervised deep learning have improved our ability to quantify cellular morphological changes in high-throughput microscopy screens, a process known as morphological profiling. However, most current methods only learn from images, despite many screens being inherently multimodal, as they involve both a chemical or genetic perturbation as well as an image-based readout. We hypothesized that incorporating chemical compound structures during self-supervised pre-training could improve learned representations of images from high-throughput microscopy screens. We introduce a representation learning framework, MICON (Molecular-Image Contrastive Learning), that models chemical compounds as treatments that induce transformations of cell phenotypes. MICON significantly outperforms classical hand-crafted features such as CellProfiler and existing deep-learning-based representation learning methods in challenging evaluation settings where models must identify reproducible effects of drugs across independent replicates and data-generating centers. We demonstrate that incorporating chemical compound information into the learning process provides small, but consistent improvements in performance and that modeling compounds specifically as treatments outperforms approaches that directly align images and compounds in a single representation space. Our findings point to a new direction for representation learning in morphological profiling, suggesting that methods should explicitly account for the multimodal nature of microscopy screening data.
arXiv:2605.15713v1 Announce Type: new Abstract: Legged manipulators extend robotic capabilities beyond static manipulation by integrating agile locomotion with versatile arm control. However, achieving precise manipulation while maintaining coordinated locomotion remains a major challenge. This work presents a hierarchical reinforcement learning framework for dynamic pick-and-place tasks using a quadruped equipped with a 6-DOF robotic arm. The framework incorporates an explicit mass estimation module enabling adaptive whole-body control for objects with varying weights. In simulation, the system achieves an 86.05% success rate with payloads up to 2.3 kg. The approach is further validated through real-world experiments across six representative scenarios with controlled variations in object physical properties (size and mass) and task heights. Specifically, within a wide vertical workspace ranging from ground level to 1.1~m-high tabletops, the system demonstrates an average success rate of 73.3% for payloads up to 1.3 kg, with an average execution time of 4.06 s. Unlike prior works that handle lightweight objects and execute pick-and-place motions with slow, piecewise motions, the proposed framework exploits concurrent locomotion and manipulation for dynamic, continuous execution. These results demonstrate the potential of quadrupedal mobile manipulators for adaptive, whole-body pick-and-place with heavier payloads and extended workspaces.
arXiv:2605.15266v1 Announce Type: cross Abstract: Preparing arbitrary logical states is a central primitive for universal fault-tolerant quantum computation and the cost of encoded-state preparation contributes directly to the overall resource overhead. This makes the synthesis of efficient general-state encoding circuits an important problem, particularly with respect to two-qubit gate count and circuit depth. Yet the synthesis of such encoders has been studied less extensively than general Clifford circuit synthesis or the preparation of specific logical Pauli-eigenstates. In this work, we develop methods for synthesizing efficient encoders for arbitrary stabilizer codes. We formulate encoder synthesis as a search over stabilizer tableaus and introduce greedy and rollout-based algorithms that exploit the freedom among stabilizer-equivalent realizations of the same encoding isometry. For code families with a modular structure, such as generalized concatenated and holographic codes, we show how large encoders can be assembled from optimized local constituent encoders, and we use SMT-based exact synthesis to obtain optimal local circuits for small instances. We further evaluate the proposed methods on a broad set of stabilizer codes, including holographic and quantum low-density parity-check (qLDPC) codes, and compare them against recent encoder-synthesis methods and existing constructions from the literature, obtaining improvements of up to 43% in two-qubit gate count and up to 70% in depth. Our results support the optimization of encoded-state preparation in several fault-tolerant quantum-computing schemes, and all methods are openly available as part of the Munich Quantum Toolkit.
arXiv:2602.14342v2 Announce Type: replace-cross Abstract: We show that high-accuracy guarantees for log-concave sampling -- that is, iteration and query complexities which scale as $\mathrm{poly}\log(1/\delta)$, where $\delta$ is the desired target accuracy -- are achievable using stochastic gradients with subexponential tails. Notably, this exhibits a separation with the problem of convex optimization, where stochasticity (even additive Gaussian noise) in the gradient oracle incurs $\mathrm{poly}(1/\delta)$ queries. We also give an information-theoretic argument that light-tailed stochastic gradients are necessary for high accuracy: for example, in the bounded variance case, we show that the minimax-optimal query complexity scales as $\Theta(1/\delta)$. Our framework also provides similar high accuracy guarantees under stochastic zeroth order (value) queries, and an improved complexity result for sampling from finite-sum potentials.
arXiv:2510.10454v2 Announce Type: replace Abstract: Large language models (LLMs) offer a generalizable approach for modeling patient trajectories, but suffer from the long and noisy nature of electronic health records (EHR) data in temporal reasoning. To address these challenges, we introduce Traj-CoA, a multi-agent system involving chain-of-agents for patient trajectory modeling. Traj-CoA employs a chain of worker agents to process EHR data in manageable chunks sequentially, distilling critical events into a shared long-term memory module, EHRMem, to reduce noise and preserve a comprehensive timeline. A final manager agent synthesizes the worker agents' summary and the extracted timeline in EHRMem to make predictions. In a zero-shot one-year lung cancer risk prediction task based on five-year EHR data, Traj-CoA outperforms baselines of four categories. Analysis reveals that Traj-CoA exhibits clinically aligned temporal reasoning, establishing it as a promisingly robust and generalizable approach for modeling complex patient trajectories. Implementation of Traj-CoA is available on https://github.com/zengsihang/Traj-CoA.
arXiv:2502.12187v3 Announce Type: replace Abstract: Hallucinations, a phenomenon where a language model (LM) generates nonfactual content, pose a significant challenge to the practical deployment of LMs. While many empirical methods have been proposed to mitigate hallucinations, recent studies established a computability-theoretic result showing that any LM will inevitably generate hallucinations on an infinite set of inputs, regardless of the quality and quantity of training datasets and the choice of the language model architecture and training and inference algorithms. Although the computability-theoretic result may seem pessimistic, its significance in practical viewpoints has remained unclear. This paper claims that those "innate" inevitability results from computability theory and diagonal argument, in principle, cannot explain practical issues of LLMs. We demonstrate this claim by presenting a positive theoretical result from a probabilistic perspective. Specifically, we prove that hallucinations can be made statistically negligible, provided that the quality and quantity of the training data are sufficient. Interestingly, our positive result coexists with the computability-theoretic result, implying that while hallucinations on an infinite set of inputs cannot be entirely eliminated, their probability can always be reduced by improving algorithms and training data. By evaluating the two seemingly contradictory results through the lens of information theory, we argue that our probability-theoretic positive result better reflects practical considerations than the computability-theoretic negative result.
arXiv:2605.03258v2 Announce Type: replace Abstract: Large language models often fail at simple counting tasks, even when items to count are in the prompt. We investigate whether this failure occurs because transformers do not represent counts internally, or because they cannot convert representations to the correct output tokens. Across three model families: Pythia, Qwen3, and Mistral, ranging from 0.4B to 14B parameters, we find evidence for the second explanation. Linear probes recover the correct count from intermediate layers with $R^2>0.99$, showing that the information is present. However, the internal directions that encode counts are nearly orthogonal to digit-token output-head rows ($|\cos| \leq 0.032$). In other words, the model stores the count in a form that the digit logits do not naturally read out. We localize this failure with two interventions. Updating only the digit rows of the output head (36,864 parameters) substantially improves constrained digit prediction (60.7--100.0% on four tasks), but it does not fix unconstrained generation (0%); we do not claim that digit-row repair fixes open-ended text. By contrast, small LoRA on attention Q/V (7.67M parameters) improves upstream routing and achieves 83.1%$\pm$7.2% in true greedy autoregressive generation (deployable fix). Logit-lens at layer 35 (entity counting; correct-digit rank): (i) median over 3 seeds drops from order-$10^4$ to 1; (ii) seed 42 shows $54{,}332 \to 838$ (median top-1 while one seed stays far below). Norm, logit-lens, and cross-task analyses generalize the bottleneck to counting, addition, and list length; nulls on MMLU and GSM8K and limited DROP transfer. These results identify counting failure as a geometric readout bottleneck, not an internal-representation failure: the model knows the count but the output pathway is misaligned with tokens needed to express it.
arXiv:2605.15372v1 Announce Type: cross Abstract: We derive an explicit formula for the intrinsic MacWilliams transform for permutation-invariant qudit codes. Such codes naturally live in symmetric power representations, where the relevant error sectors are determined by the irreducible decomposition of the conjugation action on the associated operator space. Using the multiplicity-free structure of this decomposition and the corresponding intertwiner algebra, we identify the intrinsic MacWilliams matrix with a finite Racah transform. The entries are given by a terminating hypergeometric series, and the rows of the matrix are Racah orthogonal polynomials with parameters determined explicitly by the block length and local dimension. Computing the spectrum of the degree-one twirl reveals that this spectrum lies on an affine quadratic lattice. Then we derive a tridiagonal multiplication rule from the representation theory of the adjoint sector. As consequences, we obtain closed-form orthogonality, detailed-balance, and involutivity identities for the transform. The resulting formula supplies an explicit MacWilliams matrix for computing linear programming bounds on permutation-invariant qudit codes.
arXiv:2605.15350v1 Announce Type: cross Abstract: Stochastic compositional optimization minimizes objectives of the form $\min_{\bm{x} \in \mathcal{X}} F(\bm{f}(\bm{x}), \bm{x})$, where $\bm{f}$ is accessible only through noisy stochastic queries. Existing methods for this problem assume that the outer function $F$ is continuously differentiable, which excludes many practically important applications such as robust max-of-losses, Conditional Value-at-Risk, and norm regularizers. We propose the Hybrid Momentum Stochastic Frank--Wolfe algorithm, which drops the smoothness assumption on $F$. By combining a momentum-based Jacobian tracker with a Taylor-corrected function tracker, the algorithm feeds an entire stochastic linearization -- rather than a single gradient -- into a generalized linear minimization oracle. We establish an $\mathcal{O}(K^{-1/4})$ convergence rate in the generalized Frank--Wolfe gap for non-convex objectives with $L_F$-Lipschitz outer functions, matching the optimal complexity for projection-free single-sample stochastic methods under expected smoothness. The analysis extends to heavy-tailed noise oracles with bounded $r$-th moments for $r \in (1, 2]$ and recovers the deterministic rates of Vladarean et al (2023) as the noise vanishes.