arXiv:2605.19193v1 Announce Type: new
Abstract: Multi-agent LLM debate improves factuality and reasoning, but most recipes pick a fixed round count, over-spending on easy items and under-spending on hard ones. We adapt Wald's Sequential Probability Ratio Test (SPRT) as a plug-in compute governor for LLM debates. After each round, an LLM judge emits a [0,1] consensus score on the latest agent positions; a Wald monitor accumulates the log-likelihood ratio of "useful convergence" vs "not yet useful" under a Beta likelihood family, and stops when either boundary is crossed or returns a capped best-effort outcome at R_max. Under i.i.d. assumptions the rule inherits SPRT type-I/type-II error guarantees; in deployment the calibration itself is the more important object, since it estimates whether the judge score actually separates useful from unhelpful convergence in a given domain. We evaluate two tracks: (i) a Monte-Carlo study under calibrated Beta models characterising working curves, error rates, capping behaviour, and sensitivity; and (ii) a real-LLM evaluation on 200 attempted MMLU and 200 attempted GSM8K items with three heterogeneous agents (gpt-5, claude-opus-4-6, gemini-2.5-pro) and a claude-opus-4-6 judge, using disjoint 40-item calibration subsets. On GSM8K the rule stops in 1.01 average rounds (4.06 LLM calls) at 97.0% accuracy vs 99.0% for fixed-5 debate at 15 calls: a 3.7x call reduction at -2pp accuracy. On MMLU the calibrated KL collapses to about 0 and the rule caps on 99.5% of items at 2.1x cost. The takeaway is not that SPRT makes debate more accurate, but that a classical sequential test serves as a cheap compute-control and failure-detection layer for multi-agent LLM systems.
Science Journals
arXiv:2605.19187v1 Announce Type: new
Abstract: The Koopman--von Neumann (KvN) formulation of spectrally truncated fluid and plasma dynamics is considered as a potential approach for quantum computation. The KvN framework embeds the Liouville equation into a Hilbert space with norm-preserving, unitary evolution. Here, we propose a Weyl-ordered KvN generator along with a summation-by-parts discretization, which ensures that the resulting operators are exactly unitary as required for quantum computers. The Weyl-ordered KvN generator is derived as the unique anti-Hermitian operator symmetrization for real velocity fields. The formulation operates directly in the physical amplitude space without phase-space doubling, so the Heisenberg uncertainty principle does not constrain the grid resolution during evolution. This limitation re-enters only at the measurement stage on a quantum computer. Exact discrete unitarity is proved as a purely algebraic identity that holds regardless of grid resolution or stencil order. To manage boundaries, a split-step Kraus absorbing layer is introduced via a Stinespring dilation requiring only one ancilla qubit. Validation on three test cases spanning dissipative and Hamiltonian regimes (a viscous Navier--Stokes triad, an incompressible Euler triad, and a Hasegawa--Mima drift-wave triad) confirms fourth-order convergence and machine-precision unitarity.
arXiv:2605.19185v1 Announce Type: new
Abstract: Sparse goal-conditioned planning with few cost-to-go labels can be viewed as a graph-PDE Dirichlet extension problem: extend sparse labels on a goal-dependent boundary to unlabelled graph vertices so that greedy rollouts reach the goal. We study which graph value extensions are planner-admissible under the operational argmin-Q planner. Our main result is a local action-gap certificate: if the surrogate value error along the rollout stays below half the true action gap, then the greedy rollout reaches the goal. Absolutely Minimal Lipschitz Extension (AMLE), the p=infinity endpoint of the graph p-Laplacian family, instantiates this certificate through a comparison-principle fill-distance bound. Harmonic extension, by contrast, can mis-rank local actions because its values reflect boundary hitting probabilities rather than shortest-path greedy order.
On 120 AntMaze layout-derived graph configurations, harmonic extension achieves 0.584 aggregate rollout success, while AMLE reaches 0.970. Finite high-p methods also enter a high-success regime, with success 0.903 for p=4, 0.973 for p=8, and 0.982 for a fixed-budget p=16 solver, though the p=16 row is not used as a converged endpoint ranking due to incomplete solver certification. Mechanism audits show that many rollout decisions occur in AMLE-compatible but harmonic-incompatible local geometry, and that AMLE corrects most harmonic inversions on the rollout-weighted decision scope.
arXiv:2605.19318v1 Announce Type: cross
Abstract: Amorphization of silicon is crucial to applications in photonics, microelectronics and solar cell technologies. Ultrafast lasers have been used to generate amorphous silicon from crystalline silicon using rapid nonthermal melting and solidification in room temperature. As material temperature can affect cooling rates significantly, adding temperature control in ultrafast laser modification of silicon may allow a new degree of freedom in ultrafast laser modification. In this work, we investigate the role of cryogenic temperature in governing ultrafast damage pathways via single-shot femtosecond laser irradiation of silicon from room temperature down to 24K at 1030nm. Across this temperature range, we observe a pronounced enhancement of amorphization at lower temperatures, revealed through optical microscopy, Raman spectroscopy, and Kelvin probe force microscopy (KPFM). Raman analysis identifies this ring as an amorphous surface layer, while complementary AFM and SEM imaging show temperature-dependent changes in surface morphology, including localized melt redistribution and refrozen material. To elucidate the physical origins of this behavior, we implement a carrier dependent two-temperature model (nTTM). The simulations reproduce the experimentally observed trends and indicate that reduced phonon population, modified absorption pathways, and altered lattice relaxation dynamics at cryogenic temperatures collectively promote amorphous freezing over recrystallization. This study represents the first detailed examination of silicon under ultrafast irradiation below the liquid-nitrogen regime and reveals temperature-governed mechanisms relevant for advanced silicon microstructuring.
arXiv:2605.19129v1 Announce Type: cross
Abstract: We extend the optimin notion of Ismail (2025) from mixed strategy profiles to correlated distributions. A correlated distribution is evaluated by the worst expected payoff each player can receive when opponents may either obey their private recommendations or make unilateral recommendation-contingent deviations that are strictly profitable under the posterior induced by the distribution. Correlated optimins are Pareto optimal with respect to this vector of guaranteed payoffs. We show that a correlated optimin exists in every finite game. In addition, for every correlated equilibrium, there exists a correlated optimin such that every player's guaranteed payoff is weakly higher than his or her correlated equilibrium payoff. In two-player zero-sum games, correlated optimin coincides with correlated equilibrium and yields the maximin value. Outside zero-sum games, correlated optimin may strictly improve upon all correlated equilibria. We illustrate this with a simple 2x2 game with a unique correlated and coarse correlated equilibrium, in which there exists a correlated optimin that strictly Pareto dominates the equilibrium payoff.
arXiv:2605.19128v1 Announce Type: cross
Abstract: The Koch snowflake is a classical example of a planar curve with infinite perimeter enclosing a finite, positive area. Although such examples are well known individually, classical treatments typically analyze each construction in isolation and classify them by similarity dimension. This paper develops a unified parameter-space representation for a class of self-similar planar constructions, organized by two integers -- the number of self-similar pieces $N$ and the inverse linear scale factor $r$ -- together with two derived growth ratios $\alpha = N/r$ and $\beta = N/r^2$, governing perimeter and area scaling respectively. The $(N,r)$ parameter space is partitioned into three regimes -- $N \le r$, $r < N < r^2$, and $N \ge r^2$ -- corresponding to qualitatively distinct asymptotic behaviors of perimeter and area jointly. Within the intermediate regime $r < N < r^2$, a construction-class refinement distinguishes additive constructions (region bounded by the iterated curve), which yield positive finite asymptotic area under a stated non-overlap assumption, from subtractive constructions (iterated set itself), which yield zero asymptotic area. This records a structural non-equivalence inside the same dimension class that is not visible from $D = \log N / \log r$ alone. Four worked examples illustrate the framework -- the Sierpinski triangle, Sierpinski carpet, Koch snowflake, and a Koch-style construction on a square invented by the author -- and four further constructions are analyzed predictively to demonstrate that diagnostic outputs follow from $(N, r, \text{construction class})$ without re-derivation. The contribution lies in formulation and synthesis: the paper consolidates several classical results into a single diagnostic representation in which, given $(N, r)$ and construction class, the asymptotic behavior of perimeter and area can be inferred directly.
arXiv:2605.17340v2 Announce Type: replace
Abstract: Time series foundation models rely on large-scale pretraining over diverse datasets across domains, yet their heterogeneity in temporal patterns could hinder the effectiveness of training and learning transferable time series representations. Inspired a fundamental concept, normalized power spectral density (PSD) in signal processing, we assume harmonizing datasets via PSDs in the spectral domain could reduce mismatches and enhance pretraining. We then go beyond the direct intractable minimization optimization and innovatively reformulate it as a principled harmonization approach. Specifically, we propose Harmonizer, a module that reshapes spectral structures and implicitly harmonizing PSDs across datasets, which theoretically corresponds to a shared reparameterization of second-order temporal correlations. Our theoretical analysis further reveals token interactions with Harmonizer can be efficiently mediated by a compact set of resonators, motivating a HarmonicAttention design that performs self-attention in a low-dimensional interaction space. Then, we propose Olivia, a novel time series foundation model built upon these harmonization mechanisms. Extensive experiments on two large-scale benchmarks (TSLib and GIFT-Eval) and extra 6 datasets from GluonTS, demonstrate Olivia consistently achieves state-of-the-art performance under zero-shot, few-shot, and full-shot forecasting scenarios. Our code is available at https://github.com/TSTS13/Olivia.
arXiv:2605.19195v1 Announce Type: cross
Abstract: The construction of models from data is a significant contributor to the energetic costs of computation. Because of this, understanding how foundational thermodynamic bounds apply to modeling algorithms will be increasingly important. Here, we study the thermodynamic costs of a basic and fundamental modeling algorithm: simple linear regression. Following Landauer, we approximate the thermodynamic lower bound on irreversibly performing both exact linear regression and linear regression via stochastic gradient descent as implemented on floating-point numbers. From this, we derive energycost aware scaling laws for the optimal dataset size for training a linear regression model given a generalization error dependent demand for inference. Additionally, we discuss a method to lower bound the entropy production from the mismatch cost for algorithms with continuous input variables.
arXiv:2605.19179v1 Announce Type: cross
Abstract: We present a cloud-based tool that uses drones and machine learning to help recover instrumentally observed meteorite falls. We showcase a collection of improvements made upon previous iterations of our system, as well as detail the successes and limitations of this technique when applied to observed meteorite falls in South and Western Australia. This tool is available to the meteoritics research community upon request at https://find.gfo.rocks.
arXiv:2605.19146v1 Announce Type: cross
Abstract: Financial options are fundamental to traditional markets, enabling strategies ranging from hedging to speculating. Yet, while the Automated Market Maker paradigm has revolutionized decentralized spot markets, no equivalent standard has emerged for on-chain options. Typical designs attempt to replicate centralized exchange mechanics, requiring high-frequency oracles and robust liquidation engines which may fail during stress events. This paper presents a design for amortizing perpetual options tailored to the operational and adversarial constraints of blockchain environments. Leveraging this primitive, we introduce a decentralized market framework with minimal consistency requirements. We demonstrate that this contract functions as a foundational risk primitive for DeFi, enabling applications such as endogenous collateralization and explicitly priced de-peg insurance, thereby showing that this design provides a layer for mutualizing tail risk across protocols without reliance on centralized clearing institutions.
arXiv:2605.16626v2 Announce Type: replace
Abstract: Since autonomous coding agents generate complex behaviors at high-volume, we may want to use other LLMs to monitor actions to reduce the risk from dangerous misaligned behavior. To better understand the limitations of such monitors against the diverse attack strategies that a coding agent could use, we present SLEIGHT-Bench (Subtle Low-itEration Insight-Guided Harmful Transcripts), a benchmark of synthetic transcripts containing 40 attacks across 11 categories, each showing a coding agent covertly pursuing a harmful objective (e.g. weight exfiltration, credential theft, unmonitored rogue deployment). We find that 20 of 40 attacks are never caught by an Opus 4.6 monitor with extended thinking across 10 trials at a 1% false-positive rate, while the overall catch rate is 32%. We identify several evasion strategies such as exploiting model priors, user instruction ambiguity, and state manipulation, which facilitate creating attacks that evade frontier monitors. We also elicit stronger monitor performance using coding agents as monitors versus regular prompted monitors, and for some evasion strategies show improved catch rates with targeted monitor prompts. Our dataset and evaluation framework are available at https://github.com/safety-research/sleight-bench and https://huggingface.co/datasets/sleightbench/SLEIGHT-Bench.
arXiv:2605.17326v1 Announce Type: cross
Abstract: We investigate the role of the noise schedule in diffusion processes on Lie groups, with particular emphasis on applications to lattice gauge theory. We show that a specific noise schedule leads to a linear decay of the expectation value of the Wilson action as a function of diffusion time. We compare this with Euclidean diffusion models, where such behavior requires an explicitly designed drift term, while in the Lie-group setting it arises naturally.
arXiv:2605.20182v1 Announce Type: new
Abstract: Learning universal representations from electroencephalogram (EEG) signals is a cutting-edge approach in the field of neuroinformatics and brain-computer interfaces (BCIs). Conventionally, EEG is treated as a multivariate temporal signal, where time- or frequency-domain features are extracted for representation learning. This paper investigates a simple yet effective EEG representation, i.e., microstates. Microstates represent the building blocks of brain activity patterns at a microscopic time scale. We build a universal microstate tokenizer from a large medical EEG dataset by clustering continuous EEG signals into sequences of discrete microstates. The microstate tokenizer is then adopted universally across a series of downstream tasks, including sleep staging, emotion recognition, and motor imagery classification. Experimental results show that EEG representation learning with microstates outperforms traditional time-domain and frequency-domain features under different models and across different tasks. Further analysis shows that microstates offer greater interpretability and scalability, thereby opening up applications in both cognitive neuroscience and clinical research.
arXiv:2501.09203v2 Announce Type: replace
Abstract: Visual-Spatial Systems has become increasingly essential in concrete crack inspection. However, existing methods often lacks adaptability to diverse scenarios, exhibits limited robustness in image-based approaches, and struggles with curved or complex geometries. To address these limitations, an innovative framework for two-dimensional (2D) crack detection, three-dimensional (3D) reconstruction, and 3D automatic crack measurement was proposed by integrating computer vision technologies and multi-modal Simultaneous localization and mapping (SLAM) in this study. Firstly, building on a base DeepLabv3+ segmentation model, and incorporating specific refinements utilizing foundation model Segment Anything Model (SAM), we developed a crack segmentation method with strong generalization across unfamiliar scenarios, enabling the generation of precise 2D crack masks. To enhance the accuracy and robustness of 3D reconstruction, Light Detection and Ranging (LiDAR) point clouds were utilized together with image data and segmentation masks. By leveraging both image- and LiDAR-SLAM, we developed a multi-frame and multi-modal fusion framework that produces dense, colorized point clouds, effectively capturing crack semantics at a 3D real-world scale. Furthermore, the crack geometric attributions were measured automatically and directly within 3D dense point cloud space, surpassing the limitations of conventional 2D image-based measurements. This advancement makes the method suitable for structural components with curved and complex 3D geometries. Experimental results across various concrete structures highlight the significant improvements and unique advantages of the proposed method, demonstrating its effectiveness, accuracy, and robustness in real-world applications.
arXiv:2604.19892v1 Announce Type: cross
Abstract: Incremental Potential Contact (IPC) guarantees intersection-free simulation but suffers from high computational costs due to the expensive Hessian assembly and linear solves required by Newton's method. While Preconditioned Nonlinear Conjugate Gradient (PNCG) avoids Hessian assembly, it has historically struggled with poor convergence in stiff, contact-rich scenarios due to the lack of effective preconditioners; simple Jacobi preconditioners fail to capture the global coupling, while advanced hierarchy-based preconditioners like Multilevel Additive Schwarz (MAS) are computationally prohibitive to rebuild at every nonlinear iteration. We present MAS-PNCG, a method that unlocks the power of hierarchical preconditioning for nonlinear optimization. Our key technical innovation is a Sparse-Input Woodbury update algorithm that incrementally adapts the fine-level MAS components to rapidly evolving contact sets. This bypasses the need for full preconditioner rebuilds, reducing maintenance cost to near-zero while capturing the complex spectral properties of the contact system. Furthermore, we replace heuristic PNCG search directions with a Hessian-aware 2D subspace minimization that optimally combines the preconditioned gradient and previous direction. We also apply a fast per-subdomain conservative CCD method that ensures penetration-free trajectories while avoiding overly restrictive global step sizes. Experiments demonstrate that our MAS-PNCG outperforms state-of-the-art Newton-PCG solvers, GIPC and StiffGIPC, both preconditioned with MAS up to 5.66$\times$ and 2.07$\times$ respectively.
arXiv:2410.18856v4 Announce Type: replace
Abstract: Frontier large language models (LLMs), such as GPT-5, Claude 4.5, Gemini 3, Llama 4, and DeepSeek-R1, represent a transformative class of AI tools capable of revolutionizing various aspects of healthcare by generating human-like responses across diverse contexts and adapting to novel tasks following human instructions. Their potential application spans a broad range of medical tasks, such as clinical documentation, matching patients to clinical trials, and answering medical questions. In this paper, we propose an actionable guideline to help healthcare professionals more effectively and efficiently utilize LLMs in their work, along with a set of best practices. The overall workflow consists of several main phases, including formulating the task, choosing LLMs, prompt engineering, fine-tuning, and model deployment. We start with the discussion of critical considerations in identifying medical tasks that align with the core capabilities of LLMs and selecting models based on the selected task and data, performance requirements, and model interface. We then review the strategies, such as prompt engineering and fine-tuning, to adapt standard LLMs to specialized medical tasks. Deployment considerations, including regulatory compliance, ethical guidelines, and continuous monitoring for fairness and bias, are also discussed. By providing a structured step-by-step methodology, this entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.
arXiv:2408.15377v4 Announce Type: replace
Abstract: We propose a framework of algorithm vs. hardness for all Max-CSPs and demonstrate it for a large class of predicates. This framework extends the work of Raghavendra [STOC, 2008], who showed a similar result for almost satisfiable Max-CSPs.
Our framework is based on a new hybrid approximation algorithm, which uses a combination of the Gaussian elimination technique (i.e., solving a system of linear equations over an Abelian group) and the semidefinite programming relaxation. We complement our algorithm with a matching dictator vs. quasirandom test that has perfect completeness.
The analysis of our dictator vs. quasirandom test is based on a novel invariance principle, which we call the mixed invariance principle. Our mixed invariance principle is an extension of the invariance principle of Mossel, O'Donnell and Oleszkiewicz [Annals of Mathematics, 2010] which plays a crucial role in Raghavendra's work. The mixed invariance principle allows one to relate 3-wise correlations over discrete probability spaces with expectations over spaces that are a mixture of Guassian spaces and Abelian groups, and may be of independent interest.
arXiv:2605.20151v1 Announce Type: new
Abstract: The proliferation of generative artificial intelligence has given rise to an interactive learning environment, where model parameters are continuously updated using not only data generated by natural processes, but also synthetic outputs produced by other models. This paradigm introduces two major challenges: (1) training data are no longer drawn exclusively from the target population, undermining a core assumption of classical statistical learning, and (2) model training processes become inherently correlated, as models interact with one another through repeated exposure to each other's synthetic outputs in a potentially complex manner. Establishing reliable statistical inference in such structured interactive learning environments therefore remains an important open problem. In particular, there is growing concern about model collapse, a phenomenon in which the performance of generative models progressively degrades as they are trained on synthetic data produced by earlier model generations.
Prior work on model collapse primarily focuses on a single model trained on its own output, failing to capture model performance in multi-model interactive settings. In this work, we fill this gap by investigating the performance of generative models in an interactive learning environment with general interaction patterns. In particular, we formalize model interactions using directed graphs and show that the occurrence of model collapse depends critically on the topology of the interaction graph. We further derive an explicit necessary and sufficient condition characterizing when model collapse occurs, and establish finite-sample results for linear regression and asymptotic guarantees for general M-estimators. We support our theoretical findings through extensive numerical experiments.
Hermitian hull-variation of vector rank-metric codes and self-orthogonal generalized Gabidulin codes
arXiv:2605.20109v1 Announce Type: new
Abstract: We study the Hermitian hull-variation problem for vector rank-metric codes. Except for one parameter pair, we show that the Hermitian hull dimension of such a code can be reduced to any smaller value within its equivalence class, and in particular every such code is equivalent to a Hermitian LCD code. We then address the existence of maximum rank distance (MRD) codes with prescribed Hermitian hull dimension. To this end, we introduce the notion of a \emph{scaled trace-self-dual basis} of a finite field extension, which exists in all cases, and use it to construct Hermitian self-orthogonal generalized Gabidulin codes for every prime power. Combined with the hull-variation theorem, this yields MRD codes attaining every admissible Hermitian hull dimension.
arXiv:2605.20091v1 Announce Type: new
Abstract: Kernel methods are one of the cornerstones of learning-based control, modern system identification, surrogate modelling, and related fields. A key advantage of this class of learning and function approximation methods is the availability of quantitative error bounds, which in turn play a key role in guaranteeing the safety of learned controllers and related learning-based algorithms. However, these error bounds rely on a particular property of the target function -- its reproducing kernel Hilbert space (RKHS) norm -- which is usually impossible to obtain in practice. Motivated by this severe shortcoming, we present a novel sampling-based RKHS norm estimation approach with a solid theoretical foundation, leveraging very recent advances in the theory of superconvergence in kernel methods. Our method is applicable to a broad range of practically relevant function classes and requires only reasonable prior knowledge about the target function. Extensive numerical experiments demonstrate the efficacy and practical applicability of the proposed method. By providing a reliable RKHS norm estimation approach, we remove a major obstacle to the practical deployment of learning-based control algorithms.
arXiv:2605.20074v1 Announce Type: new
Abstract: Distillation transfers knowledge from a large model trained on broad data to a smaller, more efficient model suitable for deployment. In structured prediction settings, prior knowledge about the task can guide the choice of a target architecture that is algorithmically aligned with the underlying problem. Building on recent learning-theoretic analyses of decision-tree (DT) distillation (Boix-Adsera, 2024), we study when distillation succeeds for combinatorial optimization tasks. We focus on the case where the target model is a graph neural network whose architecture is aligned with a dynamic programming (DP) algorithm for the task. Assuming that the source model is sufficiently rich, formalized through the linear representation hypothesis (LRH) (Elhage et al., 2022; Park et al., 2024), we show that the distillation problem can be solved efficiently in the complexity parameters of the DP transition function, represented as a DT. Our results provide a rigorous sufficient condition for successful distillation in the flavour of algorithmic alignment.
arXiv:2605.20066v1 Announce Type: new
Abstract: Knowledge graph question answering seeks to translate natural language questions into executable queries over knowledge graphs, but existing approaches often rely on large models or full supervision in the form of gold query annotations. This study examines whether reinforcement learning with outcome-based rewards can train a small instruction-tuned language model to perform zero-shot Text-to-SPARQL generation in the scholarly domain. Group-Relative Policy Optimization (GRPO) is applied to the Qwen3-1.7B model on DBLP-QuAD, using prompts that combine natural language questions with symbolic hints about entities and relations. Training relies on execution feedback, structural constraints, and answer-level rewards, with an additional variant that incorporates gold-query-based shaping. The resulting models are compared to the unmodified zero-shot baseline and to a supervised DoRA-finetuned baseline across answer-level accuracy, execution accuracy, category-wise scores, and generalization to held-out templates. GRPO substantially improves over the zero-shot baseline and exhibits competitive generalization, while supervised DoRA finetuning achieves higher overall accuracy on the same model scale. Ablation analyses indicate that execution-based rewards account for most gains, with additional shaping yielding limited additional benefit, suggesting that outcome-based reinforcement learning is a viable training strategy when gold queries are unavailable for token-level supervision.
arXiv:2605.20061v1 Announce Type: new
Abstract: Reinforcement learning from verifiable rewards (RLVR) is a promising paradigm for improving large language model (LLM) agents on long-horizon interactive tasks. However, in partially observable environments, incomplete observations cause agent beliefs to drift over time, while delayed rewards obscure the causal impact of intermediate decisions, exacerbating temporal credit assignment challenges. To address this, we propose ReBel (Reward Belief), a process-level reinforcement learning algorithm that explicitly models structured belief states to summarize interaction history and guide subsequent policy learning. ReBel introduces belief-consistency supervision, converting discrepancies between predicted beliefs and observed feedback into dense self-supervised signals without requiring external step-wise annotations or verifiers. It also employs belief-aware grouping to compare trajectories under similar belief states, yielding more robust and lower-variance advantage estimates. We evaluate ReBel on challenging long-horizon benchmarks, including ALFWorld and WebShop. ReBel improves task success by up to $20.4$ percentage points over the episode-level baseline GRPO and increases sample efficiency by $2.1\times$. These results suggest that belief-aware self-supervision is a promising direction for reliable long-horizon decision-making under partial observability. Code is available at: https://github.com/Fateyetian/Rebel.git.
arXiv:2605.19667v1 Announce Type: cross
Abstract: In this paper, we study a consensus-based optimization method for nonconvex bi-level optimization, where the objective is to minimize an upper-level function over the set of global minimizers of a lower-level problem. The proposed approach is derivative-free, and constructs its consensus point via smooth quantile selection combined with a Gibbs-type Laplace approximation. We establish convergence guarantees for both the associated \textit{mean-field} dynamics and its \textit{finite-particle} approximation. In particular, under suitable assumptions on smooth quantile localization, error bounds, and stability, we show that the mean-field law reaches any arbitrary prescribed Wasserstein neighborhood of the target bi-level solution with an explicit exponential rate up to the hitting time. Numerical experiments on a two-dimensional constrained problem and neural network training further support the theoretical results.
arXiv:2605.20051v1 Announce Type: new
Abstract: AI infra has become a shared execution layer for model training, deployment, and agent orchestration. Because many projects reimplement similar model-centric workflows, a vulnerability disclosed in one repository can recur as a variant in another repository with a related design. Yet the prevalence and detectability of these variants remain poorly understood. This paper presents a measurement study of vulnerability variants in AI infra. Analyzing 688 GitHub repositories and 251 publicly disclosed vulnerabilities, we find that AI infra projects frequently share overlapping functionality and recurrent vulnerable patterns, creating a concrete basis for cross-repository variants. Building on this finding, we study how to automatically identify such variants from known disclosures. We propose INFRASCOPE, a reference-driven multi-agent framework that extracts transferable vulnerability semantics from known cases and uses them to locate and validate variants in new repositories. Evaluating INFRASCOPE on 20 real-world AI infra repositories, we uncover over 20 vulnerabilities, including 11 acknowledged cases and 4 cases that have been assigned CVEs so far.