arXiv:2607.14651v1 Announce Type: new
Abstract: Persistent external memory enhances agent continuity but introduces persistent security vulnerabilities: adversarial content can be injected via standard interaction channels, retained across turns, and later distort downstream behavior. To address this challenge, we propose MemPoison, a comprehensive benchmark and analysis framework featuring 1227 hand-validated cases across four attack types, three injection channels, and three representative memory substrates, evaluated on seven open-weight and three closed-weight model families. We introduce a three-tier taxonomy: (L1) direct single-record corruption, (L2) compositional multi-record corruption and (L3) context-triggered dormant corruption. Our evaluations reveal a distinct defense frontier: while baseline write-time defenses, such as consistency checks, substantially suppress direct L1 attacks, they fail to reliably suppress L2 and L3 attacks. Through mechanistic influence decomposition (MID), we demonstrate structural blind spots in write-time defenses, which admit seemingly benign records that later become harmful through joint retrieval composition or trigger-conditioned activation. Our findings advocate for shifting from static filtering to adaptive, context-sensitive memory defense strategies.
Science Journals
arXiv:2603.05353v3 Announce Type: replace
Abstract: Retrieval-augmented generation (RAG) for long-context question answering is bottlenecked by inference-time prefilling over large retrieved contexts. A common strategy is to precompute key-value (KV) caches for individual documents and selectively recompute a small subset of tokens to restore global causal dependencies, but existing methods rely on heuristics or representation discrepancies without modeling whether selected tokens can effectively influence generation. We cast selective KV recomputation as an information flow problem and show that a simple attention-norm signal from the query reliably identifies tokens that are both semantically relevant and structurally positioned to propagate information, when computed under an inference-consistent RoPE geometry. We therefore reconstruct global positional assignments for retrieved chunks and introduce an information-flow-guided chunk reordering strategy. Experiments on Large Language Model and Vision-Language Model benchmarks demonstrate consistent gains over prior methods under comparable latency.
arXiv:2603.05754v2 Announce Type: replace
Abstract: Current Vision-Language-Action (VLA) models rely primarily on RGB perception, preventing them from capturing modalities such as thermal signals that are imperceptible to conventional visual sensors. Moreover, end-to-end generative policies lack explicit safety constraints, making them fragile when encountering obstacles and novel scenarios outside the training distribution. To address these limitations, we propose Safe-Night VLA, a multimodal manipulation framework that enables robots to see the unseen while enforcing rigorous safety constraints for thermal-aware manipulation in unstructured environments. Specifically, Safe-Night VLA integrates long-wave infrared thermal perception into a pre-trained vision-language backbone, enabling semantic reasoning grounded in thermodynamic properties. To ensure safe execution under out-of-distribution conditions, we incorporate a safety filter via control barrier functions, which provide deterministic workspace constraint enforcement during policy execution. We validate our framework through real-world experiments on a Franka manipulator, introducing a novel evaluation paradigm featuring temperature-conditioned manipulation, subsurface target localization, and reflection disambiguation, while maintaining constrained execution at inference time. Results demonstrate that Safe-Night VLA outperforms RGB-only baselines and provide empirical evidence that foundation models can effectively leverage non-visible physical modalities for robust manipulation.
arXiv:2603.23667v3 Announce Type: replace
Abstract: We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provider-diverse conditions. Echoes comprises 4,468 tracks (131 hours of audio) spanning multiple genres (pop, rock, electronic), and includes content generated by ten popular AI music generation systems. To prevent shortcut learning and promote robust generalization, the dataset is deliberately constructed to be challenging, enforcing semantic-level alignment between spoofed audio and bona fide references. This alignment is achieved by conditioning generated audio samples directly on bona-fide waveforms or song descriptors. We evaluate Echoes in a cross-dataset setting against three existing AI-generated music datasets using state-of-the-art Wav2Vec2 XLS-R 2B representations. Results show that (i) Echoes is the hardest in-domain dataset; (ii) detectors trained on existing datasets transfer poorly to Echoes; (iii) training on Echoes yields the strongest generalization performance. These findings suggest that provider diversity and semantic alignment help learn more transferable detection cues.
arXiv:2603.29759v3 Announce Type: replace
Abstract: Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing benchmarks suffer from three fundamental limitations: (1) heavy reliance on synthetic datasets constructed via simulation software, creating a significant domain gap with real-world environments; (2) oversimplified safety tasks with artificial constraints on hazard and scene types, thereby limiting model generalization; and (3) absence of rigorous evaluation protocols to thoroughly assess model capabilities in complex home safety scenarios. To address these challenges, we introduce TSHA (\textbf{T}rustworthy \textbf{S}afety \textbf{H}azards \textbf{A}ssessment), a comprehensive benchmark comprising 66,668 validated question-answer pairs, including 64,961 carefully curated training QA pairs drawn from existing indoor datasets, internet frames/images, AIGC images, newly captured images, and Hunyuan panoramic images. This benchmark also includes a highly challenging test set with 1,707 QA pairs, comprising not only a carefully selected subset from the training distribution but also newly added Sora-generated videos and Hunyuan panoramic images containing multiple safety hazards, used to evaluate the model's robustness in complex safety scenarios. Extensive experiments on 22 popular VLMs demonstrate that current VLMs lack robust capabilities for safety hazard assessment. Importantly, models trained on the TSHA training set achieve a significant performance improvement of up to +18.3 points on the TSHA test set and also exhibit enhanced generalizability across other benchmarks, underscoring the substantial contribution and importance of the TSHA benchmark.
arXiv:2604.09870v2 Announce Type: replace
Abstract: We investigate how looped transformers encode human preference, training lightweight evaluator heads on frozen Ouro-2.6B loop-iteration states on Anthropic HH-RLHF.
v2: an erratum is prepended; the original manuscript is unchanged. A post-publication audit found the three headline results inflated by two independent evaluation errors. The 95.2% pairwise evaluator accuracy is a canonical-ordering artifact: the data were correctly split, but the evaluator learned to prefer the first-presented argument; its strict antisymmetrized accuracy on the full 8,552-pair test set is 63.9%. The 84.5% pairwise probe and the below-chance 21.75% pointwise probe were source-item leaks (orientation rows and pair partners crossing the train/test split); corrected pair-disjoint values are 56.5% and 54.2% -- above chance, so the "inverted polarity" finding is withdrawn. The central finding survives at much smaller magnitude: preference is decoded more accurately relationally than pointwise (paired +2.3 points, 95% CI [+1.3, +3.3]), and the antisymmetrized evaluator still beats the linear probe, but no corrected readout rivals end-to-end reward models. The methodological findings stand (constant-output degeneracy, flip test, swap-protocol metric deflation), with one correction: antisymmetrized accuracy, not antisymmetry correlation, certifies relational discrimination. The two errors are mutually invisible -- a split audit cannot see an ordering prior, antisymmetrization cannot see a leak -- so both checks are required. Full audit in the follow-up work (Kirin, 2026, in preparation).
arXiv:2607.15046v1 Announce Type: new
Abstract: Topological concepts are frequently used to describe structured optical fields, including plasmonic near fields. Topological descriptions in terms of skyrmion numbers implicitly assume the compactness of the underlying manifold. Even when skyrmion-like textures appear locally, the compactness is usually not fulfilled in extended optical fields. Here, we use photoemission electron microscopy to investigate a plasmonic nano-focus that exhibits a sequence of radially extending alternating skyrmion and antiskyrmion textures. The full spatio-temporal reconstruction of the electric field vectors and their topology is accessible by vector polarimetry. The experiments confirm the expected oscillatory behavior of the skyrmion number and demonstrate that a global skyrmion number cannot be assigned in such non-compact fields.
arXiv:2607.15220v1 Announce Type: new
Abstract: Unsupervised visible-infrared person re-identification (USVI-ReID) is challenging due to the large modality gap and the lack of cross-modal identity annotations. Progressive association paradigms have been proposed to gradually bridge the gap, but they suffer from two critical bottlenecks: reliance on ambiguous global representations and unchecked propagation of pseudo-label noise in an open-loop manner. To address these issues, we propose Structural-Semantic Reciprocal Learning (SSRL), a framework that transforms open-loop association into a self-correcting closed-loop system. Structurally, we introduce Fine-grained Structural Decoupling (FSD) to extract discriminative body-part primitives as reliable spatial anchors, complementing ambiguous holistic silhouettes with spatially consistent structural details. Semantically, we design a Closed-loop Semantic Calibration (CSC) mechanism that reconstructs shared semantic prototypes at each epoch and feeds them back into the training loop, effectively filtering pseudo-label noise before the next clustering cycle. Through the reciprocal interaction between structural and semantic learning, SSRL achieves robust cross-modal representation. Extensive experiments demonstrate the competitive performance of SSRL against state-of-the-art USVI-ReID methods on both SYSU-MM01 and RegDB, notably surpassing several supervised counterparts on RegDB.
arXiv:2607.14684v1 Announce Type: new
Abstract: AI-generated image (AIGI) detectors achieve strong accuracy on clean benchmarks, but their performance drops sharply after images are propagated through real-world channels. We trace this fragility to what these detectors actually learn: they overfit to local artifacts left by generators in small spatial neighborhoods, which are easily destroyed by common propagation degradations such as JPEG compression and blur. Instead, we shift the discriminative cue from fragile local artifacts to more robust global structure. Building on this, we propose GlobalForge, a framework with two complementary modules. The Local Information Bottleneck (LIB) suppresses local components to block shortcut learning, while the Global Structural Reasoning (GSR) module forces every token to gather evidence from distant regions. Both modules are trained jointly under a contrastive structural loss based on degradation that keeps the resulting features stable under degradation. To support fine-grained robustness evaluation, we further introduce RealDeg-Bench, covering 7 common degradation operators and multi-step compound chains. GlobalForge improves average BAcc on 8 in-the-wild benchmark groups by $\mathbf{5.89\%}$ over the previous state-of-the-art, and is clearly ahead of representative baselines on RealDeg-Bench under both single and compound degradations. Code is available at https://anonymous.4open.science/r/GlobalForge-BE0F/.
arXiv:2604.18155v2 Announce Type: replace
Abstract: Scaling the photon-detection area of superconducting nanowire single-photon detectors (SNSPDs) has traditionally been achieved by nanowire meandering. However, material inhomogeneities and fabrication-induced defects, such as line-edge roughness, increase with nanowire length, leading to reduced internal photon-detection efficiency and elevated dark-count rates. This trade-off becomes increasingly pronounced as nanowires are scaled to sub-100 nm widths and sub-5 nm thicknesses required for mid- to far-infrared sensitivity. Here, we demonstrate an antenna-coupled SNSPD architecture that enhances the effective photon-detection area without increasing nanowire length. A crossed bowtie antenna integrated with an 80 nm-wide, 3 nm-thick WSi nanowire yields 15.7$\times$ increase in effective detection area at 7.4 $\mu$m compared to a bare nanowire of identical geometric footprint, while maintaining the same internal detection efficiency and dark-count rate. Antenna coupling provides a scalable approach to increasing photon-detection area while reducing the noise-equivalent power, offering performance benefits for applications in astronomy, biological imaging, and molecular spectroscopy.
arXiv:2604.22280v3 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) have emerged as a promising foundation for universal multimodal embeddings. Recent studies have shown that reasoning-driven generative multimodal embeddings can outperform discriminative embeddings on several embedding tasks. However, Chain-of-Thought (CoT) reasoning tends to generate redundant thinking steps and introduce semantic ambiguity in the summarized answers in broader retrieval scenarios. To address this limitation, we propose Rewrite-driven Multimodal Embedding (RIME), a unified framework that jointly optimizes generation and embedding through a retrieval-friendly rewrite. Meanwhile, we present the Cross-Mode Alignment (CMA) to bridge the generative and discriminative embedding spaces, enabling flexible mutual retrieval to trade off efficiency and accuracy. Based on this, we also introduce Refine Reinforcement Learning (Refine-RL) that treats discriminative embeddings as stable semantic anchors to guide the rewrite optimization. Extensive experiments on MMEB-V2, MRMR and UVRB demonstrate that RIME substantially outperforms prior generative embedding models while significantly reducing the length of thinking. Code is available at https://github.com/PeppaWu/RIME.
arXiv:2604.22301v2 Announce Type: replace
Abstract: We demonstrate a suspended thin-film aluminum nitride (AlN) microbolometer for narrowband very long-wave infrared detection. The device uses a 100-nm-thick AlN membrane suspended above a Pt back reflector by a 1-um air gap. Resonant absorption is set by the AlN transverse optical phonon near 15.4 um and is strengthened by suspension above the reflector. A periodic perforation pattern reduces membrane thermal mass and enhances absorption without further thinning the film. DC resistance measurements under tunable infrared illumination verify bolometric operation, and the measured spectral response follows the absorption profile expected from spectroscopic measurement of passive devices. Narrowband response is observed in the 14--18 um range, with peak responsivity of 920.8 ppm/mW at 15.48 um. This platform can enable compact wavelength-selective thermal detectors for multispectral imaging, on-chip infrared spectroscopy, and chemical sensing.
arXiv:2607.15221v1 Announce Type: new
Abstract: We demonstrate a method to measure energy shifts of the top level in a four-level ladder setup induced by atom interactions in thermal vapors. It utilizes the observation of two transmission minima corresponding to a split electromagnetically induced absorption (EIA) effect. We apply this method to measure mean Rydberg atom interactions in a hot vapor. We believe this approach could provide a valuable tool for accurately modeling mean-field Rydberg atom interactions, as well as sensing the occurrence of strong interactions.
arXiv:2607.15111v1 Announce Type: new
Abstract: Vehicle coordination at unsignalized intersections relies on accurate real-time vehicle state acquisition and reliable command-and-control (C&C) signal delivery. However, existing studies typically treat sensing, communication, and control separately, which may lead to redundant transmissions, outdated state information, and unreliable vehicle coordination. In this paper, we investigate a new scenario of distributed integrated sensing and communication (ISAC)-enabled vehicle coordination at intersections, where multiple roadside units (RSUs) collaboratively transmit sensing signals for vehicle state acquisition and C&C signals for vehicle movement control under the management of a central base station (BS). To improve signaling efficiency, we propose a unified goal-oriented semantic communication (GSC) framework, which transmits sensing and C&C signals only when they are semantically important for improving intersection traffic throughput. Specifically, an extended Kalman filter (EKF) is adopted to predict vehicle states and fuse distributed sensing measurements. A masked hybrid proximal policy optimization (MHPPO) framework is then developed to jointly determine sensing transmission decisions, C&C transmission decisions, and C&C signal contents based on a value-of-information (VoI) reward. Furthermore, we propose an uncertainty-aware transmission design (UTD), including robust beamforming and VoI-based time-division power allocation, to improve sensing and communication reliability under vehicle state uncertainty and inter-RSU interference. Simulation results show that our proposed framework achieves 100% collision-free vehicle coordination with significantly reduced signaling overhead compared with predictive ISAC baselines adapted from state-of-the-art related studies and several ablation baselines.
arXiv:2607.15118v1 Announce Type: new
Abstract: Side-channel attacks pose a significant security threat for modern computing platforms, because they exploit subtle discrepancies in CPU behaviors to leak sensitive information. To model the information leaked by a CPU via microarchitectural side-channels, recent work proposed leakage contracts: an ISA-level security abstraction that provides the foundations for secure CPU programming. Unfortunately, due to the complexity of current microarchitectures, devising a leakage contract for a CPU requires extensive manual effort and thus modern CPUs lack dedicated leakage contracts.
We present a methodology to extract instruction-centric leakage contracts for major CPU architectures with minimal manual intervention. We implemented this technique in malcos, the first template-free tool that automates the synthesis of leakage contracts for black-box CPUs. We evaluate malcos on x86 and ARM CPUs, and show that the contracts it synthesizes are precise and sound with respect to all leaks observed during synthesis. Our results demonstrate that learning leakage contracts from black-box CPUs is feasible.
arXiv:2605.07742v2 Announce Type: replace
Abstract: Inter-manufacturer plug-and-play communication in agricultural machinery is currently based on the ISO 11783 standard series, which specifies a 250 kbit/s CAN bus communication layer. To support higher-bandwidth use cases, the ISO~23870 series is being developed for next-generation Ethernet-based agricultural machine-to-machine communication. Modern Ethernet/IP-based architectures often make use of a middleware for discovery, data exchange, quality of service configuration, and security. This paper evaluates the Data Distribution Service (DDS) as a candidate middleware for secure, plug-and-play agricultural machinery networking. DDS-based proof-of-concept communication design is presented for a representative Task Controller (TC) and implement scenario, including implement-description topics and separate best-effort and reliable topics for runtime process data. Design was implemented in C++ using the FastDDS library and benchmarked on embedded hardware representative of agricultural machinery. Runtime throughput was evaluated for one-to-one and one-to-two TC-implement scenarios under four DDS security configurations. The results show that DDS security mechanisms substantially reduce maximum throughput on embedded hardware. In the tested best-effort scenarios, signing and encryption reduced mean throughput by approximately 70-84% compared with the unsecured configuration. Encrypted one-to-one best-effort case achieved approximately 4980 received process data updates per second on both the TC and implement, corresponding to about 50 process data updates per second per simulated section for 100 rate-controllable sections. These results indicate that DDS is a technically plausible middleware candidate for secure Ethernet-based agricultural machinery interoperability, while further work is required to evaluate latency, scalability, vendor interoperability, and lower-power devices.
arXiv:2605.09028v3 Announce Type: replace
Abstract: Machine learning-based Android malware detectors often fail in real-world deployment due to domain shift, where models trained on one data source perform poorly on applications from another. This paper presents a comprehensive study on the generalizability and interpretability of permission-based detectors under cross-domain conditions. Using two complementary datasets (PerMalDroid and NATICUSdroid) and five ensemble classifiers, we first establish an intra-domain baseline, where models achieve over 92% accuracy, and then quantify a severe asymmetric performance drop. While models trained on PerMalDroid generalize well to NATICUSdroid (86% accuracy), the reverse direction sees a drastic drop to 73% accuracy. Explainable AI analysis reveals bimodal feature distributions and shows that feature importance is highly unstable, with key permissions losing or gaining influence across domains. The predictive feature sets for different domains are fundamentally mismatched, as models rely on different, dataset-specific permissions. Most importantly, an ablation study demonstrates that for most models, training on a noisy feature set leads to poor generalization, confirming that domain-specific artifacts are a greater obstacle than missing features. To mitigate this, we validate a hybrid training strategy based on the intersection of common features and successfully recover cross-domain performance, achieving 88% accuracy on PerMalDroid and maintaining 97% on NATICUSdroid. These findings highlight the importance of explainable, cross-domain-robust malware detection systems and provide a practical pathway toward improving real-world deployment of permission-based Android malware detectors.
arXiv:2605.09835v3 Announce Type: replace
Abstract: We report a comprehensive measurement of the environmental $\gamma$-ray flux in Hall C of the Gran Sasso National Laboratory. A spatial mapping of the radiation was carried out using a high-purity germanium detector mounted on a movable cart and deployed at eight locations within the hall. The detector response function and full-energy-peak efficiencies were determined through Geant4 simulations validated with calibrated $\gamma$-ray sources, with particular attention devoted to the efficiency modeling and associated systematic uncertainties. In the energy range of 57-2800 keV, the average $\gamma$-ray flux is measured to be $(\mathrm{0.46} \pm \mathrm{0.06}_{stat} \pm \mathrm{0.03}_{syst})$ $\mathrm{cm}^{-2}$ $\mathrm{s}^{-1}$. The radon level was monitored for about a month using a radon detector mounted on the same cart, and a clear correlation is observed between the environmental $\gamma$-ray rate and the ambient radon concentration, consistent with the short-lived daughters of $^{222}\mathrm{Rn}$. This result represents the first high-precision and efficiency-corrected mapping of the $\gamma$-ray flux in Hall C, substantially improving its radiological characterization and providing key input for future rare-event experiments operating in this hall.
arXiv:2605.10669v2 Announce Type: replace
Abstract: The one-centre Coulomb-Sturmian convergent close-coupling method is applied
to proton collisions with the boron atom and singly charged carbon ion.
Here we report an update to our target-structure implementation, in which
configuration state functions are constructed using the method of
coefficients of fractional parentage. To assess the quality of the
structure models for the two targets, we present the excitation energies,
oscillator strengths, and dipole polarisabilities obtained from the present
configuration interaction calculations. Cross sections for total and state-selective target
excitation and electron loss are calculated from 10 keV to 1 MeV. For both
systems, the total excitation cross section is found to be dominated by
excitation of the $2s$ subshell. This emphasises the importance of a
multi-electron description of the target in such scattering calculations.
Comparisons with previous theoretical and experimental data are presented and
discussed. In particular, we find that the present calculation for the
electron-loss cross section in $p$ + C$^{+}$ collisions is in good agreement
with the available measurements across the entire overlapping incident-energy
range.
arXiv:2605.10821v2 Announce Type: replace
Abstract: Diffusion-based vision-language-action (VLA) models have emerged as strong priors for robotic manipulation, yet adapting them to real-world distributions remains challenging. In particular, on-robot reinforcement learning (RL) is expensive and time-consuming, so effective adaptation depends on efficient policy improvement within a limited budget of real-world interactions. Noise-space RL lowers the cost by keeping the pretrained VLA fixed as a denoising generator while updating only a lightweight actor that predicts the noise. However, its performance is still limited due to inefficient autonomous exploration. Human corrective interventions can reduce this exploration burden, but they are naturally provided in action space, whereas noise-space finetuning requires supervision over noise variables. To address these challenges, we propose UniSteer, a Unified Noise Steering framework that combines human corrective guidance with noise-space RL through approximate action-to-noise inversion. Given a human corrective action, UniSteer inverts the frozen flow-matching decoder to recover a noise target, which provides supervised guidance for the same noise actor that is simultaneously optimized via reinforcement learning. Real-world experiments on diverse manipulation tasks show that UniSteer adapts more efficiently than strong noise-space RL and action-space human-in-the-loop baselines, improving the success rate from 20% to 90% in 66 minutes on average across four real-world adaptation tasks.
arXiv:2605.13348v3 Announce Type: replace
Abstract: As neural components are increasingly embedded in existing symbolic software -- including safety-critical systems -- the question arises of how to specify and enforce the safety of the newly introduced neural parts. Unlike traditional logical specifications, these must be amenable not only to the standard Boolean interpretation, but also to training and optimisation. The latter calls for a quantitative interpretation of the logical syntax, subject to further requirements such as smoothness and differentiability. Moreover, the qualitative and quantitative sides of the logic must share a unifying proof-theoretic and categorical semantics. Finally, the new logic should link cleanly to the substructural and program logics that underpin the verification of existing symbolic programs. In this paper, we present a logic that ticks all of these boxes. We introduce a family of calculi, pQLL, indexed by a hardness degree $p$, prove a cut-elimination theorem for them, and establish completeness with respect to enriched residuated `soft' lattices. At $p = \infty$, \pQLL reduces to multiplicative additive linear logic (MALL), and provability in pQLL converges to provability in MALL as $p \to \infty$. We express optimisation objectives in the syntax of this logic and prove the quantitative adequacy of neuro-symbolic loss functions -- a result that has eluded the neuro-symbolic machine learning community for nearly a decade.
arXiv:2605.17644v2 Announce Type: replace
Abstract: The particle representation model (PRM) and interacting particle representation model (IPRM) describe homogeneous turbulence through orientation-conditioned structural states. In their original form, the conditional state is organized by the unit spectral direction, while the radial spectral coordinate is integrated out. We introduce a scale-conditioned Ray-Column extension in which the spectral vector is decomposed into orientation and radial wavenumber, and the conditional structure state is projected onto finite radial bands.
The formulation starts from the continuum spectral tensor and is then reduced to the ray-packet ensemble sums used in the implementation. The bands are projections of an orientation-wavenumber tensor density and retain scale-conditioned structural populations for closure evaluation. The rapid dynamics remain ray-packet resolved, while the nonlinear slow and terminal closure coefficients are evaluated from band-aggregate structure tensors formed by integrating over orientation and wavenumber within each band. The present reference closure omits conservative cascade modeling among bands.
A reference closure is built from PRM rapid kinematics, band-local effective-gradient response, slow rotational randomization, and an active large-scale enstrophy (LSE) terminal-drain map. In the active-LSE closure, the misalignment-sensing factor Psi_fd regularizes the LSE structure-to-dissipation map; the Ray-Column formulation evaluates this map on band-aggregate structural populations. The model is assessed in irrotational strain, homogeneous shear, elliptic-streamline, and rotating-shear configurations. The rotating-shear comparison with filtered LES data illustrates the payoff of retaining band information: filtered or low-pass observables can be formed before scale information is lost in the one-point reconstruction.
arXiv:2605.19302v3 Announce Type: replace
Abstract: We introduce Coherent Utility Measure Games (CUMGs) in which players' uncertainty about the distribution of payoffs is modeled using coherent utility (risk) measures. Such measures, including mean semideviation risk and conditional value-at-risk, allow for interpretable notions of players' risk aversion while retaining formal equivalence to distributionally robust games. While CUMGs, which are a subclass of distributionally robust games, are continuous games in general, they can be viewed as finite games ``lifted'' to the mixed strategy space, which illustrates computational challenges. Prior results extend to guarantee equilibrium existence in data-driven CUMGs. We show that the computation of approximate equilibria for CUMGs parameterized by several risk measures lies in PPAD. Consequently, we obtain finite multilinear complementarity programs for the computation of equilibrium for these games, which grow with $K$, the number of data samples. Unlike standard games, these programs are not linear in a two-player setting. Next, we establish the existence of approximate equilibria in finite data-driven CUMGs with small supports in the pure actions for the players, together with sparse data subsamples that guide the search for such equilibria. We also develop a stochastic first-order approach for smoothed CUMGs using data mini-batches, with bounds linking first-order error to approximate equilibrium. We include numerical experiments comparing the sparse-support search algorithm with complementarity-program solvers.
arXiv:2605.19844v5 Announce Type: replace
Abstract: Many decision processes run for a long and unknown duration: in each round new requests arrive, an irrevocable choice must be made immediately, and the system is judged by ongoing fairness requirements. Examples include food banks allocating donations, computing systems repeatedly scheduling scarce resources across users, and institutions making repeated decisions while remaining fair over time. We propose a general approach based on \emph{deficits}, which measure how far the current outcome is from satisfying each fairness requirement. The goal is to keep all deficits small at each time step, without knowing the horizon or future agent valuations. This viewpoint also highlights a natural modeling question for long-running systems: how much of the past should be counted when fairness is evaluated? We first study the full-history model, where all past rounds count equally. We propose an efficient fully-online rule. For $n$ agents, we prove anytime guarantees: after any $t$ rounds, all requirements remain satisfied up to a slack of order $\tilde O(\sqrt{t/n})$. We instantiate the rule for online allocation of indivisible goods, yielding natural relaxations of proportionality and envy-free, and for online public decision-making. We show that this slack is tight even for weak proportionality. For unrestricted classical $\mathrm{EF}c$, the exact worst-case parameter at horizon $T$ is $\lceil T/n\rceil$. We then study discounted-memory fairness, where older deficits carry smaller weight. The same fully-online rule applies to these discounted deficits, and the resulting threshold is controlled by the discount function. In particular, the time dependence is never worse than the full-history $\sqrt t$ dependence. Overall, our results show that memory is a central part of perpetual fairness. The question is not only which requirement to impose, but also how the system should count past unfairness.
arXiv:2605.21751v2 Announce Type: replace
Abstract: Text-to-optimization requires two separable capabilities: modeling -- choosing the right optimization structure -- and binding -- grounding every coefficient, index, and parameter in the concrete problem data. We study this via Text2Opt-Bench, a scalable benchmark of solver-verified optimization problems spanning 12 categories, from textbook linear programs to stochastic and multi-objective formulations with up to thousands of variables. Across 10+ models, we find that accuracy collapses as instance data grows, even when the formulation itself is simple. We call this the effective binding limit. We study it with a family of techniques, BIND, that externalize numeric data to structured files so the model binds data programmatically rather than transcribing from the prompt. When using an oracle for externalizing data, we recover between 12 and 27 accuracy points, confirming binding as a key -- but recoverable -- failure mode. In a deployable setting without oracle access, we validate our hypothesis by finetuning a model exclusively on binding and show that it outperforms end-to-end SFT and RL across three structurally distinct optimization categories, with a 1.5B binding specialist alone matching a 7B end-to-end baseline.