Forskningsradar

Science Journals

Peer-reviewade publikationer — 63965 artiklar

PolicyGuard: From Organizational Policies to Neuro-SymbolicCompliance Review Engines
arXiv:2606.32004v1 Announce Type: new Abstract: Policy-grounded document review requires determining whether a target document complies with organization-specific policies, guidelines, or playbooks. While large language models can assist with policy interpretation and document analysis, end-to-end prompting leaves the applied policy logic implicit, making compliance decisions difficult to inspect, update, and test. We present PolicyGuard, a neuro-symbolic framework for policy-grounded document compliance review. PolicyGuard converts organizational policy guidance into an executable review engine consisting of typed relational logic rules and atom-level extraction questions. During review, LLMs answer these local questions using retrieved document evidence, and a symbolic evaluator applies the formal rules to detect non-compliance. We instantiate and evaluate PolicyGuard on company-specific NDA compliance review, where contract clauses must be checked against organization-specific negotiation policies. By separating policy formalization, local document interpretation, and symbolic compliance evaluation, PolicyGuard makes document review more explicit, maintainable, and systematically testable.
Probabilistic Inversion with Flow Matching
arXiv:2606.31288v1 Announce Type: new Abstract: We demonstrate the application of Flow Matching, a technique originating from generative Artificial Intelligence, to probabilistic inversion in geophysical settings, such as seismic Full-Waveform inversion. We adapt the well-established mathematical theory of Flow Matching from generative Artificial Intelligence to the context of probabilistic inversion. We evaluate the approach with two case studies: a simple 2D velocity model to illustrate the general features of the method, and the OpenFWI dataset to show its capabilities for probabilistic inversion of more complex seismic velocity models.
PEERS: A Parallel and Exact Effective Resistance Solver via Implicit Inversion and Augmented Symbolic Analysis
arXiv:2606.31535v1 Announce Type: new Abstract: High-precision effective resistance computation is a cornerstone of Electronic Design Automation (EDA) sign-off, yet it remains a fundamental bottleneck in large-scale power grid analysis, spectral sparsification, and circuit reliability. Existing approaches face a prohibitive "precision-memory impasse": approximate methods lack the stringent accuracy required for high-stakes industrial sign-off, while exact methods either suffer from redundant query overheads or trigger $O(n^2)$ memory explosions. To resolve this, we propose PEERS, a Parallel and Exact Effective Resistance Solver powered by an implicit inverse computing model of the Cholesky factor. By integrating a state-inherited augmented depth-first search (DFS) with a dynamic query update mechanism, PEERS eliminates numerical redundancy and evaluates all-edge resistance queries in a single parallel sweep. We provide a rigorous Work-Span analysis, proving that for graphs satisfying an $O(n^\alpha)$ separator theorem, PEERS achieves a theoretically optimal parallel span of $O(n^\alpha)$ while strictly maintaining $O(nnz(L))$ space complexity. Numerical evaluations on industrial benchmarks demonstrate that PEERS achieves an average speedup of 83.3x over state-of-the-art parallel solvers under identical memory constraints. Notably, PEERS processes a 1-million-node industrial graph in just 18.8 seconds and scales to 17 million nodes in under an hour, providing the first computationally feasible path for exact all-edge resistance analysis in multi-million-gate designs.
Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments
arXiv:2606.32009v1 Announce Type: new Abstract: Vision-language-action (VLA) models across robot embodiments require high-quality observation--action supervision to learn deployable action distributions, yet scaling such robot data remains difficult, especially for high-DoF humanoids. Teleoperation provides controller-aligned supervision, while human egocentric videos capture diverse bimanual manipulation but do not directly provide executable robot actions. We introduce Human-as-Humanoid, a human-to-humanoid supervision framework that enables near-real-time human-centric action generation, making human demonstrations usable for high-DoF humanoid VLA training by jointly aligning the robot embodiment, the sensing setup, and the action-label interface. Built on PrimeU, a human-aligned 60-DoF upper-body humanoid, Human-as-Humanoid uses synchronized ego-exo videos to pair deployment-aligned egocentric observations with exocentric motion recovery, retargets the recovered human motion through staged Inverse Kinematics (IK) into controller-aligned 60-DoF action chunks, and trains the VLA model with Forward Kinematics (FK)-aware supervision to preserve wrist and fingertip task-space geometry. This converts large-scale human demonstrations from visual observations into executable observation--action supervision for the target humanoid. Experiments validate the conversion chain at the motion-recovery, robot-action-space, and real-robot deployment levels. Human-as-Humanoid yields a 4.8--7.2x raw demonstration-throughput gain over humanoid teleoperation in our data-collection analysis, and on several downstream tasks, policies post-trained only with the converted human labels generalize to real-robot deployment without target-task robot demonstrations. The official project website is available at https://zgc-embodyai.github.io/Human-as-Humanoid.
Cesium Based Laser-Atomic Oscillator
arXiv:2606.32019v1 Announce Type: new Abstract: We report the first demonstration of a laser-atomic oscillator with cesium (Cs) atoms. A laser-atomic oscillator (LAO) is analogous to an active mode-locked laser with a self-excited modulator, i.e. atoms, at a ground-state hyperfine transition frequency. Therefore, a LAO can be configured as the simplest active atomic clock or a self-oscillating, earth-field atomic magnetometer that delivers oscillation signals both optically and electrically. With the current experimental Cs-LAO setup, when it is configured as an atomic clock using the 0--0 hyperfine transition, the short-term fractional frequency instability is around 10$^{-10}$ level. When it is configured as a self-oscillating magnetometer using a magnetically-sensitive hyperfine transition, the magnetic field sensitivity is around 100 fT/$\sqrt{\rm{Hz}}$ at 60 Hz. The presented Cs-LAO uses a cavity length from $\sim6.5$ cm to $\sim11.4$ cm. Ultimately, the minimal length of a Cs-LAO device can be $\leq1.63$ cm. Our new efforts unlock the potential of building truly chip-scale atomic clocks and magnetometers.
MediEncoder: Nonlinear Representation Learning for High-Dimensional Causal Mediation Analysis
arXiv:2606.30648v1 Announce Type: cross Abstract: Causal mediation analysis decomposes a treatment effect into indirect pathways through mediators and direct pathways not operating through them. Modern biomedical studies often involve high-dimensional covariates and mediators that are noisy proxies for lower-dimensional latent biological processes. Existing methods typically rely on sparsity, linear factor models, or ignore the connection among variables in the learned representations, which can be restrictive when measurements are nonlinear and covariate and mediator factors are structurally dependent. We propose MediEncoder, a representation-learning framework for nonlinear high-dimensional mediation analysis. MediEncoder jointly learns low-dimensional covariate and mediator representations using a coupled encoder-decoder architecture with a cross-factor network that links treatment and covariate representations to mediator representations. The learned features are then used in a cross-fitted efficient influence function-based estimator of natural direct and indirect effects. The resulting estimator is multiply robust and asymptotically normal under suitable regularity conditions. Simulations show that MediEncoder improves estimation accuracy over competing dimension-reduction approaches, and an application to Alzheimer's Disease Neuroimaging Initiative data illustrates its utility in high-dimensional biomedical causal mediation analysis.
Cascading Impacts of the USA--China Trade War on Global Oilseed Supply Chain
arXiv:2606.30685v1 Announce Type: cross Abstract: Global supply chains are highly interconnected, making them vulnerable to cascading disruptions induced by trade policy shocks. Understanding how such disruptions propagate through production networks, and how mitigation mechanisms such as trade reallocation and production adjustment can alleviate their impacts, remains a central challenge. In this work, we develop a linear programming formulation of an Input-Output (IO) system that captures cascading supply-chain disruptions together with trade reallocation and production expansion. Our formulation yields a system-level equilibrium characterization that enables the joint analysis of disruption propagation and mitigation within a unified framework. We propose an efficient algorithm for computing approximate equilibrium solutions by minimizing total unmet demand in large IO systems. We apply our approach to tariff-induced disruptions in the global oilseeds supply chain arising from the U.S.-China trade war. Our results show that a localized 70% disruption to flows from the U.S. oilseeds sector to China leads to a 3.27% loss in global output, with China experiencing a disproportionate loss of 14.02%. As a counterfactual mitigation strategy, allowing a 20% reallocation from Brazil's oilseed sector to China significantly reduces global output losses to 1.36%, although pressure remains high on final-demand flows. We further investigate production expansion as an additional mitigation mechanism and show that it introduces tradeoffs between reducing global final-demand losses and protecting Brazil's domestic flows. Domestic reallocation disproportionately shifts losses toward smaller economies, while globally sourced expansion redistributes losses more broadly across the network.
HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)
arXiv:2606.31325v1 Announce Type: new Abstract: We present HistoriQA-ThirdRepublic: a French-language dataset of multi-hop historical questions derived from parliamentary debates and newspapers of the French Third Republic. Designed in collaboration with a historian, the corpus captures complex reasoning patterns typical of historical inquiry, including cross-source synthesis, temporal reasoning, and the integration of sparse evidence. The dataset is made of 1782 questions and emphasizes multi-hop connections across heterogeneous historical documents, providing a resource for evaluating retrieval-augmented and large language model systems in domain-specific contexts. We describe the methodology for constructing the corpus, including the selection and alignment of sources, question validation, and metadata integration. While the dataset focuses on French historical documents, our methodology can be readily adapted to other languages and national corpora. Finally, we demonstrate how the corpus can support realistic evaluation scenarios for multi-hop question answering, bridging the gap between NLP benchmarks and the needs of historical scholarship.
BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding
arXiv:2606.31315v1 Announce Type: new Abstract: Speculative decoding accelerates inference by using a lightweight draft model to generate candidate tokens in parallel, and are then verified by the target model, enabling lossless acceleration. Recently, diffusion-based speculative decoding further improves parallelism by generating multiple tokens per forward pass via block-level diffusion, achieving state-of-the-art (SOTA) performance. However, existing methods adopt a fixed inference block size and assume a uniform optimal decoding strategy across all inputs. In this paper, we show that this assumption is suboptimal, as the optimal block size varies across samples and plays a critical role in speculative decoding performance. Moreover, these values exhibit a clear local structure, concentrating around the training block size, which reduces the problem to a low-dimensional and structured decision space. Based on these insights, we propose BlockPilot, a sample-adaptive policy that predicts the optimal block size from the prefilling representation. Specifically, we formulate block size selection as a lightweight policy learning problem and propose an instance-adaptive decision mechanism that predicts the optimal block size based on the representation of the prefilling stage. The prediction is performed only once after prefilling, allowing for seamless integration. Extensive experiments demonstrate that our method is plug-and-play, introduces minimal overhead, and consistently improves efficiency, achieving an acceptance length of 5.92 and a 4.20$\times$ speedup on Qwen3-4B under temperature $T=1$.
Computed materials proposals depart from the structural memory of experimental discovery
arXiv:2606.30967v1 Announce Type: cross Abstract: Generative AI and high-throughput DFT pipelines propose millions of inorganic crystal structures, but lack a calibrated reference frame against experimentally realized chemistry. Here we embed 167,500 Inorganic Crystal Structure Database entries in a continuous structural-similarity space, partition it into graph communities, and replay them in time. Experimental discovery shows strong structural memory: 82.9% of new formulas enter pre-existing communities; new-community formation falls from 40.2% (1930s) to 2.6% (2010s). The communities are chemically meaningful, positively identifying nine textbook field-defining renaissances, including cuprates, colossal-magnetoresistance manganites, MAX phases, and Li-ion battery cathodes. Projecting GNoME, MatterGen-public, Materials Project, JARVIS-DFT, and Alexandria-PBE into frozen historical maps yields a cutoff-robust ordering: held-out ICSD > MatterGen > {GNoME ~ MP-theoretical} > JARVIS > Alexandria. Structural departure from experimental basins is not specific to generative AI but general across the tested computed sets. Combining structural proximity with reduced-formula precedent defines a historical synthesizability prior for triaging computed materials.
Hierarchical Clustering As a Novel Solution to the Notorious Multicollinearity Problem in Observational Causal Inference
arXiv:2606.30992v1 Announce Type: cross Abstract: Multicollinearity is a long lasting challenge in observational causal inference, especially in regressions -- highly correlated independent variables make it hard to isolate their individual impacts on outcomes of interest. While common solutions such as shrinkage estimators and principal component regressions are helpful in prediction problems, a crucial limitation hinders their applicability to causal inference problems -- they cannot provide the original causal relationships. To fill the gap, we present an innovative and intuitive solution, by employing hierarchical clustering to aggregate data in a way that effectively alleviates collinearity. This method is generally applicable to causal problems featuring multicollinearity. We use a marketing application to demonstrate how and why it works. Expenditures on different advertising channels often exhibit correlations, making it exceedingly difficult to separately measure their impact. Many previous studies proposed to leverage granular cross-sectional data for better identification but, to our knowledge, none explicitly addressed multicollinearity, which undermines causal identification even with granular data. We propose to hierarchically cluster geographic units based on marketing spend correlation to reduce collinearity, and to implement a Bayesian Marketing Mix Model with cluster-level data. Such clustering happens in two steps -- we first normalize and demean geo-level data to establish a common scale and to eliminate the common trends; we then calculate pairwise distance to summarize marketing spend correlation between geos and cluster the ones with moderate to strong correlation. Both descriptive evidence and regression analysis affirm that such hierarchical clustering effectively mitigates collinearity and facilitates the separate identification of the impact of different marketing channels.
Robust Aggregation of Calibrated Forecasts
arXiv:2606.31020v1 Announce Type: cross Abstract: Decision-makers often rely on multiple probabilistic forecasts that are individually calibrated but need not be fully informative. We develop a framework for aggregating such forecasts when the decision-maker knows only that experts satisfy calibration. We show that the joint distribution of calibrated forecasts can contain decision-relevant information that is unavailable from any single expert, so the standard optimal-in-hindsight (OIH) benchmark may substantially understate attainable performance. To formalize this idea, we introduce a robust max-min benchmark: the best payoff a decision-maker can guarantee against all profile-wise conditional-mean mappings compatible with calibration. This benchmark is tractable, admits a linear-programming formulation, and dominates the OIH benchmark up to calibration error. It can nevertheless be strictly below the Bayesian benchmark, clarifying the value of knowing experts' information structures. Finally, we provide online algorithms that attain the robust benchmark under forecast-only feedback and stronger contextual benchmarks under state feedback.
A Self-Negotiation Framework for Ethical Decision-Making during Task Interruptions in Service Robots
arXiv:2606.31357v1 Announce Type: new Abstract: Service robots operating in public environments frequently encounter interruptions when multiple users request service simultaneously. Resolving such conflicts requires ethical decision-making, as prioritizing one user request can disadvantage another. Current approaches rely on static rules or centralized arbitration and do not support autonomous, ethics-based conflict resolution. This paper addresses the question of how a single robot can arbitrate between multiple users during task interruptions and make ethically aligned decisions without relying on external coordination. We introduce a self-negotiation framework that represents each user by an ethical profile that captures their contextual ethical preferences and conditions, and resolves conflicts through an internal negotiation process. The framework is implemented in a modular ROS-based implementation and evaluated in simulation with a realistic interruption scenario. The results show that the system consistently produces user ethical preference-aligned outcomes, supports multilateral negotiation among users, and responds within 1.5 seconds, with near-linear runtime growth under increasing user input.
Beyond binary scission: a generalized three-species cascade breakage model for wormlike micellar solutions
arXiv:2606.31059v1 Announce Type: cross Abstract: Wormlike micellar fluids exhibit complex rheological behavior driven by the continuous breakage and recombination of self-assembled micellar networks. Existing two-species models provide a coarse binary representation of the micellar population, limiting their ability to resolve intermediate structural states and broad relaxation spectra. To address this limitation, we develop a three-species cascade breakage model consisting of gel-network, long chains, and short chains. By introducing an intermediate micellar state, the model links the rapid relaxation of short fragments to the slow recovery of the gel-network within a unified kinetic framework. This additional structural pathway gives rise to a three-mode viscoelastic response, improves the high-frequency description of the dynamic moduli, and produces a non-monotone constitutive curve that evolves into a stress plateau with coexisting shear bands in Couette flow. This cascade mechanism also governs the transient response, including stress overshoot, hysteresis, and multistep relaxation after shear cessation. Overall, the proposed three-species model provides a physically interpretable framework for worm-like micellar shear banding, capturing the connection between cascade microstructural evolution, broad relaxation dynamics, and macroscopic flow localization.
A Spectral Solver for Acoustic Scattering by Multiple Quasi-Axisymmetric Structures
arXiv:2606.31380v1 Announce Type: new Abstract: Acoustic scattering arises in a wide range of applications, including medical imaging, geophysical exploration, acoustic metamaterials, etc. In this paper, we develop a fast and highly accurate algorithm for acoustic scattering by multiple quasi-axisymmetric objects, whose axis of rotation is an arbitrary curve. The method is based on a Nystr\"om discretization that combines Gauss-Legendre quadrature with the trapezoidal rule. To treat the singular integrals that occur when target points are close to or coincide with source points, we reformulate them as evaluations of the modal Green's function and its derivatives, which are computed efficiently using the fast Fourier transform and convolution. The multiple scattering solver is then constructed by coupling the single scatterer discretizations through inter-body boundary integral interactions. We also present a convergence analysis for scattering problems with smooth geometries. Numerical examples demonstrate the efficiency and accuracy of the proposed method for solving multiple scattering problems involving up to 1000 quasi-axisymmetric structures.
Conditionals and Modalities in Constructive Quantum Logics
arXiv:2606.31853v1 Announce Type: cross Abstract: We investigate logics that generalize both intuitionistic logic and quantum logic. In earlier work, we introduced Ex-logic, an extension of Holliday's fundamental logic that coincides with the intersection of orthologic and the implication-free fragment of intuitionistic logic. In this paper, we add an implication connective to Ex-logic and axiomatize iEx-logic, the intersection of full intuitionistic logic and orthomodular logic with the implication connective interpreted as the Sasaki hook. As a consequence, we obtain a characterization of the lattice of logics extending iEx-logic as the product of the lattice of intermediate logics and the lattice of orthomodular logics. We also explore the robustness of our algebraic approach by briefly discussing extensions of iEx-logic with modal operators.
Quantum Derivative Pricing for SPDEs via BDSDE Representation
arXiv:2606.31076v1 Announce Type: cross Abstract: We study quantum speedups of derivative pricing for stochastic partial differential equation (SPDE) models through their backward doubly stochastic differential equation (BDSDE) representations. We develop conditional and nested quantum-accelerated multilevel Monte Carlo (QA-MLMC) methods for estimating the resulting conditional and nested expectations, improving the sampling complexity of classical Monte Carlo methods from $\widetilde{O}(\epsilon^{-2})$ to $\widetilde{O}(\epsilon^{-1})$ within additive error $\epsilon$. We apply the framework to derivative pricing and sensitivity analysis, providing quantum-accelerated estimators for prices as well as first-order and second-order Greeks, likelihood-ratio and Malliavin-weight representations for Greeks, and Heston-type stochastic-volatility models. To enable efficient multilevel coupling, we construct a family of Forward--Backward Taylor discretization schemes for the stochastic integrals arising in the BDSDE representations and establish global strong-error order one convergence for pricing and Greek estimators. Numerical experiments showcase our schemes for first-order and second-order Greeks can reach the required orders for the full quadratic quantum speedups.
Towards Inclusive Mobility Modeling: Characterizing and Evaluating Elderly Trajectory Patterns in Urban Systems
arXiv:2606.31207v1 Announce Type: new Abstract: The rapid advance of smart cities increasingly depends on trajectory data mining, yet underrepresented demographic groups, particularly the elderly, are often sparsely represented in public mobility datasets. This underrepresentation can introduce systematic bias into mobility modeling and downstream urban planning. Using the 2016-2020 Jersey City subset of the Citi Bike System Data, this study quantitatively examines how the absence of underrepresented subgroups' mobility signatures affects mobility modeling, using synthetic trajectory generation as a case study. The analysis reveals that elderly riders exhibit a structurally distinct mobility signature, including localized activity spaces (958 m vs. 1,189 m for young riders), lower mobility entropy (1.82 vs. 4.15), and asymmetric off-peak temporal patterns. To demonstrate that relying on majority-dominated training data yields biased synthetic outcomes, we further evaluate both a first-order Markov chain and a Qwen3-4B model fine-tuned with QLoRA across three demographic training settings: the full population, young riders only, and elderly riders only. Results show that models trained on majority-dominated populations systematically misrepresent elderly mobility behavior, particularly for spatial mobility metrics. The Markov model trained on the full population overestimates elderly step length by 4.5% and dwell time by 8.9%, whereas the elderly-specific model achieves substantially lower errors across most metrics. Comparisons between the Markov and LLM-based frameworks further show that higher-capability models do not necessarily improve subgroup-level fidelity under limited demographic data. These findings underscore the importance of demographic representation in mobility modeling and its downstream applications for underrepresented populations.
Full-Wave Green's-Function Modeling of Collective Single-Photon Emission in Non-Markovian Open-System QED with Finite-Bandwidth Compensation of Dispersive Interactions
arXiv:2606.31317v1 Announce Type: cross Abstract: This work presents a full-wave Green's function framework for modeling collective and coherent single-photon emission from multiple quantum emitters embedded in complex electromagnetic structures. Starting from a transverse modal completeness relation of modified Langevin noise formalism, we derive a closed set of coupled equations for population dynamics and frequency-resolved field amplitudes in the single-excitation regime. Since the electromagnetic reservoir is not traced out at the level of the dynamical amplitudes, the emitted single-photon dynamics can be modeled within the same closed set of equations without Markovian approximation in open and dissipative environments. We demonstrate that finite-bandwidth truncation of the spectral density leads to systematic deviations in coherent dispersive interactions, even when dissipative rates appear converged. To restore causal consistency, we introduce a counter-term compensation scheme that restores the missing dispersive contributions without modifying the retained non-Markovian memory kernel. To validate the scheme and demonstrate the practicality of the proposed framework, we present numerical examples ranging from benchmark configurations to a three-dimensional dispersive ring-resonator structure via finite element method. These capabilities provide a practical route for rigorously incorporating full-wave electromagnetic simulations into non-Markovian multi-emitter quantum electrodynamics, enabling predictive modeling of collective emission, coherent energy exchange, and single-photon radiation in realistic open structures.
Robustness of neural networks to random noise perturbations of their inputs
arXiv:2606.31581v1 Announce Type: new Abstract: We investigate the problem of the robustness of a trained neural network to the perturbation of its input values. More specifically, we examine the interplay between the accuracy of the network, as measured by the mean squared error, and robustness. Accordingly, we present a robustness measure, which, with high probability, suggests an upper bound on the mean squared error of the network, with respect to an input data set, for a given perturbation of the input values of the network. The measure we propose is both simple and efficient to compute, treating the neural network as a black box. We provide experimental results on several real-world data sets showing the efficacy of the proposed method. We also introduce the concept of robustness curves, which allows us to further analyse robustness within and between data sets.
Relativistic magnetohydrodynamics from kinetic theory
arXiv:2606.31327v1 Announce Type: cross Abstract: This thesis develops a kinetic-theory framework for relativistic dissipative magnetohydrodynamics under strong electromagnetic fields, motivated by quark-gluon plasma in heavy-ion collisions. Starting from the relativistic Boltzmann-Vlasov equation and using the method of moments within the 14-moment approximation, it derives causal second-order hydrodynamic equations for relativistic plasmas with increasing generality. The work first review relativistic dissipative hydrodynamics and its kinetic foundations, emphasizing the need for Israel-Stewart-type transient theories to preserve causality and stability. Electromagnetic fields are then introduced at the microscopic level, where the Lorentz force modifies the moment hierarchy and produces anisotropic transport effects absent in field-free fluids. Next, it develops relativistic dissipative magnetohydrodynamics for a non-resistive two-component plasma of oppositely charged particles. Here, the magnetic field couples the dissipative sectors of the two species, generating relative dissipative currents and coupled shear dynamics. For Bjorken expansion, the theory predicts damped oscillations in the transverse shear sector associated with cyclotron motion. Finally, the thesis treats the resistive two-component case, where the electric field evolves dynamically and couples to charge diffusion and shear stress. The resulting theory reveals current-shear feedback, transient electromagnetic generation of momentum anisotropy, and underdamped dissipative oscillations. Applications to homogeneous and Bjorken-expanding plasmas show how resistive and electromagnetic effects modify evolution beyond standard hydrodynamics. Overall, the thesis extends relativistic dissipative hydrodynamics to magnetized and resistive plasmas, providing a microscopic foundation for future studies of strongly magnetized quark-gluon plasma and astrophysical systems.
CooperScene: Multi-Modal Cooperative Autonomy Benchmark with C-V2X Communication Characterization
arXiv:2606.31219v1 Announce Type: new Abstract: Cellular vehicle-to-everything (C-V2X) enables cooperative perception, prediction, and planning beyond the field of view of individual agents. However, existing datasets often overlook the complexities of real-world deployment, such as limited communication bandwidth and its dynamics, heterogeneous sensing modalities, and scalability beyond a single cooperative partner. In this paper, we introduce CooperScene, a high-fidelity cooperative autonomy dataset with real-world C-V2X communication characterization. The dataset is organized into diverse scenes, including intersections, highway ramps, and parking lots. These scenes involve three connected and autonomous vehicles (CAVs) and one infrastructure roadside unit (RSU), all equipped with multi-modal sensors and commercial off-the-shelf C-V2X communication radios. All scenes are annotated with globally consistent 3D labels at 10 Hz, totaling 344K objects across 59K frames, underpinned by tight sensor- and agent-synchronization, centimeter-level localization and spatial alignment, precise cross-modality calibration, and 3GPP-standard-compliant C-V2X communication. CooperScene establishes a rigorous benchmark for evaluating multi-agent scaling and actual performance in real-world deployable settings. Project website for data and benchmark: https://cisl.ucr.edu/CooperScene
Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck
arXiv:2606.31411v1 Announce Type: new Abstract: Rapid advancements in generative speech technology have compromised the reliability of voice biometrics. While current spoofing detectors excel when assessed under in-domain conditions, generalisation to out-of-domain settings is often poor. We show that this can be due to linguistic bias. A reliance on linguistic cues observed in training data can then compromise robustness to cross-data. We propose a linguistic-invariant spoofing detection framework utilizing teacher-student adversarial learning. The linguistic-aware teacher model, pre-trained on linguistic content of an external dataset, guides the student detector via gradient reversal to minimize the linguistic information. To prevent the inadvertent removal of non-linguistic cues, we incorporate a Variational Information Bottleneck to enable suppression of principal cues. Across nine DF Arena datasets, our method achieves up to a 36.2% relative reduction in the EER compare to the baseline.
Beyond Cross-Reconstruction: Probing-Based Disentanglement Evaluation for Acoustic Teleportation Codecs
arXiv:2606.31365v1 Announce Type: cross Abstract: Some neural audio codecs disentangle speech into latent subspaces encoding content, speaker identity, and acoustics, enabling acoustic teleportation and voice conversion. Existing evaluations rely on cross-reconstruction quality, which cannot reliably detect leakage across partitions. We extend a probing based framework to assess disentanglement by regressing room-acoustic parameters (reverberation time, clarity, and direct-to-reverberant ratio) and classifying speaker identity, using the gap between intended and unintended partitions as the disentanglement measure. Applied to an acoustic teleportation codec, we find speaker identity is largely confined to its partition, while acoustics leak into the speech embeddings due to the training objective. Acoustic embeddings blindly estimate room parameters within 0.02 s of supervised baselines, indicating physically meaningful structure emerges without explicit supervision.
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model
arXiv:2606.31247v1 Announce Type: new Abstract: Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame rates (e.g., 25 or 12.5 Hz), ignoring the time-varying information density of speech and offering no flexibility to trade off quality for speed at inference time. Recent audio tokenizer research has proposed dynamic frame rate speech coding, which exploits this non-uniformity and enables two new capabilities: very low average frame rates and frame rate controllability. However, this technique has not yet been applied to SLMs. We introduce Flexible Spoken Language Model (FlexiSLM), the first SLM that supports dynamic and controllable frame rates on both speech input and output. Using dynamic frame rate representations, FlexiSLM outperforms fixed-frame-rate 7B models including Qwen2.5-Omni and Kimi-Audio at its high-quality operating points. We further verify that FlexiSLM can be accurately steered down to 4.0 Hz; at 6.25 Hz, it roughly halves inference time relative to 12.5 Hz while retaining strong speech-to-speech quality. Audio samples are available at https://flexislm.github.io .