Forskningsradar

Science Journals

Peer-reviewade publikationer — 53899 artiklar

CART Random Forests as Sequential Allocation over Random Opportunity Sets: A Stochastic-Control Theory of Ensemble Risk
arXiv:2605.26675v1 Announce Type: cross Abstract: CART random forests are among the most widely used modern predictive methods, with well-documented empirical success. Yet, at the mechanistic level, the algorithm is often treated as a black box because of its complexity. In this paper, we develop a stochastic-control perspective on feature-subsampled CART random forests, named CART random opportunity-set allocation (CART-ROSA). At each node, the random subset of features is interpreted as a random feasible action set, and the CART split rule as a masked-action allocation policy. This policy induces a controlled stochastic process over informative split-count states, whose terminal law determines both single-tree error and cross-tree interaction terms in the forest mean squared error (MSE). Such representation opens the black box of CART-forests by separating two design levers: the informative-opportunity rate induced by feature subsampling, and the contraction strength from the within-mask split policy. We establish that the CART policy is locally stabilizing: it contracts imbalances in informative split allocations and concentrates terminal tree geometry. At the system level, however, it can be globally suboptimal for the forest objective. Specializing to the linear model, we derive the MSE risk expansion explicitly. Our results show how an operations-research perspective makes tractable a theoretical gap difficult to access from the standard algorithmic description of CART forests.
Integrated squeezed light sources for two-mode entanglement in thin-film lithium niobate
arXiv:2605.26583v1 Announce Type: cross Abstract: Scalable generation of nonclassical light sources on an integrated platform is a key requirement for photonic quantum information processing. In particular, realizing multiple indistinguishable squeezed light sources on a single chip is an essential step toward continuous-variable quantum computing. Here, we demonstrate the fabrication of two indistinguishable and independently controllable optical parametric oscillators on a thin-film lithium niobate (TFLN) platform. The device design focuses on reproducibility, independent tunability, and compatibility with larger telecom-wavelength continuous-variable photonic circuits. We observe up to 0.5 dB of directly measured squeezing below the shot-noise level from each source. By interfering the two modes on a beam splitter, we generate an EPR-type two-mode squeezed state and verify continuous-variable entanglement through violation of the Duan-Simon inseparability criterion. This is the first demonstration of two independently tunable squeezed-light sources on a single TFLN chip and their use for generating continuous-variable entanglement.
Target-Oriented Statistical Compression: Sufficiency, Reverse Martingales, and Sequential Monitoring
arXiv:2605.26568v1 Announce Type: cross Abstract: Statistical procedures rarely retain all features of the observed data. A sufficient statistic removes information irrelevant to a parameter; a maximum likelihood estimate compresses an empirical objective into an optimizing point; and a hidden state in a sequential model compresses past observations into a learned representation. This article develops these practices under the unified notion of \emph{target-oriented statistical compression}: a useful summary preserves what matters for an inferential, predictive, or decision-relevant target, rather than every detail of the realized data path. The central object is the conditional target process \(M_n=\E(Z\given\G_n)\), where \(Z\) is the target and \(\G_n=\sigma(T_n)\) is the information retained by the compression map \(T_n\). When \((\G_n)\) is a decreasing filtration, \((M_n)\) is a reverse martingale with limit \(M_\infty=\E(Z\given\G_\infty)\). Exact sufficiency corresponds to lossless compression, while approximate summaries such as penalized estimators, principal components, and neural-network hidden states produce reverse quasi-martingale defects measuring coherence loss across compression levels. The diagnostic \(r_n=|M_n-M_{n-1}|\) is treated as an observable stability proxy, not as an unbiased estimator of the theoretical defect. Boundary degeneracy in sequential binary problems is developed as a central application. Practical boundary claims require joint assessment of boundary closeness, uncertainty control, and trajectory stability. The companion paper \citet{chang2025rm} develops the corresponding stopping procedures, finite-sample bounds, and numerical evidence; the present paper provides the broader theoretical infrastructure and extends the framework to Gaussian, Poisson, and quasi-martingale monitoring problems.
Random neural networks match observed dimensionality of neural population recordings and motivate stronger experimental tests
arXiv:2605.26551v1 Announce Type: cross Abstract: Randomly connected neural networks have long served as a theoretical tool for studying collective dynamics in neural populations, yet quantitative comparisons to experiments remain limited. Recent technological advances have made it possible to resolve population-wide correlations across neurons, and minimal models such as random neural networks predict their generic structure. Whether the two agree quantitatively remains untested. In this work, we examine whether a minimally structured random neural network can account for the low dimensionality of activity in neural population recordings by building on recent developments in Dynamical Mean-Field Theory and incorporating two additional experimentally relevant features into the model: finite measurement time and variability across behavioral contexts. We show that, when these factors are included, the dimensionality measured from large-scale recordings is consistent with the values predicted by random models. However, current recording durations make it difficult to use dimensionality to discriminate among connectivity structures. We further show that analytically predicted dimensionality varies non-monotonically with external input strength, and that the orientation similarity between neural manifolds recorded under different behavioral contexts can be more sensitive to network structure than dimensionality is. Together, these results provide quantitative guidance for experimental design to infer the connectivity structure underlying population activity.
Conditions for domain-free negative capacitance
arXiv:2605.26536v1 Announce Type: cross Abstract: While negative capacitance has been demonstrated in ferroelectric-dielectric heterostructures in the form of capacitance enhancement, all experimental evidence, to date, suggests the existence of domains therein. Here, we address the question: what are the conditions to achieve ideal, domain-free negative capacitance in ferroelectric-dielectric heterostructures? Our main claim is that for given thicknesses of the ferroelectric and the dielectric layers, there is a critical value of domain wall energy parameter -- above which the system would be stabilized in an ideal and robust domain-free negative capacitance state and would be robust against domain formation. Our analyses suggest that to achieve ideal negative capacitance, efforts should lie in understanding the means to control the domain wall energy on all fronts, both theory and experiments via high throughput design, discovery, and engineering of ferroelectrics.
Foundations of a Time-Consistent Counterfactual Actuarial Runtime for Autonomous AI Agents
arXiv:2605.26508v1 Announce Type: cross Abstract: We propose a foundational runtime actuarial layer for autonomous AI agents in which every side-effect-bearing action carries a time-consistent, counterfactual risk toll computed against a contractually fixed safe default, inside an explicit underwriting boundary. The framework treats per-action insurance as the primary unit of analysis and replaces post-hoc annual liability cover with a pre-action transaction layer. The paper establishes four structural results: (i) a well-defined counterfactual toll under a chosen safe-default mapping and continuation policy, with explicit non-uniqueness; (ii) a no-splitting property within an underwriting boundary that telescopes path-decomposed actions into a boundary potential, with a corollary tying gaming-resistance to boundary design; (iii) an irreversible-authority premium, split into a strictly positive action-level component and an if-and-only-if characterisation of the set-level robust capital increase; and (iv) a conservative runtime gating theorem that translates high-probability toll envelopes into an executed-action budget guarantee. The result is the mathematical base layer for a broader program: an empirical companion instantiates the runtime through an Actuarial Action Interface and authority-frontier experiments; a mechanism-design companion studies strategic operator incentives and cross-boundary aggregation; and a dynamic-underwriting companion studies experience rating and audit-replay calibration. The present paper states the primitive contract, the toll identity, the within-boundary no-arbitrage result, and the budget guarantee on which those later layers depend.
Negligible in Size, Significant in Effect: On Scale Vectors in Large Language Models
arXiv:2605.26895v1 Announce Type: new Abstract: Normalization layers in modern large language models (LLMs) consist of a deterministic normalization operation and a learnable scale vector. While the normalization operation has been extensively studied, the scale vector remains poorly understood despite its ubiquitous use. In this work, we present a systematic study of scale vectors in LLMs from the perspectives of expressivity, optimization, and architectural structure. First, we show empirically that although scale vectors constitute only a negligible fraction of model parameters, removing them substantially degrades LLM pre-training. Our theory further shows that, in Pre-Norm architectures, scale vectors do not increase expressivity; instead, they improve optimization through a self-amplifying preconditioning effect on subsequent linear mappings. Second, we investigate the role of weight decay for scale vectors. By distinguishing Input-Norm and Output-Norm layers, we theoretically show that weight decay is beneficial for the former but harmful for the latter, due to their distinct roles in optimization and expressivity. Third, motivated by this understanding, we propose three lightweight and complementary improvements to scale vectors: branch-specific heterogeneity, improved placement around linear mappings, and magnitude-direction reparameterization. Both theory and experiments show that each improvement yields consistent gains. Finally, we combine these improvements into a unified scale-vector strategy and evaluate it through extensive LLM pre-training experiments on dense and mixture-of-experts models ranging from 0.12B to 2B parameters, across multiple optimizers and learning rate schedules, under industrial-scale token budgets. The unified strategy consistently achieves lower terminal loss than well-tuned baselines and exhibits more favorable scaling behavior, while adding negligible parameter and computational overhead.
Megakernel vs Wavefront GPU Path Tracing
arXiv:2605.27323v1 Announce Type: new Abstract: Over the last decade, advances in GPU hardware have been driven in large part by the demands of real-time graphics, culminating in dedicated hardware ray tracing cores (RT cores). These units accelerate ray scene intersection queries directly in hardware, making physically based ray tracing algorithms increasingly practical for interactive applications. This paper compares and analyzes the performance of two ray-based rendering algorithms: forward path tracing (PT) and wavefront path tracing (WPT). GPU-based PT computes the color of each pixel by having each thread trace a single path to completion, naturally leading to a megakernel approach - while WPT maintains state buffers between specialized kernel invocations to trace path stages simultaneously. We find that WPT affords a ~16% speedup over PT in our implementation. By analyzing traces from NVIDIA Nsight Graphics, we attributed this speedup to WPT's improved cache locality compared to PT. We also find that our implementation does not achieve maximum GPU throughput across any of its units, suggesting that communication and memory latency, as well as synchronization, are the limiting factors. Finally, we address potential algorithmic improvements and future work for real-time path tracing implementation for practical applications.
TPS-Drive: Task-Guided Representation Purification for VLM-based Autonomous Driving
arXiv:2605.27038v1 Announce Type: new Abstract: Vision-Language Models (VLMs) provide a promising foundation for autonomous driving planning, yet bridging semantic reasoning and precise 3D spatial forecasting remains a critical challenge. Existing representation strategies generally follow two paths: text-aligned methods flatten continuous spatial states into symbols, which compromises geometric structure and induces "spatial hallucinations"; dense visual methods preserve spatial topology but overwhelm standard tokenizers with redundant background textures, leading to "representation interference". To address these limitations, we introduce TPS-Drive, a novel framework centered on Task-Guided Representation Purification that empowers VLMs to Think in Purified Space. At its core, an Agent-Centric Tokenizer utilizes a task-guided vector quantization mechanism supervised by a frozen 3D detection head, which explicitly reallocates limited codebook capacity from pervasive static backgrounds to critical dynamic agents and effectively isolates spatial redundancy. Leveraging this purified spatial vocabulary, TPS-Drive employs a decoupled reasoning pipeline that sequentially performs scene understanding, future forecasting, and action generation. The framework is optimized via a progressive three-stage training paradigm, culminating in reward-driven refinement that surpasses pure imitation learning. Extensive experiments validate our approach: TPS-Drive achieves accurate agent spatial state forecasting and reduces collision rates in open-loop nuScenes evaluations, while establishing new safety records on the rigorous closed-loop NAVSIMv1 and NAVSIMv2 benchmarks.
A collocation scheme that is equivalent to discontinuous Galerkin discretizations
arXiv:2605.27327v1 Announce Type: new Abstract: A spectral collocation operator with the summation-by-parts property was introduced by Chan to develop entropy-stable discontinuous Galerkin (DG) semi-discretizations (https://doi.org/10.1016/j.jcp.2018.02.033). The present work shows that semi-discretizations based on this collocation operator produce solutions that are equivalent to solutions of a DG semi-discretization using the same underlying quadrature. The equivalence holds regardless of the number of degrees of freedom in the collocation scheme and when the quadrature is not strictly positive. Extraneous degrees of freedom in the collocation scheme are associated with the nullspace of the operator and remain zero throughout an unsteady simulation. If necessary, nullspace consistency can be recovered by introducing projection-based numerical dissipation that targets only the extraneous modes. The equivalence between collocation and DG solutions is verified for the constant-coefficient advection equation and Burgers' equation on triangular meshes. The numerical results show that equivalence breaks down for entropy-stable semi-discretizations of Burgers' equation based on a skew-symmetric splitting, but that equivalence can be recovered by projecting the collocation scheme's residual onto the relevant polynomial space. In addition to investigating equivalence, the results demonstrate that the collocation operator produces semi-discretizations with favorable spectral radii compared with a commonly used summation-by-parts operator construction.
Governed Evolution of Agent Runtimes through Executable Operational Cognition
arXiv:2605.27328v1 Announce Type: new Abstract: Recent advances in agentic systems increasingly treat code as an executable operational substrate rather than as a disposable output artifact. Prior work such as \emph{Code as Agent Harness} frames validated agent-generated artifacts as runtime entities that can be created, executed, revised, persisted, and reused within long-running cognitive loops. However, the governance, lifecycle management, and operational evolution of such artifacts remain under-specified. This paper proposes a framework for governed runtime evolution in multi-agent systems through executable operational cognition. We formalize agent-generated artifacts as persistent runtime capabilities that progressively become part of the operational substrate rather than transient intermediate outputs. Building on this perspective, we introduce \emph{HarnessMutation} as a governed mechanism for lifecycle-aware runtime adaptation operating under explicit validation, traceability, evaluation, and rollback constraints. Rather than treating runtime adaptation as unrestricted self-modification, the proposed framework models evolution as a bounded and observable process over persistent operational memory. It further shows how these ideas can be operationalized over modern agent runtimes and governance-oriented orchestration systems, providing a conceptual foundation for adaptive infrastructures whose evolution remains explicit, auditable, and constrained.
Maat: The Agentic Legal Research Assistant for Competition Protection
arXiv:2605.27331v1 Announce Type: new Abstract: Competition law experts conducting legal research must review extensive volumes of cases, decisions, and judicial reports to identify precedents and assess key elements in competition and merger cases. Although general research assistants such as Claude and ChatGPT and legal assistants such as SaulLM-7B and LegalGPT are increasingly used to assist legal research, they remain inadequate for competition law analysis: they lack specialized domain expertise, provide insufficient official citations, or hallucinate competition law cases. We propose Maat, a ReAct agent that orchestrates tools corresponding to different tasks of the research process. Designed iteratively with competition law experts, Maat grounds cases and findings in official sources using RAG for reliability, provides rich in-line citations, falls back to web search when database coverage is insufficient, and prompts the user for clarification when queries are ambiguous. Maat significantly outperforms all baseline assistants on case-specific tasks and performs within range of the top baseline on theoretical question tasks. The dataset used is available on GitHub.
Muon-Catalyzed Nuclear Fusion: Physical Mechanism, Bottleneck Breakthroughs, and an Engineering Pathway
arXiv:2605.26432v1 Announce Type: cross Abstract: Muon-catalyzed nuclear fusion (\mucf) replaces atomic electrons with negative muons, compressing atomic orbitals by about two orders of magnitude and enabling deuterium--tritium (D--T) fusion under near-room-temperature conditions. This paper reviews the physical principles of \mucf{} and formulates its essential dynamics as a four-step cycle: muonic-atom formation, muon transfer, resonant \dtmu{} molecular formation, and D--T fusion with muon release and recycling. A kinetic model is used to quantify the number of catalysis cycles per muon and the corresponding energy gain. We focus on the central limitation of catalytic efficiency, namely the alpha-sticking effect, and discuss possible breakthrough routes including nuclear-spin and muon dual polarization, in-flight muon-catalyzed fusion, and heavy-ion-driven magneto-inertial fusion. Within the idealized assumptions of the present model, a four-dimensional synergistic scheme combining dual polarization, high-density confinement, electric-field-assisted muon recovery, and resonant enhancement may increase the number of catalysis cycles per muon from the present experimental record of about 150 to more than 500, potentially enabling an energy gain \(Q>2\). On this basis, we propose a conceptual fusion--fission fuel-breeding hybrid reactor, denoted as \mucf-FBR, which exploits the 14.1-MeV neutron yield of \mucf{} to breed \({}^{239}\mathrm{Pu}\) from a \({}^{238}\mathrm{U}\) blanket in a decoupled fusion--fission operating mode. This concept may offer advantages in engineering robustness, radiation-damage tolerance, and natural-uranium utilization.
FAB-Bench: A Framework for Adaptive RAG Benchmarking in Semiconductor Manufacturing
arXiv:2605.26476v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has become critical for knowledge-intensive applications, yet evaluating its performance in vertical domains remains difficult due to domain complexity, diverse context scales, and heavy reliance on expert assessments that are costly, inconsistent, and non-scalable. We introduce FAB-Bench, an end-to-end framework for adaptive benchmarking of RAG systems in semiconductor manufacturing. FAB-Bench defines six diagnostic metrics measuring factual accuracy, contextual utilization, completeness, retrieval relevance, technical depth, and reasoning consistency. The framework couples retriever diagnostics with generator-level reasoning analysis across context windows of 4K-32K tokens, quantifying how retrieval precision and generative fidelity co-evolve as contextual scope expands. From over 1,300 generated candidates, we curated a high-quality benchmark of 200 query-answer pairs spanning three synthesis strategies: needle-in-haystack, intra-document multi-topic, and cross-document multi-hop. Systematic evaluation across four LLMs and four RAG frameworks reveals three distinct context-scaling behaviors: logarithmic growth, early saturation, and cold-start dynamics, and identifies attention dilution as the primary mechanism behind performance degradation at extreme context lengths. Cross-framework validation on three additional production RAG systems confirms evaluation portability.
ProDebug: An Automated Debugging System for Prolog
arXiv:2605.27124v1 Announce Type: new Abstract: Prolog is a well-known declarative programming language commonly used in introductory courses on logic and reasoning. However, many students find Prolog challenging because it lacks the familiar debugging mechanisms found in imperative languages. In large classes, this difficulty is exacerbated by the challenge of providing timely and personalized feedback to students. In this work, we introduce ProDebug, the first tool to combine Large Language Models (LLMs) with spectrum-based and mutation-based techniques for automated debugging of Prolog assignments. ProDebug automatically identifies faults and proposes bug repairs for student Git submissions. Faults are detected using three approaches--spectrum-based, mutation-based, and LLM reasoning--while repairs are generated using mutation-based techniques and LLMs. Our evaluation on 1499 buggy student submissions from a bachelor's level programming class demonstrates the potential of automated, LLM-augmented feedback systems to scale support for declarative programming education.
NeR-SC: Adapting Neural Video Representation to Screen Content
arXiv:2605.27024v1 Announce Type: new Abstract: Implicit neural representations have emerged as a promising paradigm for video compression, with recent methods achieving competitive performance on natural video. However, screen content video -- common in remote desktop, online education, and cloud gaming -- exhibits distinct statistics: sharp edges, limited color palettes, and strong temporal redundancy. Existing neural representation methods, designed for natural scenes, lack mechanisms to exploit these properties, leaving substantial room for improvement. In this paper, we propose NeR-SC, a neural representation framework tailored for screen content video. Building on the SNeRV backbone, NeR-SC introduces three screen-content-specific modules: (i) a learnable color palette that models the discrete color structure of screen content by restricting the low-frequency sub-band to a learned color set; (ii) a multi-gate dense fusion module that replaces sequential feature fusion with dense, attention-gated cross-stage interaction; and (iii) an embedding-level frame skip strategy that bypasses redundant decoder invocations for static frames, with zero training overhead. Experiments on DSCVC and VCD show that NeR-SC achieves 40.32~dB and 41.73~dB average PSNR, outperforming representative neural video representation methods and, at low bitrates, surpassing H.264 and H.265. The skip strategy enables real-time decoding with no loss in quality.
PILOT: A Data-Free Continual Learning Approach for Real-Time Semantic Segmentation via Boundary Guidance
arXiv:2605.27128v1 Announce Type: new Abstract: Real-time semantic segmentation models offer an excellent balance between accuracy and inference speed. However, deploying these models in dynamic real world environments often requires the ability to learn novel classes incrementally without retraining on the entire dataset. This capability is known as continual learning. In this regard, the standard fine-tuning methods in deep learning often fail due to catastrophic forgetting, where the model learns new information but forgets previously trained and learned classes. Contributing to this crucial domain, the current paper proposes a novel continual learning framework tailored for PIDNet, which is a widely cited state-of-the-art real-time semantic segmentation model. Our method, PILOT(Parallel Incremental Learning Over Time), introduces a real-time and lightweight strategy by implementing a parallel Derivative-branch (D-branch) designed to capture the high frequency boundary information of novel classes while freezing the trained parameters of the original segmentation network. This novel setup allows the model to adapt to new semantic categories while preserving the knowledge of previously learned classes. By using only data associated with the new class, our model significantly reduces training overhead. Experimental results demonstrate that our approach successfully segments new classes while maintaining high mean Intersection over Union (mIoU) on the original base classes, thereby comfortably outperforming all major continual learning approaches in this domain. Overall, PILOT is shown to effectively mitigate catastrophic forgetting with minimal impact on inference latency, thus maintaining real-time performance.
ORCA: An End-to-End Interactive Copilot for Optimized Root Cause Analysis
arXiv:2605.27022v1 Announce Type: new Abstract: Causal analysis is a crucial task in many domains, including manufacturing, social science, and medicine. However, despite recent progress, the conceptual and methodological complexity of causal methods makes them largely inaccessible to domain experts. This gap prevents experts from leveraging these advances and hinders researchers who lack access to real-world data for validation. To bridge this divide, we introduce ORCA, a copilot for end-to-end causal analysis. ORCA orchestrates agents to understand the user's goals and guide them through the most appropriate causal analysis workflow, from fully automatic to highly user-guided execution. It features causal discovery, causal effect estimation, explainability and Root-Cause-Analysis (RCA). ORCA evaluates and compares performance, generates key metrics and diagrams, and generates insights through structured reports. We highlight its effectiveness across several real-world use-cases.
In-Orbit Intelligence or Ground Offloading? Inference Freshness under Intermittent Satellite Connectivity
arXiv:2605.27021v1 Announce Type: new Abstract: This paper studies how to balance onboard and ground computation under intermittent LEO connectivity for optimized inference freshness. As connectivity varies in time, the system switches among the actions of onboard computation, cached semantic transmission, raw-data offloading, and waiting. We define Age of Inference (AoInf) as the performance metric, where the age resets only upon successful task-valid updates. We formulate long-run average AoInf minimization as a finite-state average-cost semi-Markov decision process whose state captures the ground AoInf, orbital contact phase, cache occupancy, and cache age. We then transform the SMDP into an equivalent average-cost MDP and compute the solution via normalized relative value iteration (RVI). Numerical results indicate that the resulting hybrid policy reduces average AoInf relative to onboard-only and offload-only baselines, while requiring less computational resources on the satellite than the former, and fewer communication resources than the latter.
Graph-Based Modeling, Control, and Optimization for Multi-Domain and Multi-Timescale Energy Systems
arXiv:2605.27017v1 Announce Type: new Abstract: Modern energy systems in vehicles and built infrastructure are governed by high-dimensional dynamics spanning multiple physical domains (e.g., electrical, thermal, mechanical) and timescales. This tutorial paper presents a graph-based modeling approach created to facilitate the modeling, analysis, control, estimation, optimization, and design of these systems. Matured and validated through more than a decade of research spanning multiple academic institutions and companies, the graph-based approach combines transient energy conservation with an explicit mathematical representation of the network by which energy is stored and transferred within a system. Following a mathematical overview of graph-based models, examples of multi-domain component and system models from the recent literature are presented, including single-phase thermal systems, two-phase thermal systems, and electro-mechanical systems. This is followed by a survey of recent applications for decentralized and hierarchical model predictive control, design optimization, and control co-design. Lastly, the paper describes an open-source toolbox created to facilitate the generation and analysis of graph-based models.
SCENT: Aligning Mass Spectra with Molecular Structure for Olfactory Perception
arXiv:2605.27009v1 Announce Type: new Abstract: Predicting human olfactory perception from molecular structure has seen remarkable progress, yet these approaches require explicit chemical structure at inference, which is not available in practical sensing settings. We address this gap by exploring direct electron ionization mass spectrometry (EI-MS), a sensing technique that acquires chemically informative fragmentation fingerprints in seconds, as an alternative input modality for olfactory prediction. We contribute Spectrum-to-Chemical Embedding alignmeNT (SCENT), a multi-modal contrastive learning framework that aligns EI-MS representations with pretrained chemical structure embeddings, while requiring only mass spectra at inference. On the multi-label odor descriptor prediction task, SCENT significantly outperforms MS-only baselines and achieves performance comparable to structure-based models, despite requiring no explicit molecular structure at test time. The learned representations also better approximate continuous human perceptual ratings and generalize to real-world lab-measured spectra, suggesting that cross-modal alignment is an effective strategy for grounding analytical spectra in chemical semantics.
Sampling Data with Chains of Forward-Backward Diffusion Steps
arXiv:2605.27006v1 Announce Type: new Abstract: Sampling from learned high-dimensional distributions is a foundational computational problem. We introduce U-turn chains: Markov chains obtained by iterating short forward-backward steps of a diffusion model, in which each step proposes a move that remains on the learned data manifold and, paired with a Metropolis-Hastings correction, samples from energy-modified targets. For synthetic languages, we show that minimal U-turn dynamics undergoes an ergodicity-breaking phase transition driven by fragmentation of the data manifold; ergodicity is restored at larger U-turn magnitude. In the non-ergodic regime, low-level features relax faster than high-level ones, an ordering that inverts only at sufficiently large U-turn magnitude. We test these predictions on natural language and natural images. In both modalities, minimal U-turns relax slowly, especially for high-level features approximated by deep representations in CNNs or LLMs. The layer-ordering inversion appears only at large noise when mixing is efficient -- signatures consistent with strongly constrained, weakly mixing local dynamics. We discuss the implications of these results for sampling with diffusion models.
Probabilistic Recurrent Intention Switching Model
arXiv:2605.26998v1 Announce Type: new Abstract: Inverse reinforcement learning (IRL) recovers reward functions from observed behavior, yet traditional methods assume a single stationary reward that cannot capture goal switching within an episode. Recent multi-intention IRL methods address this by segmenting trajectories, but model intention transitions as either a memoryless Markov chain or via manual state augmentation with a fixed history window. We propose the Probabilistic Recurrent Intention Switching Model (PRISM), which replaces both mechanisms with a lightweight recurrent network that maps observation history to a per-step intention distribution. We prove that the resulting EM objective decomposes exactly into independent per-intention reward subproblems, each solvable in closed form, yielding an $\mathcal{O}(nK)$ E-step with no variational approximation. We evaluate PRISM on a non-Markovian gridworld, a mouse labyrinth, and BridgeData~V2 robotic manipulation, the first large-scale robotic application of multi-intention IRL. Across all settings PRISM achieves the highest held-out log-likelihood while recovering nameable, temporally coherent intentions from unlabeled demonstrations, suggesting that discrete goal switching is present in both biological and artificial agents.
On the Robustness of Machine Unlearning for Vision-Language Models
arXiv:2605.26992v1 Announce Type: new Abstract: Vision-language models (VLMs) may memorize undesirable information from training data, motivating growing interest in machine unlearning. In this work, we present the first systematic survey and robustness analysis of VLM unlearning. We provide a comprehensive taxonomy and review of existing VLM unlearning methods, together with unified evaluations under multiple prompt settings. We then propose three attack paradigms to examine whether forgotten multimodal knowledge can be reactivated through contextual prompting or downstream retraining. Extensive experiments show that many existing methods remain vulnerable under these attacks, indicating that current approaches often hide rather than fully remove target knowledge. Our study provides new insights into the robustness and limitations of current VLM unlearning methods and highlights the need for more reliable multimodal unlearning strategies. Code is available at https://github.com/XMUDeepLIT/VLM-UnL-Attack.
Towards Shared Embodied Intelligence in Humanoid Robots through Optimization Development and Testing of the Human Aware ergoCub Robot
arXiv:2605.26991v1 Announce Type: new Abstract: Collaboration is central to human behavior, enabling tasks beyond individual capability. This ability arises from coordinating actions through internal representations of others, a concept known as shared intelligence. Additionally, humans are characterized by physical bodies and cognitive abilities that are optimized in response to their environment, a phenomenon referred to as embodied cognition. Designing humanoid robots that collaborate safely and effectively with people requires unifying these principles. Here we propose an architecture that integrates shared intelligence and embodied cognition to enable robots to physically collaborate with humans, where robot hardware and control are optimized for human metrics, using representations of the human body and motion intelligence. The ultimate goal is to achieve a form of shared embodied intelligence. Specifically, our architecture optimizes robot hardware and physical intelligence parameters with respect to human ergonomic metrics. This is accomplished by modeling human-robot interaction as a function of hardware configurations and embedding human models into the robot's physical intelligence. As a concrete implementation, we present the humanoid robot ergoCub, whose morphology and control have been optimized for collaborative tasks with humans. Our approach provides a framework for designing humanoid robots that prioritize human ergonomics at both the hardware and physical intelligence levels, with applications in industrial and assistive robotics.