Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Unified ab initio quantum-electrodynamical density-functional theory for cavity-modified electron-phonon-photon coupling in solids
arXiv:2603.24095v2 Announce Type: replace-cross Abstract: Quantum-electrodynamical density-functional theory (QEDFT) provides a first-principles framework for describing materials coupled to quantized electromagnetic fields. While QEDFT has successfully captured cavity-induced modifications of electronic structures in atoms and molecules, a fully self-consistent and accurate framework to simulate and predict the structural, phonon-related, polarization and optical response of periodic solids in optical cavities has remained elusive. Here, we introduce a unified QEDFT approach that combines collective light-matter coupling parameter in the electronic ground state, density functional perturbation theory for phonons, and real-time time-dependent QEDFT for optical excitations. This framework enables ab initio calculations of cavity-modified electronic and phononic dispersions, Born effective charges, dielectric tensors, and both resonant and non-resonant optical absorption spectra. Using wurtzite gallium nitride (GaN) in an optical cavity as a case study, we demonstrate that the quantized vacuum field reshapes electronic, phononic and polarization properties, producing experimentally accessible signatures in the dielectric function and absorption spectra. These results establish QEDFT as a general first-principles platform for predicting and exploring cavity-modified quantum materials.
The Language of Security: How Prompt Syntax Shapes Secure Code Generation in Open LLMs
arXiv:2607.15937v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for source code generation despite their outputs often exhibiting security vulnerabilities. Prior work shows that prompt engineering can mitigate such risks, yet (1) they focused on high-level prompting strategies, neglecting recent evidence that fine-grained syntactic variations can substantially alter model behavior; and (2) predominantly evaluate proprietary LLMs, limiting the applicability of their findings in industrial settings where self-hosted, open models are preferred for privacy, compliance, and deployment control. In this paper, we study how fine-grained syntactic constituents of prompts influence the security of open LLM-generated code. Using a parser-driven approach, we systematically generate syntactic variants of security-relevant code generation prompts and evaluate their impact on code security across multiple open LLMs and programming languages. Our results show that specific syntactic elements, such as constraints, guards, conditions, and concept bindings, and their position within the prompt consistently affect the likelihood of generating insecure code. These findings identify prompt syntax as a concrete security control surface and provide actionable guidance for reducing vulnerability risk in LLM-assisted development.
MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction
arXiv:2607.16192v1 Announce Type: new Abstract: Humans can infer how objects are likely to move from passive observation: a cup may be lifted, a drawer may slide, and a lid may rotate shut. Such predictions expose the physical consequences of interaction needed to act in the real world. We study how to learn this anticipation from ordinary monocular videos of human-object interaction. Given a short observed video context, MotionForesight predicts future 3D trajectories for points on the manipulated object. This casts interaction prediction as object-centered 3D motion forecasting without any assumptions on the object properties. Our key insight is that video prediction models already encode rich priors about how objects move during human interactions. We redirect these priors from pixel prediction toward future 3D scene flow. We start from a dense 3D tracker built on a pretrained video model, generate pseudo-ground-truth tracks from complete clips, and train the forecaster using only the observed frames. We replace future RGB and geometry with learned mask latents and train a lightweight adapter to turn the retrospective tracking representation into a forward predictor, while freezing the large video and tracking components. Using just 40k human videos and no auxiliary inputs such as language, MotionForesight generalizes across diverse out-of-distribution objects, environments, viewpoints, and interactions. It also outperforms substantially larger models that use over a million training videos. These results show that we can efficiently re-purpose video priors into explicit geometric forecasts for embodied intelligence. https://motionforesight.github.io/
The Minimum Subgraph Complementation Problem
arXiv:2512.23687v2 Announce Type: replace Abstract: Subgraph complementation is an operation that toggles all adjacencies inside a selected vertex set. Given a graph \(G\) and a target class \(\mathcal{C}\), the Minimum Subgraph Complementation problem asks for a minimum-size vertex set \(S\) such that complementing the subgraph induced by \(S\) transforms \(G\) into a graph belonging to \(\mathcal{C}\). While the decision version of Subgraph Complementation has been extensively studied and is NP-complete for many graph classes, the algorithmic complexity of its optimization variant has remained largely unexplored. In this paper, we study MSC from an algorithmic perspective. We present polynomial-time algorithms for MSC in several nontrivial settings. Our results include polynomial-time solvability for transforming graphs between bipartite, co-bipartite, and split graphs, as well as for complementing bipartite regular graphs into chordal graphs. We also show that MSC to the class of graphs of fixed degeneracy can be solved in polynomial time when the input graph is a forest. Moreover, we investigate MSC with respect to connectivity and prove that MSC to the class of disconnected graphs and to the class of 2-connected graphs can be solved in polynomial time for arbitrary inputs.
Atomic Design Transformer: Scaffold-Conditioned 3D Molecule Generation via xTB-Reward Reinforcement Learning
arXiv:2607.15918v1 Announce Type: new Abstract: We present an SE(3)-invariant transformer for 3D-molecule generation, the Atomic Design Transformer (ADT). ADT places atoms one at a time, autoregressively. SE(3) invariance is achieved by tokenization: each new atom's position is encoded in the local coordinate frame of a previously placed atom. The backbone is a plain causal transformer. The token stream fully specifies a 3D structure together with its chemical-bond graph G, without any bond-order assignment. To score generated molecules we introduce the xTB topology-preservation rate (XTP): the fraction of molecules for which an xTB GFN2 relaxation preserves G specified by the token stream. For XTP-accepted molecules we also report the relaxation energy and the root mean square of the atomic displacement (RMSD). We evaluate two ADT models. The first is ADT pretrained on the GEOM-Drugs $\le\!30$-heavy-atom dataset; we benchmark scaffold-conditioned 3D generation across seven drug-like scaffolds from the model. It reaches an XTP of ${\sim}54\%$ and a valid-molecule yield $N^{\mathrm{gen}}/N$ of ${\sim}50\%$, where $N^{\mathrm{gen}}/N$ is the fraction of samples that are distinct, topology-preserving, and chemically valid. The second model continues from the first by reinforcement learning against the verifiable xTB reward (RLVR), using no external molecules. RLVR raises XTP to ${\sim}98\%$ and $N^{\mathrm{gen}}/N$ to ${\sim}95\%$, while approximately preserving the GEOM-Drugs size and composition distributions. Finally, we present an Inverse-Kinematics Transformer that recovers XTP for large molecules, where discretization error accumulates. ADT thus enables direct 3D generation.
PsSource: an installable Geant4 extension for transport-coupled positronium annihilation with validated Ore-Powell three-photon physics
arXiv:2604.21173v3 Announce Type: replace Abstract: PsSource is an installable Geant4 extension for configurable positronium-aware terminal positron annihilation. Geant4 transports the positron normally, after which PsSource replaces only terminal at-rest annihilation, preserves the terminal position and global time, resolves a fixed or host-defined environment, samples the annihilation class and delay, and returns ordinary two- or three-photon secondaries. Supported outcomes are direct two-photon, para-positronium two-photon, ortho-positronium two-photon, and ortho-positronium three-photon annihilation. Fixed and exponential delays and approximate phase-space, Geant4 Ore-Powell, and polarized Geant4 Ore-Powell backends are available. Regression tests verified terminal-state preservation, timing decomposition, all four classes, photon multiplicity and parentage, equivalent environment pathways, and unchanged photon physics when PsSource-specific truth recording was disabled. The Ore-Powell backend was evaluated using 100,000 three-photon events. Comparison with an analytic joint-energy distribution yielded chi-square = 1236.91 for 1274 degrees of freedom, a standardized score of -0.735, and a maximum marginal empirical CDF difference of 0.00208. Ordered photon-energy spectra agreed with native Geant4 and GATE references, with no pairwise empirical CDF difference greater than 0.00452. Photon directions and event-plane normals were isotropic with no detectable Cartesian-axis bias. The polarized backend satisfied normalization and transversality to within 1.5 x 10^-12 and preserved the ordinary Ore-Powell energy spectrum. PsSource provides a validated integration layer for positronium-aware terminal annihilation in Geant4.
A hierarchical memory architecture overcomes context limits in long-horizon multi-agent computational modeling
arXiv:2607.07666v3 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate remarkable reasoning capabilities, yet their stateless architecture fundamentally limits deployment in long-horizon research workflows requiring multi-session continuity and quantitative rigor. Here we present Ensemble QSP, a multi-agent framework featuring a three-layer hierarchical memory architecture that keeps injected context bounded and constant in project duration (mid-term project state: median 301 tokens, max 4,050, across 104 runs) by capping each state category and evicting completed work, enabling continuous autonomous operation without context degradation. The system orchestrates five specialist worker agents under domain-expert principal investigators, enforcing physical constraints through physics-based checklists and structured-domain knowledge. Comprehensive benchmarking demonstrates robust autonomous pharmacokinetic-pharmacodynamic model selection without human intervention, consistent result quality across both lower-cost and frontier LLMs, improved PK parameter recovery relative to single-agent baselines, and stable model selection across linguistically diverse prompts of the same task. Feature-level ablation across physiologically based pharmacokinetic (PBPK) models spanning a broad complexity range shows that PI-agent oversight improves debugging efficiency while preserving final accuracy across conditions. The architecture is structurally domain-agnostic, adding a new scientific domain requires only a new PI agent configuration.
Diffusion-induced instabilities promote cooperation in eco-evolutionary networks
arXiv:2607.15989v1 Announce Type: new Abstract: Understanding how cooperation persists despite the advantage of selfish behavior remains a central challenge in evolutionary dynamics. Classical models of public goods dilemmas predict dominance of defectors, yet natural and social systems often sustain cooperation. We study an eco-evolutionary public goods game on complex networks where cooperators and defectors diffuse at different rates. When the isolated system is in a defector-dominated coexistence regime, faster dispersal of defectors than cooperators leads to a symmetry-breaking transition that produces localized clusters of cooperators. In heterogeneous networks, nodes with higher connectivity become significantly more likely to exhibit cooperative dominance. A degree-based mean-field reduction supports this result by showing that network connectivity controls an effective coupling strength proportional to node degree, thereby producing a bifurcation that separates defector-dominated and cooperative states. We also address why not all hubs become cooperative by means of a multistability analysis. These results reveal how asymmetric mobility and heterogeneous connectivity jointly promote cooperation in structured populations.
The Endpoint Cardinality of Discrete Cube Skeleta
arXiv:2607.15502v1 Announce Type: cross Abstract: We determine the minimum order of a finite lattice set that contains a filled axis-parallel cube skeleton about every point of some $N$-point set of centers. For fixed integers $0\leq k<n$, the answer for $N$ centers is $N^{1-(n-k)/(2n^2)}$, up to constants depending on $n$ and $k$. Thornton proved every smaller exponent and gave a construction of this order; the endpoint lower bound was left open when $k\geq1$. Our proof combines a midpoint estimate, a labelled form of Shearer's projection inequality, and a strong induction that balances large and small radii without a dyadic pigeonhole loss. In particular, a lattice set containing a square boundary about each of $N$ centers has at least a constant times $N^{7/8}$ points.
Analysis of Semi-Supervised Learning on Hypergraphs
arXiv:2510.25354v3 Announce Type: replace Abstract: Hypergraphs provide a natural framework for modeling multiway interactions. We analyze a class of variational semi-supervised learning problems posed on random geometric hypergraphs and establish asymptotic consistency in the large-data limit. In particular, we identify scaling regimes that ensure well-posedness--yielding nontrivial label propagation rather than collapse to a constant labeling--and show that discrete minimizers converge, in the continuum, to solutions of a density-weighted p-Laplacian equation. We also propose Higher-Order Hypergraph Learning (HOHL), a multiscale regularization scheme based on powers of Laplacians associated with hypergraph-induced subgraphs. For geometric point clouds, we analyze an efficient multiscale Laplacian surrogate for HOHL and prove convergence to a higher-order Sobolev-type seminorm. Numerical experiments on standard benchmarks support the practical utility of the resulting higher-order regularization.
Entropy Geometry and Condensation in Wealth Allocation
arXiv:2602.03676v2 Announce Type: replace Abstract: We develop a statistical framework for wealth allocation in which equilibrium-like statistics follow from unbiased counting of admissible configurations rather than postulated exchange rules. Each agent is described by a value--wealth map $V_i(w)$, whose local resolution fixes the microscopic weight through a Jacobian relation. In a closed system, the microcanonical marginal and a reservoir expansion yield an emergent canonical distribution for the regular sector. Its partition sum gives a general condensation criterion: if this sector has finite wealth capacity, excess wealth concentrates on a small subset of agents. We extend the construction to open systems with variable wealth and agent number and to weak quasistatic driving. The global constraint determines an evolution equation for the common parameter $\lambda(t)$, while simultaneous changes in total wealth and value--wealth geometry produce a unified first-order response. The susceptibility $\chi_W=\sum_i\mathrm{Var}_i(w)$ equals the Fisher information of the joint canonical family, and Legendre duality gives $\mathrm{d}s^2=\chi_W\mathrm{d}\lambda^2=\chi_W^{-1}\mathrm{d}W^2=-\mathcal{S}''(W)\mathrm{d}W^2$. For power-law critical tails, finite capacity requires $p>2$; within this regime, $\chi_W$ diverges for $2<p\leq3$ and remains finite for $p>3$, while the critical boundary lies at finite Fisher--Rao distance. We also derive a qualified Cram'er--Rao duality and an open-system mixed-response relation. Contact-geometric, Airy-scaling, and stochastic-dynamical interpretations are identified only as conjectures or future work. The time-dependent results are quasistatic and do not determine microscopic relaxation times, while the information geometry describes the canonical family.
Indirect data-driven predictive control and the state-space predictor
arXiv:2602.10936v2 Announce Type: replace Abstract: We define trajectory predictive control (TPC) as a class of indirect data-driven predictive control (DDPC) methods that represent future outputs as linear in past inputs/outputs and future inputs. TPC unifies many DDPC variants with different predictor structures. We introduce a predictor with a state-space representation and show that with it, TPC inherits the mature theory of linear model predictive control. In numerical experiments, the state-space predictor outperforms existing predictors, especially for small training datasets.
Speculative Decoding with a Speculative Vocabulary
arXiv:2602.13836v2 Announce Type: replace Abstract: Speculative decoding has rapidly emerged as a leading approach for accelerating language model (LM) inference, as it offers substantial speedups while yielding identical outputs. This relies upon a small draft model, tasked with predicting the outputs of the target model. State-of-the-art speculative decoding methods use a draft model comprising a single decoder layer and output embedding matrix, with the latter dominating drafting time for the latest LMs. Recent work has sought to address this output distribution bottleneck by reducing the vocabulary of the draft model. While this can improve throughput, it compromises speculation effectiveness when the target token is out-of-vocabulary. In this paper, we argue for vocabulary speculation as an alternative to a reduced vocabulary. We propose SpecVocab, an efficient and effective method that selects a vocabulary subset per decoding step. Across a variety of tasks, we show that SpecVocab can achieve a higher acceptance length than state-of-the-art speculative decoding method, EAGLE-3. Notably, this yields up to an 8.1% increase in average throughput over EAGLE-3.
Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D
arXiv:2607.16072v1 Announce Type: new Abstract: While large language models (LLMs) can solve advanced reasoning problems in seconds, we show that even frontier models fail to perform a much simpler operation: exactly copying an input string that lies well within their context windows. We attribute this failure to positional encodings in Transformer architectures, whose inductive bias favors copying through a shortcut based on matching local contexts rather than carefully locating the corresponding input positions. To address this issue, we introduce 2D-RoPE, which organizes text into a 2D grid rather than a 1D sequence and assigns each token a row ID and a column ID. Under this view, copying becomes simply retrieving input tokens at a fixed column offset, which makes the task easy to learn. In synthetic copy experiments, shallow Transformers with 2D-RoPE achieve perfect copying at input lengths hundreds of times longer than those seen during training, whereas standard positional encodings fall far behind. We further show that the advantage of 2D-RoPE language models on copy tasks consistently holds in large-scale pretraining on DCLM with model sizes up to 1.4B parameters. Overall, our results suggest that viewing text in 2D can benefit language modeling, and we hope this encourages future work to further explore the potential of 2D positional encodings.
DyneTrion: A Spatio-temporally Coherent Generative Emulator for Protein Dynamics Across Timescales
arXiv:2607.15309v1 Announce Type: cross Abstract: Proteins function through coordinated motion across multiple spatial and temporal scales, underpinning processes such as ligand binding, allostery, and catalysis. However, accessing long-timescale conformational change through molecular dynamics (MD) simulations remains prohibitively expensive for systematic exploration across diverse systems. Here, we present DyneTrion, a generative protein dynamics emulator that jointly enforces geometric symmetry, structural consistency and temporal coherence within a single framework. DyneTrion uses a tri-attention architecture that integrates invariant point attention (IPA) for SE(3)-robust geometric updates, spatial attention anchored to a reference conformation to preserve structural integrity, and temporal attention to model correlated evolution across time frames. Across 100-ns MD trajectory simulation benchmarks, DyneTrion reproduces MD-derived flexibility, ensemble distributions and interaction observables while maintaining stereochemical validity during extrapolation. To evaluate long time-scale generalization, we introduce dynamicPDB, a dataset of over 10,000 proteins with up to 1-$\mu$s all-atom trajectories at 10-ps resolution and accompanying physical annotations. On microsecond trajectories, DyneTrion preserves free-energy landscapes and metastable-state populations, and it supports large conformational propagation in apo-to-holo transitions and fast folders. Together, DyneTrion provides a scalable path from static structure prediction toward time-resolved, ensemble-faithful protein modeling. The code is publicly available at https://github.com/fudan-generative-vision/DyneTrion
Perception-Aligned AI Outputs: End-to-End Visual Prediction for Uncertainty Communication in Clinical Decision-Making
arXiv:2205.04599v2 Announce Type: replace Abstract: Explainable Artificial Intelligence (XAI) is essential for trustworthy AI in healthcare, yet many existing methods rely on technical explanations that are difficult for clinicians and patients to interpret. We introduce Visualized Learning for Machine Learning (VL4ML), a human-centered explainability framework that communicates model predictions and uncertainty through intuitive visual representations rather than numerical or post-hoc explanations. By encoding diagnostic information in colors, patterns, and spatial structures, VL4ML enables users to interpret predictions without requiring knowledge of model internals or statistical expertise. We demonstrate the framework across multiple clinical tasks, including classification, regression, longitudinal prediction, and multimodal analysis. Its effectiveness was evaluated through a human-centered study involving 158 participants (39.2% clinical professionals) and an expert interpretability assessment. More than 79% of participants positively rated the visual explanations across evaluation dimensions, 84.0% found them more memorable than numeric outputs, and 76.9% reported faster decision-making. Over 82% successfully perceived uncertainty embedded in the visual representations without prior statistical training. No significant differences were observed between clinicians and non-clinicians or between male and female participants, indicating broad accessibility. These results suggest that VL4ML complements existing XAI and uncertainty quantification methods by providing intuitive, universally interpretable visual explanations that support transparent and trustworthy clinical decision-making.
Learn to Memorize: Scalable Continual Learning in Semiparametric Models with Mixture-of-Neighbors Induction Memory
arXiv:2303.01421v2 Announce Type: replace Abstract: Semiparametric language models (LMs) have shown promise in various Natural Language Processing (NLP) tasks. However, they utilize non-parametric memory as static storage, which lacks learning capability and remains disconnected from the internal information flow of the parametric models, limiting scalability and efficiency. Based on recent interpretability theories of LMs, we reconceptualize the non-parametric memory represented by $k$NN-LM as a learnable Mixture-of-Neighbors Induction Memory (MoNIM), which synergizes the induction capabilities of attention heads with the memorization strength of feed-forward networks (FFN). By integrating into the model's information flow, MoNIM functions as an FFN-like bypass layer within the Transformer architecture, enabling effective learning of new knowledge. Extensive experiments demonstrate that MoNIM is a retentive and scalable continual learner in both data- and model-wise, enhancing the scalability and continual learning performance of semiparametric LMs.
Why do CNNs excel at feature extraction? A mathematical explanation
arXiv:2307.00919v2 Announce Type: replace Abstract: Over the past decade deep learning has revolutionized the field of computer vision, with convolutional neural network models proving to be very effective for image classification benchmarks. However, a fundamental theoretical questions remain answered: why can they solve discrete image classification tasks that involve feature extraction? We address this question in this paper by introducing a novel mathematical model for image classification, based on feature extraction, that can be used to generate images resembling real-world datasets. We show that convolutional neural network classifiers can solve these image classification tasks with zero error. In our proof, we construct piecewise linear functions that detect the presence of features, and show that they can be realized by a convolutional network.
SC-JEPA: Stabilizing Latent Predictive Learning for Time-Series Anomaly Prediction
arXiv:2602.04643v2 Announce Type: replace Abstract: Time-series anomaly prediction aims to forecast future system failures before they fully emerge, making latent predictive models such as JEPA a promising framework for capturing precursor dynamics. However, directly applying continuous self-distillation to time-series data is often unstable and can lead to representation collapse, while also struggling to model precursors evolving at different temporal scales. To address this, we propose \textbf{SC-JEPA}, a new JEPA-based framework to model time-series anomaly prediction in a discretized predictive state space. It introduces a soft codebook bottleneck to stabilize latent predictive learning and encourage regime-level structure in the learned representations. Building on this stabilized latent space, we further design a multi-resolution predictive objective to capture precursor patterns at different temporal scales. Experiments on five real-world benchmarks show that SC-JEPA achieves strong and consistent early-warning performance.
On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
arXiv:2601.18929v2 Announce Type: replace Abstract: Vision foundation models (VFMs) have emerged as powerful tools for surgical scene understanding. However, current approaches predominantly rely on unimodal RGB pre-training, overlooking the complex 3D geometry inherent to surgical environments. Although several architectures support multimodal or geometry-aware inputs in general computer vision, the benefits of incorporating depth information in surgical settings remain underexplored. We conduct a large-scale empirical study comparing eight ViT-based VFMs that differ in pre-training domain, learning objective, and input modality (RGB vs. RGB-D). For pre-training, we use a curated dataset of 1.4 million robotic surgical images paired with depth maps generated from an off-the-shelf network. We evaluate these models under both frozen-backbone and end-to-end fine-tuning protocols across eight surgical datasets spanning object detection, segmentation, depth estimation, and pose estimation. Our experiments yield several consistent findings. Models incorporating explicit geometric tokenization, such as MultiMAE, substantially outperform unimodal baselines across all tasks. Notably, geometric-aware pre-training enables remarkable data efficiency: models fine-tuned on just 25% of labeled data consistently surpass RGB-only models trained on the full dataset. Importantly, these gains require no architectural or runtime changes at inference; depth is used only during pre-training, making adoption straightforward. These findings suggest that multimodal pre-training offers a viable path towards building more capable surgical vision systems.
Variable Aerodynamic Damping Actuation via Co-Contraction: A Structural Analogy with Variable Stiffness Actuation
arXiv:2605.07292v2 Announce Type: replace Abstract: This work identifies a passive aerodynamic damping effect induced by co-contraction in antagonistic redundant propulsion. Complementing prior work on aerodynamic promptness, which addressed active wrench-rate authority along constant-wrench fibers, we study the passive side: the local derivative of aerodynamic force with respect to air-relative velocity at a trim. This derivative defines an incremental aerodynamic damping coefficient. We prove that it increases monotonically along constant-force fibers under a mild aerodynamic hardening condition, and derive this property from a first-order Blade Element Theory model exposing the relevant speed-inflow coupling. The resulting mechanism, Variable Aerodynamic Damping Actuation (VADA), is formulated as an antagonistic aerodynamic actuation module and allocation principle, structurally analogous to variable-stiffness actuation at the level of fiber motions and incremental impedance modulation. An impedance-form interpretation clarifies common- and differential-mode roles, while a propeller-data-based assessment using the UIUC Propeller Database shows that the identified damping has practical small-UAV magnitude, is comparable to ordinary low-speed body-drag damping, and depends strongly on low-advance-ratio thrust sensitivity.
DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning
arXiv:2607.16090v1 Announce Type: new Abstract: Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and target domains. In this paper, we consider the setting of online dynamics adaptation, where policies are trained in the source domain with sufficient data, while only limited interactions with the target domain are allowed. There are a few existing works that address the dynamics mismatch by employing domain classifiers, value-guided data filtering, or representation learning. Instead, we study the domain adaptation problem from a generative modeling perspective. Specifically, we introduce DADiff, a diffusion-based framework that leverages the discrepancy between source and target domain generative trajectories in the generation process of the next state to estimate the dynamics mismatch. Both reward modification and data selection variants are developed to adapt the policy to the target domain. We also provide a theoretical analysis to show that the performance difference of a given policy between the two domains is bounded by the generative trajectory deviation. More discussions on the applicability of the variants and the connection between our theoretical analysis and the prior work are further provided. We conduct extensive experiments in environments with various shifts to validate the effectiveness of our method. The results demonstrate that our method provides superior performance compared to existing approaches, effectively addressing the dynamics mismatch. We provide the code of our method at https://github.com/hanyang-chen/DADiff-release
How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA
arXiv:2607.16094v1 Announce Type: new Abstract: Compositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations like object selection, spatial relation resolution, and attribute verification. Despite strong aggregate performance, the mechanistic basis of VLM failures on this task remains underexplored. To address this gap, we analyze vision-operation misalignment in VLMs by examining how failures relate to specific reasoning operations and the internal computational pathways through which they arise and propagate. We introduce an Operation-centric mechanistic framework that decomposes VLM failures by both the reasoning operation where they originate and the internal computational pathway through which they propagate. Our analysis reveals four mechanistically distinct failure modes: grounding failure, reasoning failure, attribute extraction failure, and language prior dominance failure. Each characterized by a unique relationship between visual grounding strength and answer correctness. Through three complementary causal interventions applied across all transformer layers, we further demonstrate a pathway dissociation: grounding failures route exclusively through the feedforward network, reasoning failures route through late-layer attention, and attribute extraction failures localize to the answer-position feedforward computation. This dissociation demonstrates that different failure types require fundamentally different corrective strategies, providing a principled foundation for targeted improvements to VLM reliability in multimedia reasoning.
Let the Body Follow: Coupled Egocentric Control for Whole-Body Robot Teleoperation
arXiv:2607.16095v1 Announce Type: new Abstract: Whole-body teleoperation requires users to coordinate perception, manipulation, posture, and mobility across multiple robot components. This coordination is difficult because users must simultaneously control the robot's head, arms, torso, and base while maintaining task awareness and avoiding kinematic or environmental constraints. In this paper, we propose coupled egocentric control, a body-following teleoperation approach in which the robot's torso and base automatically respond to the operator's head and arm motions. Rather than requiring explicit touchpad commands for every torso or base adjustment, the system lets users focus on gaze and hand control: head pitch adjusts torso height, head yaw drives base rotation, end-effector height adjusts torso motion, and end-effector workspace boundaries trigger base translation. We evaluate this approach in a user study on whole-body teleoperation of a TIAGo mobile manipulator for home-care-inspired tasks. Compared with a baseline hybrid interface, coupled egocentric control improves object manipulation efficiency, reduces button-based control effort and arm singularities, lowers mental demand and overall workload, and increases ease of use, ease of learning, confidence, and user preference for torso and base control.
ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning
arXiv:2607.16131v1 Announce Type: new Abstract: Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tables, charts, and textual context. However, existing methods often fail because they struggle to locate decisive visual evidence, accurately read structured scientific visuals, and integrate multimodal observations into reliable reasoning. We introduce ToolSciVer, the first tool-augmented framework for MSCV to our knowledge. ToolSciVer equips a VLM with three type-aware visual tools, table row/column focus, chart-to-structure parsing, and high-resolution region zoom, which convert dense scientific visuals into explicit, claim-facing evidence, and trains the policy with Group Relative Policy Optimization (GRPO) under a composite reward of answer correctness, format validity, length control, tool-use efficiency, and tool-validity penalties. Experiments on SciVer and MuSciClaims datasets on five VLMs from three model families (Qwen, InternVL, Gemma) demonstrate that our method achieves superior performance compared to four competitive baselines including prompting-based and RL-based tool-use methods, highlighting the effectiveness of learned, type-aware tool use for scientific claim verification.