Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term Convergence
arXiv:2510.16657v3 Announce Type: replace-cross Abstract: Synthetic data has been increasingly used to train frontier generative models. However, recent studies raise key concerns that iteratively retraining a generative model on its self-generated synthetic data may keep deteriorating model performance, a phenomenon often coined model collapse. In this paper, we investigate ways to modify the synthetic retraining process to avoid model collapse, and even possibly help reverse the trend from collapse to improvement. Our key finding is that by injecting information through an external synthetic data verifier, whether a human or a better model, synthetic retraining will not cause model collapse. Specifically, we situate our theoretical analysis in the fundamental linear regression setting, showing that verifier-guided retraining can yield near-term improvements, but ultimately drives the parameter estimate to the verifier's "knowledge center" in the long run. Our theory further predicts that, unless the verifier is perfectly reliable, these early gains will plateau and may even reverse. Indeed, our experiments across linear regression, Variational Autoencoders (VAEs) trained on MNIST, and fining-tuning SmolLM2-135M on the XSUM task confirm these theoretical insights.
Entropy-Driven Initiation and Cellular Uptake Mediated by Viscoelastic Cytoskeleton: A Kinetic Phase Diagram from Onsager Variational Principle
arXiv:2607.12766v2 Announce Type: replace-cross Abstract: A fundamental question in receptor-mediated endocytosis remains unanswered: what initial driving force brings ligands and receptors into close proximity? While previous models assume pre-existing contact and overlook this initiation problem, we propose that entropic forces from nanoscale biomolecules in crowded cellular environments provide the essential driving mechanism. We develop a unified continuum model rooted in the Onsager variational principle, where engulfment depth serves as the generalized coordinate and the driving force derives from a free energy landscape of entropic, binding, membrane, and cytoskeleton contributions. The framework naturally incorporates: (i) entropy-driven adhesion as initiation; (ii) ligand-receptor binding as the sustaining force; (iii) membrane deformation via the Helfrich-Canham Hamiltonian; and (iv) cytoskeleton viscoelasticity through the elastic-viscoelastic correspondence principle. The kinetic phase diagram predicts a critical biomolecule concentration for initiation, a lower bound of ligand density for complete engulfment, a finite size window for engulfable particles, and an optimal virus radius of 30--60 nm that decreases with increasing binding energy. The Onsager solubility condition naturally yields the phase boundaries. The model exhibits asymptotic consistency with the classic Asakura-Oosawa result in the large-particle flat-surface limit. Stiffer cells lead to longer engulfment times and narrower size windows. Strikingly, the optimal size matches HIV-1 dimensions under physiologically realistic parameters. This work provides a variational foundation for cellular uptake with implications for virology, nanotechnology, and drug delivery.
Single-shot laser-pulse-induced magnetization reversal in CoFeB/MgO-based magnetic tunnel junctions
arXiv:2510.25102v2 Announce Type: replace-cross Abstract: We demonstrate single-shot laser-pulse-induced magnetization reversal in rare-earth-free CoFeB/MgO magnetic tunnel junctions (MTJs), a material system widely adopted in spin-transfer torque magnetic random-access memory (STT-MRAM). By tuning the Ru capping layer thickness, we modify the laser energy absorption profile and observe magnetization reversal from the parallel (P) to antiparallel (AP) state, with switching observed for $t_\text{Ru} \geq 2.0\,$ nm. Furthermore, we detect magnetization reversal in a micro-scale MTJ device via the tunnel magnetoresistance (TMR) effect. Our findings suggest that ultrafast spin transport, dipolar interactions, or a combination of both may contribute to the switching process, although the precise mechanism remains to be clarified. This work represents a significant step toward integrating ultrafast optical control with MTJ technology.
A Noise-Robust Elicit-to-Optimize Framework for Distortion Riskmetrics via Inverse Reinforcement Learning
arXiv:2607.14373v1 Announce Type: new Abstract: We propose a noise-robust elicit-to-optimize framework that integrates inverse reinforcement learning (IRL) and reinforcement learning (RL) for eliciting agents' risk preferences and optimizing policies under a broad class of risk objectives characterized by distortion riskmetrics. On the elicitation side, we propose an adaptive Bayesian IRL method that infers agents' latent risk objectives from their noisy observed decisions, explicitly allowing agents to take stochastic and suboptimal actions. We establish the existence of a finite set of distinguishing questions that identifies the preferred distortion riskmetric within the candidate class and prove that the convergence rate of the algorithm is of order $O(\exp(-cm+O(\sqrt{m\log m})))$ under general settings, where $c>0$ is a constant and $m$ denotes the number of algorithm iterations. On the optimization side, we develop a model-free RL algorithm for optimizing policies under conditional distortion riskmetrics. By representing the objective as an integral of the conditional cost quantile function with respect to the distortion function, the method unifies distortion-riskmetric objectives. We optimize diverse risk objectives by extending the Proximal Policy Optimization (PPO) algorithm with policy, value, and quantile neural networks, where the quantile network estimates the full conditional cost quantile function and enables numerical evaluation of general risk objectives. A comprehensive empirical study demonstrates the framework's elicitation accuracy and effectiveness in complex financial environments.
Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment
arXiv:2607.14682v1 Announce Type: new Abstract: Efficient multimodal document question answering with explicit visual grounding, locating the precise document region that supports each answer remains an open challenge. Current approaches bifurcate into Supervised Fine-Tuning (SFT), which requires large annotated datasets and reaches optimization plateaus, and reasoning-centric Reinforcement Learning (RL), which depends on verbose intermediate traces that inflate inference token cost without clear benefit. We introduce Perception-RFT, a training framework that applies Group Relative Policy Optimization (GRPO) to multimodal document QA, bypassing intermediate reasoning tokens to directly align visual features with structured grounding outputs. To rigorously evaluate the necessity of reasoning, we construct a reasoning variant under identical reward settings. We find that reasoning-enabled models suppress their reasoning traces during training, converging to direct perception-based policies at the 4B parameter scale, reducing per-query inference token length by more than 60%, while reasoning-enabled RL underperforms perception-only training. Through a fine-grained analysis of Qwen3-VL-4B optimization dynamics, we confirm that SFT saturation and cold-start RL instability established in text-domain post-training extend to multimodal, and identify a previously uncharacterized Grounding Divergence: a selective trade-off between semantic robustness and geometric precision on two out of distribution (OOD) benchmarks (4,828 samples) under joint RL optimization. We further show that an early SFT$\rightarrow$RL transition achieves comparable precision with 65% less training data.
Assisting Mission-Critical Traffic Flows with Active Queue Management in Industrial Internet of Things
arXiv:2607.14478v1 Announce Type: new Abstract: Mission-critical Industrial Internet of Things (IIoT) traffic flows require bounded network latency and jitter guarantees to ensure the safe functioning of critical industrial infrastructure. These flows are typically communicated via commodity network routers with conventional First-In-First-Out (FIFO) buffers. FIFO has proven to be the culprit of the well-known bufferbloat phenomenon, and the deployment of Active Queue Management (AQM) schemes have demonstrated significant performance improvements for latency-sensitive applications over the Internet in the IT domain. However, the bufferbloat phenomenon and the efficacy of AQM schemes have not been studied in IIoT-based OT domain. In this paper, we propose the use of AQM as a lightweight and non-intrusive mechanism for assisting mission-critical traffic flows in IIoT networks. Our experimental results demonstrated that multi-queue AQM schemes provide substantial flow isolation and capacity sharing benefits, and significantly improve the performance of mission-critical traffic flows under network pressure. We further provide deployment recommendations based on our experimental insights.
Entangled criticality and irreversibility in random Markov dynamics
arXiv:2602.04905v2 Announce Type: replace-cross Abstract: We introduce a two-parameter ensemble of random discrete-time Markov models that simultaneously captures critical slowing down and broken detailed balance. Extending a previously studied heterogeneous Markov ensemble, we incorporate correlations between forward and backward transition rates through a single asymmetry parameter $\gamma$, while heterogeneity is controlled by $\epsilon$. Using results from random matrix theory, we identify a critical locus $\epsilon_c(\gamma,N)$ at which relaxation times diverge and spectral universality breaks down, in Markov models with $N$ states. We characterize the behavior of entropy production, predictive information, and relaxation dynamics across the ensemble, showing that many observables depend strongly on heterogeneity but only weakly on asymmetry, except near the symmetric limit. Applying maximum-likelihood inference to human fMRI and EEG data, we find that both modalities operate near the predicted critical locus and occupy a similar region of the $\epsilon-\gamma$ plane, supporting a super-universality of human brain dynamics. While ensemble averages are well captured by the null model, empirical data exhibit substantially enhanced variability, indicating subject-specific structure beyond random expectations. Our results unify criticality and nonequilibrium measures within a single framework and clarify their intertwined role in the analysis of complex biological dynamics.
Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
arXiv:2603.23723v2 Announce Type: replace-cross Abstract: Deep spatially selective filters achieve high-quality enhancement with real-time capable architectures for stationary speakers of known directions. To retain this level of performance in dynamic scenarios where only the speakers' initial directions are given, accurate, yet computationally lightweight tracking algorithms become necessary. Assuming a frame-wise causal processing style, temporal feedback allows for leveraging the enhanced speech signal to improve tracking performance. In this work, we investigate strategies to incorporate the enhanced signal into lightweight tracking algorithms and autoregressively guide deep spatial filters. Our proposed Bayesian tracking algorithms are compatible with arbitrary deep spatial filters. To increase the realism of simulated trajectories during development and evaluation, we develop a synthetic data generation framework based on the social force model. Results validate that the autoregressive incorporation significantly improves the accuracy of our Bayesian trackers, resulting in superior enhancement with none or only negligibly increased computational overhead. Real-world recordings complement these findings and demonstrate the generalizability of our methods to unseen acoustic conditions.
Frobenius orbit codes and finite-field Fourier spectra: exact distance and coset distributions
arXiv:2605.20062v2 Announce Type: replace-cross Abstract: Let $q$ be a prime power, let $m\ge1$, put $Q=q^m$, and let a permutation $\pi$ of a finite index set $I$ satisfy $\pi^m=1$. We study the $\F q$-linear code \[ \mathcal C(\pi)=\{x\in\Lfield^I:x_{\pi(i)}=x_i^q\}, \] which includes the Fourier spectra of $\K$-valued functions on finite abelian groups split by $\Lfield$. Our structural starting point is a $\F q$-linear Hamming isometry taking each cycle of $\pi$ of length $\ell$ to the subfield repetition code $\{(b,\ldots,b):b\in\F{q^\ell}\}$. From this normal form we obtain exact symbol weight and received-word distance enumerators, all list sizes, and the covering radius. An occupancy generating function gives the complete coset-leader distribution, the exact mean distance, the probability of a unique nearest word, and the number of deep holes. For the full cyclic family of length $q^m-1$, the deficit of a uniformly random ambient word from the covering radius is asymptotically Poisson with mean $(m-1)/2$, with an additional $1$ when $m$ is even; a central limit theorem follows. We also determine the deep-hole probability and the second-order gap to the sphere-covering bound. Classical Fourier--Galois descent and cyclotomic orbit enumeration enter only to identify the Fourier specialization.
Immediate 3D Gaussian Splat Reconstruction of Unordered Input with Global Consistency
arXiv:2607.14481v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) has become the method of choice for reconstructing and real-time rendering of captured scenes. To capture a scene with good visual quality, continuous image sequences are usually combined with out-of-order shots for better scene coverage. Structure from motion can reconstruct such captures, but only after they are all available and often with high computational cost. Incremental reconstruction methods -- often derived from SLAM solutions -- provide immediate feedback, but cannot handle the out-of-order capture we require. We provide the first immediate feedback solution for such radiance field capture that provides global consistency. We first introduce a method for fast matching in out-of-order sequences, by repurposing visual place recognition models and a covisibility graph, and provide an efficient way to find highly connected keyframes, improving quality even for ordered sequences. We show how these steps -- together with GPU optimization and careful Gaussian primitive placement -- provide fast local reconstruction, in our challenging radiance field reconstruction case. We then introduce a novel cluster-based method, again using the covisibility graph, to provide efficient loop closure that does not require sequential input. Finally, to handle large scenes in our context, we introduce a progressive hierarchy that allows our method to scale to large environments, without compromising efficiency. Our results show we provide immediate feedback 3DGS reconstruction with good visual quality in several datasets, with up to thousands of input images.
Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models
arXiv:2607.14315v1 Announce Type: new Abstract: In this paper, we present a comprehensive framework for assessing the explainability of various XAI methods, such as LIME and SHAP, across multiple datasets and machine learning models, with the ultimate goal of creating a unified multidimensional explainability score. Our methodology focuses on three key aspects of explainability: fidelity, simplicity, and stability. We leverage benchmarking experiments to systematically evaluate these aspects and use the insights gained to construct an offline knowledge base. This knowledge base captures the explainability scores for each registered model and serves as a valuable resource for context-dependent evaluation of explainability. By analyzing the complementary characteristics and metadata of AI models, datasets, and XAI methods, the knowledge base will enable the estimation of explainability scores for previously unseen datasets and models. Properties like fidelity, simplicity, and stability may vary significantly based on the dataset, underlying model, and domain expertise of the end user. We demonstrate our framework by applying it to three open-source datasets, discussing the implications of the obtained results in relation to the characteristics of the datasets. Our work contributes to the growing field of XAI by providing a robust and versatile tool for evaluating and comparing the explainability of various XAI methods, ultimately supporting the development of more transparent and trustworthy AI systems.
Uncertainty-Aware Multi-Source Retinal Fluid Segmentation in OCT
arXiv:2607.12212v2 Announce Type: replace-cross Abstract: Measuring retinal fluid from optical coherence tomography (OCT) drives treatment decisions in macular disease, but manual annotation is slow and segmentation models trained on one scanner degrade on another. We present an attention-guided TransUNet that segments three fluid types across four independent OCT sources, combining a domain-adaptive normalisation scheme with an uncertainty estimate that flags unreliable pixels. The model reaches a mean fluid Dice of 0.78, and -- most usefully for clinicians -- its uncertainty is 1.34x higher exactly where expert graders disagree (p<10^-4), turning a raw segmentation map into an actionable clinical triage signal.
Penny: Transition Network Analysis of Learner-Chatbot Interactions in Scaffolded EFL Writing
arXiv:2607.14575v1 Announce Type: new Abstract: Generative AI chatbots promise to transform English as a Foreign Language (EFL) writing by providing immediate, personalised feedback. However, their pedagogical value depends on how learners engage with them - a process often treated as a "black box." This study uses Transition Network Analysis to model the temporal dynamics of Japanese EFL learners using "Penny," an LLM-powered writing chatbot. Analysis of over 4,500 writing sessions and 21,000 chatbot interactions reveals two dominant behavioural loops: a "Revision Loop," where feedback leads directly to successful error correction, and a "Chat Loop," where learners engage in sustained dialogue with the chatbot following feedback. Crucially, EFL proficiency significantly shapes interaction: high-proficiency learners engage more in open dialogue and negotiation with the chatbot, while low-proficiency learners rely more heavily on repetitive corrective feedback cycles. The findings demonstrate that AI-scaffolded writing is a non-linear, dialogic process and highlight the need for differentiated chatbot design to move beyond simple error correction and foster deeper cognitive engagement for all learners.
On-Policy Delta Distillation
arXiv:2607.15161v1 Announce Type: new Abstract: On-policy distillation is an alternative post-training method in reinforcement learning that alleviates the constraints imposed by reward models by providing token-level supervision from a teacher model. Although on-policy distillation has been studied and applied across various settings, its fundamental design remains underexplored. In this paper, we introduce a new distillation reward, termed the delta signal, instead of directly imitating the teacher's output distribution. The delta signal is defined as the difference between the teacher model and its base model prior to instruction tuning for reasoning capability. It therefore captures the changes induced by reasoning tuning and provides a more direct signal for transferring reasoning capabilities. Using extensive empirical evidence, we show that the delta signal substantially improves on-policy distillation and refer to the new distillation method as On-Policy Delta Distillation (OPD$^2$). Experiments across mathematics, science, and code-reasoning benchmarks demonstrate that OPD$^2$ consistently outperforms conventional on-policy distillation, enabling reasoning LLMs to achieve strong performance with only a short post-training period. Code will be available at https://github.com/naver-ai/opd2
Unified framework of optical thermodynamics and optical pressure
arXiv:2607.15162v1 Announce Type: new Abstract: Optical thermodynamics is a newly developed framework that applies principles from statistical mechanics to describe the intricate behavior of weakly nonlinear, multimode photonic systems. Utilizing this theory, the collective dynamics of complex optical arrangements can be systematically uncovered and understood. The purpose of this work is to examine fundamental aspects of optical thermodynamics, including optical pressure, and provide a unified framework that can be applied to effectively any optical setting within the domain of validity of optical thermodynamics. We find that in addition to the conservation laws, the remaining extensive and intensive parameters of the system are naturally provided by the parameters of the propagation constants. At this point, several thermodynamic approaches exist for analyzing optical forces in multimode settings. Here, we develop a new theoretical methodology that unifies these perspectives in a variety of different configurations, irrespective of whether they are discrete or continuous. We apply our theory in four different settings. By studying Su-Schrieffer-Heeger lattices, we elucidate the thermodynamics of polyatomic chains and show that intercell and intracell bonds can display different optical forces. In addition, we provide a thermodynamic formalism to predict and understand the optical pressure at equilibrium arising in arrangements characterized by a continuous index variation and apply our results to graded-index fibers.
Sharp Stability Threshold and Certification for Designing Stable Residual Architectures
arXiv:2607.14576v1 Announce Type: new Abstract: We propose \emph{the sublinear-growth principle} for deep residual architectures -- a sharp stability threshold on the input-magnitude exponent of every residual block's velocity field: $$\|v(x, t)\| \leq c\,\|x\|^q + b, \qquad q \in [0, 1].$$ The threshold $q = 1$ is established via two independent arguments. Classical ODE theory gives a global forward flow on $[0, T]$ at $q \le 1$ and exhibits divergent velocity fields at any $q > 1$. The optimal-control analysis, via the Hamilton-Jacobi-Bellman equation, sharpens this to a selection statement: the training optimum is bang-bang on the boundary of the admissible class, so the optimum at $q > 1$ blows up while the optimum at $q \le 1$ is safe by construction. The exponent criterion $q \le 1$ is thereby a necessary and sufficient condition for stable training. It clarifies architectural placements that ensure the stability of training and inference, explaining, for instance, the stabilizing role of layer normalization. The sublinear-growth velocity fields form \emph{the right function space} on which forward dynamics, adjoint sensitivity, and architectural composition are all well-controlled. An arithmetic of input-magnitude exponents under the five operations that build residual blocks enables efficient certification of $q_k \le 1$ at the level of architectural primitives, in place of ad hoc trial and error in the search for stable neural architectural designs. A parameter-free modification reduces the supercritical Mamba block from $q = 5$ to $q = 1$ without layer normalization, demonstrating this point. Experiments on Mamba and PatchTST confirm that the $q \le 1$ variants train stably: the criterion is the input-magnitude exponent, not the presence of a normalization layer.
CoTu at EXACT 2026: Neuro-Symbolic Reasoning for Transparent Educational QA
arXiv:2607.14735v1 Announce Type: new Abstract: Transparent educational question answering asks for answers that are not only correct but explainable, and doing so with small models rules out the reasoning power of the largest proprietary systems. The EXACT 2026 competition poses this problem concretely: open-weight language models of at most 8B parameters, self-hosted, with a natural-language explanation for every answer. It pairs two tasks: logical reasoning over university regulations, and multi-step physics problem solving. We describe the system that team \cotu{} developed to address both, a neuro-symbolic Program-of-Thought pipeline in which a 4B backbone writes a program rather than stating an answer directly: for regulation queries it emits a Z3 encoding whose entailment verdict grounds the deduction, and for physics it emits numerical Python, both wrapped in a shared self-correction loop and a unified explained-JSON output. Answer-type routing, distillation-based task fine-tuning, and a latency-aware serving stack -- SGLang with speculative decoding -- keep the system within the 60-second per-query limit. The system achieved a \textbf{perfect score} on the physics task in both automated selection rounds and obtained the \textbf{highest final-round technical score} of any team -- $13.44/15$, combining automated answer evaluation with expert-judged reasoning depth -- with the equally weighted presentation score included, \cotu{} placed 3rd overall. Grounding answers in a symbolic solver yields correct, verifiable deductions at the 4B scale, and the residual difficulty lies in premise selection rather than the deduction itself.
The climates and thermal emission spectra of prime nearby temperate rocky exoplanet targets
arXiv:2504.00978v2 Announce Type: replace-cross Abstract: Over the course of the past decade, advances in the radial velocity and transit techniques have enabled the detection of rocky exoplanets in the habitable zones of nearby stars. Future observations with novel methods are required to characterize this sample of planets, especially those that are non-transiting. One proposed method is the Planetary Infrared Excess (PIE) technique, which would enable the characterization of non-transiting planets by measuring the excess infrared flux from the planet relative to the star's spectral energy distribution. In this work, we predict the efficacy of future observations using the PIE technique by potential future observatories such as the MIRECLE mission concept. To do so, we conduct a broad suite of 21 General Circulation Model (GCM) simulations with ExoCAM of seven nearby habitable zone targets for three choices of atmospheric composition with varying partial pressure of CO$_2$. We then construct thermal phase curves and emission spectra by post-processing our ExoCAM GCM simulations with the Planetary Spectrum Generator (PSG). We find that all cases have distinguishable carbon dioxide and water features assuming a 90$^\circ$ orbital inclination. Notably, we predict that CO$_2$ is potentially detectable at 15 $\mu\mathrm{m}$ with MIRECLE for at least four nearby known non-transiting rocky planet candidate targets in the habitable zone: Proxima Cenaturi b, GJ 1061 d, GJ 1002 b, and Teegarden's Star c. Our ExoCAM GCMs and PSG post-processing demonstrate the potential to observationally characterize nearby non-transiting rocky planets and better constrain the potential for habitability in our Solar neighborhood.
Reinforcement Learning in Switching Non-Stationary Markov Decision Processes: Algorithms and Convergence Analysis
arXiv:2503.18607v2 Announce Type: replace Abstract: We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite set of MDPs governed by a latent Markov chain while the agent observes only the external state. We show that the long-term effect of this switching is equivalent to stationary dynamics parameterized by the stationary distribution of the hidden Markov chain. For fixed policies, we derive a closed-form expression for the SNS value function and prove that standard temporal-difference (TD) learning converges to it almost surely despite persistent non-stationarity. We further establish that policy iteration converges to the optimal policy of the equivalent averaged environment, and prove that tabular Q-learning converges almost surely to the optimal Q-function. The framework is validated on a wireless communication network with Markovian channel noise, demonstrating its practical efficacy for decision-making in rapidly time-varying systems.
Scaling Evaluation-time Compute with Reasoning Models as Evaluators
arXiv:2503.19877v3 Announce Type: replace Abstract: As language model (LM) outputs get more and more natural, it is becoming more difficult than ever to evaluate their quality. Simultaneously, increasing LMs' "thinking" time through scaling test-time compute has proven an effective technique to solve challenging problems in domains such as math and code. This raises a natural question: can an LM's evaluation capability also be improved by spending more test-time compute? To answer this, we investigate employing reasoning models-LMs that natively generate long chain-of-thought reasoning-as evaluators. Specifically, we examine methods to leverage more test-time compute by (1) using reasoning models, and (2) prompting these models to evaluate not only the response as a whole (i.e., outcome evaluation) but also assess each step in the response separately (i.e., process evaluation). In experiments, we observe that the evaluator's performance improves monotonically when generating more reasoning tokens, similar to the trends observed in LM-based generation. Furthermore, we use these more accurate evaluators to rerank multiple generations, and demonstrate that spending more compute at evaluation time can be as effective as using more compute at generation time in improving an LM's problem-solving capability.
FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models
arXiv:2607.14739v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapping current observations and language to actions without explicit forward prediction of world dynamics. Existing visual foresight methods predict future visual states but lack explicit motion guidance: they show where to go but not how to get there. We argue that future feature prediction and sparse point tracking are naturally complementary: the former provides the goal state, while the latter captures the continuous motion path toward it. We propose FoMoVLA, a framework that augments VLA representations with explicit spatio-temporal supervision by jointly learning future feature foresight and sparse 2D point tracking, enhancing the continuous action policy. FoMoVLA introduces compact foresight tokens to decode future feature states, decodes sparse temporal 2D point trajectories to model compact geometric motion, and couples both through a lightweight future-conditioned cross-attention module that enables consistent reasoning between anticipated states and point dynamics. Extensive experiments on LIBERO, RoboCasa GR-1 Tabletop, and LIBERO-Plus demonstrate state-of-the-art performance and strong zero-shot generalization. Project page is available at https://liauto-research.github.io/FoMoVLA.
DINE: Distance Is Not Enough -- Learning Global Deformation Priors for Robust Soft-Tissue Point Cloud Registration
arXiv:2607.14946v1 Announce Type: new Abstract: Non-rigid point cloud registration is central to soft-tissue shape analysis, but large deformations, noise, and outliers make correspondence estimation challenging. Most learning-based methods rely on local objectives such as Chamfer distance, which encourage point-wise proximity but do not constrain the global plausibility of the predicted deformation field. We address this limitation with DINE, a maximum a posteriori framework that augments distance-based registration with a learned statistical prior over displacement vector fields. DINE is applied to two registration backbones, Robust-DefReg and DefTransNet, using a two-stage strategy: a first-stage model is trained with Chamfer distance, its predicted deformation fields are used to estimate a prior, and the model is then refined with a combined distance and negative log-prior objective. We compare a full-field PCA Gaussian prior with a per-vector normalizing-flow prior. Experiments on DeformedTissue and SynBench show lower mean Chamfer distance under deformation and corruption. On DeformedTissue, DINE-PCA reduces Chamfer distance by approximately 27--69\% relative to the corresponding Stage-1 backbone across deformation levels, and improves robustness by up to 66\% for outliers and 83\% for Gaussian noise. On SynBench, improvements are modest at the smallest deformation levels and reach approximately 59--79\% from moderate to severe deformation. These results suggest that global deformation plausibility is an important constraint for reliable soft-tissue point cloud registration. (The code will be published soon.)
DialogueVPR: Towards Conversational Visual Place Recognition
arXiv:2607.14115v1 Announce Type: new Abstract: Inspired by how humans communicate spatial information, language-guided geo-localization has gained significant traction for its intuitive and practical value. Despite this progress, most methods still rely on a static, one-shot retrieval paradigm, which fails to handle the ambiguity and incompleteness inherent in real-world natural language descriptions. We propose a paradigm shift to reasoning retrieval and introduce Dialogue Place Recognition (DlgPR), which casts localization as an interactive, dialogue-driven reasoning process. To support this new task, we present DlgQuest-Cities, the first large-scale dialogue-based benchmark for place recognition, and a unified reasoning framework that couples a cross-modal multi-level retriever with an intelligent questioner, DQ-pilot. DQ-pilot is trained in a curriculum: supervised fine-tuning on a curated DQ-cities-20k subset followed by reinforcement refinement on a harder DQ-cities-10k split via GRPO. Two task-aligned metrics guide learning: a Discriminative Difficulty Index (DDI) for curriculum sampling and a Positional Retrieval Gain (PRG) reward that directly measures retrieval improvement induced by a question. Experiments show this reasoning-based approach significantly outperforms baselines. The code and model are available at https://github.com/Graysonggg/DlgPR.
On the Disagreement in Perturbation-based xAI -- Benchmarking Perturbation Choices for Flood Detection from SAR Images
arXiv:2607.14743v1 Announce Type: new Abstract: Perturbation-based xAI methods are widely used to analyze the behavior and predictions of deep learning models. By altering input regions and measuring the resulting changes in class probabilities with respect to the original image, they assign relevance scores and generate heatmaps that reflect each region's contribution to the prediction. Despite their apparent simplicity, however, perturbation-based methods are sensitive to parameter choices. In this work, we focus on two key parameters of the perturbation pipeline, namely the patch geometry, including the size and shape of the perturbed regions, and the perturbation type, defined by the replacement scheme. Grounded in the use case of flood detection from Synthetic Aperture Radar imagery, we conduct a comprehensive investigation of how relevance estimation changes under different perturbation settings. Beyond visual inspection of the resulting relevance maps, we evaluate their consistency across perturbation strategies and their faithfulness to the model's reasoning. We demonstrate how different perturbation choices can steer the resulting relevance maps, yielding ambiguous and even contradictory explanations. Our findings emphasize the importance of methodological settings in perturbation-based xAI. They underscore the need to carefully inspect and evaluate perturbation choices and to treat them as an integral part when interpreting explanations, ensuring a robust understanding of both the explanations and model predictions.
Routing Ceilings Are Domain-Independent: Structural Prior Injection in Code Security Vulnerability Detection
arXiv:2607.14628v1 Announce Type: new Abstract: Large language models (LLMs) exhibit a well-documented gap between latent capability and consistent activation: the router hypothesis posits that models possess the knowledge to solve a task but lack reliable internal routing to activate it. Prior work in formal mathematical reasoning (SAIR, C\'azares 2026) reports that structural priors (cheatsheets) raise in-distribution performance dramatically, yet collapse below the zero-shot baseline out-of-distribution (OOD) -- and that iterative recalibration amplifies rather than corrects the collapse. We test whether this phenomenon is cross-domain by reproducing the SAIR design in source-code security vulnerability detection, evaluating three LLMs (GPT-OSS-120B, Llama-3.3-70B, Gemma-4-31B) across three vulnerability categories (CWE-798, CWE-284, and the non-CWE N+1 anti-pattern) spanning syntactic, contextual, and semantic complexity, then transferring cheatsheet-augmented prompts to real-world CVE data from VUDENC (CWE-89, CWE-22). Our findings replicate and extend SAIR: (F1) structural priors lift semantic-vulnerability recall from 20.0% to 100.0% across all models; (F2) zero-shot performance degrades along a semantic complexity gradient; (F3) the same cheatsheets that saturate synthetic performance amplify distribution-shift collapse on real CVE data (CWE-89: 100% synthetic F1 to 48.9% on VUDENC, -51.1pp); (F5) iterative recalibration produces a v2 cheatsheet that performs worse than v1 on real data, mirroring SAIR's AN45c-vs-AN38 finding. These results provide evidence that the cross-distribution trade-off surface documented in SAIR generalises to code security, and that the router hypothesis is cross-domain. We argue the structural nature of the collapse motivates distribution-aware training over prompt calibration. Code and evaluation scripts: https://github.com/bytepro-ai/bitcoder-v2-research