Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Data Alchemy: Mitigating Cross-Site Model Variability Through Test Time Data Calibration
arXiv:2407.13632v2 Announce Type: replace Abstract: Deploying deep learning-based imaging tools across various clinical sites poses significant challenges due to inherent domain shifts and regulatory hurdles associated with site-specific fine-tuning. For histopathology, stain normalization techniques can mitigate discrepancies, but they often fall short of eliminating inter-site variations. Therefore, we present Data Alchemy, an explainable stain normalization method combined with test time data calibration via a template learning framework to overcome barriers in cross-site analysis. Data Alchemy handles shifts inherent to multi-site data and minimizes them without needing to change the weights of the normalization or classifier networks. Our approach extends to unseen sites in various clinical settings where data domain discrepancies are unknown. Extensive experiments highlight the efficacy of our framework in tumor classification in hematoxylin and eosin-stained patches. Our explainable normalization method boosts classification tasks' area under the precision-recall curve(AUPR) by 0.165, 0.545 to 0.710. Additionally, Data Alchemy further reduces the multisite classification domain gap, by improving the 0.710 AUPR an additional 0.142, elevating classification performance further to 0.852, from 0.545. Our Data Alchemy framework can popularize precision medicine with minimal operational overhead by allowing for the seamless integration of pre-trained deep learning-based clinical tools across multiple sites.
Unconstrained Scheme for Geometrically Constrained Gradient Flows
arXiv:2607.08577v1 Announce Type: new Abstract: In this paper, we study the approximation of gradient flows of harmonic maps, which serve as model problems for applications in micromagnetics, liquid crystals, and nonlinear plate bending. Harmonic maps are vector fields that are critical points of the Dirichlet energy subject to the constraint that the vector field be unit length pointwise. Most existing time-stepping schemes for gradient flows deal with the constraint by linearizing the unit length constraint at every step, which involves solving for the solution increment in the tangent space of the constraint. These schemes lead to robust control over the violation of the constraint, but require solving degenerate saddle point systems at every step that may be difficult to precondition. In this paper, we propose a scheme that first computes the unconstrained increment and then projects this increment pointwise onto the tangent space. With an additional stabilization, this scheme is energy stable under mild step size restrictions and provides robust control of the unit length constraint violation. Our new scheme only requires the solution of decoupled symmetric positive definite systems at every step, which translates to a large increase in computational efficiency. We also propose a computable a posteriori criterion and a variable time-stepping procedure that guarantee the stability of the scheme. We conclude with computational examples demonstrating the efficacy of the scheme, and present a computational extension of the scheme to nonlinear plate bending.
Robust Dynamic Operating Envelopes in Unbalanced Three-Phase Distribution Systems
arXiv:2607.08578v1 Announce Type: new Abstract: This paper proposes a robust optimization formulation to calculate dynamic operating envelopes (DOEs) to safely operate unbalanced three-phase distribution systems. Unlike conventional formulations that satisfy network constraints only at the envelope bound, the robust formulation covers the entire envelope range. We formulate a robust non-linear programming (NLP) problem with the full AC power flow equations, as well as an approximate linear programming (LP) model. Numerical simulations are run with real-world data from Belgium and two different distribution test feeders. The paper compares the conventional approaches with their robust counterparts and examines the trade-off between constraint violation and envelope size as well as accuracy and solve time aspects.
HeadRoom: Lightweight, Edge-deployable Pipeline for Adaptive Notification Routing
arXiv:2607.08083v1 Announce Type: new Abstract: Emerging wearables, such as smart glasses, can deliver notifications through multiple sensory channels, but there is still a limited understanding of how to choose the right channel at the right moment. We propose HeadRoom, a lightweight, edge-deployable pipeline that estimates the availability of visual and auditory channels in real time from egocentric video and audio. Our controlled user study (N=25) shows that, under high perceptual load, routing notifications to the more available channel reduces response time relative to routing them to the less available channel. This work opens up a new possibility for adaptive routing of notifications in wearable and immersive systems.
Wat3R: Underwater 3D Geometry Learning without Annotations
arXiv:2607.08772v1 Announce Type: new Abstract: Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-quality 3D annotations. Pioneering methods rely on massive dense annotations that are impractical in underwater settings. In this paper, we propose Wat3R, a cross-domain semi-supervised learning framework designed to adapt feed-forward 3D reconstruction models from air to underwater scenes. Uniquely, our method eliminates the need for any annotated underwater data following a teacher-student architecture, that learns robust geometry representations merely on abundant unlabeled real underwater video footage. We also design a cross-view consistency loss that leverages geometric cues from other views to compensate for the information degradation in the current view caused by water attenuation and scattering. Furthermore, considering the lack of comprehensive evaluation benchmarks, we construct Water3D, a diverse dataset covering various water bodies and underwater scenarios, designed for geometric task evaluation. Experimental results demonstrate that Wat3R outperforms current state-of-the-art methods in underwater multi-view depth estimation and point cloud reconstruction. The dataset and code are available at https://github.com/LSXI7/Wat3R .
Attention-Based Segmentation of WMHs and Differentiation of Vascular vs. Demyelinating Lesions
arXiv:2607.08171v1 Announce Type: new Abstract: White Matter Hyperintensities (WMHs) are commonly observed in brain Magnetic Resonance Imaging (MRI) scans. They are associated with various neurological conditions, including vascular and inflammatory demyelinating diseases. Despite differing in etiology, WMHs from these conditions often appear similar on Fluid Attenuated Inversion Recovery (FLAIR) images. This similarity makes differential diagnosis challenging. In this work, we highlight the potential of combining attention-based segmentation with feature-driven classification. This approach supports more accurate and efficient classification between vascular and demyelinating white matter pathologies. For segmentation, we evaluate the effectiveness of attention mechanisms, specifically the Bottleneck Attention Module (BAM) and the Convolutional Block Attention Module (CBAM). We also test different architectures, particularly Attention U-Net. In addition, we explore advanced training strategies, such as patch-based learning and a 2.5D approach, to enhance lesion detection. After segmentation, we extract morphological features from the lesion masks. We then use them to classify WMHs based on their underlying cause. Our experiments utilize five publicly available datasets with diverse imaging protocols to promote model generalizability, despite limited sample sizes. The results suggest that attention-based segmentation and feature-driven classification offer a promising direction for discriminating vascular and demyelinating white matter lesions. Further validation in larger clinical cohorts is still needed.
Probing individual phonon-polaritonic nanoparticle-on-mirror cavities by infrared nanospectroscopy
arXiv:2607.07941v1 Announce Type: new Abstract: Nanoparticle-on-mirror (NPoM) cavities enable extreme light confinement and strong light-matter interactions, but their realization with phonon-polariton materials in the mid-infrared spectral range remains largely unexplored. Here, we use nano-FTIR spectroscopy to study the near-field response of individual phononic NPoM cavities formed by gold nanoparticles on a quartz substrate supporting phonon-polaritons. By placing a metal tip on top of the NPoM and recording the tip-scattered field, we observe two reproducible cavity resonances, identified as the fundamental and a second-order antenna modes by numerical simulations. The calculations show that, in absence of the tip, the NPoM cavity exhibits ultrasmall mode volumes ($V \sim 10^3$ nm$^3$) and high quality factors ($Q \sim 100$), resulting in extraordinary field intensity enhancements ($F \sim 10^4$) and Purcell factors ($P_F \sim 10^9$). They also indicate that the nano-FTIR tip enables efficient excitation and readout of these phononic NPoM modes without perturbing them, while enhancing the intrinsic local field intensity in the NPoM gap by two orders of magnitude. Our results establish phononic NPoM cavities as a promising platform for mid-infrared nanophotonics and pave the way for ultrasensitive vibrational spectroscopy based on nano-FTIR measurements of individual cavities.
Electronic Strong Coupling of Gas-Phase Molecular Iodine
arXiv:2602.09243v2 Announce Type: replace Abstract: Molecular polaritons, hybrid light-matter states formed from the strong coupling of molecular transitions and discrete photonic modes, are a compelling platform for optical control of chemical reactivity. Despite the origins of the field of polaritonics in atomic gases, strong coupling of molecular gases remains underexplored. The pristine, solvent-free gas-phase environment may prove ideal for gaining mechanistic understanding of molecular behavior under strong light-matter coupling. In this work, we achieve electronic strong coupling of the B-X, $\nu_1$ = 0$\rightarrow$32, J = 53$\rightarrow$52 and B-X, $\nu_1$ = 0$\rightarrow$34, J = 103$\rightarrow$102 rovibronic transitions of gas-phase iodine (I$_2$) lying near 532.2 nm. We access a range of coupling strengths and detuning conditions with fine control over molecular number density and cavity length stabilization. This effort represents the first demonstration of electronic polaritons in a molecular gas and opens a new platform for polariton photochemistry and photophysics.
KronQ: LLM Quantization via Kronecker-Factored Hessian
arXiv:2607.07964v1 Announce Type: new Abstract: Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining. Existing second-order PTQ methods, including GPTQ, construct quantization objectives exclusively from input activation statistics, effectively assuming that all output channels contribute equally to the layer-wise reconstruction objective. We propose KronQ, a PTQ framework that challenges this assumption by introducing the gradient covariance into the quantization pipeline. Under the Kronecker-factored Hessian approximation, the quantization loss depends jointly on both the activation and gradient covariances, and KronQ exploits this at two complementary levels. (1) KronQ introduces bidirectional incoherence processing, extending the existing input-side random rotation to the output dimension using the gradient covariance, reducing weight magnitude variance across both input and output dimensions. (2) KronQ derives a new sensitivity metric for inter-layer mixed-precision allocation, driven by the gradient and activation Hessian traces. Notably, in the case of 2-bit weight-only quantization on LLaMA-3-70B, while GPTQ and GPTAQ diverge or produce degenerate quantizations (>2000 perplexity on WikiText-2), KronQ achieves 7.93 perplexity.
ASMR: Agentic Schema Generation for Ship Maintenance Report Writing
arXiv:2607.08177v1 Announce Type: new Abstract: In this paper, we study the automatic schema generation problem: given a collection of historical ship maintenance and operational reports across multiple form categories, automatically discover compact and informative schemas that capture the essential information requirements of each report type. To address this challenge, we propose ASMR, a modular agentic framework consisting of two specialized agents. A Field Generation Agent extracts semantic concepts from historical narratives and generates candidate schema fields through adaptive multi-granularity clustering, while a Structural Optimizer Agent employs reinforcement learning to identify compact, informative, and non-redundant schema representations. The resulting schemas can guide report authors toward producing more complete, consistent, and actionable reports. Preliminary results demonstrate the promise of the proposed approach and highlight several open research challenges at the intersection of data management, agentic AI, and human-centered AI.
ShannonProver: Towards Automating Formal Cryptographic Proofs
arXiv:2607.02847v2 Announce Type: replace Abstract: Cryptographic proofs are produced at a scale that increasingly exceeds the community's ability to verify them manually. Machine-checked proofs offer a path toward scalable proof verification, but writing proof scripts for expressive proof assistants such as EasyCrypt remains a major bottleneck: even when the high-level proof plan is known, converting it into proof tactics requires substantial reasoning effort. This paper presents ShannonProver, an agentic framework for automating cryptographic proofs. ShannonProver targets the setting in which a cryptographer provides the security model and a decomposition of the target theorem into lemma-level proof obligations, while the system automatically constructs EasyCrypt proof scripts for those obligations. We evaluate ShannonProver on a dataset of formal cryptographic proofs in EasyCrypt. The dataset spans textbook primitives, deployed protocols, and standardization efforts such as NIST proposals, and includes expert case studies drawn from a corpus that has not previously been available online. We show that ShannonProver can automate substantial portions of cryptographic proof engineering for case studies such as ChaChaPoly1305 and MEE-CBC. More broadly, this work suggests a path toward accelerating cryptographic research: as agents automate the proof-engineering burden, cryptographers can iterate more quickly on new constructions, obtain machine-checked assurance earlier, and bring trustworthy protocols from design to deployment faster.
Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG
arXiv:2607.02966v2 Announce Type: replace Abstract: Cross-lingual retrieval-augmented generation (RAG) is often deployed in an English-evidence regime, where users query in diverse languages but retrieved passages remain English. In this setting, generation can fail despite strong base models: English evidence induces language drift (English or code-switching outputs) and models use evidence unreliably when producing non-English answers. We attribute these failures to two post-training challenges: (i) errors are prefix-dependent, so fixed-trajectory supervision suffers from prefix mismatch; and (ii) sequence-level (partly discrete / judge-based) rewards yield noisy credit assignment and high-variance updates. We propose TR-RAG, a teacher-regularized RL recipe that couples reward optimization with on-policy distillation on student-visited prefixes. A compact student samples on-policy answers, while a stronger frozen teacher is queried only on those prefixes and provides a prefix-wise student-to-teacher reverse-KL anchor. We further introduce a reward decomposition for English-evidence multilingual generation, combining language consistency, character 3-gram recall, and an LLM-judge score for evidence-grounded correctness. Across three benchmarks (BioASQ-ENKB5, Hotpot-ENKB5, and naturally multilingual MKQA) and two backbones, TR-RAG improves the composite of language adherence and evidence-grounded correctness over strong baselines. Crucially, the teacher anchor acts as a safety net: on in-domain languages it prevents the large language-consistency collapses (up to ~27 percentage points) that reward-only RL can suffer by drifting below even the base model, while on distant out-of-distribution languages, where reward-only RL stalls at the base model's ceiling, it still improves evidence grounding; and on character 3-gram recall the compact student sometimes surpasses its 70B teacher.
Neural non-canonical Hamiltonian dynamics for long-time simulations
arXiv:2510.01788v2 Announce Type: replace Abstract: This work focuses on learning non-canonical Hamiltonian dynamics from data, where long-term predictions require the preservation of structure both in the learned model and in numerical schemes. Previous research focused on either facet, respectively with a potential-based architecture and with degenerate variational integrators, but new issues arise when combining both. In experiments, the learnt model is sometimes numerically unstable due to the gauge dependency of the scheme, rendering long-time simulations impossible. In this paper, we identify this problem and propose two different training strategies to address it, either by directly learning the vector field or by learning a time-discrete dynamics through the scheme. Several numerical test cases assess the ability of the methods to learn complex physical dynamics, like the guiding center from gyrokinetic plasma physics.
TOPO-Bench: An Open-Source Topological Mapping Evaluation Framework with Quantifiable Perceptual Aliasing
arXiv:2510.04100v2 Announce Type: replace Abstract: Topological mapping offers a compact and robust representation for navigation, but progress in the field is hindered by the lack of standardized evaluation metrics, datasets, and protocols. Existing systems are assessed using different environments and criteria, preventing fair and reproducible comparisons. Moreover, a key challenge - perceptual aliasing - remains under-quantified, despite its strong influence on system performance. We address these gaps by (1) formalizing topological consistency as the fundamental property of topological maps and showing that localization accuracy provides an efficient and interpretable surrogate metric, and (2) proposing the first quantitative measure of dataset ambiguity to enable fair comparisons across environments. To support this protocol, we curate a diverse benchmark dataset with calibrated ambiguity levels, implement and release deep-learned baseline systems, and evaluate them alongside classical methods. Our experiments and analysis yield new insights into the limitations of current approaches under perceptual aliasing. All datasets, baselines, and evaluation tools are fully open-sourced to foster consistent and reproducible research in topological mapping.
SoK: Market Microstructure for Decentralized Prediction Markets (DePMs)
arXiv:2510.15612v4 Announce Type: replace Abstract: Decentralized prediction markets (DePMs) allow open participation in event-based wagering without fully relying on centralized intermediaries. We review the history of DePMs which date back to 2011 and includes hundreds of proposals. Perhaps surprising, modern DePMs like Polymarket deviate materially from earlier designs like Truthcoin and Augur v1. We use our review to present a modular workflow comprising eight stages: underlying infrastructure, market topic, share structure and pricing, market initialization, trading, market resolution, settlement, and archiving. For each module, we enumerate the design variants, analyzing trade-offs around decentralization, expressiveness, and manipulation resistance. We also identify open problems for researchers interested in this ecosystem.
Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator
arXiv:2607.07993v1 Announce Type: new Abstract: Identifying faithfulness hallucinations in LLM-generated outputs remains challenging due to the scarcity of high-quality annotated data. Recent work relies on advanced LLMs to synthesize training data, including rationales, labels, and hallucinated claims. However, these methods treat the generator as a static component, limiting iterative improvement of the detector. To address this limitation, we introduce Hallucination Self-Play (HSP), a novel framework that enables the detector to bootstrap with an evolved generator. HSP involves two roles initialized from the same base model, a detector that assesses the faithfulness of model outputs, and a generator that produces increasingly hard-to-detect hallucinated responses. Specifically, the detector is first fine-tuned on human-labeled data and then employed as a reward model to train the generator via reinforcement learning from AI feedback (RLAIF). In turn, the evolved generator synthesizes hallucination data to further optimize the detector through rule-based reinforcement learning. Experiments on RAGTruth benchmark and two model families demonstrate that the proposed framework can progressively enhance a small LLM to match or even outperform advanced LLMs without external supervision. Our code is available at https://anonymous.4open.science/r/Hallucination-Self-Play-50B5 .
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing
arXiv:2607.08646v1 Announce Type: new Abstract: As available training data approaches its physical limit, gains from Scaling Laws have begun to diminish. Consequently, improving Large Language Models (LLMs) now depends less on data expansion and more on higher-quality data utilization. However, in the context of large-scale corpora, existing refinement methodologies face significant limitations in quality, efficiency, and reliability: Rule-based approaches are constrained by fixed heuristics and struggle with instance-level variations; LLM-based approaches improve quality but fail to meet the efficiency and reliability requirements of large-scale data processing. To address these challenges, we propose UltraX, a function-calling refinement framework for large-scale pre-training data that completes the editing function space by introducing insertion in addition to deletion and modification, enabling fine-grained instance-level editing. Specifically, UltraX builds a reliable program-supervision generation pipeline. In this pipeline, dataset-adaptive prompt optimization first guides an expert LLM to produce high-quality end-to-end refined texts, and Line Alignment Mapping and Dynamic Context Replacement then convert original-refined text pairs into structured program supervision. Meanwhile, UltraX improves supervision quality and stabilizes the training distribution with low-confidence example filtering and ratio-controlled sampling by operation combination. During inference and execution, it normalizes and validates model outputs through sliding-window prediction, global operation aggregation, and systematic post-processing, improving the stability and reliability of large-scale execution. Experiments show that UltraX achieves the highest average performance across all corpora and also matches or surpasses baselines with fewer training tokens, demonstrating stronger data efficiency and refinement reliability.
EdgeRefine: Privacy-Utility Balance for Graphs via Jaccard Sampling under Edge Differential Privacy
arXiv:2607.08659v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) have shown considerable success in learning from graph-structured data, but their use in privacy-sensitive areas remains difficult because graph structure can leak sensitive link information. To satisfy edge-level differential privacy, a common approach is to inject noise into all elements of the graph's adjacency matrix, thereby obfuscating the existence of any single edge. However, stronger privacy requires more noise, and excessive noise reduces utility, making the privacy-utility balance a major barrier to practical privacy-preserving graph learning. To address this issue, we propose EdgeRefine, a local differential privacy framework that improves this trade-off through adaptive edge refinement. EdgeRefine first estimates edge-existence probabilities using Jaccard similarity and ranks edges for noisy edge removal. To ensure the sparsity and reliability of the final graph, it uses the privacy budget $\epsilon$ to determine the ratio of true to false edges, samples them separately based on this probability ranking, and controls the total number of edges with a separate sampling rate $k$. Extensive experiments show that EdgeRefine achieves accuracy comparable to the noise-free baseline and substantially outperforms other privacy-preserving methods across datasets and GNN architectures. Under privacy budget $\epsilon = 2.5$, EdgeRefine improves node classification accuracy over state-of-the-art baselines by 17.8\% on ACM under GAT and 19.7\% on Cora under GCN. In graph classification, it achieves an average accuracy degradation of around 5\% compared to the noise-free baseline. Under graph reconstruction attacks, EdgeRefine maintains relative absolute error levels above 1 across all privacy budgets, averaging 1.962 on Cora and 1.472 on AMAP, indicating strong resilience against privacy leakage.
Structure-Guided Gauss-Newton Method: Linear Advection-Reaction Equation
arXiv:2607.07506v2 Announce Type: replace Abstract: The least-squares neural network (LSNN) method introduced in [5] for linear advection-reaction equations is capable of accurately approximating discontinuous solutions without a priori knowledge of the interface location. However, the resulting discretization is a non-convex optimization problem that is computationally intensive and complex. In this paper, we propose a structure-guided Gauss-Newton (SgGN) method that alternates between the linear (output) and the nonlinear (hidden layer) parameters. At each outer iteration, the linear parameters are computed by a linear solver, and the nonlinear parameters are updated by a modified Gauss-Newton (GN) method that explicitly removes the singularities of the GN matrix. Numerical experiments for all test problems presented in [5] show that the SgGN method is superior to the Adam optimizer [13], the commonly used first-order optimization algorithm, not only in computational cost but, more importantly, in accuracy
Trustworthy Machine Learning through the Lens of Combinatorial Optimization: Survey and Research Perspectives
arXiv:2607.07762v1 Announce Type: new Abstract: Modern machine learning (ML) increasingly relies on complex models whose behavior is difficult to characterize beyond empirical performance metrics. Across a wide range of tasks, including prediction, generation, and decision-making, models with similar empirical performance can exhibit markedly different properties in terms of their transparency, interpretability, robustness, fairness, privacy, and certifiability. This survey highlights how optimization- and certification-oriented reasoning can provide a useful framework for reasoning about such differences, supporting tasks ranging from model training and selection to auditing and certification. We review and synthesize recent advances at the intersection of combinatorial optimization (CO) and trustworthy ML, covering both training and post-training tasks, including interpretable model learning, explanation generation, robustness analysis, fairness auditing, model compression, and privacy attacks and protections. Across these domains, CO formulations offer additional capabilities over purely heuristic approaches, e.g., gradient-based ones, notably global guarantees, formal certificates, and explicit treatment of trade-offs. While scalability remains an important challenge, continued progress in solvers and hybrid algorithms suggests a growing role for CO in the design and deployment of trustworthy ML systems.
Buffy versus Bella: An archetypometric analysis and comparison
arXiv:2607.07826v1 Announce Type: new Abstract: Fictional stories and characters embody and encode social norms, and their study is a powerful tool through which to understand culture and society. Vampire stories and folklore, in particular, have long both reflected and refracted people's preoccupation with disease, sexuality, death, and immortality. Here, we explore female main characters from two popular vampire franchises of the 21st century: Buffy Summers from the eponymous Buffy the Vampire Slayer and Bella Swan from the Twilight series. We employ the archetypometrics framework, built from 2,000 characters assesed across 464 semantic differential traits, to understand Buffy's and Bella's archetypes compared to one another and characters in their own stories, as well as within a larger societal context. While Buffy and Bella are female protagonists who share focus on love and romance, they differ broadly on their underlying traits and overall archetypes. Buffy -- presented as a prototypical high school cheerleader -- largely bucks traditional gender norms as an strong Adventurer-Hero. Bella -- stylized as ``not like the other girls'' -- largely conforms to traditional gender norms as a weak Outcast archetype. In each instance, our use of archetypometrics offers a detailed, character-based lens for assessing female protagonists in contemporary vampire narratives, with clear potential for broader application across other storytelling forms.
Shift & Drift: A Zero-Shot Benchmark for Generalizable and Robust Autonomous Driving Motion Planning
arXiv:2607.07844v1 Announce Type: new Abstract: While closed-loop motion planners trained on large-scale, object-level datasets, e.g., nuPlan, demonstrate strong in-distribution (ID) performance, their generalization to novel urban topologies and recovery mechanisms following execution perturbations remain under-explored. To address this, we present Shift & Drift, a novel dual-track benchmark designed to rigorously stress-test motion planners across two critical axes of distribution shift: (1) The Semantic Shift Track leverages a novel conversion pipeline that transforms the aerial, DeepScenario Open 3D dataset into the nuPlan simulation framework. This enables zero-shot evaluation of planners trained on North American and Singaporean data against 1,182 scenarios spanning four German cities and the US city of San Francisco featuring dense pedestrian-cyclist interactions. (2) The State-Distribution Drift Track injects stochastic perturbations into the ego vehicle's dynamics to quantify robustness against compounding execution errors. Based on this, we systematically evaluate the failure modes of diverse planning paradigms under semantic and state-distribution shifts. While imitation learning methods achieve high scores in ID benchmarks, they exhibit significant failures under semantic shift, particularly in pedestrian-dense environments, and suffer from persistent drift when subjected to temporally correlated actuation noise. In contrast, the evaluated reinforcement-learning-based planner demonstrates more graceful degradation, maintaining higher safety and progress metrics across both tracks. Our findings reveal an empirical trade-off between imitation fidelity and closed-loop resilience, providing the community with a rigorous benchmark to evaluate progress toward reliable deployment.
SAM-MT: Real-Time Interactive Multi-Target Video Segmentation
arXiv:2607.08688v1 Announce Type: new Abstract: Modern Video Object Segmentation (VOS) involves tracking and segmenting user-specified targets. While recent approaches have achieved remarkable performance in single-target scenarios, extending them to multi-target settings typically involves replicating the single-target processing for each individual object, resulting in reduced frame rates (FPS) with unbounded latency as target count increases. Built upon Segment Anything 2 (SAM2), we propose SAM-MT, which addresses this by transforming the model into an interactive framework for real-time Multi-Target video segmentation. SAM-MT uses explicit queries to represent different individual targets, in parallel with a shared representation for global context. It employs decoupled masked attention to keep individual identities distinct from cross-target interference, and sparse memory for stable temporal evolution, along with specialized strategies for occlusion handling and overlap prevention. SAM-MT successfully decouples latency from the number of targets, achieving real-time speed on par with single-target baselines (>36 FPS for 10 targets) while maintaining SAM2's robust video segmentation performance.
Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima
arXiv:2607.08380v1 Announce Type: new Abstract: An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian. Classical analyses typically require the step size to be uniformly smaller than twice the reciprocal of the sharpness, but this condition is frequently violated in the training of deep neural networks. Recent work bridges this gap in the setting of overparametrised least-squares with a \emph{single scalar output}, providing a normal form for large-step GD in a neighbourhood of an \emph{isolated} flat minimum and establishing three corresponding convergence results. In this paper, we extend this theory in two directions: (1) to overparametrised least-squares with \emph{vector-valued outputs} (including regression with arbitrarily many observations), and (2) to a neighbourhood of a \emph{manifold} of flat minima (which we show is essential for applications such as matrix factorisation). We generalise both the normal form and all three convergence theorems of \cite{macdonaldeos} to this broader setting, overcoming several technical challenges, including the solution of a singular partial differential equation via a novel method that may be of independent interest. We further show that our framework applies to deep matrix factorisation under mild assumptions, yielding several new structural results. In particular, we prove that the set of flat minima forms a fibre bundle over a product of spheres, and that the sharpness is Morse-Bott along this manifold.
Analysis of collisional and facility effects in a magnetic nozzle plasma expansion
arXiv:2607.07861v1 Announce Type: new Abstract: An axisymmetric, quasineutral three-fluid model is proposed to study the plasma expansion in a magnetic nozzle under the presence of neutrals coming either from the plasma source or as an homogeneous background. As a difference with other models, electron cooling in the plume is achieved by treating the electron energy flux as mainly convective and without the need to postulate any anomalous resistivity. Solutions are presented for the electron high-magnetization limit, in which the electron main magnitudes can be integrated along magnetic lines. Ionization, elastic and charge-exchange collisions with neutrals do not change the main qualitative features of the plasma expansion, known from previous collisionless models. Ionization enhances the plasma flow in the nozzle, and leads to additional electron cooling, which decreases the electric potential fall along the nozzle. The efficiency of the nozzle is quantified in terms of the gain of magnetic thrust and the plume divergence angle. Two types of boundary conditions are discussed for the electron flow: local current ambipolarity conditions at the nozzle throat and global current-free conditions at the outer boundary (i.e., metallic vacuum chamber walls). These last ones are shown to be physically more reliable: they introduce the influence of the chamber walls on the plasma expansion by shaping the ambipolar electric field; they permit the extrapolation to undisturbed free space conditions; and they approximate better experimental trends with the background pressure.