Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA
arXiv:2606.03728v1 Announce Type: new Abstract: Retrieval-augmented generation systems for legal question answering typically retrieve passages based on semantic similarity and provide them to a language model, which then generates cited answers. Prior work assumes that highly ranked passages are most likely to be usefully cited by the model. Perturbation-based attribution methods, such as C-LIME, have been used exclusively for post-hoc explanation. However, on the AQuAECHR benchmark, semantic similarity does not correlate with passage attribution. Within a retriever's candidate pool, similarity-based ranking performs worse than random selection at surfacing gold citation paragraphs. To address this limitation, a lightweight cross-encoder is trained on continuous perturbation-based attribution scores to re-rank passages prior to generation. This approach is evaluated on the AQuAECHR benchmark, using two language models and five-fold cross-validation. The re-ranker substantially improves citation faithfulness and alignment with gold expert answers. Notably, two re-rankers trained independently on different models converge beyond their raw attribution agreement. This finding indicates that the cross-encoder reduces model-specific noise and produces a shared relevance signal that partially transfers across models, although same-model re-ranking remains more effective. These results demonstrate that perturbation-based attribution provides a practical, model-agnostic training signal for citation-aware retrieval.
Wave-mean decomposition of scale-dependent kinetic energy from surface drifters
arXiv:2606.03744v1 Announce Type: new Abstract: Separating waves and mean flows is a fundamental challenge in ocean dynamics. Lagrangian filtering of passive-tracer time series into high-frequency wave and low-frequency mean-flow components provides a practical route, as the relevant time scales are often cleanly split in the Lagrangian frame. Here we show that Lagrangian filtering can be applied to surface drifter observations, providing a powerful approach to quantify wave and mean-flow contributions to surface kinetic energy statistics. A key methodological choice is to implement the filtering in a generalized Lagrangian mean (GLM) framework, attributing filtered velocities to mean rather than particle trajectories; this produces more physically interpretable diagnostics. Using Gulf of Mexico drifter data, we compute second-order velocity structure functions (SF2s) for waves and mean flow components across spatial scales. With these filtered SF2s as a benchmark, we illustrate that Helmholtz decomposition of unfiltered SF2s alone should not be interpreted as a dynamical wave-mean decomposition. Applying Helmholtz decomposition to the filtered SF2s further illuminates seasonal dynamics. Mean-flow surface kinetic energy is rotationally dominated at scales larger than O(1) km, while at and below O(1) km, divergent and rotational contributions are approximately equipartitioned in both summer and winter, suggesting low-frequency divergent motions and possible associated vertical exchange. Winter mean flows are more active than summer mean flows over 500 m-10 km. Super-inertial motions are broadly consistent with linear waves. In winter, wave kinetic energy is concentrated at smaller spatial scales than in summer, possibly reflecting enhanced downscale transfer by stronger submesoscale mean flows.
A Voxel-Based Quantum Computing Method (VBQC) for Solid Mechanics Problem
arXiv:2606.03515v1 Announce Type: new Abstract: Quantum computing presents a promising method to overcome the efficiency and memory constraints in large-scale mechanical problems, with numerous successful applications demonstrated in fluid mechanics. However, solid mechanics problems usually require irregular grids for spatial discretization, due to the Lagrange formulations and complex boundaries, which makes the quantum simulation of the system matrix, e.g., the mass or stiffness matrix which is often referred to as the Hamiltonian in quantum computing, difficult to be effectively conducted. This study proposes a voxel-based quantum computing method (VBQC) for the quantum simulation of Hamiltonians in solid mechanics. VBQC applies voxel grids to discretize the spatial domain, thereby enabling the system matrix to exhibit the tridiagonal fractal property. Based on this property, the system matrix can be decomposed into three groups of fundamental matrices, $\mathbf{k}_{n}$, $\mathbf{c}_{n}$, and $\mathbf{q}_{n}$. This decomposition process is referred to as the KCQ decomposition. By integrating the KCQ decomposition with the quantum Fourier transform and the quantum multiplexer, VBQC enables efficient quantum simulation of Hamiltonians in solid mechanics. Three specific solid problems with different dimensions and numbers of variables are applied to preliminarily verify the correctness of the proposed VBQC for solid mechanics problems.
Emergent Ordinal Geometry in Transformers Trained on Local Comparisons
arXiv:2606.01269v2 Announce Type: replace Abstract: Transitive inference is the challenge of inferring that A < C from knowing only adjacent relations (A < B, B < C). It is solved by humans and animals not through logical chaining but via an analogue mental number line, whose signature is the symbolic distance effect: distant comparisons are easier than nearby ones. We ask whether Transformers acquire the same primitive, training small models exclusively on adjacent comparisons from a hidden total order and evaluating generalization to unseen distant pairs. We find that out-of-distribution generalization emerges alongside a striking geometric reorganization: entity embeddings collapse onto a one-dimensional manifold whose principal axis recovers the hidden rank order with near-perfect fidelity, and this structure is sensitive to optimization in ways that produce grokking-like transient dynamics. Critically, even when accuracy is at ceiling, decision confidence and geometric separation both scale monotonically with rank distance, directly mirroring the symbolic distance effect observed across decades of behavioural experiments on humans, primates, and rodents. We further show the same rank-aligned geometry in a pretrained large language model, where it tracks the topology of each ordinal relation: linear for sizes and digits, cyclic for months. These results ground a 50-year-old behavioural regularity in the geometry of learned representations, offering a mechanistic account of transitive inference that bridges cognitive science and modern neural networks.
Finite-Temperature de Bruijn Identities: Fisher Information as the Spectral Gap of Blahut--Arimoto Dynamics
arXiv:2606.03813v1 Announce Type: new Abstract: We uncover a finite-temperature extension of de Bruijn's identity -- the classical relation $\frac{d}{dt}h(X+\sqrt{t}Z)=\frac{1}{2}J(X)$ connecting differential entropy and Fisher information. Our framework is the spectral theory of Blahut--Arimoto (BA) dynamics, recently developed by Wang~\cite{Wang2026} for the analysis of rate-distortion optimization. The central observation is elementary yet profound: for Gaussian sources, the spectral gap $\lam$ of the BA relaxation kernel $\G$ satisfies $\lam = 1/(2\beta\sigma^2)$~\cite{Wang2026}, while the Fisher information of the source is $J = 1/\sigma^2$. Hence \[ {\lam = \frac{J}{2\beta}} \] for all inverse temperatures $\beta > 1/(2\sigma^2)$. This identifies the BA spectral gap as a \emph{finite-temperature regularization of Fisher information}. From this observation we derive an exact finite-temperature de Bruijn identity: \[ \frac{\partial F_\beta}{\partial \sigma^2} = \frac{1}{2\beta\sigma^2} = \lam, \] where $F_\beta$ is the BA free energy. This identity holds for all finite $\beta$ without any limit procedure. The classical de Bruijn identity follows as the exact consequence $\beta\,\partial F_\beta/\partial\sigma^2 = J/2$. The significance is structural: classical de Bruijn is not an isolated fact about Gaussian convolutions, but the $\beta\to\infty$ shadow of a one-parameter family of exact identities living in the spectral geometry of rate-distortion optimization. We discuss implications for the entropy power inequality, the $\chi^2$-dissipation structure of BA dynamics, and the geometric unification of information inequalities.
Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation?
arXiv:2606.03837v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) and probing enable adaptation of foundation models using only a small number of trainable parameters, making it attractive for video understanding where annotation and computation are expensive. However, video PEFT has focused on adapting image-pretrained models, while standard PEFT methods can also be applied to video representations. These settings are rarely compared and both confine temporal reasoning to a single component of the model, leaving open how temporal context should be distributed across backbone, PEFT and probe. In this work we provide a systematic study of model adaptation strategies for video understanding. We evaluate methods across appearance-focused, motion-focused and spatially dense settings, with a particular focus on scenarios with limited data where parameter-efficiency is most beneficial. Our results provide new insights into PEFT and probing across settings and demonstrate the importance of temporal context allocation for effective video adaptation
Uncovering Turbulent Dynamics in Stenotic Flows from 4D-flow MRI Measurements via Resolvent Analysis and Data Assimilation
arXiv:2606.03838v1 Announce Type: new Abstract: This study presents a hybrid experimental and computational framework that couples in vitro 4D phase-contrast magnetic resonance imaging (4D-flow MRI) measurements with data assimilation and linear modeling to characterize the flow linear amplification mechanisms. We manufacture an idealized stenosis phantom with a cosine-shaped contraction and acquire three-dimensional (3D) mean velocity measurements at Reynolds number 3960 using 4D-flow MRI. To overcome the inherent displacement artifact, we perform data assimilation via a two-step optimization strategy using physics-informed neural network (PINN). This approach first corrects measurement artifacts before extracting the unknown mean pressure and eddy viscosity fields. The RANS-compatible mean flow then serves as the base state for global linear stability analysis (LSA) and resolvent analysis. The global LSA reveals stationary eigenmodes located in the recirculation bubble that exhibit a positive growth rate for azimuthal wavenumbers m=2 and m=3. The forced dynamics of this eigenmode dominates the low-frequency dynamics. Resolvent analysis identifies a broadband pseudo-resonance associated with the convective instability of the separated shear-layer, with maximal amplification for m=0. This methodology demonstrates how integrating sparse experimental MRI data with physics-based modeling enables the identification of mean fields and coherent structures. By leveraging the capabilities of 4D-flow MRI to non-invasively measure 3D velocity fields without requiring physical or optical access, this approach is a first step in the application of linear analysis to cardiovascular flows.
Clustered Self-Assessment: A Simple yet Effective Method for Uncertainty Quantification in Large Language Models
arXiv:2606.03846v1 Announce Type: new Abstract: Large language models (LLMs) demonstrate remarkable performance across diverse tasks, but they often generate responses that appear plausible while being factually incorrect. This problem is compounded by the lack of explicit uncertainty estimates, which makes it difficult for users to judge the reliability of model outputs. Existing uncertainty quantification methods typically rely on indirect signals, such as entropy across sampled generations. These signals can be difficult to interpret and do not fully leverage the model's ability to assess its own uncertainty. We propose a simple yet effective self-assessment method for uncertainty quantification in LLMs. Our approach groups sampled generations into semantically distinct clusters, converts them into answer options in a structured multiple-choice question, and uses the probability assigned by the LLM to each option as a confidence estimate. Experiments across multiple models and datasets show that our method consistently outperforms baseline approaches. Notably, it achieves competitive performance with as few as two additional samples, demonstrating both its effectiveness and efficiency.
Approximation by short exponential sums with geometric error decay based on Gauss quadrature
arXiv:2606.03855v1 Announce Type: new Abstract: We present new short exponential sum approximations of length $N$ for $f_1(x)=\frac{1}{a+x}$ with $a>0$ on $[0, \infty)$ and for $f_2(x)= {\mathrm e}^{-x^2/2\sigma}$ with $\sigma>0$ on ${\mathbb R}$ with geometric error decay ${\rho}^{-2N}$ for user-defined $N \ge 2$ and $\rho >1$. The approximations are built over consecutive intervals $[b_j, \, b_{j+1}) \subset [0, \infty)$, $j \in {\mathbb N}_{0}$, with interval lengths that depend on $\rho$ and grow exponentially for $f_1$ and are equidistant for $f_2$. All parameters determining the exponential sum approximations on $[b_j, \, b_{j+1})$ are easily computed from the initial parameters on $[b_0, \, b_{1})$, ensuring numerical stability. Our method is based on Gauss-Laguerre and Gauss-Hermite quadrature, respectively, applied to suitable parametric integral representations of $f_1$ and $f_2$. This technique ensures consistent relative errors across all intervals. Using the obtained exponential sum approximations, we achieve highly accurate approximations of $\log(x)$ on $[1,\infty)$ and of the error function $\mathrm{erf}(x)$ with predictable geometric error decay. Numerical examples for $N=8$ and $N=10$ clearly illustrate the theoretical error estimates.
PHASOR: Phase-Anchored Universal Action Representations for Humanoid Embodiments
arXiv:2606.01851v2 Announce Type: replace Abstract: Learning a good action embedding space is fundamental to scalable robot policy learning, yet existing methods treat action latents as task-specific intermediates rather than first-class representations. The resulting latents are unstructured, embodiment-specific, and weakly tied to motion semantics, limiting interpretability, controllability, and transferability across robots. We position the action embedding space itself as a first-class design target, with downstream policy quality emerging from representation quality. Exploiting motion's intrinsic periodicity, we factorize it into a phase manifold that captures cyclic structure via FFT-parametric coefficients, together with a pose branch that conditions the manifold on non-periodic configuration detail. Combined with motion-semantic distillation, this factorized structure yields a cross-embodiment motion manifold that is interpretable and embodiment-agnostic by design. Anchoring multiple humanoid robots to a shared human-pretrained manifold then produces a unified action embedding space across diverse platforms, achieving strong cross-embodiment retrieval and consistent gains on downstream robot tasks.
Attend to Anything: Foundation Model for Unified Human Attention Modeling
arXiv:2606.03540v1 Announce Type: new Abstract: Existing human attention (saliency) modeling methods persist as highly fragmented across modalities, scenes, and task formulations. Consequently, even with increasing model capacity and data scale, current models predominantly remain scene-dependent and task-specific, failing to practically generalize in real-world applications. To address the fundamental limitations, we present the Attend to Anything Model (AAM), a multi-modal foundation model that unifies attention modeling across various image, video, and audio-visual tasks and scenes. AAM reformulates attention as a cognitive entailment relationship organized in a general-to-specific hierarchy, implemented through language prompts with hierarchical embeddings in hyperbolic space. Furthermore, to unify static image and dynamic video attention, we adopt a fluid-dynamics perspective, formulating video-frame attention as a diffusive temporal evolution governed by the Fokker--Planck equation. Extensive experiments on 16 benchmarks demonstrate that AAM consistently outperforms state-of-the-art methods by an average of 6\% across various scenarios, while achieving approximately a 4$\times$ speedup in video inference. Overall, these results demonstrate that AAM provides a principled foundation for future research on attention and saliency-related tasks. The dataset and code will be available at https://github.com/wz-zhao/Attend-to-Anything.
D2MDT: Department-aware Multidisciplinary Team Consultation with Deliberation for Efficient Clinical Prediction
arXiv:2606.03543v1 Announce Type: new Abstract: Electronic health records (EHRs) are central to clinical prediction, but existing methods either rely on correlation-driven deep models or use single large language models (LLMs), making it difficult to support multidisciplinary clinical reasoning. Recent multi-agent systems (MAS) provide a promising alternative, yet current EHR-grounded MAS methods still suffer from weak evidence differentiation across agents and redundant multi-round interaction. We propose D2MDT, a Department-aware MultiDisciplinary Team Consultation with Deliberation for Efficient clinical prediction. D2MDT first constructs structured EHR evidence and consultation-ready semantic evidence for multi-agent consultation. It then assigns patient-specific department perspectives to doctor agents and retrieves complementary evidence for collaborative consultation. To improve efficiency, D2MDT further introduces residual deliberation, which updates only unresolved consensus rather than replaying the full discussion history. Finally, D2MDT fuses the refined consensus report with structured EHR representations for prediction. Experiments on mortality prediction show that D2MDT improves both predictive performance and consultation efficiency. We release the code online to ease the reproducibility of this paper.
From Prompt to Service: An SLM-Based Agent Orchestration Gateway for AI-Driven Virtual Worlds
arXiv:2606.03557v1 Announce Type: new Abstract: As generative AI capabilities expand, AI-driven virtual worlds face a growing architectural challenge. Users interact through in-world interfaces in multimodal ways, yet their requests demand fundamentally different AI backend models and computational resources. Embedding these capabilities directly into virtual world systems reduces extensibility, complicates maintenance, and limits the ability to coordinate services distributed across edge and cloud infrastructure. This paper presents an SLM-based Agent Orchestration Gateway, a lightweight runtime coordination mechanism that decouples a virtual world client from heterogeneous AI backends through intent-driven service routing. An edge-deployed SLM classifies the semantic intent of each user prompt, a configurable service registry validates and resolves the routing decision, and the selected backend is invoked transparently, enabling new AI capabilities to be introduced in the virtual world without modifying the client application. The gateway is implemented and evaluated within the InterwovenXR virtual museum testbed. The evaluation shows that compact SLMs can serve as reliable intent routers on edge hardware, and that task-specific fine-tuning can transform sub-billion-parameter models into practical, low-latency routers. A layered configuration pairing a fine-tuned sub billion-parameter model as router with a larger SLM for conversational response generation is shown to be deployable on mid-range edge hardware and more efficient than delegating both responsibilities to a single model. The findings show that SLMs can support practical AI service orchestration in virtual worlds and the work contributes an evaluated architecture for scalable, extensible, and edge-supported AI interaction, enabling virtual agents become access points to distributed generative AI services.
Magnet-Free Proton Therapy with 4D Pencil Beam Delivery Optimisation
arXiv:2606.03562v1 Announce Type: new Abstract: Objective. Motion management is a critical challenge in proton therapy for mobile tumours. This study aims to develop and evaluate a novel four-dimensional (4D) pencil beam delivery strategy that incorporates respiratory motion into a dynamic treatment plan to improve dose conformity and treatment efficiency. Approach. To assess this 4D pencil beam delivery strategy, a mobile phantom was used. The generated 4D treatment plans were assessed with various scanner configurations, including gantry-free and magnet-free scanner heads. For each setup, the treatment time, dose conformity, and robustness against irregular breathing patterns were quantified. The influence of scanner head design and patient-specific motion irregularities on overall plan quality was evaluated. Main Results. The 4D planning tool generated treatment plans that achieved clinically acceptable dose distributions across all configurations. In magnet-free configurations, static beam operation substantially increased treatment time and reduced dose conformity. In contrast, configurations using a single scanner magnet, without a gantry, maintained acceptable conformity within practical treatment times. Significance. The proposed 4D delivery strategy demonstrates feasibility for treating mobile targets with simplified, gantry-free and magnet-free scanner designs. Further improvements could be achieved by synchronising the patient's breathing with 4D delivery, which may enhance dose accuracy during irregular or interrupted breathing. By reducing system complexity while preserving dosimetric performance, this approach offers a pathway toward more accessible and cost-effective proton beam therapy for motion-affected tumours.
SparseStreet: Sparse Gaussian Splatting for Real-Time Street Scene Simulation
arXiv:2606.03909v1 Announce Type: new Abstract: While 3D Gaussian Splatting has shown promising results in street scene reconstruction, existing methods require massive numbers of Gaussian primitives to capture fine details, leading to prohibitive storage costs and slow rendering speeds. We observe that dynamic objects (e.g., vehicles and pedestrians) demand high-fidelity representations to maintain temporal consistency, while static background regions often contain substantial redundancy. Motivated by this, we propose SparseStreet, a general compression framework specifically designed for street scenes. First, we introduce a node-based learnable pruning strategy that systematically removes low-contributing Gaussian primitives while preserving visually critical regions. Second, after the scene representation stabilizes, we apply background compression, further reducing redundancy in static regions. Our method effectively preserves the geometry and appearance of dynamic objects while significantly reducing the total number of Gaussian primitives. Extensive experiments on the Waymo and nuScenes demonstrate that SparseStreet achieves up to 80% compression ratio with minimal quality degradation, enabling resource-efficient, high-fidelity dynamic scene reconstruction. Project website: https://sparsestreet.github.io/.
Learned Non-Maximum Suppression for 3D Object Detection
arXiv:2606.03568v1 Announce Type: new Abstract: Post-processing is a critical stage in LiDAR-based 3D object detection, where dense and overlapping proposals must be filtered for compact and reliable perception. This work introduces two learned filtering modules that replace heuristic non-maximum suppression (NMS) by leveraging relations among detections. D2D-Rescore employs transformer-based detection-to-detection (D2D) attention, while GossipNet3D adapts the 2D GossipNet concept to 3D through localized message passing in bird's-eye view. A metric-aware matching strategy aligned with the nuScenes evaluation protocol ensures consistent training and validation behavior, improving overall detection performance. Both approaches improve mean average precision (mAP), nuScenes detection score (NDS), and true positive quality compared to CircleNMS, particularly for small and infrequent classes, while adding minimal computational overhead. These results demonstrate that learned, detection-level filtering can enhance 3D detector reliability without modifying the base network, offering a principled alternative to heuristic suppression. Code is available at https://github.com/rst-tu-dortmund/learned-3d-nms .
When Attention Collapses: Stage-Aware Visual Token Pruning from Structure to Semantics
arXiv:2606.03569v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have demonstrated remarkable capabilities but suffer from significant computational overhead during inference. While visual token pruning offers a promising solution, existing methods predominantly rely on initial attention scores. This single-metric paradigm presents a critical flaw: high attention scores inherently collapse onto semantically similar regions, thereby severely reducing feature diversity and discarding vital contextual details. To address this, we introduce Structure-to-Semantics (STS), a novel two-stage visual token pruning framework that explicitly decouples the pruning process. The first stage employs a repulsion-based sampling mechanism to maximize spatial and structural diversity. The second stage leverages instruction-aware cross-attention to precisely filter out prompt-irrelevant tokens. This two-stage synergy constitutes the core of STS, first ensuring geometric coverage and then refining the retained tokens according to semantic relevance. Extensive evaluations demonstrate that STS mitigates the redundancy caused by attention-based selection, improving both structural diversity and fine-grained task alignment of the preserved visual tokens.
Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching
arXiv:2606.03577v1 Announce Type: new Abstract: Wide-baseline matching (WBM) requires integrating geometric understanding, viewpoint changes, fine-grained perception, and occlusion reasoning, making it a challenging testbed for spatial reasoning in multimodal large language models (MLLMs) deployed in physical environments. However, current MLLMs lack systematic evaluation and training frameworks for these capabilities. We introduce ReasonMatch-Bench, a benchmark stratified by viewpoint displacement and matching granularity across indoor, outdoor, and object-centric scenarios, and show that current MLLMs still struggle with fine-grained wide-baseline correspondence: on a difficult 90-sample subset, human annotators achieve 84.0 F1, while the best existing baseline reaches 37.2. To bridge this gap, we build a scalable data-generation pipeline that automatically extracts wide-baseline view pairs from large-scale video-3D corpora, including RGB-D videos and SfM reconstructions, yielding diverse and verifiable supervision. We further propose Dynamic Correspondence Reinforcement Learning (DCRL), which combines Image-Level Viewpoint Progression and Point-Level Correspondence Curriculum to improve WBM training through verifiable rewards without explicit CoT supervision. Extensive experiments show that DCRL substantially improves ReasonMatch-Bench and transfers to related spatial benchmarks, while maintaining general visual understanding performance with modest gains on several benchmarks.
Diffusing in the Right Space: A Systematic Study of Latent Diffusability
arXiv:2606.03578v1 Announce Type: new Abstract: Latent diffusion models leverage visual tokenizers to compress images into latent spaces for efficient generative modeling. However, better reconstruction quality of a tokenizer does not necessarily translate into better generation quality, suggesting that latent representations should be evaluated not only by fidelity but also by their diffusability. Recent studies have proposed diverse explanations for diffusion-friendly latent spaces, including semantic separability, affine equivariance, distribution uniformity, spatial structure, spectral smoothness, and manifold continuity. Yet these properties are often validated on a limited set of tokenizers, leaving it unclear which factors are most predictive of downstream generation quality and whether such conclusions hold beyond the specific settings in which they are introduced. In this work, we conduct a systematic study of latent diffusability by training a large collection of tokenizers with diverse regularization strategies, architectures, and latent configurations, and evaluating them with multiple downstream diffusion backbones. Our analysis identifies several latent properties that consistently correlate with generation quality and exhibit strong generalization across experimental settings. Beyond existing metrics, we introduce Velocity Irreducible Variance (VIV), a measure of velocity ambiguity induced by trajectory crossings. Extensive experiments show that VIV is one of the most stable predictors of generation quality.
Beyond Gradient Descent: Adam for Analog Ising Machines
arXiv:2606.03917v1 Announce Type: new Abstract: As Moore's law reaches its limits, Ising machines offer a promising alternative computing approach for difficult optimization problems. However, many analog, time-continuous Ising machines rely on gradient-descent-like dynamics to find solutions, which can limit speed and robustness. We investigate whether momentum and Adam optimization can improve these systems. Since these optimizers are traditionally formulated in discrete time, we derive continuous-time versions suitable for analog, time-continuous Ising-machine dynamics. On Max-Cut benchmarks, we find that Adam-based dynamics substantially reduce time-to-target and improve solution quality compared with gradient-descent- and momentum-based dynamics. We further introduce a first-order continuous-time approximation of Adam that is intended as a simpler starting point for future physical implementations and while performing better than the full Adam formulation in a continuous-time setting. We also study a purely algorithmic discrete-time setting, where the performance gap is reduced on easier problem instances, while the Adam-based update rule performs best on harder weighted problem instances. These results identify continuous-time Adam dynamics as a powerful design principle for analog Ising machines.
Correcting Neural Operator Spectral Bias via Diffusion Posterior Sampling with Sparse Observations
arXiv:2606.03936v1 Announce Type: new Abstract: Neural operator surrogates (NO) approximate PDE solutions orders of magnitude faster than numerical solvers, but suffer from spectral bias: high-frequency content is systematically attenuated, limiting reliability where fine-scale structure matters. Sparse sensor measurements of the field are often available too, offering pointwise accuracy without spectral distortion but covering only a small fraction of the domain. We address this by treating NO predictions as auxiliary observations in a diffusion posterior sampling framework. Our method, FreqNO-DPS (https://github.com/niccoloperrone/FreqNO-DPS), combines an unconditional score-based diffusion prior, trained on high-fidelity simulations, with diffusion posterior sampling (DPS) conditioned on sparse observations and guided by a frozen neural operator. Naive integration reintroduces the surrogate's spectral bias; we resolve this with a closed-form, spectrally shaped guidance score that weights the surrogate by its frequency-dependent accuracy and needs no denoiser backpropagation. A distribution-free analysis bounds the approximation error across the frequency-diffusion-time plane and shows the guidance's frequency dependence is preserved regardless of distributional assumptions. On 3D elastic wavefield prediction at 5% and 2% sensor coverage, the method reaches near-zero spectral bias across all bands, where both the surrogate and sensor-only DPS show systematic high-frequency attenuation. Isotropic guidance, the natural baseline, improves pointwise accuracy but carries the bias into the posterior nearly intact, confirming that frequency-dependent calibration is essential, not merely beneficial. The framework needs only paired surrogate/reference data and exploits no problem-specific structure beyond the residual's approximate spectral diagonality, verifiable for new surrogates via the coherence diagnostic we provide.
Training a Predictive Coding Network on ImageNet using Equilibrium Propagation
arXiv:2606.03584v1 Announce Type: new Abstract: Equilibrium Propagation (EP) is a physics-based training framework that has primarily been employed in energy-based models, including continuous Hopfield networks, nonlinear resistive networks and coupled phase oscillators. However, EP's practical applications have so far remained limited to relatively small-scale problems. Predictive coding networks (PCNs), another class of energy-based models rooted in computational neuroscience, are typically trained with a specialized algorithm and have likewise not yet been demonstrated at large scale. In this work, we develop an EP-based training method for PCNs which combines the centered variant of EP with a novel equilibration scheme for PCNs. Using this approach, we train a 10-layer convolutional PCN (VGG10) on full-size ImageNet, achieving 13.23\% test error rate on the top-5 classification task, close to the 12.2\% backpropagation baseline. To our knowledge, this is the first demonstration of both PCNs and EP-based training at ImageNet scale. These results significantly extend the scalability of both approaches and suggest that the primary challenges in scaling EP in other physical systems may come more from the computational properties of these systems than from inherent limitations of the EP framework.
A Comparison of Multirate Co-Simulation Techniques for Field-Circuit Coupled Problems
arXiv:2606.03594v1 Announce Type: new Abstract: This paper compares three different multirate splitting approaches for the application on field-circuit coupled magnetoquasistatic simulations. For these methods, again three different variants for exchanging values between the field and circuit are tested, namely voltages, currents and flux correction terms. All scenarios are applied on two different benchmark problems, i.e. a coil inductor and transformer model coupled to different circuits. The convergence behavior of different time steppers (Implicit Euler and Trapezoidal Rule) is determined for all possible settings, and guidelines for practical applications are derived.
An Efficient Parity-Blocked Method for Band-Structure Computation of 3D Anisotropic Phononic Crystals
arXiv:2606.03599v1 Announce Type: new Abstract: Band-structure calculations for three-dimensional anisotropic phononic crystals require the repeated solution of large elastic generalized eigenvalue problems along Bloch paths. In standard staggered-grid discretizations, anisotropic coupling may involve derivative components located at incompatible grid positions, so additional interpolation or averaging closures are often introduced. This paper proposes a parity-blocked rotated staggered discretization based on four Bloch-periodic body-diagonal differences. The directional derivatives are reconstructed from these diagonal differences, leading to a Hermitian $B_hC_hB_h^H$ generalized eigenvalue formulation that incorporates anisotropic derivative coupling without separate interpolation closures. On even grids, when the stiffness and mass matrices are nodewise local multiplication matrices, the body-diagonal shifts preserve two independent parity invariants. The discrete velocity space is then decomposed exactly into four mutually independent block subspaces, and the full discrete spectrum can be recovered by solving the four smaller eigenvalue problems and merging their spectra. The full and block formulations are further organized in a unified Fourier SVD framework, which supports $\Gamma$-point zero-mode treatment, shift-invert Krylov iteration, inner PCG solves, and GPU matrix-vector products. Numerical experiments for a three-dimensional two-phase anisotropic phononic crystal show that the block implementation preserves the full-space spectrum while substantially reducing the wall-clock time. The results demonstrate that the proposed method provides a structured and efficient solver for large-scale band-structure computations of three-dimensional anisotropic phononic crystals.
A Sparse Bayesian Learning Algorithm for Estimation of Interaction Kernels in Motsch-Tadmor Model
arXiv:2505.07068v2 Announce Type: replace-cross Abstract: In this paper, we investigate the data-driven identification of asymmetric interaction kernels in the Motsch-Tadmor model based on observed trajectory data. The model under consideration is governed by a class of semilinear evolution equations, where the interaction kernel defines a normalized, state-dependent Laplacian operator that governs collective dynamics. To address the resulting nonlinear inverse problem, we propose a variational framework that reformulates kernel identification using the implicit form of the governing equations, reducing it to a subspace identification problem. We establish an identifiability result that characterizes conditions under which the interaction kernel can be uniquely recovered up to scale. To solve the inverse problem robustly, we develop a sparse Bayesian learning algorithm that incorporates informative priors for regularization, quantifies uncertainty, and enables principled model selection. Extensive numerical experiments on representative interacting particle systems demonstrate the accuracy, robustness, and interpretability of the proposed framework across a range of noise levels and data regimes.