Forskningsradar

Science Journals

Peer-reviewade publikationer — 53080 artiklar

Emergent heavy-tailed distributions from a Markovian random walk
arXiv:2605.22933v1 Announce Type: cross Abstract: The emergence of heavy-tailed statistics in complex systems is conventionally attributed to non-local stochastic jumps or non-Markovian memory. Here, we present a one-dimensional random walk where power-law behaviors arise instead from a strictly local, discrete-time Markovian mechanism. The step length is governed by a deterministic function of the walker's position, establishing a positive feedback loop that induces strong effective correlations along the trajectories. Through analytical derivations in the continuum limit and extensive numerical simulations, we show that this rule yields a robust, non-Gaussian stationary state. The exact analytical solution is obtained in the closed form of a symmetric, Lorentz-like distribution, $\rho_{\text{st}}(x) \propto (|x|/l+r\Delta x)^{-2}$, confirming asymptotic power-law tails that decay as $|x|^{-2}$ over six decades. Furthermore, by employing the Onsager-Machlup path-integral formalism, we demonstrate that effective velocity and acceleration acquire physical meaning along the shortest fluctuation trajectories. Crucially, we find that a non-zero initial acceleration acts as the fundamental mechanism driving the walker away from the origin, ensuring both the emergence of scale-free statistics and the normalizability of the stationary distribution. This minimal pathway provides a new microscopic foundation for the widespread $-2$ power law observed across multidisciplinary complex systems.
Construction of EAQECCs with imperfect ebits
arXiv:2605.23119v1 Announce Type: cross Abstract: We generalize the stabilizer formalism for entanglement-assisted quantum error-correcting codes with noisy ebits (EAQECCs-Ne) from the binary case to the general $q$-ary case, where $q$ is a prime power. By leveraging the structure of the generalized Pauli group over $\mathbb{F}_q$ and symplectic geometry over $\mathbb{F}_q^{2n}$, we establish a unified framework for constructing EAQECCs-Ne for qudit systems. Equivalent formulations in terms of symplectic geometry over $\mathbb{F}_q$ and additive codes over $\mathbb{F}_q^{2n}$ are derived. We further construct several families of $q$-ary EAQECCs with noise ebits and analyze their performance compared to optimal stabilizer codes. Our results demonstrate that under certain noise conditions, the proposed EAQECCs-Ne can outperform standard stabilizer codes with equivalent error-correcting capability, offering a promising approach for fault-tolerant quantum computation in high-dimensional quantum systems.
Chirality-sensitive mobility and dissipation of Brownian motion on a helical landscape
arXiv:2605.23803v1 Announce Type: cross Abstract: We study the Brownian dynamics and linear response of a particle with inertia moving in a 2-dimensional helical landscape imprinted on a cylindrical surface. In the harmonic well approximation, the deterministic motion separates into free propagation along the screw direction and harmonic motion in the transverse screw-normal direction. We show that for isotropic damping this simplification survives in the Langevin description, whereas anisotropic damping along the axial and angular directions couples the stochastic dynamics and destroys separability. The resulting anisotropic model is formulated as a linear Ornstein-Uhlenbeck process in phase space with a zero mode associated with diffusion along the screw coordinate, so that in an infinite system the full phase-space dynamics does not relax to a stationary distribution. To treat transport in this setting, we construct the stationary dynamics in the stable subspace obtained after projecting out the zero mode. This leads to a linear response theory for this system and yields closed analytical expressions for stationary time-correlation functions and the dynamical mobility tensor in both the time and frequency domains. The off-diagonal elements of the mobility tensor describe cross-response between axial forcing and angular motion, and between applied torque and axial transport. Consistent with time reversal symmetry, these cross mobilities are equal and provide a direct dynamical signature of the helical geometry. In addition, a simultaneous application of driving in both the axial and angular direction reveals asymmetry in energy dissipation rate due the helical landscape.
YASPS: A Symbolic Framework for Extensible, High-Performance IPC Simulation
arXiv:2605.23088v1 Announce Type: new Abstract: Incremental Potential Contact (IPC) enables robust, contact-rich simulation by casting elasticity and contact as a single energy minimization problem, but high-performance IPC pipelines are typically built from specialized kernels and assembly logic tied to fixed energies, primitive types, and parameterizations, making extensions costly and combinatorial. We present YASPS, a GPU-oriented framework that removes this extensibility bottleneck by making structure explicit in a differentiable intermediate representation. YASPS introduces two first-class relational operators: JOIN, which composes dependent quantities across user-declared relations (e.g., element-to-vertex connectivity), and UNION, which represents alternative parameterizations within a relation (e.g., mixing free vertices with affine-body or other parameterizations without fragmenting the program). Because JOIN and UNION are part of the symbolic program, YASPS differentiates through them using dedicated rules and an efficient second-order procedure that reuses intermediate Jacobians and reduces Hessian-projection cost. From the same relational description, YASPS derives the global gradient/Hessian sparsity and block layout, enabling structure-aware block-sparse storage and compression, and JIT-compiles CUDA kernels for evaluation, derivatives, assembly, and solving. Across IPC-style examples, including layered cloth-on-bunny, mixed rigid/deformable bunnies, and a caged deformation model, YASPS supports rapid front-end extensions with minimal back-end changes while achieving competitive end-to-end performance; its Hessian compression yields near 10x faster CG iterations in our benchmarks.
Training-Free Rate-Distortion-Perception Traversal With Diffusion
arXiv:2603.04005v2 Announce Type: replace Abstract: The rate-distortion-perception (RDP) tradeoff characterizes the fundamental limits of lossy compression by jointly considering bitrate, reconstruction fidelity, and perceptual quality. While recent neural compression methods have improved perceptual performance, they typically operate at fixed points on the RDP surface, requiring retraining to target different tradeoffs. In this work, we propose a training-free framework that leverages pre-trained diffusion models to traverse the entire RDP surface. Our approach integrates a reverse channel coding (RCC) module with a novel score-scaled probability flow ODE decoder. We theoretically prove that the proposed diffusion decoder is optimal for the distortion-perception tradeoff under AWGN observations and that the overall framework with the RCC module achieves the optimal RDP function in the Gaussian case. Empirical results across multiple datasets demonstrate the framework's flexibility and effectiveness in navigating the ternary RDP tradeoff using pre-trained diffusion models. Our results establish a practical and theoretically grounded approach to adaptive, perception-aware compression.
Weak wave turbulence as a precursor to universal coarsening in a homogeneous Bose gas
arXiv:2605.22906v1 Announce Type: cross Abstract: Relaxation and condensation of an isolated low-energy Bose gas provide an ideal setting for the study of the universal features of far-from-equilibrium many-body dynamics and the emergence of long-range order. Conceptually, the emergence of such order involves two steps: the formation of local coherence, on a system-specific microscopic lengthscale, and the spreading of coherence, over lengthscales much larger than any microscopic scale. The latter is understood in terms of universal phase-ordering kinetics, or coarsening, characterized by an algebraic growth of the coherence length. Here, for a homogeneous Bose gas with tunable interactions, we show that the former also has a universal description, within the framework of weak wave turbulence (WWT). Specifically, the initial transport of particles to low momenta corresponds to an inverse turbulent cascade that is, in agreement with the WWT theory, characterized by a power-law momentum distribution, with exponent $\gamma = 2.4(1)$, and transport times ${\propto} (na)^{-2}$, where $n$ is the gas density and $a$ the $s$-wave scattering length.
LiveFigure: Generating Editable Scientific Illustration with VLM Agents
arXiv:2605.23527v1 Announce Type: new Abstract: Scientific illustrations are essential for depicting conceptual designs, methodologies, and experimental workflows in research, playing a pivotal role in communicating complex academic insights. However, creating high-quality scientific illustrations remains a labor-intensive task for human scientists. While recent generative image models have advanced prompt-based editing, the synthesis of fully editable figures remains a fundamental challenge. Valid editability involves structured transformations of graphical elements, scales, attributes, and text, rather than simple pixel-level changes. Existing models generate raster outputs that do not support manual correction or layout adjustment, limiting their utility in scientific publishing, where editable vector figures are typically required for submission. To address this challenge, we introduce LiveFigure, an agentic framework driven by VLM agents that imitates the multi-step drawing workflow of human researchers. It first plans figure blueprints by drawing inspiration from high-quality references in previous works, then generates executable scripts that produce figures via the PowerPoint interface based on skills and experience, and finally refines the outputs with targeted visual diagnostics, producing fully vectorized, editable figures that meet publication standards. Extensive experiments demonstrate that LiveFigure generates inherently editable figures, achieving 80% publication-readiness in only 17 manual edits, far surpassing the 24% rate of the strongest baseline, NanoBanana. Human preference studies further validate this advantage, with LiveFigure securing a 60% win rate against NanoBanana. Our code is available at https://github.com/tsinghua-fib-lab/LiveFigure.git.
Do Synthetic Brain MRIs Reliably Improve Tumour Classification? A StyleGAN2-ADA Class-Plane Augmentation Study on BRISC 2025
arXiv:2605.23094v1 Announce Type: cross Abstract: Generative augmentation is often proposed as a remedy for small medical-image datasets, but synthetic images are only useful when they improve downstream task performance. "Augmentation" here means synthetic supplementation: GAN-generated samples added to the real training pool, not geometric or photometric transforms of existing images. Twelve class-plane StyleGAN2-ADA generators were trained on constrained BRISC 2025 partitions to test whether their output, with or without InceptionV3 feature-space filtering, improves held-out tumour classification across three classifier families: a random forest (RF) on InceptionV3 features, a compact two-headed convolutional neural network (CNN), and MobileViTV2, a mobile hybrid convolutional-transformer. Each was evaluated at 1:1 and 1:2 real-to-synthetic ratios. An independent GPT-5.5 blind test placed gated real-versus-synthetic discrimination at 57.73% (95% CI: 54.48--60.92%) on the model-legible subset -- modestly above chance. The RF classifier did not benefit from the synthetic MRIs. The CNN showed consistent mean gains that did not survive Holm correction. MobileViTV2 showed the clearest benefit: filtered 1:1 augmentation improved tumour classification accuracy by 1.02% absolute (95% CI: 0.54--1.54%; Holm-corrected p = 0.0104). A secondary efficiency analysis found that every augmented CNN condition selected its checkpoint 42--64% earlier than baseline, while compute-matched MobileViTV2 runs reached selection after 50--67% fewer real-data epochs. Overall, augmentation utility was found to be architecture- and ratio-dependent, not guaranteed by visual fidelity alone.
Atmosphere as a steam engine
arXiv:2605.23875v1 Announce Type: new Abstract: Earth's atmosphere operates a steam cycle in which water vapor evaporates from the surface, expands, condenses, and returns as precipitation. The Clausius-Clapeyron law relates the incremental expansion work of saturated water vapor to latent heat converted at a Carnot efficiency corresponding to the temperature difference between evaporation and condensation. We generalize this relation to an atmospheric column with condensation occurring over a range of heights and derive the expansion work per mole of precipitated water. This includes the gravitational work associated with lifting moist air to the mean condensation height, the expansion work generated by condensation, and a correction for incomplete condensation. Using GPCP v3.3 precipitation and observational constraints on condensation height, we estimate the global steam-engine power as $W_v=4.4\pm0.9$ W/m2, close to an independent estimate of total atmospheric power, $W=W_P+W_K\simeq4.3\pm0.6$ W/m2, obtained from the gravitational power of precipitation and kinetic energy generation by horizontal pressure gradients diagnosed from MERRA-2. Kinetic energy generation is $W_K\simeq3.2\pm0.3$ W/m2, of which at least two thirds is generated in the lower atmosphere. The smaller upper-atmospheric contribution, dominated by temperature-related pressure gradients, is comparable to Lorenz available potential energy generation. The agreement between steam-engine and atmospheric power is linked to condensation and precipitation fallout. By removing water from the atmospheric gas phase and enabling column-mass redistribution, precipitation maintains surface pressure gradients that drive cross-isobaric flow in the frictional lower atmosphere. The steam-engine framework thus provides a thermodynamic basis for condensation-induced atmospheric dynamics and identifies a major lower-atmospheric power pathway associated with water phase transitions.
Soft Mobility Theory
arXiv:2605.23869v1 Announce Type: new Abstract: Predicting how a deformable body moves and deforms in a viscous flow underlies problems ranging from microorganism locomotion to soft microrobotics, yet existing frameworks are either problem-specific or ill-suited to inverse design. We propose the soft mobility theory: applying the principle of virtual power and the Lorentz reciprocal theorem to a hyperelastic body in a background Stokes flow yields a configuration-dependent ordinary differential equation for the generalized coordinates of the body. This soft mobility equation extends classical rigid-body mobility theory in that the mobility, elastic, body-force, and flow-coupling tensors all depend explicitly on the instantaneous deformation. We specialize the framework to assemblies of hydrodynamically interacting spheres connected by elastic springs, using the Rotne-Prager-Yamakawa approximation to compute the mobility, and validate it on canonical problems spanning rigid and flexible bodies in quiescent and shear flows. An open-source JAX implementation makes entire simulations end-to-end differentiable. This allows efficient gradient-based inverse design: as proofs of concept, we recover the asymptotic optimum of a three-sphere swimmer and design a soft gyrotactic "surfer" that exploits passive deformation to ascend faster than its rigid counterpart in a Taylor-Green flow.
Strong Teacher Not Needed? On Distillation in LLM Pretraining
arXiv:2605.23857v1 Announce Type: new Abstract: Knowledge distillation generally assumes a strong-to-weak relationship where stronger teachers yield better students. In this work, we examine this assumption about distillation in large language model pretraining. By varying architecture sizes and training token budgets, we create strong-to-weak, same-level, and weak-to-strong teacher-student relationships, and study distillation's effectiveness under each. We find that the teacher need not be strong: with proper mixing of the language modeling and knowledge distillation losses, even small and undertrained teachers improve larger students. At the same time, a stronger teacher is not always better: pushing the teacher further, through more parameters or more training tokens, can saturate or even reverse the distillation gains. We further observe that distillation improves generalization (out-of-distribution and downstream performance) more readily than in-domain fitting. Together, these results challenge the common belief that distillation pretraining always requires a strong teacher.
IntentionNav: A Benchmark for Intent-Driven Object Navigation from Implicit Human Instruction
arXiv:2605.23187v1 Announce Type: new Abstract: Existing object navigation benchmarks usually tell an embodied agent which object category to find, such as microwave or chair. Human-facing embodied AI is often asked something less direct: "I need something to warm this food" or "the room feels stuffy." The agent must infer the object that can satisfy the need, find a scene-grounded instance, and decide whether the goal has been reached. We study this setting as intent-driven object navigation and introduce IntentionNav, a diagnostic benchmark for active object search from implicit human instructions. Each episode provides a free-text intent, RGB-D observations, and pose, but withholds the target object name. IntentionNav contains 500 intents over 176 Isaac Sim scenes and 64 target categories. Each intent is rewritten in four controlled instruction styles and annotated with one of four intent modes, separating surface phrasing from semantic cue type under matched geometry. This paired design supports analysis of target inference, language robustness, neighborhood reachability, and terminal success rather than only aggregate success. We evaluated three VLMs using a fixed active-navigation agent. Models identify the intended target in 48.3 percent of episodes and enter its 2 m neighborhood in 68.7 percent, but terminate successfully in only 24.9 percent and achieve grounded 1 m success in 5.5 percent. Success is highest for event-script intents (28.7 percent) and lower for physical-state and affordance intents (19.2 percent and 18.5 percent), showing that indirect human intent remains a bottleneck for target selection, visual verification, and terminal localization in active embodied search.
Edge Assisted Multi-Camera Vehicle Tracking Framework for Real-Time and Scalable Deployment
arXiv:2511.13904v2 Announce Type: replace Abstract: Cameras are a core sensing modality in modern intelligent transportation systems (ITS), providing rich visual information on road-user activities. Multi-Camera Vehicle Tracking (MCVT) uses this data to reconstruct vehicle trajectories across camera networks, supporting applications such as traffic flow prediction and optimisation. However, most existing MCVT studies emphasise tracking accuracy while paying limited attention to real-time performance and scalability, both essential for real-world and city-scale deployment. To address this gap, we propose Edge-Assisted, Scalable and Efficient MCVT (EASE-MCVT), a distributed edge--server framework designed for real-time throughput and scalable operation. On the edge side, each camera stream is processed through object detection, single-camera tracking, geo-mapping and feature extraction, while only lightweight metadata, including vehicle locations and appearance features, is sent to the central server for cross-camera association. To improve both tracking accuracy and system efficiency, EASE-MCVT is optimised from algorithmic and system perspectives. Algorithmically, it introduces a dynamic workload scheme for tracklet-level feature extraction, a server-side re-match module to reconnect fragmented tracklets, and a self-supervised camera link model that learns spatio-temporal constraints to accelerate and stabilise cross-camera association. Systemically, it integrates production-oriented data engineering components to standardise deployment and data exchange for large-scale operation. To the best of our knowledge, EASE-MCVT is the first MCVT framework explicitly designed to address both real-time performance and scalability in a distributed edge--server setting. Experiments on the RoundaboutHD and CityFlow datasets demonstrate real-time throughput with competitive tracking accuracy, paving the way for city-wide real-time traffic management.
Reflections on the design, applications and implementations of the normative specification language eFLINT
arXiv:2511.12276v2 Announce Type: replace Abstract: Checking the compliance of software against laws, regulations and contracts is increasingly important and costly as the embedding of software into societal practices is becoming more pervasive. Moreover, the digitalised services provided by governmental organisations and companies are governed by an increasing amount of laws and regulations, requiring highly adaptable compliance practices. A potential solution is to automate compliance using software. However, automating compliance is difficult for various reasons. Legal practices involve subjective processes such as interpretation and qualification. New laws and regulations come into effect regularly and laws and regulations, as well as their interpretations, are subjected to constant revision. In addition, computational reasoning with laws requires a cross-disciplinary process involving both legal and software expertise. This paper reflects on the domain-specific language eFLINT developed to experiment with novel solutions to these challenges. Specifically, the language has been developed to experiment with the abstract syntax and semantics of a language supporting different types of reasoning for various applications. The language combines declarative and procedural elements, formalises connections between legal concepts and computational concepts, and is designed to automate compliance checks before, during and after a software system runs. The various design goals and applications areas for the language give rise to (conflicting) requirements. This paper presents and reflects on the current design of the language by recalling applications and requirements. As such, this paper reports on results and insights of an investigation that can benefit language developers within the field of automated compliance.
DELICATE: Diachronic Entity LInking using Classes And Temporal Evidence
arXiv:2511.10404v2 Announce Type: replace Abstract: In spite of the remarkable advancements in the field of Natural Language Processing, the task of Entity Linking (EL) remains challenging in the field of humanities due to complex document typologies, lack of domain-specific datasets and models, and long-tail entities, i.e., entities under-represented in Knowledge Bases (KBs). The goal of this paper is to address these issues with two main contributions. The first contribution is DELICATE, a novel neuro-symbolic method for EL on historical Italian which combines a BERT-based encoder with contextual information from Wikidata to select appropriate KB entities using temporal plausibility and entity type consistency. The second contribution is ENEIDE, a multi-domain EL corpus in historical Italian semi-automatically extracted from two annotated editions spanning from the 19th to the 20th century and including literary and political texts. Results show how DELICATE outperforms other EL models in historical Italian even if compared with larger architectures with billions of parameters. Moreover, further analyses reveal how DELICATE confidence scores and features sensitivity provide results which are more explainable and interpretable than purely neural methods.
One-Forcing: Towards Stable One-Step Autoregressive Video Generation
arXiv:2605.23458v1 Announce Type: new Abstract: Recent advances have substantially improved real-time interactive video generation in the autoregressive regime. However, most existing few-step autoregressive video generation methods, often distilled from a corresponding many-step teacher, default to a 4-step sampling configuration, which still incurs considerable latency during deployment and suffers from severe quality degradation when the number of sampling steps is further reduced, particularly in the one-step setting. Trajectory-style consistency distillation methods often produce videos with weak dynamics, while DMD-based approaches, such as Self-Forcing, tend to yield blurry frames. To address this challenge, we propose One-Forcing, a simple yet effective approach which augments the DMD objective with an auxiliary GAN loss for high-quality and efficient one-step video generation. Experiments on VBench show that One-Forcing achieves a total score of 83.76, establishing state-of-the-art performance among one-step causal video generation methods and remaining competitive with strong many-step approaches. We further demonstrate that one-step framewise autoregressive generation can be achieved stably with merely one-third of the training cost of the chunkwise model, a setting that prior methods have failed to achieve successfully.
Human Decision-Making with Persuasive and Narrative LLM Explanations
arXiv:2605.23867v1 Announce Type: new Abstract: Large language models (LLMs) have the potential to aid and improve human decision-making in classification tasks, not only by providing fairly accurate predictions, but also in their ability to generate cogent narrative explanations of those predictions. Prior work has demonstrated that people generally find AI narrative explanations to be understandable, trustworthy, and convincing for changing beliefs and opinions; however, less is known about the impact of narrative explanations on objective human decision-making performance. Here we conduct a large-scale human behavioral experiment to evaluate decision-making performance with LLM-generated narrative explanations of varying persuasiveness. We found the degree of persuasiveness, or lack thereof, for LLM-based explanations did not meaningfully impact decision accuracy over a simple AI prediction alone, in agreement with typical results with explainable AI based on feature importance. We found evidence that narratives increased reliance on AI, but both when the AI prediction was correct and incorrect. Exploratory analyses also indicated that the more persuasive narratives may have had a detrimental effect on decision response times and the ability to discriminate between a correct and incorrect AI prediction. Overall, this work indicates that including narrative explanations with AI predictions may involve tradeoffs for decision-making performance, and more work is needed to determine how and when narrative explanations impact human decision-making.
Stochastic simulation of partial discharge inception
arXiv:2511.04356v2 Announce Type: replace Abstract: We present a Monte Carlo method for simulating the inception of electric discharges in gases. The input consists of an unstructured grid containing the electrostatic field. The output of the model is the estimated probability of discharge inception per initial electron position, as well as the estimated time lag between the appearance of the initial electron and discharge inception. To obtain these quantities electron avalanches are simulated for initial electron positions throughout the whole domain, also including regions below the critical electric field. Avalanches are assumed to propagate along field lines, and they can produce additional avalanches due to photon and ion feedback. If the number of avalanches keeps increasing over time we assume that an electric discharge will eventually form. A statistical distribution for the electron avalanche size is used, which is also valid for gases with strong electron attachment. We compare this distribution against the results of particle simulations. Furthermore, we demonstrate examples of inception simulations in 2D Cartesian, 2D axisymmetric and 3D electrode geometries.
A blueprint for constructing 3-pass AKE protocols under commitment-based models
arXiv:2605.23843v1 Announce Type: new Abstract: The commitment-based AKE model provides a formal security framework for key exchange protocols that avoid long-term cryptographic material, achieving authentication through a final out-of-band verification of session-derived values. Within this model, secure KA-based and KEM-based protocols were previously constructed via a commitment-based MT compiler, yielding optimized 4-pass protocols. In this work, we show that 3-pass protocols secure under this model exist for both primitives. These protocols are constructed ad hoc, following the core ideas of the commitment-based MT authenticator, and their SK security in the unauthenticated model is proved using the same game-based techniques, achieving bounds of the same form as those previously achieved. The resulting protocols provide one-way authentication in three message exchanges.
Leveraging Foundation Models for Causal Generative Modeling
arXiv:2605.23861v1 Announce Type: new Abstract: Causal generative modeling is essential for developing reliable and transparent AI systems capable of counterfactual reasoning. While existing approaches focus on integrating causal constraints during the training of generative models, they often lack a unified framework to leverage the zero-shot reasoning capabilities of pretrained foundation models. We introduce FM-CGM, a modular framework for end-to-end visual causal reasoning using pretrained foundation models. FM-CGM formalizes the causal pipeline through three core components: a concept extractor, a concept manipulator, and a counterfactual generator. By leveraging a large reasoning model for causal inference and a text-to-image diffusion model for generation, our approach enables zero-shot causal discovery, intervention, and counterfactual generation. We then develop Causal Semantic Guidance (CSG), a cross-attention-based mechanism that ensures semantic interventions propagate to descendant concepts while preserving invariant regions. We empirically show that our approach can identify plausible causal structures and is suitable for faithful counterfactual image generation.
Anatomy-Guided Vision-Language Learning with Angular Prototype Separation for Multi-Label Video Capsule Endoscopy Classification Under Class Imbalance
arXiv:2603.17879v2 Announce Type: replace Abstract: This work presents a multi-label temporal event detection framework for video capsule endoscopy (VCE) that addresses the extreme class imbalance inherent in the Galar dataset by combining two principal contributions: an Angular Separation Loss on class prototypes and a Biological State Machine temporal decoder. The backbone remains BiomedCLIP, a biomedical vision-language foundation model. Three consecutive frames are fused through a Local Differencing Attention module that amplifies transient pathological signals by suppressing static temporal redundancy. An Anatomy Context Head then conditions pathological predictions on soft anatomical activations, exploiting the known spatial co-occurrence structure of GI findings. Learnable text-feature prompts and prototype-based logit augmentation are trained alongside an Angular Separation Loss that penalizes off-diagonal cosine similarity between class prototypes, preventing the prototype collapse that afflicts rare classes under extreme imbalance. To counteract the skewed label distribution, the training regime combines asymmetric focal loss, inverse-frequency weighted sampling, temporal Mixup, Exponential Moving Average, and per-class threshold calibration. The Biological State Machine decoder replaces naive gap merging with a physiologically grounded forward-only state transition over anatomy labels, eliminating the fragmentation artefact that produced hundreds of spurious anatomy events per video in the prior approach and reducing per-video anatomy output to 2--3 clinically realistic events. On the held-out RARE-VISION test set comprising three NaviCam examinations (161,025 frames), the updated pipeline achieves an overall temporal mAP@0.5 of 0.3597 and mAP@0.95 of 0.3399, representing a relative improvement of 46% and 44% respectively over the prior submission, with total inference completed in approximately 21 minutes on a single GPU.
X-TRACK: Physics-Aware xLSTM for Realistic Vehicle Trajectory Prediction
arXiv:2511.00266v2 Announce Type: replace Abstract: Accurate trajectory prediction is crucial for safe and reliable autonomous driving systems, requiring models that capture long-term temporal dependencies while accounting for social interactions among neighboring vehicles in highway driving scenarios. While Long Short Term Memory (LSTM) networks have been widely used in the domain of trajectory prediction, they have limitations such as limited memory capacity and scalar cell state. The recently introduced Extended Long Short Term Memory (xLSTM) addresses these limitations of traditional LSTMs by introducing exponential gating and enhanced memory structures, making them better suited for modeling long-term temporal dependencies. Despite their potential, xLSTM-based models remain underexplored in the context of vehicle trajectory prediction. This paper introduces a novel xLSTM-based highway trajectory prediction framework, X-TRAJ, as the first application of xLSTM, and its physics-aware variant, X-TRACK (eXtended LSTM for TRAjectory prediction Constraint by Kinematics), which explicitly integrates vehicle motion kinematics into the model learning process. By introducing physical constraints, the proposed model generates realistic and feasible highway trajectories. A comprehensive evaluation on the publicly available highway datasets, highD and NGSIM, demonstrates that X-TRACK outperforms state-of-the-art baselines on highD and is among the state-of-the-art models on the NGSIM dataset.
Sparser Block-Sparse Attention via Token Permutation
arXiv:2510.21270v2 Announce Type: replace Abstract: Scaling the context length of large language models (LLMs) offers significant benefits but is computationally expensive. This expense stems primarily from the self-attention mechanism, whose $O(N^2)$ complexity with respect to sequence length presents a major bottleneck for both memory and latency. Fortunately, the attention matrix is often sparse, particularly for long sequences, suggesting an opportunity for optimization. Block-sparse attention has emerged as a promising solution that partitions sequences into blocks and skips computation for a subset of these blocks. However, the effectiveness of this method is highly dependent on the underlying attention patterns, which can lead to sub-optimal block-level sparsity. For instance, important key tokens for queries within a single block may be scattered across numerous other blocks, leading to computational redundancy. In this work, we propose Permuted Block-Sparse Attention (\textbf{PBS-Attn}), a plug-and-play method that leverages the permutation properties of attention to increase block-level sparsity and enhance the computational efficiency of LLM prefilling. We conduct comprehensive experiments on challenging real-world long-context datasets, demonstrating that PBS-Attn consistently outperforms existing block-sparse attention methods in model accuracy and closely matches the full attention baseline. Powered by our custom permuted-FlashAttention kernels, PBS-Attn achieves an end-to-end speedup of up to $2.75\times$ in long-context prefilling, confirming its practical viability. Code available at https://github.com/xinghaow99/pbs-attn
D2 Actor Critic: Diffusion Actor Meets Distributional Critic
arXiv:2510.03508v3 Announce Type: replace Abstract: We introduce D2AC, a new model-free reinforcement learning (RL) algorithm designed to train expressive diffusion policies online effectively. At its core is a policy improvement objective that avoids the high variance of typical policy gradients and the complexity of backpropagation through time. This stable learning process is critically enabled by our second contribution: a robust distributional critic, which we design through a fusion of distributional RL and clipped double Q-learning. The resulting algorithm is highly effective, achieving state-of-the-art performance on a benchmark of eighteen hard RL tasks, including Humanoid, Dog, and Shadow Hand domains, spanning both dense-reward and goal-conditioned RL scenarios. Beyond standard benchmarks, we also evaluate a biologically motivated predator-prey task to examine the behavioral robustness and generalization capacity of our approach. Code: https://github.com/d2ac-actor-critic/d2ac-public
TCAD + Allpi$\text{x}^2$ Simulation study of MALTA2, a Depleted Monolithic Active Pixel Sensor for future tracking
arXiv:2605.23860v1 Announce Type: new Abstract: In this work, a hybrid simulation framework combining TCAD and Allpi$\text{x}^2$ is presented to investigate the sensor properties of MALTA2, a depleted monolithic active pixel sensor designed for future tracking. The study starts from 3D modeling and transient simulations in TCAD, with generic doping profiles and simple well structures. The resulting doping profiles and electric field are extracted and fed into Allpi$\text{x}^2$ for high-statistics Monte Carlo simulations in both DUT-only and full-telescope mode. Simulations reveal a strong dependence of sensor performance, specifically the detection efficiency and cluster size, on the doping concentration of the N-type blanket at the sensor surface. The doping concentration is then optimized by comparing simulations with measurement data. The active depth of the depleted region of the MALTA2 sensor is estimated in both simulations and measurements using a grazing angle method, in which the sensor is positioned at various inclinations relative to the beam, covering angles from 0 to 60 degrees. Excellent agreement on active depth is obtained with the optimal doping concentration, showing a deviation of 2\% from the measured value at a threshold of 450\,$\text{e}^-$. Consequently, the framework offers a generic toolkit for sensor studies without requiring proprietary information.