Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Rethinking the UI of GenUI: A Tale of Two Designs
arXiv:2606.13843v2 Announce Type: replace Abstract: GenUI is an emergent class of AI tools that use large models to generate UI mock-ups based on users' high-level descriptions, promising to democratize UX design exploration for a broader audience. Most GenUI designs to date tend to inherit the conventions of conversational large models, such as ChatGPT and Gemini, where a user describes their design needs primarily via an unstructured prompt, and the tool then takes a depth-first approach, delving into the design right away and producing a high-fidelity prototype. In this research, we rethink how well this unstructured, depth-first, and high-fidelity GenUI design can support early-stage, 0-to-1 design exploration. To probe this question, we propose a contrastive design with structured input, breadth-first exploration, and low-fidelity generation. We then conducted a comparison study with 24 UX designers and product managers who conducted mini design exploration exercises using an existing GenUI tool and our contrastive GenUI tool. Findings reveal participants' perceived benefits and trade-offs of the two GenUI designs: structured input surfaces key facets but requires more work, raising entry barriers to start exploration; breadth-first workflow reveals more possibilities, but previewing UX ideas spanning many screens remains hard; and though low fidelity has value, professionals favor high fidelity because it fits practice and GenAI heightens fidelity expectations. We conclude with design implications for GenUI and similar AI-powered creativity support tools.
Gefen: Optimized Stochastic Optimizer
arXiv:2606.13894v2 Announce Type: replace Abstract: AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to training memory, increasing the already substantial cost of large-scale pretraining. We propose Gefen, a memory-efficient optimizer that automatically shares second-moment estimates across parameter blocks and quantizes the first moment using a learned codebook, thereby reducing AdamW's memory footprint by ~8x while maintaining the same performance, corresponding to a reduction of 6.5 GiB per billion parameters. The method is motivated by a theoretical result showing that large mixed Hessian entries constrain the ratio of squared gradients toward one, suggesting that Hessian-aligned parameters are natural candidates for sharing second-moment statistics. Since computing Hessians is impractical at scale, Gefen infers block structure from the initial squared gradients, requiring no architecture-specific metadata or hyperparameters beyond AdamW defaults. Gefen learns an exact histogram-based dynamic-programming quantization codebook and reuses the same blocks for first-moment scaling. Across diverse pretraining experiments, Gefen achieves the lowest peak optimizer memory among the compared AdamW-like methods while maintaining AdamW-level performance. In single-machine or distributed training, the reduced memory footprint enables larger microbatches and improves throughput significantly over AdamW, providing a practical drop-in replacement with lower memory usage that can increase throughput and enable training larger models or using larger global batch sizes. We provide the complete Python implementation, including fused CUDA kernels at https://github.com/ndvbd/Gefen
Contact-Consistent Interaction Dynamics Normalization for Predictive Physical Human--Robot Interaction
arXiv:2606.14617v2 Announce Type: replace Abstract: Safe physical human--robot interaction on floating-base robots requires interaction regulation under changing contact constraints. We develop a contact-consistent normalization in which the residual end-effector channel is represented as a linear double integrator in acceleration coordinates. Both discrete prediction matrices are independent of configuration and support mode; posture and contact enter only through task-inertia force recovery and constraints. The controller combines a constant-Hessian receding-horizon QP, an acceleration-disturbance observer, and a priority-consistent realization. Classical operational-space impedance is shown to be the unconstrained infinite-horizon limit. MuJoCo experiments on a 17-DOF biped and a Menagerie-derived Unitree G1 model evaluate sustained forces, transmitted shocks, and scheduled contact-model changes. Disturbance estimation is the dominant source of fixed-stance accuracy, while covariance inflation gives only scenario-dependent transient benefit. Dynamic walking and hardware validation remain outside the present evidence.
RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos
arXiv:2606.16278v2 Announce Type: replace Abstract: Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect and reproduce at scale. Editable 3D Gaussian Splatting (3DGS) simulation offers a promising alternative by reconstructing real driving scenes and supporting controllable scene editing. However, edited 3DGS-rendered videos still suffer from a significant Sim-to-Real gap, including rendering artifacts, degraded foreground assets, inconsistent illumination, and temporal flickering. Existing restoration and video generation methods are insufficient for this task, as they often fail to jointly repair 3DGS-specific artifacts, improve visual realism, and ensure temporal consistency. To fill this gap, we propose RealityBridge, a structure-preserving and asset-aware Sim-to-Real framework for edited 3DGS driving videos. RealityBridge uses multimodal controls, including rendered videos, foreground masks, edge maps, and semantic masks, together with a lightweight GateNet for adaptive condition allocation across backbone layers. We further construct targeted training data and introduce autoregressive long-video training with reward-guided post-training to improve restoration quality, temporal stability, and hallucination suppression. Extensive experiments on internal and public driving datasets show that RealityBridge outperforms existing methods in artifact removal, illumination harmonization, and long-sequence temporal consistency.
KVEraser: Learning to Steer KV Cache for Efficient Localized Context Erasing
arXiv:2606.17034v2 Announce Type: replace Abstract: Post-hoc context erasing over the KV cache is challenging because a local edit has a global consequence: once a span has been processed, its influence propagates into the cached states of all subsequent tokens. This issue arises naturally in long-context LLM applications, where stale retrieved facts, incorrect tool observations, retracted user preferences, or harmful prompt injections may be identified only after prefill. Exact erasing must then recompute all tokens after the deleted span, making its computational cost depend on suffix length rather than erased-span length. We introduce KVEraser, a learned KV-cache editing method for efficient localized context erasing. Given a processed context and a span to remove, KVEraser replaces only the KV states of the erased interval with learned steering states while reusing the remaining cache unchanged. To learn a transferable erasing mechanism, we build a two-stage training pipeline: generic span-neighbor pre-training teaches the eraser to suppress the influence of the erased span, while task-specific fine-tuning adapts this capability to downstream scenarios. Experiments show that KVEraser nearly matches full recomputation in post-erasure performance on in-domain tasks across 1K--32K context lengths, while its latency increases by only 24% compared with a 17.6x increase for full recomputation. KVEraser also generalizes to unseen long-document QA tasks with harmful factual distractors, achieving the best performance among approximate baselines with a 3--4x speedup over full recomputation.
Optical Emission Spectroscopy Measurements of keV Apparent Ion Temperatures in Avalanche Energy's Centrifugal Mirror Machine
arXiv:2606.17195v2 Announce Type: replace Abstract: Newly formed ions in $E \times B$ devices are rapidly accelerated by strong radial electric fields and execute large cycloidal orbits in the presence of an axial magnetic field. At locations where these orbits intersect, ions originating from different birth radii arrive with substantially different velocities, producing a non-Maxwellian velocity distribution with a large velocity variance. Through Coulomb collisions and collective interactions, this distribution relaxes toward a drifting Maxwellian in the rotating frame. Here, we present the first optical emission spectroscopy (OES) measurements of the line-of-sight-convolved ion-velocity distribution, from which an apparent ion temperature is determined, in Avalanche Energy's centrifugal mirror machine. High-resolution $H_\alpha$ spectra obtained along five chordal lines-of-sight spanning the plasma radius are analyzed using two complementary models representing limiting cases of the ion dynamics: a collisionless cycloidal model based on the ion-velocity distribution arising from deterministic single-particle orbits, and a rotating Gaussian model based on collisions and collective processes that fully randomize the cycloidal motion into a drifting Maxwellian in the rotating frame. Combined, these approaches bracket the possible degree of velocity-space relaxation and provide a stringent test of the inferred ion energies. Both models reproduce the measured spectra relatively well and yield density-weighted apparent ion temperatures of $1.40\pm0.43$ keV for the rotating Gaussian model and $1.55\pm0.24$ keV for the cycloidal model. These results provide direct spectroscopic evidence that strong $E \times B$ rotation in a device only a few centimeters in size can generate ion populations with keV energy spreads.
REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation
arXiv:2607.10147v1 Announce Type: new Abstract: Automated chest X-ray report generation has recently benefited from reinforcement learning (RL) and large language models. However, RL training often suffers from instability or limited exploration due to fixed Kullback-Leibler (KL) regularization and a static reference policy that accumulates KL pressure over time. We propose Response-Weighted and Validation-Anchored Policy Optimization (REVA-PO), a RL framework that stabilizes long-term training via Response-Weighted Regularization (RER) and Validation-Anchored Policy Reset (VAPR). RER dynamically adjusts per-response KL weights based on advantage and reference-policy entropy, relaxing constraints for high-quality responses while tightening them for low-quality ones. Complementarily, VAPR periodically synchronizes the reference and current policies to the best validation checkpoint, resetting accumulated regularization pressure to expand the viable exploration space. To ensure a robust starting point, we employ a three-stage pipeline consisting of warm-up training, classifier-guided supervised fine-tuning, and RL. Extensive evaluations on MIMIC-CXR and IU-Xray demonstrate that REVA-PO sets new state-of-the-art benchmarks in both linguistic quality and clinical accuracy. Notably, BLEU-4 improves by 5.1% on MIMIC-CXR and 3.6% on IU-Xray, while CheXpert F1 and RadGraph F1 scores increase by 4.5% and 12.8%, respectively, over prior leading methods. The code is publicly available at https://github.com/LiGuo12/REVA_PO/.
JAX-FEM-ANISO: Differentiable GPU-Accelerated Finite Element Framework for Inverse Identification of Finite-Strain Anisotropic Plasticity
arXiv:2606.17390v2 Announce Type: replace Abstract: We present a fully differentiable, GPU-accelerated finite element framework JAX-FEM-ANISO, for forward simulation and inverse parameter identification of finite-strain anisotropic plasticity. Built on JAX-FEM, the framework exploits modern accelerator architectures by parallelizing the three major computational bottlenecks in nonlinear FEM: elemental weak-form and tangent-stiffness evaluation, global sparse matrix assembly, and sparse linear solution. For a large-scale forward problem with 3 million degrees of freedom, JAX-FEM-ANISO on a single NVIDIA H100 GPU achieves up to 9.4$\times$ speed-up over a 24-core CPU Abaqus baseline. Automatic differentiation is applied through the constitutive update and solver workflow, providing consistent Jacobians for complex constitutive models without manual derivation and accurate gradients for PDE-constrained inverse analysis. Compared with finite differences, the JAX-AD gradients avoid step-size sensitivity and provide the required sensitivities at substantially lower computational cost. For inverse characterization, we combine information-rich, topology-optimized heterogeneous specimens with full-field displacement data to identify advanced constitutive model parameters from a single test, replacing what would otherwise require many conventional experiments. We demonstrate accurate recovery of anisotropic yield and hardening parameters in progressively challenging settings, including uniform and spatially varying material properties. The resulting AD-based formulation enables efficient optimization in high-dimensional parameter spaces where finite-difference approaches are computationally infeasible. These results establish differentiable, GPU-accelerated FEM as a practical high-throughput engine for simulation, characterization, and optimization workflows in advanced manufacturing.
Distinguished Scaling and UTSD Structure in Weak Shock Reflection at Nearly Glancing Incidence
arXiv:2606.18618v2 Announce Type: replace Abstract: We study weak shock reflection from a rigid wall in the joint limit of weak shock strength and nearly glancing incidence. In the distinguished scaling $\Mach=1+\lambda\alpha^2$, the inner reflection region is governed by the unsteady transonic small-disturbance (UTSD equation and is controlled, to leading order, by the single parameter $a_0=1/(2\sqrt{\lambda})$, independent of the ratio of specific heats $\gamma$. Thus the known UTSD detachment value $a_d=\sqrt2$ corresponds in this scaling to $\lambda_d=1/8$, with Guderley--Mach reflection for $\lambda>1/8$. The physical trajectory angle is obtained by multiplying the canonical UTSD trajectory function $g(a)$ by the Mach-number strength scale $\delta=\sqrt{2(\Mach^2-1)}$, so that $\chi_{\rm phys}=\delta g(a)+O(\delta^2)=2\sqrt{\lambda}\,\alpha g(a_0)+O(\alpha^3)$. We rederive the self-similar UTSD reduction, sonic parabola, and shock polar in order to make the convention and the detachment map self-contained. We also record a formal adjoint solvability expression for the first correction $H(a;\gamma)$, while specifying the free-boundary data required to evaluate it. Finally, a time-marching solver for the full leading-order canonical UTSD system is benchmarked at $a_0=0.5$: retaining the transverse compression $u>1$ gives a $u=0.5$ contour location consistent with the Hunter--Tesdall triple-point benchmark. This computation is used only as a leading-order benchmark, not as a substitute for an adaptive self-similar Guderley free-boundary solver.
Ricci flow for the Bures--Helstrom qubit metric
arXiv:2606.19493v3 Announce Type: replace Abstract: The Bures--Helstrom metric is the minimal monotone Riemannian metric on the state space of a qubit. With the quantum Fisher normalization used here, it identifies the Bloch ball with a geodesic hemisphere of the unit round three--sphere. We describe its Ricci flow explicitly. In a general rotationally symmetric gauge the flow is a coupled system for the radial lapse and warping factor; a single scalar equation appears only after a Hamilton--DeTurck gauge choice. In the corresponding moving DeTurck frame the squared warping function $\Psi=\Phi^2$ satisfies the linear forced heat equation \begin{equation*} D_t\Psi=\Psi_{ss}-2, \end{equation*} while the fixed-lapse coordinate form contains the associated transport term. Since the Bures--Helstrom metric is Einstein, the geometric flow itself is the homothetic shrinker \begin{equation*} g(t)=(1-4t)g_{\mathrm{BH}}, \end{equation*} with scalar curvature $6/(1-4t)$ and extinction time $T=1/4$. Thus the metric remains inside the monotone cone for all $t<T$ and leaves the cone of nondegenerate Riemannian metrics only through the collapsed limit. We also record the volume--normalized flow, for which the Bures--Helstrom metric is a fixed point. Its linearization is the shifted round--sphere Laplacian $\Delta_{\mathbb S^3}+3$, with spectrum \begin{equation*} \sigma_\ell=-(\ell-1)(\ell+3), \end{equation*} and spectral gap $5$ after removal of the scaling mode.
Prismriver: Formalization of Music Theory and Algorithmic Composition in Lean 4
arXiv:2606.19936v2 Announce Type: replace Abstract: Music theory obeys a rich set of mathematical rules and symmetries. These symmetries follow mathematical structures which can be verified and expressed in the precise language of a proof assistant. In this paper, we present Prismriver, a formalization library of music theory in Lean 4. We use Prismriver to generalize beyond existing work that assumes equal temperament tuning. We also discuss modelling counterpoint music theory with Prismriver. By formalizing music theory in Lean 4, we open the door to verifiable algorithmic composition and accompaniment generation. Prismriver also has a custom DSL integrated with MusicXML exports to interoperate with other music software. Prismriver can be used to compose music with Lean, using monadic composition primitives.
Deep-Ocean Application-Specific Neutrino Experiment
arXiv:2606.20780v2 Announce Type: replace Abstract: This report introduces the concept, prototype design, projected costs, and scientific goals of a mobile experiment for detecting geoneutrinos originating from uranium and thorium decay chains in the Earth's mantle. This will constrain the planet's radiogenic heat production and unearth its geochemical makeup. This design of a deep-ocean mobile neutrino experiment, which is not mirrored by any active or planned experiments, supports physics and geoscience's goal of multi-modal data on the Earth's internal composition and structure. Based on geoscientific studies, this design is expected to achieve a 50--100-fold reduction in crustal background compared to similarly sized continental detectors, thereby enabling direct measurements of mantle geoneutrinos. The multiple stereoscopic projections enabled by the detector's unique mobility can map spatial variations in heat-producing elements within the mantle. Beyond discussing the design, we report on our collaboration's most recent hardware developments in the active prototyping of this detector. We briefly highlight the potential multiuse and interdisciplinary nature of this detector.
From Embedding Geometry to Spectral Search: Energy Dispersion Networks For Vector Retrieval
arXiv:2606.21535v2 Announce Type: replace Abstract: High-dimensional vector spaces, particularly embedding spaces with dense semantic structure, are often interpreted primarily leveraging solely geometric relationships. In this work, we show that they can also be viewed as spectral energy networks induced by the topology of their underlying feature-space manifold with relevant improvements for downstream tasks. Building on this perspective, we introduce Graph Wiring, a general framework for exploiting feature-space spectral structure, together with Spectral Indexing, its task-specific instantiation for vector search. By coupling geometric similarity with spectral information, the proposed method improves Head-Tail coherence and semantic alignment relative to purely geometric retrieval methods. It further supports adaptive search behavior through tau-modulation, providing the flexibility increasingly required by modern Retrieval-Augmented Generation (RAG) pipelines. We present the complete algorithmic pipeline, establish its theoretical foundation through epiplexity, and evaluate the approach across benchmark and industrial settings using the open-source arrowspace library.
FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving
arXiv:2606.21587v2 Announce Type: replace Abstract: Deep reinforcement learning is pivotal for closed-loop autonomous driving yet remains constrained by severe bottlenecks in sampling efficiency. Standard parallel sampling mitigates this but suffers from the straggler effect, where the premature termination of a single environment necessitates a synchronized batch re-initialization, leading to suboptimal sample utilization and prohibitive re-initialization latency. To address this, we propose FAST, a synchronous parallel framework tailored for closed-loop simulation. Specifically, FAST employs Dynamic Parallel Sampling Alignment (DPSA) to maintain vectorization synchronization by extending terminated episodes via virtual continuation, thereby decoupling the sampling loop from individual terminations. By dynamically triggering global truncation based on the termination rate of parallel clips, FAST effectively eliminates the bottleneck of premature resets without sacrificing data diversity. Furthermore, to strictly preserve theoretical consistency, we incorporate a Scaled Mask-Padding Optimization (SMPO) that leverages validity masking and adaptive loss normalization to nullify the bias from auxiliary padding data. Empirical evaluations demonstrate that FAST achieves at least a 1.78 times wall-clock speedup over the single-clip baseline while preserving statistical unbiasedness.
Small edits, large models: How Wikipedia advocacy shapes LLM values
arXiv:2606.24890v3 Announce Type: replace Abstract: Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia appears in nearly every major language model training dataset and is weighted more heavily than web-crawled text. The Pro-Animal Wikipedians (PAW), a group of advocates who add sourced animal welfare content to relevant articles, have made 125 edits across 115 pages. Using gradient-based data attribution (Bergson; MAGIC), we traced how these edits influence language model behavior. TrackStar retrieval attribution on Llama 3.1 8B found that PAW-edited sections made up 68 percent of the highest-attributed documents for animal welfare queries (p < 0.0001) but only 52 percent for unrelated queries about the same companies (p = 0.53): the model links PAW content specifically to animal welfare topics, not to the entities in general. MAGIC counterfactual influence estimation on Llama-3.2-1B, run across five random training-order seeds, gave the same picture even more sharply: in every seed, the top-10 most influential documents on animal welfare queries were all PAW edits (10 of 10, 5 of 5 seeds), while on general queries the same top-10 sat at chance (4 to 6 of 10). Mean PAW influence exceeded mean control influence on animal welfare queries with p < 0.0001 in every seed, an effect 6 to 30 times larger than on general queries. Leave-subset-out validation gave Spearman rho = 1.00 for all 10 runs. When we fine-tuned separate models on PAW content versus control content, each model performed better specifically on the type of text it was trained on: the PAW-trained model cut perplexity on animal welfare text from 12.4 to 8.4, while the control-trained model cut perplexity on control text from 16.1 to 11.4. A small, coordinated Wikipedia editing campaign therefore measurably shapes how language models handle the topics those edits address.
RWGBench: Evaluating Scholarly Positioning in Related Work Generation
arXiv:2606.24894v3 Announce Type: replace Abstract: Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited. Existing RWG evaluations largely inherit summarization-oriented metrics, using lexical or semantic similarity to reference sections as proxies for quality. However, related work writing is fundamentally a citation-level scholarly positioning task: it requires selecting, organizing, and framing prior work to clarify how a target paper relates to, differs from, and contributes beyond existing research.As a result, models may generate coherent and semantically-relevant text while exhibiting academically critical failures, such as inappropriate citation selection or misplaced references, that conventional metrics do not capture.To this end, we introduce \textbf{RWGBench}, a benchmark that evaluates RWG from the perspective of citation decision-making rather than text similarity. RWGBench is constructed from a large-scale collection of 40,108 computer science papers and a retrieval corpus of 1.09 million documents, with a carefully curated test set comprising 100 papers and their corresponding published related work sections.We propose a multi-dimensional evaluation framework that assesses citation selection, contextual appropriateness, organization, and discourse structure.Experiments reveal systematic limitations in current systems that are obscured by standard evaluations, while Oracle studies further disentangle retrieval-level and generation-level bottlenecks. Human evaluation further shows that our citation-centric metrics align substantially better with expert judgment than surface-level text metrics. RWGBench offers a citation-centric testbed for developing and evaluating related work generation systems that are better aligned with scholarly writing practices.
What Does It Mean to Break a Distillation Defense?
arXiv:2606.25059v2 Announce Type: replace Abstract: Black-box LLMs (accessible only via API) are vulnerable to distillation attacks, in which an attacker queries the model and trains a student on its outputs. A recent line of work proposes output perturbation defenses that modify the teacher's output to reduce student performance while preserving utility for legitimate users. As a relatively new family of approaches, output perturbation defenses lack a shared threat model, making it difficult to compare them, reason about composing them with other attacks, or evaluate their robustness against realistic adversaries. This underspecification matters beyond technical evaluation: when defenses are deployed to protect intellectual property or justify regulatory compliance, an imprecise threat model can create a false sense of security. We propose a threat model framework that describes attackers along three dimensions: a query budget, a data budget, and an interface profile that captures how attackers interact with the API. Using antidistillation sampling as a case study, we show that whether the defense is considered effective depends on the assumed threat model. We argue that future work on distillation defenses, along with any governance or policy frameworks built around them, should explicitly specify and stress-test attacker capabilities along our three dimensions.
Average-Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning
arXiv:2606.25680v2 Announce Type: replace Abstract: Underwater vehicles operate from a fixed onboard energy budget that propulsion rapidly depletes, so a controller that completes its task while drawing less thruster power directly extends mission range and endurance. Reinforcement learning yields capable model-free controllers for station-keeping and trajectory tracking, but optimizing task accuracy alone drives the policy toward oscillatory, energy-wasting actuation. The established remedy subtracts an energy penalty from the reward, yet this sets the task-power trade-off through a single weight with no physical units: a target power level cannot be specified, the weight must be re-tuned for every vehicle and task, and a mismatched weight can even raise power. This paper instead formulates energy-efficient underwater control as a constrained Markov decision process in which average thruster power is subject to an explicit budget, solved with a PPO-Lagrangian algorithm. The power level is set by declaring a budget in physical units, and a single dual variable is updated online to meet it for each vehicle and task, without manual weight search. Across three vehicles and four tasks in the MarineGym simulator, the energy-constrained policy draws the least power in all twelve settings, reducing it by 14--65\% (up to 64.9\%) over a task-only baseline and below an energy-reward baseline everywhere, while remaining the smoothest in ten settings and preserving task accuracy except in one deliberately power-limited regime. Imposing energy as an explicit constraint thus offers a tuning-free route to energy-efficient underwater control that needs no per-vehicle, per-task weight search.
Collisions and Stopping of Fast Charged Particles in Matter
arXiv:2606.25847v3 Announce Type: replace Abstract: This text is intended to offer a consistent presentation of the theory of collisions and stopping of charged particles in matter, limited to the range of intermediate kinetic energies where atomic aggregation effects are relatively unimportant and processes such as the creation of particle-antiparticle pairs are not likely to occur. The first three Chapters contain introductory material on the classical description of electromagnetic fields in matter, an overview of quantum wave equations for a particle in a central potential, and an account of elementary atomic-structure models. Chapters 4 and 5 are devoted to the classical and quantum theories of elastic collisions of charged particles with atoms. The theory of inelastic collisions and stopping is split into two parts: first, collisions with atoms are considered within the plane-wave Born approximation in Chapter 6, which includes a derivation of the Bethe stopping power formula; second, the theory of inelastic collisions in dense materials is based on the dielectric formalism, which is formulated for the electron gas, and extended to arbitrary materials by means of optical-data models in Chapter 7. Chapter 8 offers a detailed review of the theory of stopping, starting with the classical study by Bohr and ending with derivations of the Bloch and Barkas corrections to the stopping power. Chapter 9 deals with general aspects of transport theory, including derivations of energy-straggling distributions and multiple-scattering distributions, which are the basis for condensed simulation schemes of charged particle transport. Finally, Chapter 10 describes the Fortran programs elastic and sbethe, which implement the main theoretical models presented in the preceding Chapters and are distributed as ancillary information.
Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training
arXiv:2606.26102v3 Announce Type: replace Abstract: Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these processes may inadvertently degrade values instilled during pre-training. We investigate whether the domain of post-training data differentially affects the retention of animal compassion values in a Llama 3.1 8B model mid-trained on compassion-oriented synthetic data, using both SFT (helpfulness via Dolly-15k vs. coding via Magicoder-110K) and GRPO (helpfulness via RLHFlow vs. coding via Magicoder), evaluated on the ANIMA 2.2 benchmark and MORU benchmark (Moral Reasoning Under Uncertainty). Helpfulness training significantly degrades animal compassion relative to coding training on ANIMA (SFT: 35.7% vs. 65.2%; GRPO: 18.7% vs. 32.0%), replicating across two independent helpfulness datasets and two training paradigms. On English MORU items, helpfulness training degrades general moral reasoning by 25.5 percentage points (46.4% vs. 71.9%), a striking gap that rivals the compassion effect in magnitude. However, this effect does not transfer cross-lingually: on the multilingual MORU benchmark, the domain effect disappears (SFT: 52.3% vs. 51.2%). In contrast, the animal compassion effect transfers consistently across languages, with Magicoder's ANIMA percentage-point gain over the base model 4.5 times larger on non-English items than English items. This divergence suggests that values instilled through mid-training are encoded more deeply and cross-lingually than reasoning improvements from domain-specific post-training. These results suggest that, for labs building on value-laden mid-training, coding-domain post-training may better preserve mid-trained values than helpfulness post-training without harming general reasoning capabilities.
Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare
arXiv:2606.26104v3 Announce Type: replace Abstract: Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about animal welfare. Using vocabulary-matched stance-contrast probes on a held-out animal-welfare benchmark, we measure how each of ten linguistic features changes Llama-3.2-1B's preference for pro-animal-welfare reasoning when used as fine-tuning data. Eight of the ten features produce statistically significant shifts. Seven move the model toward stronger pro-animal-welfare reasoning: assertive certainty, explicit moral vocabulary, emotion words, evaluative claims, narrative structure, depicted harm severity, and immediate temporal framing. Two move it the other way: hedged language and concrete sensory description both dilute the pro-animal-welfare stance. First-person perspective has no statistically significant effect. The practical recommendation for anyone writing animal-welfare text that may end up in LLM training corpora: assert a position rather than describe a scene neutrally. The features that shift the model are the ones that make the writer's position explicit; the features that dilute it hold animal-welfare content but withhold stance.
Forget, Anticipate and Adapt: Test Time Training for Long Videos
arXiv:2606.26515v3 Announce Type: replace Abstract: Test Time Training (TTT) is a mechanism in which a model adapts to an incoming test-sample by performing some self-supervised (SSL) task and updating its weights even during inference. This procedure does not require labels at test-time. This paper focuses on TTT for long-videos. A major concern with existing approaches is: 1) they perform TTT updates using a sliding window containing frames in the past, whose compute increases linearly with the size of window. This becomes computationally intractable when the videos are hours long. 2) TTT is performed even when temporally close frames look similar, thereby consuming a lot of compute. We present the Frame Forgetting Network (FFN) that: 1) operates on only three frames within the sliding window, namely the frame that exits, the current frame and the frame after that. The model still manages to retain temporal context and work for hours long-videos; 2) mathematically define a surprise metric: how much new information the incoming frame contains with respect to the past seen frame. This facilitates determining how to modify the effective window size during TTT and constitutes the core mechanism of an adaptive windowing algorithm. Additionally, we curate a dataset EpicTours containing up to 3 hour long videos of walking city-tours, whereas earlier datasets on this problem were only 5 min long. We demonstrate FFNs empirical effectiveness on dense-segmentation, video classification tasks, generalization to depth-estimation, and multi-hour long videos.
NaviCache: Test-Time Self-Calibration Caching for Video Generation
arXiv:2606.26795v2 Announce Type: replace Abstract: Video Diffusion Models (VDMs) is constrained by immense computational costs. While offline calibration-based acceleration suffers from calibration data dependency, prohibitive calibration duration, and susceptibility to distribution shifts, offline calibration-free methods eliminate these hurdles. However, since they rely on instantaneous zero-order approximations where the mapping between input and output differences varies in real-time, they are susceptible to observational noise and ignore the intrinsic momentum within the diffusion trajectory. In this paper, we propose NaviCache, a plug-and-play test-time self-calibration method re-conceptualizing feature evolution as an Inertial Navigation System (INS) problem. NaviCache bridges the fundamental domain gap and the non-stationary nature of diffusion by modeling the relative coupling between input and output variations. We introduce a dual-state estimation architecture that adaptively tracks the feature change ratio and its latent drift, initialized via a specialized Initial Alignment phase. By integrating a time-dependent noise schedule with an uncertainty-aware Measurement Update mechanism, NaviCache provides a theoretically grounded mechanism for error-bounded computation skipping. Extensive experiments on the HunyuanVideo, Wan, and Open-Sora series demonstrate that NaviCache exhibits more accurate error judgment for computation skipping and achieves outstanding comprehensive performance.
FlowArk: Boosting Agentic Data-flow Analysis for Android Apps via Context-Aware Knowledge Reuse
arXiv:2607.11308v1 Announce Type: new Abstract: Data-flow analysis is foundational to Android app privacy and security auditing. Recent coding agents can assist with non-trivial source-to-sink data-flow analysis tasks by searching, reading, and reasoning over repository code. However, when these tasks are executed as a batch workload, current agentic analysis setups incur substantial re-analysis cost. Agent instances assigned to different taint sources may inspect shared code fragments, because code reuse in the target app can cause different data-flow paths to converge on shared program logic. Since these agent instances are context-isolated, analysis of these shared code fragments can be repeated within a batch, unnecessarily consuming API budget and limiting scalability. We propose FlowArk, a knowledge-reuse system that reduces re-analysis cost in batch agentic data-flow analysis by making knowledge from completed analyses available to later agent instances. Specifically, FlowArk distills completed analysis histories into reusable knowledge candidates, packages these candidates into matchable knowledge entries, and injects matched entries into a later agent instance's context. We implement FlowArk on OpenCode and evaluate it on 4,685 source-to-sink data-flow analysis tasks from 50 open-source Android apps. Compared with standard OpenCode, FlowArk-enabled OpenCode maintains comparable analysis quality while reducing end-to-end API cost by 26.83%. In addition, under a USD 100 budget, FlowArk completes 36.66% more tasks (1,060 vs. 776).
EDAR: Learning Environment-Dependent Action Representations for Robotic Manipulation
arXiv:2607.11427v1 Announce Type: new Abstract: Learning effective action representations is critical for robotic manipulation, where raw control trajectories are often noisy, redundant, and difficult to model directly. Existing methods mainly encode the structure of the action stream itself, treating the role of actions in the environment as implicit. Yet manipulation is about changing the world: the same action segment can induce different outcomes under different scene contexts, making action semantics inherently environment-dependent. We propose EDAR, an Environment-Dependent Action Representation that grounds action tokens in both executable control structure and expected visual consequences. By coupling motor commands with their environment-conditioned effects, EDAR encourages the learned action space to capture interaction semantics rather than merely command-level patterns. Experiments on simulated and real-robot manipulation benchmarks demonstrate that EDAR improves downstream policy learning, especially in long-horizon manipulation. These results highlight the importance of grounding action representations in executable control structure and environment-conditioned visual change.