Forskningsradar

Science Journals

Peer-reviewade publikationer — 60005 artiklar

Collaborating with Artists in the Search for Life
arXiv:2606.24943v1 Announce Type: cross Abstract: Art and science collaborations that go beyond outreach and advertisement in service of science have the potential to unlock new ways of seeing and understanding the Universe that science alone cannot reach. In this white paper for the NASA Decadal Astrobiology Research and Exploration Strategy (DARES) request for information, we outline examples and benefits of artscience and research-creation methods for astrobiology. The search for life and its origin is inherently interdisciplinary and requires novel approaches that could benefit from the training artists receive in design thinking, contextualization, speculation, and community building. We take a look at this process in action through the work of Robert Irwin during the 1970 NASA Habitability Symposium, Carl Sagan's approach to mixing art and science, and the Transition Design framework of creativity-led problem solving. Each example underscores a specific advantage of deeper art-science collaborations: Irwin's creative approach to problem-solving broke scientists from conventional thought patterns, Sagan's contextualization helped align scientific work with ethical and societal considerations, and design-led research is shown to improve planning and efficiency, even for problems as complex as searching for life. Specific implementation recommendations include specifically allowing funding for artist consultations in research grants, reviving NASA's artist-in-residence program, and supporting artscience training initiatives within the astrobiology community.
Curvature-induced smectic-C order of tangentially anchored hard spherocylinders on a sphere with a rigidly locked director field
arXiv:2606.24961v1 Announce Type: cross Abstract: We study the strict locked-orientation limit of hard spherocylinders on a sphere, in which the rod axes are rigidly locked to a prescribed tangential director field and cannot reorient. Because the bulk hard-rod phase diagram contains no smectic-C phase, any coherent tilt isolates a geometric curvature mechanism rather than a finite-stiffness equilibrium effect. A ratio-symmetric recognition cost fixes the layer spacing at the bulk close-contact value and yields a hierarchy of geometric statements: the lower edge of the smectic-area window at $45^\circ$ follows from reciprocal symmetry; the upper edge at $58.3^\circ$ is a falsifiable channel-saturation hypothesis; the smectic-A to smectic-C boundary is a closed-form prediction; and the rod tilt angle is set by the rod-to-radius ratio, modulated by a chirality envelope peaking near $24^\circ$. Locked-orientation Monte Carlo across fifteen geometries confirms these predictions with no fitted elastic constants: the smectic area peaks at $55^\circ$, and a coherent smectic-C window is detected.
Dual Agreement Consistency Learning for Semi-Supervised Fetal Ultrasound Segmentation
arXiv:2606.25254v1 Announce Type: cross Abstract: Maternal-fetal US is the primary imaging modality for monitoring fetal development, yet accurate automated segmentation remains challenging due to the scarcity of pixel-level annotations. To address this issue, we propose DACL, a semi-supervised framework for robust fetal US image segmentation. DACL jointly trains a deployment-oriented lightweight convolutional network (1.47\thinsp\mathrm{M} parameters) and a Transformer-based network, leveraging labeled data for supervised learning and unlabeled data via CPS. To enhance prediction stability, we introduce a dual-agreement consistency loss that couples pixel-wise probabilistic divergence with entropy-guided confidence alignment. Unlike conventional CPS methods that enforce agreement only at the prediction level, DACL explicitly regularizes both distributional alignment and uncertainty, thereby suppressing unreliable pseudo-labels and enabling stable cross-architecture pseudo-label learning under extreme annotation scarcity. Furthermore, an interpolation-based consistency strategy using mixup is applied to unlabeled samples to enhance robustness. Under 5% labeled data, DACL improves Dice by up to 2.77% and reduces HD95 by up to 14.69 mm compared with the strongest recent semi-supervised methods, demonstrating significant improvements in boundary accuracy on both fetal head and abdomen datasets. These results demonstrate the effectiveness of agreement-based consistency learning for annotation-efficient fetal US segmentation. Our code is on GitHub.
Rapid and robust laser-frequency auto-locking using Bayesian-optimization and discrete-wavelet-transformation algorithms
arXiv:2606.25267v1 Announce Type: cross Abstract: Rapid and robust laser-frequency auto-locking is essential for the field deployment of quantum communications, quantum computing, and precision-measurement technologies; however, achieving this remains a considerable challenge. Here, we propose and demonstrate an auto-locking scheme employing Bayesian optimization and discrete biorthogonal wavelet transformation. First, the reference is rapidly sought by making intelligent use of historical observations, eliminating the inherent blindness of the traditional parameter-scanning method. Second, the frequency reference is robustly identified by pinpointing transition signals with the discrete biorthogonal wavelet transformation and analyzing their immutable frequency differences and relative magnitudes, which are determined by the inherent atomic structure and remain resistant to environmental disturbances. This proposed approach achieves a fivefold acceleration in reference searching compared to conventional scanning methods in the case where the laser frequency drifts far away from the reference. Crucially, it achieves an identification accuracy of more than 99.5 %, even under severe 50 % laser-intensity fluctuations, $9.95^\circ$ photodiode misalignment, and $18^\circ$C Rb cell temperature elevation. Finally, locking the laser frequency to the identified reference with a lead zirconate titanate-current double-servo loop narrows the linewidth to 20 kHz. We believe that this rapid, robust, and high-performance auto-locking technique will be pivotal towards the deployment of the next generation of practical quantum technologies in demanding field environments.
Three-Dimensional Positive-Cone Oldroyd-B Flows:Geometric Continuation and Residual-Work Criteria
arXiv:2606.25438v1 Announce Type: cross Abstract: We prove a three-dimensional positive-cone continuation criterion for the stress-diffusion-free Oldroyd-B system on the periodic torus. Writing the positive conformation tensor as A = exp(B), we show that finite-time breakdown of a strong H^s solution, s > 5/2, can occur only through loss of the logarithmic spectral envelope of A or divergence of the endpoint vorticity clock given by the time integral of the B^0_{infty,1} norm of curl u. The proof combines compact positive-cone envelopes, endpoint Biot-Savart estimates, and high-order logarithmic conformation estimates, without using stress diffusion. We also derive a positive-cone Reynolds admissibility criterion with an exact residual-work cost. The least L^2 conformation residual needed to pay positive pressure-free residual work is determined by the entropy-dual lever G = I - A^{-1}, and this cost degenerates quantitatively near the equilibrium A = I. Together, the two criteria identify the same positive-cone obstruction in the strong and relaxed regimes: before breakdown one must control the endpoint flow clock on a compact logarithmic cone, while after passage to a relaxed description positive residual work must be paid for by an exact entropy-dual conformation defect.
A Differentiable DFT-Based Framework for Inverse Materials Design
arXiv:2606.25502v1 Announce Type: cross Abstract: Discovering solid-state materials with target properties remains a central challenge in computational materials science. Existing approaches -- high-throughput screening, surrogate optimization, and generative models -- require extensive evaluations or training data and extrapolate poorly to unseen compositions. Here we develop a first-principles inverse-design framework, integrating reverse-mode automatic differentiation (AD) into KKR-CPA -- the Korringa--Kohn--Rostoker method with the coherent potential approximation -- where atomic compositions are continuous variables to be optimized. Reverse-mode AD yields gradients of objective functions with respect to composition at a cost independent of the number of candidate elements, enabling gradient-based optimization to identify materials from compositional spaces spanning dozens of elements. In this framework, any computable quantity can serve as the objective. We demonstrate this generality through two contrasting applications, magnetic alloys and half-metals, yielding candidates such as (Lu$_{0.553}$Yb$_{0.447}$)(Co$_{0.759}$Fe$_{0.241}$)$_2$Fe$_3$ and FeZr(Sb$_{0.94}$Te$_{0.06}$). Our framework offers a physically grounded route from a target property to the material that realizes it.
Preparing two-mode magnonic Schr\"odinger cat states in a cavity-magnon-qubit system
arXiv:2606.25511v1 Announce Type: cross Abstract: The cavity-magnon-qubit system has recently been demonstrated as a new platform for preparing macroscopic quantum states in magnonic systems. Here, we propose to prepare a two-mode magnonic cat state, which is also a non-Gaussian entangled state, based on this practical system involving two yttrium-iron-garnet (YIG) spheres and a superconducting qubit coupled to a common microwave cavity. By adiabatically eliminating the cavity and resonantly driving the qubit, an effective magnon-qubit conditional-displacement interaction is achieved. Further working in the magnon-magnon strong-coupling regime and considering two identical magnon frequencies and coupling strengths to the cavity, two hybridized magnon modes are formed, of which the bright mode is prepared in a cat state after a projective measurement on the qubit, while the dark mode remains in its initial vacuum state. Such a state corresponds to a two-mode cat state of two original magnon modes, which share strong non-Gaussian entanglement. We also discuss practical dissipation and dephasing effects on the cat state. The results indicate that strong nonclassicality and non-Gaussian entanglement are present in the two-mode cat state using fully feasible parameters.
A topology-tuned pressure valve across the isoreticular RHO zeolite family
arXiv:2606.25557v1 Announce Type: cross Abstract: The isoreticular index of the eight-member embedded RHO zeolite hierarchy operates as a phenomenological design knob that tunes the mechanical critical pressure of the framework valve from $\sim 0.94$ GPa for parent RHO down to a predicted $\lesssim 0.03$ GPa for the largest member PST-28, more than an order of magnitude lower across a single family of synthesizable nanoporous solids. The molecular-valve effect that defines zeolite RHO (a reversible centric-to-acentric phase transition triggered by water, gas pressure, or mechanical loading) is shown here to be not a peculiarity of the smallest member but a generic property of the whole hierarchy, with a critical pressure that decays exponentially with the isoreticular order $k$. Combining lattice dynamics, full elastic-constant tensors, and finite-temperature free-energy reconstructions within a classical core-shell description of the pure-silica frameworks (independently validated against r$^2$SCAN+rVV10 density-functional theory on $G_1$ and $G_2$), we find that the entire family is well described by an effective mean-field Landau picture in which the framework distortion couples quadratically to the volumetric strain. We emphasise that $p_c$ is a mechanical (hydrostatic) critical pressure of the bare pure-silica framework, used as a proxy for the intrinsic framework softness and for the cation- and dehydration-driven response of the real aluminosilicates; it is not a gas-adsorption pressure. The exponential extrapolation places $p_c(G_6\text{-}G_8) \lesssim 0.1$ GPa (model-dependent band $0.03$-$0.15$ GPa), identifying the higher-order members as the softest, most stimuli-responsive frameworks of the hierarchy; whether this intrinsic softness translates into guest-driven switching at low gas activities will depend on the Al distribution, extra-framework cations and adsorbed molecules, and remains to be tested experimentally.
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
arXiv:2602.06566v3 Announce Type: replace Abstract: Despite recent successes, test-time scaling -- i.e., dynamically expanding the token budget during inference as needed -- remains brittle for vision-language models (VLMs). Unstructured visual reasoning chains entangle perception and reasoning, leading to long, disorganized contexts where small perceptual mistakes may cascade into completely wrong answers. Reasoning also requires expensive reinforcement learning with hand-crafted rewards. Here, we introduce SPARC (Separating Perception And Reasoning Circuits), a modular framework that explicitly decouples visual perception from reasoning. Inspired by sequential sensory-to-cognitive processing in the brain, SPARC implements a two-stage pipeline where the model first performs explicit visual search to localize question-relevant regions, then conditions its reasoning on those regions to produce the final answer. This separation enables independent test-time scaling with asymmetric compute allocation (e.g., prioritizing perceptual processing under distribution shift), and supports selective optimization (e.g., improving the perceptual stage alone when it is the bottleneck for end-to-end performance). It also accommodates compressed contexts by running global search at lower image resolutions and allocating high-resolution processing only to selected regions, thereby reducing visual token count and compute. SPARC outperforms monolithic baselines and strong visual-grounding approaches across challenging visual reasoning tasks, such as improving Qwen3VL 4B on the $V^*$ VQA benchmark by 6.7 points and surpassing "thinking with images" by 4.6 points in an OOD setting with a $200\times$ lower token budget.
The Neumann problem for a multivalued p-Laplace equation of Allen-Cahn type with a multiplicative stochastic force
arXiv:2606.25615v1 Announce Type: cross Abstract: In this paper, we consider a parabolic problem with constraint written as a differential inclusion, driven by a multiplicative colored noise and involving a p-Laplace operator (for $p \geq 2$), nonlinear random source terms and subject to Neumann boundary conditions on a bounded Lipschitz domain of $R^d$ with $d \geq 1$. This contribution aims at proving existence and uniqueness of a solution for such a multivalued problem. On one hand, the existence result is proved by the analysis of a semi-implicit time discretization scheme constructed on a smoother version of our problem, itself obtained by a regularization "\`a la Moreau-Yosida" of the subdifferential term. The key point of our approach consists in finding a clever relation between the time step denoted $\tau$ and the Moreau-Yosida regularization parameter denoted $\epsilon$ in view to pass simultaneously to the limit with respect to $\tau$ and $\epsilon$. On the other hand, the uniqueness of the solution is proved by standard arguments.
Quantum Detectability in Invisibility Cloaks
arXiv:2606.25666v1 Announce Type: cross Abstract: Classical invisibility cloaks are designed to suppress selected scattering signatures and thereby make an object appear absent to external electromagnetic probes. However, the suppression of a classical scattering observable does not, by itself, establish that all information about the concealed object has been removed from the detected quantum state of light. Here we formulate the detectability of classically cloaked objects as a quantum-state distinguishability problem. Treating a linear passive cloak as an effective Gaussian quantum channel acting on the accessible detected modes, we show that local quantum undetectability requires the detected first and second moments to be independent of the hidden-object parameter. In this framework, quantum Fisher information provides an operational criterion for whether the concealed parameter remains estimable from the detected output state. We derive displacement- and covariance-level detectability conditions and show that a nonzero parameter imprint surviving in the detected Gaussian state leads to a nonzero accessible quantum Fisher information. To connect the criterion with a physical cloaking model, we analyze a regularized cylindrical transformation-optical cloak in the Born limit and compare the scaling of the classical scattering response with the derivative-based quantum sensitivity. The analysis shows that reducing a scattering amplitude is not equivalent to eliminating local quantum-state sensitivity. Loss, environmental noise, and finite numerical aperture degrade the accessible information, but quantum undetectability is reached only when the parameter imprint is removed from the detected state or projected entirely outside the accessible subspace. These results provide a Gaussian-channel framework for assessing when classical cloaking does, and does not, imply quantum-state undetectability.
Sp(2N, R) interferometry in multi-mode Gaussian bosonic systems for optimal metrology and quantum control
arXiv:2606.25768v1 Announce Type: cross Abstract: Multi-mode interferometers for bosons in Gaussian states are important systems for quantum metrology with precision beyond the standard quantum limit and for bosonic quantum computing. However, there is a lack of theoretical foundation for generic $N$-mode Gaussian interferometry. In this work, we study quantum metrology and quantum control in multi-mode bosonic systems with quadratic Hamiltonians, exploiting the fundamental Sp$(2N,R)$ symmetry of such interferometers. We show that the optimal quantum control to maximize sensitivity requires aligning squeezing and displacement in the same direction. We propose Sp$(2N,R)$ echo, a multi-mode generalization of the SU$(1,1)$ interferometry, to achieve the sensitivity of phase estimation set by the quantum Fisher information. In addition, we introduce a geometrical means for reversing many-body dynamics with Sp$(2N,R)$ dynamical symmetry, such as dynamics of the bosonic Kitaev chain. Our schemes are readily realizable in optical, atomic, and mechanical platforms.
Generating Input Distributions for Explaining Portfolio Optimization Pipelines
arXiv:2606.25808v1 Announce Type: cross Abstract: We propose a predict-optimize-explain framework that uses gradient-based sample generation to interpret various portfolio models by identifying macroeconomic conditions that induce specified portfolio outcomes. Unlike traditional feature-importance methods, this approach directly probes decision pipelines (predictive models coupled with portfolio optimization) by constructing economically meaningful what-if questions. We focus on four such questions: under what macroeconomic conditions a predict-then-optimize pipeline closes or reverses its return gap with a predict-and-optimize pipeline; what conditions lead a pipeline to diversify rather than concentrate its allocation; when a pipeline trained on calm markets overtakes one trained through crises; and what conditions would let a pipeline match a benchmark return. These examples illustrate how our framework uncovers key behavioral differences between various decision pipelines. Beyond these cases, the proposed framework is flexible and can support a wide range of probing questions tailored to specific portfolio objectives. Our findings highlight the value of integrating prediction, optimization, and explanation to produce more robust and transparent portfolio strategies.
DADF: A Distribution-Aware Debiasing Framework for Watch-Time Regression in Recommender Systems
arXiv:2605.17863v2 Announce Type: replace Abstract: Watch-time prediction is a central regression task in short-video recommender systems, where labels are highly long-tailed and residual errors vary systematically across observed watch-time regions. In practice, a model may appear globally calibrated while still overestimating short views and underestimating long views, because opposite errors cancel out in aggregate. Existing methods mainly improve the first-stage watch-time predictor, but often leave such residual distributional bias insufficiently corrected. We propose DADF, a distribution-aware debiasing framework for watch-time regression. Instead of replacing a deployed predictor, DADF performs second-stage multiplicative residual correction on top of it. DADF combines three complementary designs: a dynamic distribution-aware transformation for stabilizing long-tailed correction targets, a debias-factor-aware module for modeling heterogeneous residual patterns using inference-time observable factors, especially video duration, and a multi-label-aware module that exploits auxiliary prediction signals from engagement heads. We evaluate DADF on public short-video benchmarks and a large-scale industrial ranking system. DADF consistently improves both pointwise accuracy and ranking quality across datasets and backbones. In the industrial setting, it achieves an aggregated 2.07 percentage-point ranking-quality gain over the production baseline, consistently reduces MAE, and yields statistically significant online lifts of 0.649% in average time spent per device and 0.656% in total app time. These results demonstrate that DADF effectively mitigates local calibration bias and provides a practical plug-in solution for debiasing long-tailed continuous targets. The source code is available at https://github.com/liuzhao09/DADF.
Metric distortion Under Probabilistic Voting
arXiv:2405.14223v5 Announce Type: replace Abstract: Metric distortion in social choice is a framework for evaluating how well voting rules minimize social cost when both voters and candidates exist in a shared metric space, with a voter's cost defined by their distance to a candidate. Voters submit rankings, and the rule aggregates these rankings to determine a winner. We extend this framework to incorporate probabilistic voting, recognizing that real-world voters exhibit randomness in how they vote. Our extension includes various probability functions, notably the widely studied Plackett-Luce (PL) model. We show that the distortion results under probabilistic voting better correspond with conventional intuitions regarding popular voting rules such as \textsc{Plurality}, \textsc{Copeland}, \textsc{Random Dictator} and \textsc{Borda} than those under deterministic voting. For example, in the PL model with candidate strength inversely proportional to the square of their metric distance from a voter, we show that \textsc{Copeland}'s distortion is at most 2, whereas that of \textsc{RandomDictator} is $\Omega(\sqrt{m})$ in large elections (i.e., number of voters $n \rightarrow \infty$), where $m$ is the number of candidates. This contrasts sharply with the classical model, where \textsc{RandomDictator} beats \textsc{Copeland} with a distortion of 3 versus 5. In the PL model where the candidate strength is inversely proportional to the distance raised to power $\theta$, the distortion under \textsc{Borda} is $\Theta(m^{1-2/\theta})$ when $\theta>2$ and $\Theta(1)$ otherwise. This generalizes the classical deterministic voting model where the distortion of \textsc{Borda} is $2m-1$. The proof uses a novel variant of asymptotic duality where we choose the Lagrange multiplier via asymptotically maximizing the derivative of the objective function. Overall, our work opens a new frontier for analyzing voting rules.
A Methodology for Integrating Life Cycle Assessment into a Multidisciplinary Design Analysis and Optimization Framework for Sustainable Launcher Development
arXiv:2606.25945v1 Announce Type: cross Abstract: The increasing number of orbital and sub-orbital launches makes it necessary to investigate the environmental impacts of launch vehicles and incorporate eco-design considerations into their development. In response, the European Space Agency has promoted Life Cycle Assessment (LCA) as a standardization methodology to mitigate environmental impacts of present and future space missions. This need is further amplified in the NewSpace, where numerous configurations and innovative technologies are explored, reinforcing the importance of integrating environmental considerations. At early design stages, launch vehicle architecture can be formalized through a multi-physics optimization problem based on Multidisciplinary Design Analysis and Optimization (MDAO) methods, where disciplines such as propulsion, aerodynamics, structure, and trajectory are coupled to obtain trade-offs among candidate configurations. This paper proposes a methodology to integrate an LCA discipline within an MDAO framework for launch vehicle design. The approach relies on parametric life-cycle inventories depending on design and coupling variables, covering component and propellant production as well as transport to the launch site. Launch emissions are evaluated from optimized trajectory profiles and characterized in terms of climate change impact. The methodology is illustrated on a representative expendable launch vehicle, where multi-objective optimizations assess trade-offs between performance and environmental indicators. Results highlight antagonistic behaviors among environmental impact categories, emphasizing the importance of carefully defining environmental objectives in eco-design studies. The generic nature of the methodology lays the foundation for integrating LCA into early-stage launch vehicle design, enabling exploration of trade-offs between performance, cost, and environmental considerations.
High-Performance Nanophononic Resonators in Self-Suspended WSe$_2$ Domes and Drums
arXiv:2606.25946v1 Announce Type: cross Abstract: Van der Waals materials are ideally suited for the implementation of high-frequency nanophononic resonators with atomically flat interfaces. Here, we present two versatile van der Waals-based nanophononic architectures: First, we introduce self-supporting nano-domes of WSe$_2$ as a scalable platform for the simultaneous generation of hundreds of high-quality nanoacoustic resonators with resonance frequencies in the 100 GHz range. Second, we engineer self-supporting nano-drums that reach record-high working frequencies for 2D-semiconductor transducers beyond 1 THz. Through optical pump-probe spectroscopy experiments and photoelastic linear chain model calculations, we gain a detailed understanding of the intricate interplay between phononic mode hybridization across heterostructures, the differences between modes close to the center and edge of the acoustic Brillouin zone, and the temporal structure of the photoelastic response. Both architectures have potential applications in low-cost nanoacoustic probing and the ultrafast modulation of quantum emitters in two-dimensional semiconductors. While nano-drums surpass the THz frequency barrier, nano-domes appear as an accessible, low-cost alternative for developing scalable nanophononic technologies.
Measurable Majorities Are Not Finitely Axiomatizable
arXiv:2606.25954v1 Announce Type: cross Abstract: This theoretical note studies the finite axiomatizability of strict majority reasoning in finite social decision frames. Moss and Pedersen (2026) introduce a coherence criterion that characterizes exactly when qualitative majority judgments are representable by a finitely additive measure. The question addressed here is whether that coherence criterion can be replaced, in the finite setting, by any bounded finite fragment. We prove that it cannot. For every $k\ge 1$, we construct a maximal standard frame whose shortest coherence violation has length exactly $2k+2$. Hence there is no uniform finite bound on the incoherence index of social decision frames, resolving Conjecture 5.7 stated by Moss and Pedersen (2026). The construction is geometric, in the sense that it proceeds via orthogonality and dimension in rational vector spaces, and self-contained: it isolates a symmetric family of half-sized voting blocs and extends it to a maximal frame in which every shorter balanced obstruction is excluded. Along the explicit infinite sequence of universe sizes obtained in the construction, this also establishes the middle-layer family predicted by Conjecture B.25 by Moss and Pedersen (2026). Together with the soundness and completeness theorem for the Moss-Pedersen minimal logic for strict majorities, this establishes that measurable social decision frames are not finitely axiomatizable in that language.
Learning to Erase Private Knowledge from Multi-Documents for Retrieval-Augmented Large Language Models
arXiv:2504.09910v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) is a promising technique for applying LLMs to proprietary domains. However, retrieved documents may contain sensitive knowledge, posing risks of privacy leakage in generative results. Thus, effectively erasing private information from retrieved documents is a key challenge for RAG. Unlike traditional text anonymization, RAG should consider: (1) the inherent multi-document reasoning may face de-anonymization attacks; (2) private knowledge varies by scenarios, so users should be allowed to customize which information to erase; (3) preserving sufficient publicly available knowledge for generation tasks. This paper introduces the privacy erasure task for RAG and proposes Eraser4RAG, a private knowledge eraser which effectively removes user-defined private knowledge from documents while preserving sufficient public knowledge for generation. Specifically, we first construct a global knowledge graph to identify potential knowledge across documents, aiming to defend against de-anonymization attacks. Then we randomly split it into private and public sub-graphs, and fine-tune Flan-T5 to rewrite the retrieved documents excluding private triples. Finally, PPO algorithm optimizes the rewriting model to minimize private triples and maximize public triples retention. Experiments on four QA datasets demonstrate that Eraser4RAG achieves superior erase performance than GPT-4o.
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
arXiv:2505.12843v2 Announce Type: replace Abstract: Reinforcement Learning from Human Feedback (RLHF) relies on reward models to align large language models with human preferences. However, RLHF often suffers from reward hacking, wherein policy learning exploits flaws in the trained reward model to maximize reward scores without genuinely aligning with human preferences. A significant example of such reward hacking is length bias, where reward models usually favor longer responses irrespective of actual response quality. Previous works on tackling length bias have notable limitations, these approaches either mitigate bias without characterizing the bias form, or simply assume a linear length-reward relation. To accurately model the intricate nature of length bias and facilitate more effective bias mitigation, we propose FiMi-RM (Bias Fitting to Mitigate Length Bias of Reward Model), a framework that autonomously learns and corrects underlying bias patterns. Our approach consists of three stages: First, we warm up by training a standard reward model which inherently contains length bias. Next, we deploy a lightweight fitting model to capture the non-linear relation between length and reward. Finally, we incorporate this learned relation into the reward model, effectively decoupling length from reward while preserving preference modeling capabilities. Experimental results demonstrate that FiMi-RM achieves a more balanced length-reward distribution. Furthermore, when applied to alignment algorithms such as Direct Preference Optimization (DPO) and Best-of-N (BoN), our debiased reward model improves length-controlled win rate and reduces verbosity without compromising its performance.
A Flow-rate-conserving CNN-based Domain Decomposition Method for Blood Flow Simulations
arXiv:2509.15900v2 Announce Type: replace Abstract: This work aims to predict blood flow with non-Newtonian viscosity in stenosed arteries using convolutional neural network (CNN) surrogate models. An alternating Schwarz domain decomposition method is proposed which uses CNN-based subdomain solvers. A universal subdomain solver (USDS) is trained on a single, fixed geometry and then applied for each subdomain solve in the Schwarz method. Results for two-dimensional stenotic arteries of varying shape and length for different inflow conditions are presented and statistically evaluated. One key finding, when using a limited amount of training data, is that incorporating a physics-aware constraint, as, in our case, flow rate conservation, into the USDS improves the prediction accuracy and convergence behavior of the Schwarz method compared to a purely data-driven USDS. As the USDS is a data-driven, inexact subdomain solver, admissible parameter ranges for the geometry and inflow configurations must be defined and tested.
Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation
arXiv:2509.22193v2 Announce Type: replace Abstract: Distilling reasoning traces from strong teacher models has become the standard recipe for building capable small language models. Yet reasoning traces are 5-20$\times$ longer than standard instruction fine-tuning (IFT) outputs, meaning every practitioner who chooses reasoning distillation implicitly forgoes training a larger IFT model on the same compute budget. Whether this trade-off is worthwhile remains unaddressed. We study it with a controlled experiment: a single teacher generates paired IFT and reasoning outputs for identical prompts by toggling only its reasoning mode, isolating supervision format as the sole variable. Training students at five scales (0.5B to 14B) and evaluating on 18 benchmarks, we find that at matched FLOPs, IFT lies on or near the Pareto frontier across the majority of configurations. Reasoning reaches the Pareto frontier only on open-ended tasks at 7B and above. Even there, a sequential curriculum mixing just 25-50\% reasoning data with IFT captures most of the accuracy benefit at far lower compute cost.
ScalingAR: Scaling Confidence for Autoregressive Image Generation
arXiv:2509.26376v3 Announce Type: replace Abstract: Test-time strategies have shown remarkable success in improving large language models, but their application to next-token prediction (NTP) autoregressive (AR) image generation remains largely underexplored. Existing test-time scaling (TTS) methods for visual autoregressive models (VAR) rely on frequent partial decoding and external reward models, which are inefficient and often ineffective for NTP-based image generation due to the inherent instability of intermediate decoding results. To address these limitations, we propose ScalingAR, a novel test-time scaling framework tailored for NTP-based AR image generation. ScalingAR introduces token entropy as a confidence signal and operates at two complementary levels: (i) Profile Level, integrates intrinsic uncertainty and conditional utilization into a unified confidence state, and (ii) Policy Level, leverages this state for adaptive trajectory pruning and dynamic guidance scheduling. Without requiring early decoding or auxiliary rewards, ScalingAR achieves significant improvements across diverse benchmarks. Experiments show that ScalingAR (I) improves base models by $12.5\%$ on GenEval and $15.2\%$ on TIIF-Bench, (II) reduces visual token consumption by $62.0\%$ while outperforming baselines, and (III) enhances robustness, mitigating performance degradation by $26.0\%$ in challenging scenarios. These results establish ScalingAR as a robust and efficient test-time scaling solution for autoregressive image generation.
Sampling Strategies for Robust Universal Quadrupedal Locomotion Policies
arXiv:2510.07094v2 Announce Type: replace Abstract: This work focuses on sampling strategies of configuration variations for generating robust universal locomotion policies for quadrupedal robots. We investigate the effects of sampling physical robot parameters and joint proportional-derivative gains to enable training a single reinforcement learning policy that generalizes to multiple parameter configurations. Three fundamental joint gain sampling strategies are compared: parameter sampling with (1) linear and polynomial function mappings of mass-to-gains, (2) performance-based adaptive filtering, and (3) uniform random sampling. We improve the robustness of the policy by biasing the configurations using nominal priors and reference models. All training was conducted using the RaiSim simulation environment, tested in simulation on a range of diverse quadrupeds, and zero-shot deployed onto hardware using the ANYmal quadruped robot. Compared to multiple baseline implementations, our results demonstrate the need for significant joint controller gains randomization for robust closing of the sim-to-real gap.
Epistemic Bias Injection: Manipulating LLM Opinion via Selective Context Retrieval
arXiv:2512.00804v3 Announce Type: replace Abstract: When answering user queries, LLMs often retrieve knowledge from external sources stored in retrieval-augmented generation (RAG) databases. These are often populated from unvetted sources, e.g. the open web, and can contain maliciously crafted data. This paper studies attacks that can manipulate the context retrieved by LLMs from such RAG databases. Prior work on such context manipulation primarily injects false or toxic content, which can often be detected by fact-checking or linguistic analysis. A more subtle threat, which we call epistemic bias injection (EBI), is where adversaries inject factually correct yet epistemically biased passages that systematically favor one side of an open-ended issue. Although linguistically coherent and truthful, such adversarial passages effectively crowd out alternative viewpoints during retrieval from the RAG and push LLM outputs towards an attack-desired stance. As a core contribution, we propose a novel characterization of the problem: We give a geometric metric that quantifies stance polarity and epistemic bias. This metric can be computed directly on embeddings of text passages. Leveraging it, we construct EBI attacks and develop a lightweight prototype defense called BiasDef for them. We evaluate them both on a comprehensive benchmark constructed from public question answering datasets. Our results show that: (1) the proposed attack induces significant stance polarity shifts, effectively evading existing retrieval-based sanitization defenses, and (2) BiasDef substantially reduces adversarial retrieval and epistemic bias in LLM's answers. Overall, this demonstrates the new threat as well as the ease of employing epistemic bias metrics for filtering in RAG-enabled LLMs.