Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation
arXiv:2607.07669v1 Announce Type: new Abstract: Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-leaning English, leaving dialectal generation, the harder half of the problem, largely unaddressed. We introduce \textbf{DiaLLM}, which continually pretrains three open-weight language model families on the International Corpus of English and applies implicit and explicit post-training paradigms, each combined with three model alignment strategies, giving the first controlled comparison of these components across Australian, Indian, and Northern British English. Our results reveal that dialectal robustness and generation are \emph{dissociated}: benchmarks are shaped by continual pretraining and SFT, while alignment visibly reshapes generation in ways benchmarks do not capture. Explicit variety-targeted adaptation produces output reliably recognised as dialectal and preferred over broad alignment, yet the method that most aggressively optimises the dialectal reward is not preferred by human evaluators. Independent linguistic analysis corroborates this reward-quality gap, most clearly on two of the three families. No single alignment method dominates, and closing the gap will require richer reward designs and continued investment in dialectal resources. We release all code, checkpoints, and preference datasets.
Predicting Multi-Order Magnetic Polariton Resonances for Radiative Properties Tailoring by Distributed Circuit Model
arXiv:2607.07657v1 Announce Type: new Abstract: Surface plasmon polaritons (SPPs) and magnetic polaritons (MPs) are fundamental resonance modes that are widely used to tailor the thermal radiation properties of micro/nanostructured metamaterials. Lumped circuit models (LCMs) are usually constructed empirically to describe the MP resonance conditions, and different LCMs have to be constructed for different orders of MPs, but these are difficult to be built for high-order MP modes due to the complex electromagnetic field distribution. This work proposes a new type of circuit model, distributed circuit model (DCM), to describe and predict multi-order MP resonances inside the structure based on the minimum total impedance condition. This allows both fundamental and high-order MP resonances to be predicted with a unified circuit, significantly simplifying the analysis of high-order MPs. More importantly, the DCM shares a similar and clear physical picture as the LCM for describing MPs. The MP resonance conditions for four typical structures are derived. Theoretical predictions based on DCMs are compared with and validated by rigorous numerical simulations. This study deepens the understanding and facilitates the design of MP-based thermal radiation metamaterials.
Hadronic vacuum polarization in hydrogen-like atoms and ions amid the interplay of recoil and finite-size effects
arXiv:2607.07658v1 Announce Type: new Abstract: Hadronic vacuum polarization (hVP) enters simple atomic systems at a level that is small yet decisive for the precision spectroscopy now underway. We evaluate the hVP contributions to the Lamb shift and the hyperfine splitting (HFS) in ordinary and muonic hydrogen (H and $\mu$H) and hydrogen-like helium-3 ions ($^3$He$^+$ and $\mu^3$He$^+$), using the dispersive data-driven approach and state-of-the-art empirical parametrizations of the $R$ ratio. At the centre of the analysis is the interplay of recoil and finite-size effects: the recoil corrections that dominate the HFS in muonium (Mu), where both constituents are pointlike, are shown to be suppressed by the nuclear elastic form factors (FFs). Our results for the leading hVP contribution to the Lamb shift agree with the literature within uncertainties. Furthermore, we present a first evaluation of the subleading $O(Z^5\alpha^6)$ hVP-finite-size correction, which is by no means negligible in $\mu^3$He$^+$. Our results for the hVP contribution to the HFS deviate significantly from all previous evaluations. For the ground-state HFS, we obtain $2.153(11)~\mu$eV in $\mu$H and $-15.19(57)~\mu$eV in $\mu^3$He$^+$, as well as $0.0860(4)~$kHz and $-0.476(17)~$kHz in ordinary H and $^3$He$^+$, respectively. Notably, our result for $\mu$H differs from previous evaluations by roughly ten times the experimental precision anticipated by the upcoming CREMA and FAMU measurements.
Dynamic Object Detection and Tracking in Construction: A Fisheye Camera and LiDAR Sensor Fusion Model
arXiv:2607.06896v1 Announce Type: new Abstract: Robust dynamic object detection and tracking are essential for enabling robots to operate safely and effectively alongside humans in complex environments such as construction sites. While LiDAR-based SLAM and occupancy grid methods offer viable solutions for detecting and tracking motion, many state-of-the-art 3D vision approaches rely heavily on pre-trained neural networks and require additional post-processing to identify moving objects. Sensor fusion techniques, combining the precision of LiDAR with the semantic richness of RGB imagery, offer a promising alternative. In this work, we present a novel framework that enhances a quadruped robot equipped with a LiDAR sensor and an upward-facing fisheye camera for real-time dynamic object detection and tracking. After identifying moving objects within a registered point cloud, our method assigns semantic labels by projecting 3D coordinates onto a 2D cylindrical panorama, aligning with real-time image-based detections for observation update of the Kalman filter. The proposed system demonstrates high precision, simplicity, and robustness, particularly in handling objects transitioning between dynamic and static states, thus it is well-suited for deployment in real-world construction environments.
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF
arXiv:2607.07693v1 Announce Type: new Abstract: Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences. However, applying RLHF to diffusion models remains highly feedback inefficient, as existing approaches typically require large amounts of human or reward model evaluations. This limitation reduces the practicality of diffusion RLHF in realworld settings where feedback is the primary bottleneck. In this paper, we propose two complementary strategies that substantially improve the feedback efficiency of diffusion RLHF while preserving generalization to unseen prompts. Our key observation is that reward information in diffusion trajectories is unevenly distributed: not all denoising timesteps or trajectories contribute equally to learning from a reward signal. By emphasizing informative timesteps and trajectories during optimization, we obtain more effective gradient updates. First, we introduce a per-timestep weighting scheme that reweights denoising steps during policy optimization. We theoretically connect this weighting to the optimal convergence properties of proximal policy optimization (PPO) and approximate the resulting weighting trend empirically. Second, we introduce a replay mechanism that prioritizes informative trajectories, enabling the model to reuse past samples instead of repeatedly querying new rewards. Together, these strategies significantly improve the feedback efficiency of diffusion RLHF. Under identical hyperparameter settings, our approach achieves up to a 6$\times$ improvement in sample efficiency compared to widely used diffusion RLHF baselines.
Hypergraph backboning
arXiv:2606.00893v2 Announce Type: replace Abstract: Hypergraphs provide a natural framework for describing complex networked systems with higher-order, non-dyadic interactions. Due to their high dimensionality and often redundant structure, a key challenge is to develop methods that simplify hypergraph representations while preserving the essential structure of interactions. Here we present a principled, efficient, and non-parametric information-theoretic method for pruning nested and/or redundant structures in hypergraphs, enabling a minimal representation of higher-order interactions in the presence of local heterogeneity. Our approach naturally extends to weighted hypergraphs, where higher-order topology and hyperedge weights combine to identify the system's structural backbone. We validate the method on controlled synthetic hypergraphs and apply it to empirical datasets from diverse domains, demonstrating substantial sparsification without loss of core structural information.
Taming nonlinear energy diffusion: The case of time-crystal energy condensates
arXiv:2607.07325v1 Announce Type: cross Abstract: We study a bulk-driven nonlinear variant of the Kipnis-Marchioro-Presutti model of stochastic energy diffusion in which local collisions are biased to induce a net energy flow, resembling the effect of an external field. Starting from the microscopic master equation, we derive the hydrodynamic description of the driven system via a local equilibrium approximation, obtaining explicit expressions for the energy current and the associated diffusivity and mobility transport coefficients, which are nonlinear functions of the local energy density. We test our findings in kinetic Monte Carlo simulations of the model and, as a proof of concept, we demonstrate the versatility of this driving mechanism to control nonlinear energy transport by inducing time-crystalline phases. In particular, we show that appropriately designed packing fields induce the spontaneous formation of traveling energy condensates, exhibiting robust long-range temporal order reminiscent of continuous time crystals. Our results provide a simple yet powerful framework to study bulk-driven nonlinear energy diffusion in stochastic many-body systems, offering a bridge between microscopic dynamics, macroscopic transport, and controlled spatiotemporal order.
VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models
arXiv:2606.27646v2 Announce Type: replace Abstract: Conventional machine-vision pipelines typically rely on high-quality optics that produce clean, human-interpretable images, and optical design has therefore been driven by image-level criteria such as resolution, aberration correction, and pixel fidelity. However, such optics are often impractical for size-, cost-, or form-factor-constrained applications, where compact meta-optics offer an attractive alternative but operate under strict physical efficiency limits. We propose CODA, a co-design framework that optimizes a continuous-density meta-optic front-end for frozen-model recognition using differentiable image formation and adjoint-gradient updates of Maxwell-based simulations. CODA directly optimizes the cross-entropy loss of a fixed zero-shot CLIP classifier without learned reconstruction, image signal processing, or image-fidelity auxiliary objectives. In a two-dimensional simulated imaging benchmark on ImageNet-100, CODA improves CLIP ViT-L/14 zero-shot accuracy from 53.75 $\pm$ 3.57$\%$ with a focal-concentration baseline to 65.41 $\pm$ 3.99$\%$. The optimized optics further transfer without re-optimization across CLIP, SigLIP, and DINOv2 on ImageNet-100, CIFAR-100, and Food-101. These results demonstrate that, under constrained meta-optic imaging, downstream recognition can be improved by aligning optical design with frozen vision-model objectives rather than conventional image-formation criteria.
CogAdapt: Adapting Clinical ECG Foundation Models for Wearable Cognitive Load Assessment
arXiv:2605.22774v4 Announce Type: replace Abstract: Assessing cognitive load continuously and at low latency would help adaptive human-computer interaction, but it remains hard because labeled data are scarce and models generalize poorly across subjects. Recent ECG foundation models, pre-trained on millions of clinical diagnostic ECG recordings, yet they do not apply directly to wearable devices when the sensor configuration and the task both differ. We present CogAdapt, a framework that adapts a clinical ECG foundation model to wearable cognitive load assessment. CogAdapt has two parts. LeadBridge is a learnable adapter that maps 3-lead wearable signals to a 12-lead-compatible representation. ProFine is a progressive fine-tuning strategy that unfreezes encoder layers in stages while limiting representational drift in the pre-trained model. On two public datasets (CLARE and CL-Drive) under leave-one-subject-out cross-validation, CogAdapt reaches macro-F1 of 0.626 and 0.768, improving over from-scratch baselines by 11.2 and 16.1 percentage points. The results show that a clinical ECG pretraining can support subject-independent cognitive load assessment from wearable sensors.
Virtual-Memory Powersort
arXiv:2605.27147v2 Announce Type: replace Abstract: We give a more space-efficient implementation of adaptive mergesort: Virtual-Memory Powersort. Using internal buffering techniques, we significantly reduce the memory consumption of the algorithm; specifically, for sorting $n$ objects the required buffer area is reduced from space for $n/2$ objects to $O(\sqrt{n \log n})$ objects. While this space-efficiency can be achieved (indeed reduced to $O(1)$) conceptually very easily with known inplace merging algorithms, using these as a drop-in replacement for the standard merge algorithm incurs a substantial slow-down. Virtual-Memory Powersort, by contrast, uses the same number of moves and comparisons as previous Powersort implementations up to an additive $O(n)$ term. We report on an empirical running-time study comparing our implementation against other Powersort variants and state-of-the-art stable sorting methods, demonstrating that almost in-place stable sorting can be achieved with negligible overhead in many scenarios.
From Noisy Traces to Root Causes: Structural Trajectory Analysis and Causal Extraction for Agent Optimization
arXiv:2607.07702v1 Announce Type: new Abstract: The optimization of long-horizon agents increasingly relies on reflection-based mechanisms, where a large language model (LLM) acts as an optimizer to diagnose agent failures and improve agent policies. However, real execution traces are difficult to use directly for optimization: large trace collections are often redundant and heterogeneous, making optimization inefficient and prone to overfitting to low-value failures; meanwhile, each individual trajectory also contains many irrelevant steps, while naive context reduction methods such as truncation or sliding windows can discard causally important evidence and produce misleading optimization signals. To resolve this dilemma, we introduce STRACE (Structural TRajectory Analysis and Causal Extraction), a framework that constructs high signal-noise optimization contexts for more precise and effective optimization. At the batch level, STRACE mines failure patterns to filter redundant traces and retain representative failures; within each selected trace, it performs causal localization over a textual dependency graph to remove non-causal steps and identify the true root-cause module for optimization. Empirical results demonstrate that STRACE significantly outperforms standard context-filtering baselines. Notably, on a challenging formal verification task (VeruSAGE-Bench), it successfully optimizes human-expert designed agents, delivering $1.4\times$ success-rate improvement (42.5% to 58.5%). The code is available at https://github.com/moomight/STRACE .
What Semivalues Cannot See: The Information Content of Anonymous Marginal Values
arXiv:2607.07013v1 Announce Type: new Abstract: The semivalue family shares a common kernel: games invisible to every anonymous marginal value at once, nonzero from four players (Kleinberg and Weiss, 1985; Amer, Derks and Gim\'enez, 2003). Crisman and Orrison (2015) ask what useful structure this kernel carries; this paper gives a concrete answer. In Harsanyi-dividend coordinates the joint information of all semivalues is exactly each player's total synergy at each coalition size, so the kernel is synergy arranged in closed circuits. We prove: order-$\le d$ mixed-difference audits recover exactly the degree-$\le d$ dividend-slice harmonics, with closed-form dimension at every rung; nonzero blind games fail superadditivity, monotonicity, and core existence, yet distinct convex games with identical values under every semivalue exist from four players, with exact perturbation thresholds; the positive weighted Shapley family attains full information $2^n-1$, so anonymity is the binding axiom within the marginal framework; and a coalition of size $c$ defeats every audit of order $\le d$ precisely when $c\ge2d+2$, within the convex class for small perturbations. An exhaustive census at $n=5$ exhibits non-isomorphic voting rules with identical values under every semivalue power index; no weighted game participates in any collision, prompting a swing-rigidity conjecture. Measured against the theory, classical cooperative games sit at $0.90$ to $1.00$ visibility to the family versus $0.089$ for a random game.
A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving
arXiv:2607.07103v1 Announce Type: new Abstract: Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic safety. These scenarios are severely under-represented in naturalistic driving data, and existing trajectory and language-augmented datasets seldom provide high-risk event labels, semantic annotations, and verifiable safety signals. Here we present K-Risk, a knowledge-augmented dataset that combines structured driving trajectories with large language model generated semantic annotations for safety-critical driving scenarios. K-Risk integrates 20 human-driven and autonomous-vehicle trajectory datasets from Europe, China, and the United States, covering highways, urban freeways, intersections, and roundabouts. Using a unified risk-centric extraction pipeline, K-Risk curates 31,398 high-risk events, together with a 1,036-event extreme subset of near-collision cases. Each event is released as a synchronized trajectory, metadata, and language triplet containing structured scenario descriptions, abnormal-behavior notifications, and, for a representative subset, causal risk analyses and action recommendations validated through a closed-loop simulator with iterative reflection. By combining multi-dimensional risk annotations, interpretable language supervision, and verifiable decisions, K-Risk bridges structured traffic trajectories, semantic reasoning, and decision supervision, providing a standardized foundation for developing and evaluating next-generation risk-aware autonomous driving agents.
CompoVista: A Composition-Graph-Based Visual Analytics System for Compositional Analysis of Traditional Chinese Paintings
arXiv:2607.07105v1 Announce Type: new Abstract: Composition in Traditional Chinese Paintings (TCPs) carries spatial, narrative, and cultural-aesthetic meaning. Systematic compositional analysis is therefore important for understanding their visual language and artistic meaning. Traditional compositional analysis is mainly qualitative and interpretation-driven. It supports close reading of individual paintings, but it is difficult to discover, compare, and verify compositional patterns across large painting collections. To better understand these challenges, we conducted a literature review and in-depth interviews with two art historians. Based on these findings, we introduce the Composition Graph, a scene-graph-based representation for TCP composition. It models a painting through four layers: entities, relations, void space, and context. Based on this representation, we develop CompoVista, a canvas-based visual analytics system for composition-oriented exploration of TCPs. CompoVista allows art historians to construct and revise format-aware painting cohorts through visual queries and context queries. It also supports cohort-level inspection of entity distributions and relations, comparison of compositional differences across cohorts, and tracing aggregate patterns back to painting-level evidence.We evaluated CompoVista through a task-based user study with 12 domain participants, two case studies, and expert interviews. The results show that CompoVista supports composition-oriented cohort construction, pattern discovery, iterative refinement, and evidence inspection. The evaluation also reveals future needs, including clearer result explanations, fuzzier composition queries, and stronger exploration history management. Our work contributes a composition-specific structured representation and an integrated visual analytics workflow for studying TCP composition at collection scale.
Reaction-Network-Level Discovery of Ammonia Synthesis Catalysts via Ten-Million-Scale Generative Exploration
arXiv:2606.22926v2 Announce Type: replace Abstract: Catalyst discovery for ammonia synthesis is inherently a reaction-network challenge because catalytic performance is governed not by a single adsorbed intermediate, but by a surface's orchestrated compatibility with multiple distinct intermediates across competing dissociative and associative pathways. However, navigating ultra-large chemical spaces under such multi-intermediate constraints remains a formidable bottleneck for conventional screening workflows. Here, we report a reaction-network-level catalyst discovery framework driven by ten-million-scale generative exploration. By coupling adsorbate-specific generative Transformers with high-throughput machine learning potentials, we systematically map the structure-property landscapes of four critical intermediates (N*, NH*, NNH*, and HNNH*). Scale-dependent overlap analysis shows that the full four-intermediate compatibility space remains strongly under-sampled at conventional 105-106 generative scales, emerging exclusively under ten-million-scale exploration. By generating approximately 15 million configurations per adsorbate, followed by structural compression and machine-learning-potential predictions, we identified 279 highly potential target materials. This sparse compatibility space successfully recovers traditional Fe- and Ru-based motifs while uncovering previously unexplored catalyst families. Representative DFT calculations validate pathway-dependent mechanisms: Fe-V emerges as a dissociative-pathway lead by significantly lowering the initial N2 dissociation barrier, whereas Al-Pd-Zr efficiently stabilizes associative intermediates as an associative-pathway lead. These findings establish multi-intermediate reaction-network compatibility as a robust criterion for discovering advanced catalysts from multi-million generative chemical spaces.
Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning
arXiv:2607.07117v1 Announce Type: new Abstract: In text-to-image in-context learning (T2I-ICL), a model has to infer a latent compositional pattern from fewshot demonstrations for generating a query image. Recent studies show that state-of-the-art multimodal large language models struggle with this setting, particularly due to limited compositional reasoning and sensitivity to prompt construction. In this work, we propose a Tree-of-Thoughts (ToT) reasoning framework for T2I-ICL that introduces a multi-stage reasoning and selection layer that generates, evaluates, and selects among multiple candidate hypotheses before constructing the final prompt for image synthesis. By exploring alternative reasoning branches and selecting a coherent interpretation, the proposed approach mitigates prompt ambiguity and compositional errors. We implement the proposed approach in a complete ToT-T2IICL inference pipeline and evaluate it on the CoBSAT benchmark. Both qualitative and quantitative results show that structured multi-branch reasoning leads to more consistent and semantically aligned image generation compared to baseline and Chain-of-Thought prompting strategies, without any additional training or fine-tuning.
Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer
arXiv:2510.24108v2 Announce Type: replace Abstract: Human demonstrations are widely considered the cornerstone of end-to-end (E2E) autonomous driving despite human demonstration's scarcity for long-tail and safety-critical scenarios. Nonetheless, current E2E autonomous driving (AD) training paradigms continue to rely on human demonstrations. Imitation learning (IL) requires human demonstrations for training, whereas reinforcement learning (RL) has emerged as a promising alternative to reduce this dependency. However, most existing RL methods for E2E AD still rely implicitly on human demonstrations. A pure rewards-based RL method can overcome the need for human demonstrations, but general RL policy gradient methods suffer from the cold-start problem. In this paper, we propose ZTRS (Zero-human demonstration end-to-end autonomous driving with TRajectory Scorer) - a complete RL-based E2E planning paradigm trained solely on real-world images and rule-based rewards, entirely without human demonstration. Through our proposed Exhaustive Policy Optimization (EPO), a policy gradient variant tailored for enumerable trajectory actions and dense supervision, ZTRS enables the model to generalize better to long-tail driving scenarios. We demonstrate this generalization through our SOTA performance against IL approaches on both long-tail Navhard and closed-loop HUGSIM datasets. Project page: https://zhenxinli.net/ZTRS/.
Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
arXiv:2511.13726v2 Announce Type: replace Abstract: We propose RT (Refine Thought), a method that can enhance the semantic reasoning ability of text embedding models. The method obtains the final semantic representation by running multiple forward passes of the text embedding model. Experiments show that RT achieves significant improvements on semantic reasoning tasks in BRIGHT and the person-job matching benchmark PJBenchmark, while maintaining consistent performance on general-purpose semantic understanding tasks such as C-MTEB. Our results indicate that RT is effective because it further activates the semantic reasoning ability learned during pretraining by decoder-only text embedding models (e.g., Qwen3-Embedding-8B). RT can be seen as a test-time inference method.
Data-Driven Soft Labeling Scales DNA Read Classification to Whole-Body Cell-Type Deconvolution
arXiv:2607.04987v2 Announce Type: replace Abstract: Cell-type deconvolution, the task of estimating the proportions of constituent cell types in a heterogeneous biological sample, is a core problem in computational biology. Methods that rely on epigenetic marks such as DNA methylation typically operate on aggregated methylation estimates, discarding the pattern-level information carried by individual DNA reads. Existing read-level approaches that exploit this information are scarce, and all remain restricted to few-class settings; scaling them further is an open problem because, at scale, non-discriminative reads dominate and hard labels conflict with the many-to-many mapping between methylation patterns and cell types, preventing classifier convergence. To overcome this, we propose data-driven soft labels that estimate the conditional cell-type distribution for each read, and integrate this scheme into Syto, a new modular framework for read-level classification-based deconvolution. On a whole-body atlas of 39 human cell types, Syto reduces MSE by 2.56$\times$ over SoTA, with gains transferring to an out-of-distribution dataset spanning 16 tissues. Syto lays the foundation for modeling increasingly large cell-type panels, with improved applications in biology and healthcare. The proposed soft-labeling scheme is further translatable to any setting with a many-to-many signal-to-label mapping.
BifrostUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation
arXiv:2605.03452v2 Announce Type: replace Abstract: High-quality demonstration data are essential for humanoid robot skill learning, especially for whole-body behaviors that require coordinated perception, locomotion, and manipulation. Existing data-collection methods largely rely on robot teleoperation, which is constrained by hardware accessibility, operator expertise, and limited efficiency. Inspired by the Universal Manipulation Interface (UMI), we propose BifrostUMI, a portable and robot-free framework for humanoid whole-body data collection. BifrostUMI uses lightweight VR devices and UMI-inspired grippers to collect sparse human keypoint trajectories, wrist-view observations, and gripper actions. These demonstrations train a high-level policy to predict future keypoints, which are retargeted to robot-native whole-body references and executed by a whole-body controller. Experiments in five real-world scenarios demonstrate the effectiveness of the proposed framework and validate the collected demonstrations for transferable humanoid whole-body skill learning.
Understanding Two-Layer Neural Networks with Smooth Activation Functions
arXiv:2507.14177v2 Announce Type: replace Abstract: This paper aims to understand the training solution, which is obtained by the back-propagation algorithm, of two-layer neural networks whose hidden layer is composed of the units with smooth activation functions, including the usual sigmoid type most commonly used before the advent of ReLUs. The mechanism contains four main principles: construction of Taylor series expansions, strict partial order of knots, smooth-spline implementation and smooth-continuity restriction. The universal approximation for arbitrary input dimensionality is proved and the explanation of training solutions is given. Through the principles proposed, the mystery of ``black box'' of the solution space is largely revealed. The new proofs employed also enrich approximation theory.
Macroscopic position-position entanglement by photon recoil in Rydberg atoms
arXiv:2607.07167v1 Announce Type: cross Abstract: Entanglement between two spatially separate matter particles can be generated via many means and often resides in the internal states of particles. Here, via Rydberg blockade in two spatially separate neutral atoms, we find that the photon recoil in Rydberg excitation can push one atom microns away provided the other atom exerts a state-dependent Rydberg-mediated blockade. When the atoms are recaptured by optical traps, a position-position entangled state between two spatially separate atoms can emerge. This realizes a Bell state of two atoms, where the entanglement exists in the position of each atom and the distance between the two possible locations of each atom can be in the hundred-micron regime.
Reinforcement Federated Learning Method Based on Adaptive OPTICS Clustering
arXiv:2306.12859v3 Announce Type: replace Abstract: Federated learning is a distributed machine learning technology, which realizes the balance between data privacy protection and data sharing computing. To protect data privacy, feder-ated learning learns shared models by locally executing distributed training on participating devices and aggregating local models into global models. There is a problem in federated learning, that is, the negative impact caused by the non-independent and identical distribu-tion of data across different user terminals. In order to alleviate this problem, this paper pro-poses a strengthened federation aggregation method based on adaptive OPTICS clustering. Specifically, this method perceives the clustering environment as a Markov decision process, and models the adjustment process of parameter search direction, so as to find the best clus-tering parameters to achieve the best federated aggregation method. The core contribution of this paper is to propose an adaptive OPTICS clustering algorithm for federated learning. The algorithm combines OPTICS clustering and adaptive learning technology, and can effective-ly deal with the problem of non-independent and identically distributed data across different user terminals. By perceiving the clustering environment as a Markov decision process, the goal is to find the best parameters of the OPTICS cluster without artificial assistance, so as to obtain the best federated aggregation method and achieve better performance. The reliability and practicability of this method have been verified on the experimental data, and its effec-tiveness and superiority have been proved.
Disturbance-aware Motion Planning for Over-actuated Underwater Vehicles Exploiting Actuation Redundancy for High-fidelity 3D Reconstruction
arXiv:2607.07139v1 Announce Type: new Abstract: Underwater robots often operate near delicate targets where high-power thrusters resuspend sediments and induce turbulence, degrading image quality at the sensor input. Conventional controllers optimize vehicle-centric objectives, such as tracking and stability, without accounting for the impact of actuation on sensing. We address this actuation-to-perception coupling by exploiting redundancy in over-actuated platforms. For an eight-thruster ROV, multiple thrust allocations can yield the same motion; we search this null space to minimize predicted disturbance in a task-relevant target region while enforcing motion constraints. Our method uses a control-oriented thruster-wake proxy derived from actuator-disk theory with directional attenuation and validated by PIV ($R^2 = 0.99$ near the wake axis; $R^2 > 0.82$ in the primary wake region), together with a real-time redundancy-resolving allocator running at 10 Hz (45 ms/solve). Across 440 trials, the approach reduces target-region particle velocity by 67% ($p < 0.001$), improves 3D reconstruction RMSE by 55% versus a disturbance-unaware baseline ($1.9 \pm 0.4$ mm vs. $4.3 \pm 1.8$ mm), and achieves a 98.5% reconstruction success rate. The framework supports autonomous scanning, which is quantitatively evaluated, and operator-assisted inspection, which is demonstrated in the supplementary materials.
Unraveling Machine Behavior by Multi-Level Bias Analysis and Detection: Methodology and Application to Computer Vision
arXiv:2607.07236v1 Announce Type: new Abstract: This study investigates the presence and propagation of bias within Neural Networks through a comprehensive multi-level analysis spanning the learned latent space, layer activations, and the network's parameters. Based on this taxonomy, we propose three bias detection approaches: 1) SpaceBias (new method), which characterizes the latent space prior to the final classification layer using neighbor-probability distributions and quantifies bias with the two-sample Kolmogorov-Smirnov test on the per-group distributions. 2) ActivationBias (extension of the existing method InsideBias), which analyzes the activations of neural network filters and quantifies bias via a Mann-Whitney U test, based on the observed fact that underrepresented groups exhibit lower activation levels in the final convolutional layers. 3) WeightBias (extension of the existing method IFBiD), which uses a secondary neural network trained to identify biased patterns directly in the parameters of task-specific models. Unlike conventional methods, which assess neural network outcomes and treat the model as a black box, our proposed techniques provide insight into how biases manifest within the network architecture itself at different levels, offering a more nuanced and detailed understanding. Experiments are conducted on two complementary applications: gender classification in the DiveFace dataset (72,000 face images) and digit classification on a colored-MNIST benchmark with controlled bias severity. In total, more than 127,000 models with varying degrees and types of bias were trained and evaluated. The severity sweep shows that the internal disparity, and with it the detection performance, decreases smoothly as the training distribution approaches balance. The results highlight the importance of methods that provide deeper insight into the behavior of AI models.