Forskningsradar

Science Journals

Peer-reviewade publikationer — 52194 artiklar

Three-Dimensional and Selective Displacement Sensing of a Levitated Nanoparticle via Spatial Mode Decomposition
arXiv:2409.08827v4 Announce Type: replace Abstract: We propose and experimentally demonstrate a novel detection method that significantly improves the precision of real-time measurement of the three-dimensional displacement of a levitated dipolar scatterer. Our technique relies on spatial mode sorting of the light scattered by the levitated object, allowing us to selectively extract the position information of all translational degrees of freedom with minimal losses. To this end, we collect all the light back-scattered from a levitated nanoparticle using a parabolic mirror and couple it into a spatial mode sorter. We measure displacement sensitivities ($\sqrt{S_{\mathrm{imp}, x}}, \sqrt{S_{\mathrm{imp}, y}}, \sqrt{S_{\mathrm{imp}, z}}$) $=$ (1.7, 2.4, 1.0) $\times$ $10^{-14}$ m/$\sqrt{\mathrm{Hz}}$ below the zero-point motion ($x_{\mathrm{zpm}}, y_{\mathrm{zpm}}, z_{\mathrm{zpm}}$) $=$ (2.2, 2.4, 1.6) $\times$ $10^{-12}$ m of the levitated particle considered here. In the regime where environmental decoherence is not limited by gas collision we estimate that our method can reach measurement efficiencies of $(\eta_{^{\mathrm{tot}}}^{_{x}}, \eta_{^{\mathrm{tot}}}^{_{y}}, \eta_{^{\mathrm{tot}}}^{_{z}}) = (0.13, 0.18, 0.33) > 1/9$, which would enable the 3D motional quantum ground state of a levitated optomechanical system.
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
arXiv:2603.16947v2 Announce Type: replace Abstract: Although vision-language navigation (VLN) has progressed rapidly, zero-shot VLN in continuous environments (VLN-CE) remains highly challenging when using lightweight vision-language models (VLMs), whose limited reasoning capacity makes long-horizon navigation unreliable. In this paper, we propose LightZeroNav to tackle the three major bottlenecks when using lightweight VLMs in zero-shot VLN-CE,i.e.,information redundancy from multi-source inputs, inaccurate progress estimation caused by noisy textual memory, and task entanglement between action execution and stage transition. Using only RGB observations and a lightweight open-source Qwen3-VL-8B backbone, LightZeroNav achieves competitive performance with GPT-4o (~200B) without task-specific training, graph search, or waypoint predictors, demonstrating its effectiveness in zero-shot VLN-CE.
Robust and Resilient Soft Robotic Object Insertion with Compliance-Enabled Contact Formation and Failure Recovery
arXiv:2509.17666v2 Announce Type: replace Abstract: Object insertion tasks are prone to failure under pose uncertainty and environmental variation, often requiring manual fine-tuning or controller retraining. We present a novel approach for robust and resilient object insertion using a passively compliant soft wrist that enables safe contact absorption through large deformations, without high-frequency control or force sensing. Our method structures insertion as compliance-enabled contact formations, sequential contact states that progressively constrain degrees of freedom, and integrates automated failure recovery strategies. Our key insight is that wrist compliance permits safe, repeated recovery attempts; hence, we refer to it as compliance-enabled failure recovery. We employ a pre-trained vision-language model (VLM) that assesses each skill execution from terminal poses and images, identifies failure modes, and proposes recovery actions by selecting skills and updating goals. In simulation, our method achieved an 83% success rate, recovering from failures induced by randomized conditions, including grasp misalignments up to 5 degrees, hole-pose errors up to 20 mm, fivefold increases in friction, and unseen square/rectangular pegs, and we further validated the approach on a real robot. Project page is available at https://omron-sinicx.github.io/compliance-enabled-failure-recovery/.
Causal Attribution via Activation Patching
arXiv:2603.13652v2 Announce Type: replace Abstract: Attribution methods for Vision Transformers (ViTs) aim to identify image regions that influence model predictions, but producing faithful and well-localized attributions remains challenging. Existing attribution methods face several limitations, with gradient-based, relevance-propagation, and attention-based methods relying on local approximations, while perturbation or optimization-based methods intervene on inputs, tokens, or surrogates rather than internal patch representations. The key challenge is that class-relevant evidence is formed through interactions between patch tokens across layers; methods that operate only on input changes, attention weights, or backward relevance signals may therefore provide indirect proxies for patch importance rather than directly testing the predictive effect of contextualized patch representations. We propose Causal Attribution via Activation Patching (CAAP), which estimates the contribution of individual image patches to the ViT's prediction by directly intervening on internal activations rather than using learned masks or synthetic perturbation patterns. For each patch, CAAP inserts the corresponding source-image activations into a neutral target context over an intermediate range of layers and uses the resulting target-class score as the attribution signal. The resulting attribution map reflects the causal contribution of patch-associated internal representations on the model's prediction. The causal intervention serves as a principled measure of patch influence by capturing semantic evidence after initial representation formation, while avoiding late-layer global mixing that can reduce spatial specificity. Across multiple ViT backbones and standard metrics, CAAP consistently outperforms existing methods in various settings and produces more faithful and localized attributions.
OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism
arXiv:2603.14371v2 Announce Type: replace Abstract: Embodied AI agents increasingly require parallel execution of multiple tasks, such as manipulation, conversation, and memory construction, from shared observations under distinct time constraints. Recent Mixture-of-Transformers (MoT) Vision-Language-Action Models (VLAs) architecturally support such heterogeneous outputs, yet existing inference systems fail to achieve efficient multi-task parallelism for on-device deployment because of redundant computation and resource contention. We identify isolated KV cache management as the root cause. To address this, we propose unified KV cache management, an inference design that treats the KV cache as a first-class shared resource across tasks and over time. This abstraction enables two key optimizations: cross-task KV sharing eliminates redundant prefill of shared observations, while cross-frame continuous batching decouples variable-length language decoding from fixed-rate action generation across control cycles. We implement this design for $\pi_{0.5}$, a popular MoT VLA, and evaluate it on both NVIDIA GeForce RTX 4090 and Jetson AGX Thor, two representative platforms for on-device VLA inference. OxyGen achieves up to 3.7$\times$ speedup over isolated execution, delivering over 200 tokens/s language throughput and 70 Hz action frequency simultaneously without degrading action quality, and we further validate the gains on a real humanoid robot with on-board Jetson AGX Thor.
CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering
arXiv:2603.16091v3 Announce Type: replace Abstract: In factual question answering, many errors are not failures of access but failures of commitment: the system retrieves relevant evidence, yet still settles on the wrong answer. We present CounterRefine, a lightweight repair layer for short-form RAG that treats the first answer as a hypothesis to test. Given a draft, CounterRefine issues answer-conditioned expansion queries to retrieve candidate-specific evidence, then applies a constrained KEEP or REVISE refinement step whose proposed revisions are accepted only after deterministic validation. The design is intentionally narrow: it adds one evidence-gathering pass and one guarded refinement call rather than replacing the retriever or building a broad agentic system. On the full SimpleQA benchmark, CounterRefine improves a matched one-pass RAG baseline by up to 5.8 correct-rate points; in the full Claude trace, it changes only 5.6% of outputs, with 180 beneficial outcome changes and 8 harmful ones. These findings suggest a simple but important direction for knowledgeable foundation models: beyond accessing evidence, they should also be able to use that evidence to reconsider and, when necessary, repair their own answers.
Fast and Reliable Gradients for Deformables Across Frictional Contact Regimes
arXiv:2603.16478v2 Announce Type: replace Abstract: Differentiable simulation establishes the mathematical foundation for solving challenging inverse problems in computer graphics and robotics, such as physical system identification and inverse dynamics control. However, rigor in frictional contact remains the "elephant in the room." Current frameworks often avoid contact singularities via non-Markovian position approximations or heuristic gradients. This lack of mathematical consistency distorts gradients, causing optimization stagnation or failure in complex frictional contact and large-deformation scenarios. We introduce our unified fully GPU-accelerated differentiable simulator, which establishes a rigorous theoretical paradigm through: Long-Horizon Consistency: enforcing strict Markovian dynamics on a coupled position-velocity manifold to prevent gradient collapse; Unified Contact Stability: employing a mass-aligned preconditioner and soft Fischer--Burmeister operator for smooth frictional optimization; Robust Material Identification: resolving FEM singularities via a derived "Within-block Commutation" condition. Our experiments demonstrate our solver efficacy in bridging the Sim-to-Real gap, delivering precise, low-noise gradients in contact-rich tasks like dexterous manipulation and cloth folding. By mitigating the gradient instability issues common in conventional approaches, our framework significantly enhances the fidelity of physical system identification and control.
CoUn: Empowering Machine Unlearning via Contrastive Learning
arXiv:2509.16391v3 Announce Type: replace Abstract: Machine unlearning (MU) aims to remove the influence of specific "forget" data from a trained model while preserving its knowledge of the remaining "retain" data. Existing MU methods based on label manipulation or model weight perturbations often achieve limited unlearning effectiveness. To address this, we introduce CoUn, a novel MU framework inspired by the observation that a model retrained from scratch using only retain data classifies forget data based on their semantic similarity to the retain data. CoUn emulates this behavior by adjusting learned data representations through contrastive learning (CL) and supervised learning, applied exclusively to retain data. Specifically, CoUn (1) leverages semantic similarity between data samples to indirectly adjust forget representations using CL, and (2) maintains retain representations within their respective clusters through supervised learning. Extensive experiments across various datasets and model architectures show that CoUn consistently outperforms state-of-the-art MU baselines in unlearning effectiveness. Additionally, integrating our CL module into existing baselines empowers their unlearning effectiveness.
AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents
arXiv:2605.17933v1 Announce Type: new Abstract: Vision-language model (VLM) agents increasingly rely on memory-augmented reinforcement learning to reuse experience across long-horizon tasks, yet most existing frameworks store memory as text and depend on proprietary teacher models to summarize or refine it. This design is poorly matched to spatial decision making: geometric priors are compressed into lossy language, and sparse interaction is often supervised through delayed textual feedback rather than dense visually grounded signals. We argue that reusable experience for VLM agents should remain visually grounded. Based on this insight, we propose \textbf{AtlasVA}, a teacher-free visual skill memory framework that organizes memory into three complementary layers: spatial heatmaps, visual exemplars, and symbolic text skills. AtlasVA further evolves danger and affinity atlases directly from trajectory statistics and lightweight grid heuristics, and reuses these self-evolving atlases as potential-based shaping rewards for reinforcement learning. This unifies perception, memory, and optimization without external LLM supervision. Experiments on \textsc{Sokoban}, \textsc{FrozenLake}, 3D embodied navigation, and 3D robotic manipulation benchmarks show that AtlasVA consistently outperforms text-centric memory baselines and competitive VLM agents, with especially strong gains on spatially intensive tasks. Homepage: https://wangpan-ustc.github.io/AtlasvaWeb
AutoVecCoder: Teaching LLMs to Generate Explicitly Vectorized Code
arXiv:2605.17978v1 Announce Type: new Abstract: Vectorization via Single Instruction, Multiple Data (SIMD) architectures is a cornerstone of high-performance computing. To fully exploit hardware potential, developers often resort to explicit vectorization using intrinsics, as compiler-based auto-vectorization frequently yields suboptimal results due to conservative static analysis. While Large Language Models (LLMs) have demonstrated remarkable proficiency in general code generation, they struggle with explicit vectorization due to the scarcity of high-quality corpora and the strict semantic constraints of low-level hardware instructions. In this paper, we propose AutoVecCoder, a novel framework designed to empower LLMs with the capability of automated explicit vectorization. AutoVecCoder integrates two core components: VecPrompt, an automated data synthesis pipeline to inject domain-specific intrinsic knowledge; and VecRL, a reinforcement learning framework that aligns code generation with execution efficiency. AutoVecCoder-8B trained by this framework achieves state-of-the-art performance on the SSE and AVX subsets of SimdBench and, in some cases, generates implementations surpassing standard -O3 optimizations, effectively overcoming the inherent bottlenecks of traditional automated vectorization.
Confidence-Gated Robot Autonomy: When Does Uncertainty Actually Help?
arXiv:2605.18045v1 Announce Type: new Abstract: Robotic systems often use predictive uncertainty to decide whether to act autonomously or defer to a fallback policy. In threshold-gated autonomy, uncertainty matters mainly through its ability to rank likely errors. Standard metrics such as expected calibration error and AUROC do not directly test whether uncertainty changes act/defer decisions. We therefore evaluate uncertainty using Spearman rank correlation, paired bootstrap equivalence testing, and act/defer agreement. Across three temporal activity-recognition benchmarks, we find a dataset-dependent competence regime below which uncertainty provides a weak and unstable error ranking. Above this regime, softmax heuristics, MC Dropout, and ensembles produce similar gating behavior, while threshold choice has a much larger effect on execution outcomes. A multi-seed embodied simulation shows the same pattern for collision rate and cost once realized autonomy is matched. Under temporal covariate shift, ranking quality remains stable, but fine grained semantic OOD detection remains near chance. These results suggest that simple uncertainty proxies can suffice for selective gating once the base model is competent, but not for semantic novelty detection.
Elastic wave propagation governs impulse enhancement in pulsed jets through flexible nozzles
arXiv:2605.17319v1 Announce Type: new Abstract: Inspired by cephalopod jet propulsion through compliant funnels, this study investigates elastic wave propagation and energy exchange in passively deforming cylindrical nozzles through three-dimensional, two-way fluid-structure interaction simulations. Flexible nozzles with varying stiffness ($Eh = 75 - 500~\mathrm{N\,m^{-1}}$, where $E$ and $h$ are Young's modulus and nozzle thickness, respectively) are subjected to a pulsatile jet inflow at $Re \sim 4000$. Increasing nozzle flexibility reduces the deformation-wave speed in accordance with Moens-Korteweg scaling, thereby prolonging the nozzle expansion phase. This delayed expansion enhances jet entrainment and elastic energy storage while suppressing early shear-layer roll-up and vortex formation. During contraction, the stored elastic energy is released, thereby enhancing jet acceleration and vortex formation. For the most flexible nozzle, the primary vortex-ring circulation increases by 52.13%, the vortex convection distance by 9.00%, and the peak outlet kinetic energy flux by a factor of 4.62 compared with a rigid nozzle. These effects collectively yield a 61.92% increase in total hydrodynamic impulse. These findings identify passive wave-speed tuning via nozzle compliance as a mechanism to enhance pulsed-jet thrust for bio-inspired underwater propulsion.
Electrolyte flows under magnetic fields: Manning-like counterion condensation in one dimension
arXiv:2605.18076v1 Announce Type: new Abstract: We present a theoretical framework for unidirectional electromagnetohydrodynamic flow of dilute electrolytes under perpendicular magnetic fields. Starting from the Navier--Stokes equation coupled with the Poisson--Nernst--Planck formulation, we show that the problem admits a sequential decoupling: the Stokes equation is solved first to obtain the velocity profile, which defines a hydrodynamic potential entering the Nernst--Planck description of ions. This Lorentz-force-induced potential competes with electrostatic attraction and significantly alters ionic distributions. We analyze this mechanism in two canonical geometries. In planar Couette shear, it produces a Manning--Oosawa-like condensation transition in one dimension, a phenomenon absent in classical electrostatics. We derive an eigenvalue equation predicting a sharp threshold between counterion enrichment and depletion at the charged wall. In cylindrical Taylor--Couette flow, the same effect shifts the classical Manning criterion by a magnetic parameter, enabling tunable control of condensation. These findings extend Manning--Oosawa phenomenology to driven, non-equilibrium systems and provide a basis for magnetic manipulation of screening in electrolytes, with implications for microfluidics, electrochemical systems, and nonlinear boundary-value theory.
DARE-EEG: A Foundation Model for Mining Dual-Aligned Representation of EEG
arXiv:2605.18298v1 Announce Type: new Abstract: Foundation models pre-trained through masked reconstruction on large-scale EEG data have emerged as a promising paradigm for learning generalizable neural representations across diverse brain-computer interface applications. However, a critical yet overlooked challenge is that EEG encoders must learn representations invariant to incomplete observations-when different masked views of the same signal have minimal overlap, existing methods fail to constrain them to a consistent latent subspace, leading to degraded transferability. To address this, we propose DARE-EEG, a self-supervised foundation model that explicitly enforces the mask-invariance property through dual-aligned representation learning during pre-training. Specifically, we introduce mask alignment that constrains representations from multiple masked views of the same EEG sample via contrastive learning, complementing anchor alignment that aligns masked representations to momentum-updated complete features for semantic stability. Additionally, we propose conv-linear-probing, a parameter-efficient strategy that adapts pre-trained representations to heterogeneous electrode configurations and sampling rates through decoupled spectro-spatial projections. Extensive experiments across diverse EEG benchmarks demonstrate that DARE-EEG consistently achieves state-of-the-art in accuracy performance while maintaining relatively low parameter complexity and superior cross-dataset portability compared to existing methods. Furthermore, DARE-EEG contributes to effectively discovering and utilizing the rich potential representations in EEG.
SD-Search: On-Policy Hindsight Self-Distillation for Search-Augmented Reasoning
arXiv:2605.18299v1 Announce Type: new Abstract: Search-augmented reasoning agents interleave internal reasoning with calls to an external retriever, and their performance relies on the quality of each issued query. However, under outcome-reward reinforcement learning, every search decision in a rollout shares the same trajectory-level reward, leaving individual queries without step-specific credit. Recent process-supervision approaches address this gap by drawing step-level signals from outside the policy, relying either on a much larger teacher model, or on sub-question annotations produced by a stronger external system. In contrast, we propose SD-Search, which derives step-level supervision from the policy itself through on-policy hindsight self-distillation, requiring neither an external teacher nor additional annotations. In SD-Search, a single model plays two roles that differ only in conditioning: a student that sees only the context available at inference time, and a teacher that additionally conditions on a compact hindsight block summarizing the search queries and final outcomes of a group of rollouts sampled from the same question. Since the teacher knows how each rollout unfolded and which ones succeeded, its query distribution implicitly marks which decisions were worth making, and the student is trained to recover this behavior by minimizing the token-level Jensen--Shannon divergence to the teacher at search-query positions. This layers a dense, step-level signal on top of GRPO's coarse trajectory reward. Crucially, this signal is produced by the policy itself within the standard RL training loop, without external model inference, auxiliary annotation pipeline, or additional training stage.
What Would GPT Click: Practical Effects of Human-AI Behavioral Misalignment and the Cost of Synthetic Participants in User Experience
arXiv:2605.18302v1 Announce Type: new Abstract: Synthetic participants represent a methodologically concerning concept that threatens the integrity of UX research. Findings from previous experiments specify how AI outputs are misaligned with the behaviors and thoughts of real humans in various ways. However, industry voices keep underestimating their severity, advocating for practical compromises where good-enough data does not need to be perfect, and all issues will be solved by future tuning. Our study tackles the lack of systematic understanding of the practical issues that arise with synthetic behavior and its use for steering decisions within real contexts. Within twelve diverse first click tests (n = 3431) obtained from real UX practice, we examine the ability of GPT to predict where humans click and how they reason about their behavior. Results (e.g., significantly different distribution from real data in 53% of tasks) demonstrate critical failures to reflect the patterns in which users click on visual elements and the underlying cognitive processes. Participant personas, chain-of-thought reasoning in GPT, and different sampling parameters fail to create sensible fidelity improvements apart from inflating believability. We expose a multitude of nuanced distortions in synthetic responses that reduce their overall analytical usefulness as a decision-making resource, compared with real data. Observed distortions can be theoretically linked to the properties categorically inherent to LLMs: their statistical nature and encoding of semantic heuristics dependent on their training on linguistic data.
Violin ''Playing-In'': Disentangling Physical Change from Player Adaptation via Physical Measurements
arXiv:2605.18305v1 Announce Type: new Abstract: It is a widespread belief among musicians that a violin's sound ``opens up'' or improves through regular playing. However, physical evidence for this ``playing-in'' effect remains elusive. This study revisited the phenomenon by testing two hypotheses: (1) the instrument undergoes physical evolution, or (2) the player undergoes behavioral adaptation. We conducted a longitudinal study centered on a seldom-played test violin played daily by a professional soloist for six months, alongside two stored control violins and a control group of ten violinists (N=10). All three violins had been rarely played prior to the experiment. Assessments performed at the beginning (``Before'' phase) and end (``After'' phase) of the period included input admittance measurements via laser vibrometry, standardized sound recordings, and acquisition of the soloist's bowing gestures via motion capture. Results revealed no acoustical evolution in the played violin exceeding the environmental drift observed in the control instruments. Furthermore, kinematic analysis showed no significant drift in the soloist's bowing strategy (bow position, velocity, contact point). These findings suggest that the ``playing-in'' phenomenon is driven neither by macroscopic physical changes in the instrument nor by a fundamental reorganization of the player's technique.
Directional Scattering-Induced Optical Forces on a Mie Particle near a Metal Interface
arXiv:2604.19428v2 Announce Type: replace Abstract: Optical manipulation of Mie-resonant dielectric nanoparticles is strongly influenced by their enhanced scattering and multipolar response, which fundamentally modifiesthe balance of optical forces. In this work, we study the optical forces acting on a resonant dielectric nanoparticle placed near a metal interface, where scattering occurs into both free-space and surface plasmon-polariton (SPP) channels. We show that the interference of electric and magnetic dipole moments leads to highly directional scattering in these channels, and the direction and magnitude of the scattering-induced force are directly linked to the angular directivity of the corresponding radiation channels. We show that in a cross-beam configuration, where the radiation-pressure contribution is suppressed, the optical force can be changed for almost 2{\pi} in a wide range of particle sizes that provides a route toward optical sorting of resonant nanoparticles.
Latent-IMH: Efficient Bayesian Inference for Inverse Problems with Approximate Operators
arXiv:2601.20888v3 Announce Type: replace-cross Abstract: We study sampling from posterior distributions in Bayesian linear inverse problems where $A$, the parameters to observables operator, is computationally expensive. In many applications, $A$ can be factored in a manner that facilitates the construction of a cost-effective approximation $\tilde{A}$. In this framework, we introduce Latent-IMH, a sampling method based on the Metropolis-Hastings independence (IMH) sampler. Latent-IMH first generates intermediate latent variables using the approximate $\tilde{A}$, and then refines them using the exact $A$. Its primary benefit is that it shifts the computational cost to an offline phase. We theoretically analyze the performance of Latent-IMH using KL divergence and mixing time bounds. Using numerical experiments on several model problems, we show that, under reasonable assumptions, it outperforms state-of-the-art methods such as the No-U-Turn sampler (NUTS) in computational efficiency. In some cases, Latent-IMH can be orders of magnitude faster than existing schemes.
RIE-Greedy: Regularization-Induced Exploration for Contextual Bandits
arXiv:2603.11276v2 Announce Type: replace-cross Abstract: Real-world contextual bandit problems with complex reward models are often tackled with iteratively trained models, such as boosting trees. However, it is difficult to directly apply simple and effective exploration strategies--such as Thompson Sampling or UCB--on top of those black-box estimators. Existing approaches rely on sophisticated assumptions or intractable procedures that are hard to verify and implement in practice. In this work, we explore the use of an exploration-free (pure-greedy) action selection strategy, that exploits the randomness inherent in model fitting process as an intrinsic source of exploration. More specifically, we note that the stochasticity in cross-validation based regularization process can naturally induce Thompson Sampling-like exploration. We show that this regularization-induced exploration is theoretically equivalent to Thompson Sampling in the two-armed bandit case and empirically leads to reliable exploration in large-scale business environments compared to benchmark methods such as epsilon-greedy and other state-of-the-art approaches. Overall, our work reveals how regularized estimator training itself can induce effective exploration, offering both theoretical insight and practical guidance for contextual bandit design.
Crosstalk-free Chiral Anomaly Bulk States in Photonic Crystals
arXiv:2605.17717v1 Announce Type: new Abstract: Ultracompact cladding-free waveguide arrays with zero inter-channel spacing and negligible crosstalk open a new avenue for high-density integrated photonic circuits. However, existing cladding-free waveguide arrays typically rely on conventional trivial bulk modes, making them highly susceptible to scattering losses at sharp bends or in the presence of obstacles and defects. To overcome this limitation, we theoretically propose and experimentally demonstrate a robust, crosstalk-free, and cladding-free photonic waveguide array based on chiral anomaly bulk states (CABSs) in photonic crystals. By interfacing distinct Dirac photonic crystals that host Dirac cones at different high-symmetry points ({\Gamma} and K) in the Brillouin zone and carefully engineering the boundary conditions, the boundary-induced CABSs in adjacent channels become effectively decoupled due to a large momentum separation, thereby eliminating inter-channel crosstalk. More importantly, we experimentally demonstrate that these crosstalk-free CABSs are robust to perturbations, including metallic obstacles, air defects, and sharp bends. We further extend the CABS-based waveguide array to two dimensions and demonstrate a cladding-free triangular resonator and a crosstalk-free waveguide crossing, both of which are previously unattainable. Our work establishes a new design paradigm for cladding-free, crosstalk-free, and ultracompact topological photonic devices, paving the way for robust, highly integrated photonic circuits.
Incentive-Aware Federated Averaging with Performance Guarantees under Strategic Participation
arXiv:2603.20873v2 Announce Type: replace Abstract: Federated learning (FL) is a communication-efficient collaborative learning framework that enables model training across multiple agents with private local datasets. While the benefits of FL in improving global model performance are well established, individual agents may behave strategically, balancing the learning payoff against the cost of contributing their local data. Motivated by the need for FL frameworks that successfully retain participating agents, we propose an incentive-aware federated averaging method in which, at each communication round, clients transmit both their local model parameters and their updated training dataset sizes to the server. The dataset sizes are dynamically adjusted via a Nash equilibrium (NE)-seeking update rule that captures strategic data participation. We analyze the proposed method under convex and nonconvex global objective settings and establish performance guarantees for the resulting incentive-aware FL algorithm. Furthermore, under a merely monotone game setting, we consider a welfare loss minimization framework and establish asymptotic convergence of the scheme. Numerical experiments on the MNIST and CIFAR-10 datasets demonstrate that agents achieve competitive global model performance while converging to stable data participation strategies.
From Documents to Segments: A Contextual Reformulation for Topic Assignment
arXiv:2605.17714v1 Announce Type: new Abstract: Traditional topic modeling assigns a single topic to each document. In practice, however, many real-world documents, such as product reviews or open-ended survey responses, contain multiple distinct topics. This mismatch often leads to topic contamination, where unrelated themes are merged into a single topic, making it difficult to identify documents that truly focus on a specific subject. We address this issue by introducing segment-based topic allocation (SBTA), a reformulation of topic modeling that assigns topics not to entire documents, but to segments: short, coherent spans of text that each express a single theme. By modeling topical structure at the segment level, our approach yields cleaner and more interpretable topics and better supports analysis of multi-theme documents. To support systematic evaluation, we construct a SemEval-STM, a new dataset inspired by aspect-based sentiment analysis. Documents are first decomposed into topical segments using large language models (LLMs), followed by human refinement to ensure segment quality. We also propose a segment-level extension of the word intrusion task, enabling human evaluation of topical coherence at the granularity where topics are actually assigned. Across multiple models and evaluation metrics, we show that SBTA improves clustering quality and interpretability. Overall, this work provides a practical, scalable framework for fine-grained topic analysis in heterogeneous text corpora where documents naturally span multiple topics. URL: https://huggingface.co/datasets/LG-AI-Research/SemEval-STM
TLS Certificate and Domain Feature Analysis of Phishing Domains in the Danish .dk Namespace
arXiv:2603.21652v2 Announce Type: replace Abstract: Phishing attacks remain a persistent cybersecurity threat, and the widespread adoption of TLS certificates has unintentionally enabled malicious websites to appear trustworthy to users. This study examines whether certificate metadata and domain characteristics can help distinguish phishing domains from benign domains within the Danish .dk namespace. A dataset was constructed by combining registry information from Punktum dk with phishing reports and popularity rankings from external sources. TLS certificate attributes were collected using Netlas, while additional domain-based features were derived from DNS records and lexical analysis of domain names. The analysis compares phishing, popular, and less frequently visited domains across several feature categories, including Certificate Authorities (CAs), validity periods, missing certificate fields, SAN structure, registrant geography, hosting providers, and lexical properties of domain names. The results indicate that several features show observable differences between phishing and highly popular domains. However, phishing domains often resemble less popular domains, resulting in substantial overlap across many characteristics. Consequently, no individual feature provides a reliable standalone indicator of phishing activity within the Danish namespace. The findings suggest that certificate and domain attributes may still contribute to detection when combined, while also highlighting the limitations of relying on individual indicators in isolation. This work provides an empirical overview of phishing-related infrastructure patterns in the Danish .dk ecosystem and offers insights that may inform future phishing detection approaches.
Sometin Beta Pass Notin (SBPN): Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation
arXiv:2605.17710v1 Announce Type: new Abstract: Although modern multilingual Automatic Speech Recognition (ASR) systems support several Nigerian languages, their performance consistently lags behind high-resource languages like English and French. Nigerian languages present unique modelling hurdles, including acute data scarcity, inconsistent orthography, tonal diacritics, diverse accents, frequent code-switching, and localized named entities. To address these challenges, we developed a multilingual ASR framework utilizing a two-stage distillation process. First, we employ student-teacher knowledge distillation from existing monolingual models, conditioned on robust language-specific N-gram language models. Second, we perform iterative self improvement using pseudo-labelled data to further refine accuracy. Our method significantly bridges the performance gap, achieving on average a relative Word Error Rate (WER) reduction of 29 % over monolingual baselines. Our models also outperform state-of-the-art multilingual models across major benchmarks, including Common Voice and Fleurs. We introduce Sometin Beta Pass Notin (SBPN), a foundational multilingual ASR model covering Yor\`ub\'a, Hausa, Igbo, Nigerian Pidgin, and Nigerian English. SBPN is released in two sizes: SBPN-Base (120 M parameters) and SBPN-Large (600 M parameters). By releasing these as open foundation models, we aim to provide ASR resources for further research into the rich phonetic and cultural landscape of the region.