Forskningsradar

Science Journals

Peer-reviewade publikationer — 60531 artiklar

SAD-LoRA: Spectral Alignment for Low-Rank Knowledge Distillation
arXiv:2607.04306v1 Announce Type: new Abstract: Distilling a fine-tuned teacher into a LoRA-adapted student is a standard recipe for parameter-efficient compression, but output-level KD does not explicitly control which rank-$r$ weight subspace the adapter occupies. We propose \textbf{SAD-LoRA} (\textbf{S}pectral \textbf{A}lignment \textbf{D}istillation), which selects this subspace from the data-weighted student-space reference update $\DWT\Sigx^{1/2}$ and maintains it during training via a differentiable principal-angle loss on $\colspan(B)$. We show that the data-weighted distillation error decomposes exactly into subspace misalignment, within-subspace coefficient mismatch, and irreducible rank residual; standard KD can affect the first term only indirectly through output gradients. On controlled synthetic problems with a flat teacher spectrum, SAD-LoRA reduces the subspace-misalignment term from $51\%$ to nearly zero and lifts final subspace alignment from $0.49$ to $1.00$. On RoBERTa-large to RoBERTa-base distillation across six GLUE tasks, SAD-LoRA improves rank efficiency: at $r{=}4$, it matches or beats the strongest included spectral baseline on five of six tasks, and at $r{=}8$ it gives the best result on SST-2 and CoLA. Ablations identify subspace alignment as the load-bearing component, while coefficient matching is auxiliary.
WinTA-GIL: Windowed Trajectory Alignment for GNSS-IMU-LiDAR Heading Refinement in Intermittent Signal Environments
arXiv:2607.04879v1 Announce Type: new Abstract: Although multi-source fusion positioning systems have achieved significant progress, accurate and reliable heading estimation remains a critical challenge due to the lack of gravitational constraints and the inherent weak observability of heading in complex environments. Most existing methodologies are specifically tailored for the startup phase, relying on a singular initial alignment to establish the heading reference. Consequently, these approaches lack the adaptability required to refine heading estimates dynamically, which renders the system highly vulnerable to accumulated drift and observation noise during prolonged navigation or immediately following GNSS signal outages. To address these limitations, this paper proposes WinTA-GIL, a novel heading refinement framework that integrates information from Global Navigation Satellite System (GNSS), Inertial Measurement Unit (IMU), and Light Detection and Ranging (LiDAR) through a temporal window-based optimization strategy. Unlike conventional alignment methods restricted to the startup phase, WinTA-GIL leverages high-precision local trajectories from LiDAR-Inertial Odometry (LIO) to register against filtered GNSS observations. This approach transforms heading estimation into a repeatable, trajectory-based consistency optimization problem. In particular, an adaptive re-estimation mechanism based on state discrimination is incorporated to trigger heading corrections whenever necessary, thereby effectively suppressing the inertial drift accumulated during challenging conditions. Extensive experiments on both open-source and self-collected datasets demonstrate that WinTA-GIL significantly outperforms state-of-the-art approaches in both estimation accuracy and system robustness.
Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees
arXiv:2607.04430v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in question answering (QA) systems, yet they may generate hallucinated or misaligned responses without reliable confidence estimates. Uncertainty quantification (UQ) offers a natural basis for selective answering, where a system answers only when its prediction is deemed reliable and abstains otherwise. However, existing uncertainty scores for LLMs are often heuristic: a threshold chosen on such scores does not, by itself, provide statistical guarantees on the error rate among accepted answers. We propose CIC, a confidence-interval-based calibration framework that converts arbitrary uncertainty scores into risk-controlled selective answering rules. Given a held-out calibration set, CIC evaluates each generated response using an application-specific alignment criterion and associates it with an uncertainty score and a binary error label. For each candidate uncertainty threshold, CIC estimates the acceptance-conditioned error rate and constructs a high-probability upper confidence bound using either Hoeffding-style or Clopper-Pearson confidence intervals. It then selects the largest threshold whose upper bound is below a user-specified risk level $\alpha$, thereby maximizing the answering rate subject to a finite-sample reliability constraint. Under exchangeability, CIC guarantees with probability at least $1-\delta$ that the selected threshold, if non-null, controls the error rate among accepted answers at level $\alpha$. We evaluate CIC on both closed-ended and open-ended QA benchmarks across seven LLMs and multiple uncertainty estimators. Experimental results show that CIC consistently achieves valid risk control while retaining strong answering efficiency, providing a practical and statistically grounded mechanism for deploying LLMs in reliability-sensitive QA workflows.
Knowledge-Informed Local Causal Discovery of Optimal Adjustment Sets
arXiv:2607.04447v1 Announce Type: new Abstract: Local causal discovery is a scalable alternative to global structure learning. However, it can struggle to identify valid adjustment sets in data-scarce settings because of finite-sample uncertainty, incomplete local neighborhoods, and unresolved Markov equivalence. Although many application domains provide structured background knowledge, its integration into local causal discovery remains limited. We propose b-LOAD, a knowledge-informed extension of the LOAD algorithm for local discovery of optimal adjustment sets. b-LOAD incorporates prior edge constraints directly into the local structure-learning procedure and uses Meek's rules to expand the discovery frontier dynamically, yielding a knowledge-constrained partially directed graph over the relevant local subgraph. This strategy helps prevent structurally relevant nodes introduced by prior knowledge from being excluded by local search. We prove that, under sound background knowledge, the procedure monotonically refines the admissible equivalence class and can enlarge the set of identifiable causal queries, enabling recovery of optimal adjustment sets that are not identifiable from observational conditional-independence information alone. Empirically, b-LOAD improves downstream causal effect estimation relative to purely data-driven and standard knowledge-augmented baselines, particularly in data-scarce and structurally complex regimes. Results on real-world biological networks show that locally targeted prior knowledge provides the largest gains and remains beneficial under moderate structural noise. These findings position b-LOAD as a scalable approach for converting fragmented domain knowledge into more reliable causal-effect estimation.
Dynamic Image-Informed Selection of Biomechanical Tumor Growth Models
arXiv:2607.04551v1 Announce Type: new Abstract: Glioblastoma progression is strongly influenced by evolving mechanical interactions between the tumor and surrounding brain tissue. However, the extent to which finite-deformation mechanics and constitutive assumptions improve subject-specific prediction as tumor burden evolves remains unclear. We introduce a sequential Bayesian inference and dynamic model selection framework that assimilates longitudinal murine magnetic resonance imaging (MRI) data to calibrate spatially varying tumor diffusivity, proliferation rate, and tissue stiffness in biomechanical tumor growth models. Competing formulations were compared at each imaging time, including reaction-diffusion without mechanics and reaction-diffusion coupled to linear elasticity or hyperelastic mechanics, using posterior model plausibility to adapt model choice for individualized one-scan-ahead prediction as new MRI scans are acquired. Across the studied animals, mechanically coupled models were consistently more plausible than the uncoupled reaction-diffusion model, and the evolution of model plausibility indicated an increasing role of mass effect and stress-mediated feedback of tumor growth during progression. While linear and hyperelastic coupled tumor growth models often produced similar tumor morphology, they yield distinct stress, deformation, and inferred stiffness fields, with the hyperelastic formulation often receiving higher posterior plausibility at later imaging times. These results indicate that, within the present longitudinal murine dataset, mechanical coupling is favored for image-informed glioma growth prediction and that constitutive assumptions should be evaluated sequentially for each subject rather than fixed a priori.
Enhancing Large Multimodal Models in Key Information Extraction via Scene-Aware Document Synthesis
arXiv:2607.04636v1 Announce Type: new Abstract: Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-server Large Multimodal Models (LMMs), while compact locally deployable models lack sufficient KIE supervision. We present SAYRE, a scene-aware document synthesis framework for generating scalable KIE training data without hand-crafted template design. Given a few exemplar documents, SAYRE captures category-specific content patterns and layout conventions to synthesize document-schema-annotation triples. It further introduces error-driven generation, which expands real-world failure cases into hard training examples while preserving their structural patterns. Experiments on constrained- and open-category KIE show that SAYRE consistently improves Qwen3-VL backbones and achieves the strongest overall performance among on-device LMMs. Data scaling experiments show an overall upward trend as more synthesized data is introduced, especially for smaller models and open-category extraction. Error analysis further shows that synthesized training reduces field-level errors by improving schema-aware extraction over dense tables, business identifiers, and contract clauses. These results establish scene-aware synthesis as an effective data-centric approach for improving practical multimodal KIE.
Learning Probabilistic Prompt for Continual Learning
arXiv:2607.04711v1 Announce Type: new Abstract: Continual learning aims to progressively learn from a sequence of tasks, each containing a disjoint subset of classes, while preserving previously learned knowledge. Prompt-based continual learning methods propose to learn a small set of parameters, i.e., prompts, by associating them with a query feature of an input image. These methods optimize the prompts, attempting to represent diverse patterns of images. However, we have observed that existing prompt-based methods suffer from a prompt collapse problem, that is, the prompts tend to be highly similar to each other, thereby failing to capture the diverse data distributions in continual learning scenarios. To address this issue, we propose in this paper a novel prompt-based continual learning framework that captures diverse patterns of images across a sequence of tasks. To this end, we model each prompt as a probabilistic distribution and construct a mixture of these distributions, from which we sample diverse prompts. This enables our model to effectively capture highly diverse image distributions in the continual learning process. We also present a distribution regularization loss to prevent abrupt changes in the prompt distributions throughout the training process. We show extensive experimental results for continual learning on standard benchmarks, including ImageNet-R, CIFAR-100, and CUB-200, demonstrating the effectiveness of our framework.
Detection of scintillation light in noble gases with wavelength-shifting optical fibers
arXiv:2607.04944v1 Announce Type: new Abstract: Wavelength-shifting (WLS) techniques enable particle detectors based on noble gases, whose scintillation light is predominantly emitted in the vacuum-ultraviolet. We investigate WLS fibers coated with tetraphenyl butadiene (TPB) for scintillation light detection in gaseous xenon and argon at pressures up to 8.5 bar, motivated by future high-pressure xenon time-projection chambers of the NEXT program. Two detector configurations are studied: an elongated high-pressure vessel with four PTFE panels equipped with WLS fibers read by temperature-stabilized SiPMs, and a compact box-shaped detector operated at 1 bar Xe with WLS fibers read out by PMTs. Both operate with continuous gas purification. The detector response is characterized using cosmic muons and alpha particles from a $^{241}$Am source. With the SiPM setup, we measure a light collection efficiency (LCE) of ${1.18 \pm 0.01~\mathrm{(sta.)}~^{+0.07}_{-0.09}~\mathrm{(sys.)}~\%}$ for xenon and ${1.07 \pm 0.01~\mathrm{(sta.)}~^{+0.06}_{-0.08}~\mathrm{(sys.)}~\%}$ for argon. With PMT readout, we measure a LCE of ${0.45 \pm 0.01~\mathrm{(sta.)} \pm 0.05~\mathrm{(sys.)}~\%}$ in xenon, in agreement with the SiPM result once photon detection efficiency is accounted for. Average scintillation waveforms in xenon and argon are studied to assess the time structure of the emitted light. Cosmic-muon measurements yield a mean energy required to produce a scintillation photon $45\pm7~\mathrm{(sta.)}~^{+4}_{-5}~\mathrm{(sys.)}~\mathrm{eV}$ at 1.5 bar, in agreement with the literature. The results demonstrate that TPB-coated WLS fiber systems can reliably detect scintillation light in high-pressure gaseous noble detectors, with a LCE representing an upper limit for realistic large-scale TPCs, where additional photon losses from materials and fiber attenuation are expected.
Performance evaluation of scheduling tasks in many-core systems utilizing processes and threads
arXiv:2607.04821v1 Announce Type: new Abstract: This study assesses the scalability of process-based and thread-based schedulers for many-core shared-memory systems using a memory-intensive row-wise quick-sort workload on large three-dimensional tensors. The process-based evaluation considers bounded prolific, bounded collective, and three pipe-based producer-consumer schedulers: one-to-one, one-to-many, and many-to-many. These pipe schedulers dynamically stream task identifiers to worker processes, exchanging increased inter-process communication overhead for enhanced runtime load balancing and flexible chunk-based task dispatching. The thread-based evaluation examines static, dynamic, guided, chunk-based, chunk-stealing, adaptive chunk, and AIMD adaptive scheduling strategies. The AIMD scheduler employs an additive-increase multiplicative-decrease policy inspired by TCP congestion control, utilizing an exponentially weighted moving average (EWMA) of CPU utilization to regulate a contention window that limits the number of concurrently active chunks. The adaptive chunk scheduler further modifies chunk size based on observed per-thread execution speed. Experimental results on a 24-core x86-64 platform indicate that thread schedulers deliver the highest overall performance, with dynamic and guided scheduling yielding the most favorable practical outcomes. Among process schedulers, pipe-based designs demonstrate the strongest scalability, with one-to-one pipes excelling for smaller workloads and many-to-many pipes preferred for larger workloads. In summary, lightweight thread scheduling is optimal for shared-memory row sorting, while AIMD/adaptive scheduling and pipe-based process scheduling remain valuable for contention-aware execution, explicit inter-process coordination, and distributed-style heterogeneous workload management.
Shifting from Discrete to Continuous Reference Data: QSM-Derived Horizontal Tree Biomass Distribution for Deep Learning Biomass Estimation
arXiv:2607.05260v1 Announce Type: new Abstract: Conventional modeling approaches for LiDAR-based above-ground biomass (AGB) estimation rely on discrete plot-level inventory aggregates. This methodology introduces boundary-effect uncertainties that may severely degrade model performance within small field plots. To solve this limitation, we evaluate a Horizontal Biomass Distribution (HBD) reference mapped continuously from Quantitative Structure Models (QSMs). We trained a sparse 3D U-Net on simulated broadleaved forest structures using three AGB reference types: a standard forest inventory (FI) plot-level aggregate, an edge-effect-free QSM plot-level aggregate, and a continuous HBD mapping. Evaluating training plot sizes scaling from 100 to 2500 $m^2$ , QSM-based models systematically outperformed FI approaches at small plot sizes. Specifically, for 100 $m^2$ plots, the HBD reference reduced the relative root mean square error (RRMSE) by 16.84 $\pm$ 4.37 % and increased $R^2$ by 0.22 $\pm$ 0.05 against the FI baseline. By replacing plot level aggregates with HBDs as AGB reference, this methodology corrects for edge-effects and shows that using an HBD-based reference enhances model performance for small plot sizes.
Representing and Detecting Label Ambiguity in IMU-Based Exercise Evaluation
arXiv:2607.04842v1 Announce Type: new Abstract: Home-based physiotherapy is performed without supervision, which leads to incorrect execution and motivates systems that assess movement automatically from inertial measurement units (IMUs). Such systems assign each repetition to a category, yet a relevant share of repetitions falls near a class boundary, where even trained raters disagree. Classifiers trained with one-hot labels collapse these borderline repetitions onto a single class and discard this ambiguity. We address this with a method that automatically generates a label distribution per repetition without a large rater pool. We train a network to reproduce the full distribution with a Kullback-Leibler objective, the ambiguity approach, and compare it against a one-hot cross-entropy baseline on four IMU exercise datasets. From the network output we further determine whether a repetition is ambiguous and which classes are relevant to it. The ambiguity approach matched or exceeded the baseline classification on all four datasets, and detected ambiguity and the relevant classes more reliably. Representing the label distribution in the training target therefore adds information about ambiguity at no cost to classification.
Pretraining Curricula Enable Selective Fine-tuning
arXiv:2607.04846v1 Announce Type: new Abstract: Transformers follow implicit curricula whereby some tasks are learned before others. However, how explicit pretraining curricula influence learning, generalization, and the selectivity of fine-tuning is unclear. This is important for AI safety, where fine-tuning is used to selectively suppress misaligned behaviors. Here, we compare curricula that pretrain tasks in a balanced (sampled uniformly) or an imbalanced (one task early, the other late) fashion. We show that imbalanced learning of two conflicting copy tasks promotes in-context learning and improves the selectivity of refusal fine-tuning. Ablations and activation patching show that this occurs because imbalanced pretraining encourages tasks to be disentangled in separable neural circuits, whereas balanced training routes both tasks through a common pathway. We extend these findings to a synthetic language learning task involving rule-consistent and rule-violating data, where imbalanced curricula similarly lead to more localized, less entangled rule representations, resulting in more robust rule-following behavior. Together, these results suggest that imbalanced pretraining curricula may be an important tool for promoting disentangled representations, with direct consequences for the precision and reliability of safety fine-tuning.
Anomalous Peak Formation in the Second Stability Zone of Quadrupole Mass Filters: Role of Non-linear Resonances Induced by Single Rod Defects
arXiv:2607.04867v1 Announce Type: new Abstract: Operating a quadrupole mass filter (QMF) within its second stability zone offers superior mass resolving power but introduces extreme sensitivity to structural imperfections. While symmetric misalignments are well-documented, this work combines analytical modeling with SIMION trajectory simulations to investigate the unaddressed electrodynamic impact of a completely localized, asymmetric defect on a single rod. Unlike diagonally symmetric perturbations, which reduce the four-fold rotational symmetry to two-fold symmetry while maintaining a smooth stability landscape, a single-rod defect breaks the remaining symmetry, giving rise to pronounced transmission ridges within the second stability zone. Spatial multipole expansion reveals that this structural breakdown injects odd-parity harmonics -- dominated by the hexapole A3 field -- which couple the orthogonal transverse equations of motion. By tracking secular indices through explicit zone-II Floquet projection mappings, we demonstrate that these asymmetric fields drive destructive, higher-order nonlinear secular resonances. Ions traversing these precise parametric coordinates undergo rapid amplitude growth and collide with adjacent electrodes, establishing critical geometric tolerance frameworks and operating conditions for high-performance mass spectrometry.
WildSplat: Feedforward Gaussian Splatting from Unposed In-the-Wild Images
arXiv:2607.05347v1 Announce Type: new Abstract: While feedforward 3D reconstruction excels at efficient novel view synthesis, it typically falters when faced with scenes under varying illumination. To this end, we introduce WildSplat, the first feedforward 3D Gaussian Splatting framework capable of appearance-conditioned novel-view synthesis for unposed in-the-wild images. To handle inconsistent photometric conditions, we propose a dual-branch architecture that explicitly decouples geometry from appearance. The geometry branch extracts an appearance-invariant 3D structure and jointly predicts camera poses. To govern the rendering appearance, the appearance branch injects target appearance cues into the content features via a globally pre-modulated cross-attention mechanism. To further prevent feature entanglement, we introduce a joint multi-reference training strategy that stabilizes the training process. Extensive experiments show that WildSplat surpasses existing optimization-based and feedforward methods, achieving state-of-the-art performance in in-the-wild novel view synthesis and appearance editing from sparse inputs in a single forward pass.
U3DWind: A Low Altitude Wind Field Dataset and Benchmark for Urban Air Mobility
arXiv:2607.04495v1 Announce Type: new Abstract: Urban Air Mobility (UAM) requires reliable assessment of low-altitude wind hazards, because winds, gusts, and building-induced turbulence have been recognized as critical factors affecting vehicle stability, route feasibility, vertiport siting, and airspace management. While wind-tunnel experiments, computational fluid dynamics (CFD), multiscale downscaling, reduced-order models, and UAV planning datasets have advanced wind-aware analysis, public resources for data-driven, city-scale UAM planning remain limited in geographic coverage, scenario diversity, vertical extent, building realism, and task-oriented benchmarking. To address this gap, we introduce U3DWind, a building-resolved low-altitude wind-field dataset generated using our GPU-accelerated Lattice Boltzmann Method--Large-Eddy Simulation (LBM-LES) framework for rapid urban flow simulation. U3DWind covers five megacities in China: Beijing, Shanghai, Guangzhou, Shenzhen, and Hong Kong. It contains 720 simulations, with 16 inflow directions, three reference wind speeds, and three seasonal atmospheric scenarios (annual, summer, and winter) for each city. At a 10 m grid resolution, the dataset provides three-dimensional three-component (3D3C) velocity, turbulent kinetic energy (TKE), flow density, and fluid--solid masks. To support operationally relevant evaluation, we further define five baseline tasks: wind-field prediction, sparse-sensor wind-field reconstruction, site wind-exposure ranking, airworthiness wind-compliance risk scoring, and noise propagation modeling. As a multi-city, building-resolved 3D urban wind-field dataset, U3DWind enables systematic evaluation of wind-induced impacts in low-altitude traffic scenarios and provides an open benchmark for urban airspace management and data-driven high-fidelity urban flow simulation.
Focusing and light collection effects on plasma-induced frequency-resolved optical switching (PI-FROSt) traces
arXiv:2607.04920v1 Announce Type: new Abstract: Plasma-Induced Frequency-Resolved Optical Switching (PI-FROSt) is a promising and recently proposed phase-matching-free technique for characterising ultrafast pulses across broad spectral ranges. We investigate the mechanisms of PI-FROSt trace formation through numerical simulations and experimental validation. The results reveal that trace characteristics are highly sensitive to the relative focusing geometry between pump and probe pulses, as well as the spatial region selected for signal collection. Depending on these conditions, the interplay between plasma defocusing and positive lens-like nonlinear effects causes either intensity depletion or enhancement in the probe beam, flipping the PI-FROSt trace. Simulations demonstrate that optimal gate stability also depends strongly on the focusing scheme and the collecting region. This study highlights that precise spatial and temporal optimisation is essential to properly exploit the benefits of this broadband pulse characterisation technique.
Multi-Robot Open Adaptive Teaming Across Unseen Environments, Partners, and Scales
arXiv:2607.04972v1 Announce Type: new Abstract: Deploying robot teams in the real world requires simultaneous adaptation to unseen environments, unknown partners, and varying team sizes, yet existing approaches often address these challenges in isolation under the closed-world assumption of fixed teammates. We formalize this as open adaptive multi-robot teaming and propose a hypergraphic-form game formulation that captures team-level cooperative relationships beyond pairwise interactions, providing a principled foundation for coordination structure inference when team composition changes dynamically within episodes. Unlike graph neural network architectures, this is a game-theoretic construct for modeling strategic interactions and payoff structures among agents. Building on this formulation, we develop the Hypergraphic Open-ended Learning Algorithm (HOLA), which progressively expands partner and environment diversity during training rather than optimizing for fixed configurations. Evaluated on cooperative pursuit with multi-drone and multi-quadruped platforms, HOLA outperforms all baselines across all three adaptability dimensions. Learned policies transfer directly to physical hardware without fine-tuning, with successful deployments on Crazyflie and Zsibot L1 platforms confirming robust real-world coordination in novel environments with unseen teammates.
Abstract Color Voronoi Diagrams and Circular Sequences of Color Permutations
arXiv:2607.05383v1 Announce Type: new Abstract: Abstract Voronoi diagrams are defined in terms of a given system of planar bisecting curves satisfying some simple combinatorial properties. They offer a unifying framework for a wide range of concrete Voronoi instances on generalized sites and metrics. In this paper, we formulate higher-order abstract color Voronoi diagrams of a set $S$ of $n$ colored abstract sites, simultaneously considering all concrete instances under their umbrella. We prove that the number of vertices in the order-$k$ abstract color Voronoi diagram is at most $4k(n-k)-2n$, and present an iterative construction algorithm. The bound directly applies to a family of $m$ disjoint simple polygons of total complexity $n$. For simple polygons the bound can further improve to $O(\min\{k(n-k),(m-k)^2n\})$. A critical ingredient of our proof is a combinatorial analysis on circular sequences of color permutations derived from the unbounded edges of these diagrams, which is interesting in its own right.
Contaminated Multi-task Learning with Heterogeneity: Fundamental Limits and Optimal Algorithms
arXiv:2607.02681v1 Announce Type: cross Abstract: Integrating information across related tasks can improve estimation and prediction in transfer, multi-task, and federated learning, but contamination and heterogeneity make robust borrowing challenging. We study a contaminated multi-task empirical risk minimization (ERM) framework in which an $\epsilon$ fraction of $K$ tasks, each with sample size $n$, may be arbitrarily contaminated while the remaining tasks are heterogeneous. Our goal is to estimate both the global minimizer of the average risk and the clean task-specific minimizers, thereby combining robustness and personalization. In the Gaussian mean model, we show that several common paradigms, including adaptive and robust regularization around a shared center, global matrix regularization, decomposition-based regularization, and score-based outlier-task detection, all suffer from a worst-case contamination error of order $\epsilon\sqrt{d/n}$, which is suboptimal compared to the lower bound $\epsilon/\sqrt{n}$. This identifies a dimension-dependent barrier for these approaches. We then establish minimax lower bounds for a general heterogeneous ERM setting and propose a computationally efficient filtering-based robust multi-task gradient descent method. Under local strong convexity, smoothness, and sub-Gaussian gradient assumptions, the proposed method attains high-probability upper bounds matching the minimax rates up to logarithmic factors over a broad regime. In particular, it removes the extra $\sqrt{d}$ contamination dependence of many regularization-based methods and score-based outlier detection, while achieving personalization to local tasks under strong heterogeneity. Simulations and a real-data analysis demonstrate strong robustness and personalization relative to a broad range of benchmark methods.
Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference
arXiv:2607.05116v1 Announce Type: new Abstract: As MoE models scale to hundreds of experts, placement and pruning decisions increasingly dictate communication volume, affecting the performance of distributed inference across GPUs and nodes. We propose CAP (Communication-Aware Assignment and Pruning), a framework that considers computation, communication and accuracy together for efficient MoE inference through expert placement and pruning. It consists of three components: (1) Co-activation driven expert placement, which groups frequently co-activated experts to reduce inter-device and inter-node communication; (2) Communicationcomputation trade-off adjustment, which generates placements with different computational load and communication volume; and (3) Communication-aware expert pruning, which selectively removes routing destinations to reduce communication with limited accuracy degradation. By combining these components, CAP selects an efficient operating strategy for different hardware configurations. Across our single-node and multi-node experiments, it achieves 1.23x - 1.86 x throughput improvement over DeepSeek EPLB and sequential placement in vLLM, and preserves better model accuracy at the same target speedup under lossy acceleration.
Small-scale dynamo saturation across magnetic Prandtl numbers using the EDQNM closure
arXiv:2607.02743v1 Announce Type: cross Abstract: Small-scale dynamos (SSDs) are believed to be the primary source of magnetic fields in all turbulent astrophysical systems, especially those with weak rotation such as elliptical galaxies and galaxy clusters. The initial kinematic phase of these dynamos is relatively well understood. Here we demonstrate analytically and numerically that, in an appropriate limit, the eddy-damped quasi-normal Markovian (EDQNM) closure for incompressible magnetohydrodynamic turbulence is strictly equivalent to the earlier models of kinematic dynamos. Moreover, it allows the extension of the kinematic dynamo framework to multi-scale turbulent flows and into the nonlinear regime. The EDQNM closure also enables us to explore a wide parameter range which is inaccessible to direct numerical simulations of the SSD. Using nonhelical EDQNM simulations, we identify several asymptotic regimes of nonlinear dynamo action when the system is highly turbulent with fluid Reynolds number $Re \gtrsim 10^6$ for magnetic Prandtl number $Pm > 1$ and magnetic Reynolds number $Rm \gtrsim 10^6$ for $Pm < 1$: 1) the kinematic growth rate approaches a value independent of $Pm$, 2) the saturated magnetic to kinetic energy ratio similarly converges to $\simeq 0.55$ across $Pm$, while the ratio of magnetic to kinetic integral wavenumbers asymptotes to $\simeq 3$. For all $Pm$, we further find strong feedback between magnetic field and velocity field largely via Alfv\'{e}nisation leading to a saturated kinetic and magnetic spectra with almost the same inertial range with a slope of $-3/2$. These findings could provide guidance for future global simulations and for modeling the nonlinear regime of astrophysical systems living in these extreme limits.
Fast counting and sampling for ferromagnetic two-spin systems
arXiv:2607.05248v1 Announce Type: new Abstract: We introduce two new models equivalent to ferromagnetic two-spin systems: a weighted subgraph model and a random cluster type model. Using these new connections, we obtain an efficient sampling algorithm and a new randomised algorithm that efficiently approximates the partition function of ferromagnetic two-spin systems in certain parameter regimes. No efficient sampling algorithms are known before in this regime, and our new estimation algorithm runs in near-quadratic time for bounded degree graphs and in polynomial time for general graphs, improving upon the previous algorithm of Guo, Liu, and Lu (2020).
Mixture-Constrained Max Pooling Improves Separation-Based Bird Species Classification
arXiv:2607.03221v1 Announce Type: cross Abstract: Bird species classification from field recordings remains challenging due to overlapping vocalizations and incomplete species labels. We study source separation as a preprocessing for bird species classification to improve multi-species detection. Specifically, we employ an ensemble of two separators, FTRNN and TF-Locoformer, both trained with mixture invariant training (MixIT). To address the false positive gain caused by separation errors in separated outputs, we propose mixture-constrained max pooling (MCM), which clips the predicted probability from each separated channel based on the corresponding species probability in the original mixture. The classifier is applied to each separated output and the original mixture independently, and MCM aggregates the predictions into a final per-species probability. Experiments on two real-world datasets show that the ensemble outperforms individual separators and MCM outperforms standard max pooling across multiple metrics, and reveal that separation leads to both true positive gain for present species and false positive gain for absent species.
High Success Probability, Fidelity, and Purity Nonlinear Optical Two-Qubit Gates on Chip
arXiv:2607.03313v1 Announce Type: cross Abstract: Optical two-qubit gate with high success probability, fault-tolerant fidelity, and high-purity outputs is a fundamental yet unsolved challenge, essential for large-scale optical quantum computing toward quantum advantage. Here, we propose a feasible scheme for such gate using thin-film lithium niobate platform, enabling \c{hi}(2) nonlinear photon-photon interaction with 100% efficiency. By decoupling photon interaction and qubit flip operations, fidelity ceiling is removed, and output state purity is recovered by spectral-phase pre-compensation based on a full-spectral photon interaction model, yielding a CNOT gate with 84% success probability, 93% purity, and unity fidelity.
CaSPECT: Discovering Causally Homogeneous Subgroups via Directed Spectral Clustering
arXiv:2607.03364v1 Announce Type: cross Abstract: We propose \textbf{CaSPECT}, a causal spectral clustering framework for discovering causally homogeneous subgroups from observational data. Rather than clustering in covariate space, CaSPECT defines similarity through the topology of a learned directed acyclic graph (DAG); a bootstrap-stabilised PC algorithm recovers the causal skeleton; a novel \emph{Orientation Validation Score} (OVS) combines PC bootstrap evidence with DirectLiNGAM to orient edges robustly; directed edges are weighted by backdoor-identified average treatment effects estimated via OLS or double machine learning. Chung's directed Laplacian provides a spectral embedding in which individuals close together share the same causal propagation pathways. We establish almost-sure consistency of the full pipeline and validate the method through a controlled simulation study and on LaLonde CPS1, IHDP, and 401(k) datasets, where CaSPECT recovers a positive and statistically significant treatment effect within the causally comparable subpopulation and corrects for severe confounding without requiring a pre-specified propensity score model.