Forskningsradar

Science Journals

Peer-reviewade publikationer — 56237 artiklar

An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge
arXiv:2607.12468v1 Announce Type: new Abstract: We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Large-s80 (WavLM-Large) segmentation, CAM++ embedding-based two-speaker clustering, and a LoRA-adapted omniASR LLM 7B v2 recognizer, with no oracle segmentation or speaker labels at test time. On the official Development set (150 conversations, 21 language/accent categories) the system attains a macro tcpMER of 29.27%, versus 79.15% for the official baseline; on the Evaluation set it scores 50.23%. We also analyze two engineering choices that substantially affect tcpMER. First, embedding-based speaker clustering outperforms an end-to-end-style alternative that assigns speakers from ASR <sc> turn markers alone. Second, overlap-aware segmentation, although intended to raise diarization recall, increases tcpMER because overlapped speech is transcribed twice.
High-Frequency Gravitational Wave Constraints from Precision Spectroscopy
arXiv:2607.12617v1 Announce Type: cross Abstract: Gravitational waves affect the propagation of electromagnetic waves in laser cavities, modulating the frequency of emitted photons. We use this effect to search for high-frequency gravitational waves between 100 kHz and 100 MHz using optical precision spectroscopy. Our limits constrain much of this frequency range for the first time. We discuss future improvements of the technique, which we expect to enhance the sensitivity by eight orders of magnitude, and to extend the frequency coverage up to at least 1 GHz.
Optimal Assembly of Repurposed Lithium-Ion Battery Packs under Cell Heterogeneity and Screening Uncertainty
arXiv:2607.12951v1 Announce Type: new Abstract: The growing supply of retired electric vehicle batteries presents an opportunity for second-life stationary energy storage, but assembling heterogeneous retired cells into reliable packs is challenging due to substantial variation in capacity, DC internal resistance (DCIR), and self-discharge. This paper proposes a robust optimization framework for cell-to-pack assembly of second-life batteries. A topology-screening stage first identifies minimum-cell series-parallel configurations satisfying inverter and energy requirements, reducing the dimensionality of the subsequent assignment problem. For each candidate topology, a mixed-integer linear program selects cells and assigns them along the series string, enforcing power, voltage, and energy requirements as hard constraints while minimizing a normalized, weighted sum of DCIR spread, capacity spread, and self-discharge imbalance. Additionally, measurement uncertainty in capacity and DCIR is modeled as bounded intervals to guarantee feasibility under worst-case parameter deviations. The framework is evaluated on four heterogeneous inventories for a 10 kW/10 kWh stationary backup application. The proposed method satisfies all feasibility requirements in every case, while single-metric sorting heuristics each fail on at least one inventory. Relative to the best single-metric baseline by objective value, it reduces the normalized mismatch objective by 76-87%, demonstrating that jointly optimizing cell matching with application-level feasibility requirements improves heterogeneous second-life pack assembly under screening uncertainty.
Traceable In Situ Microwave Power Measurement at the Cryogenic Device Plane in a Dilution Refrigerator
arXiv:2607.12751v1 Announce Type: new Abstract: Accurate knowledge of the microwave power delivered to a cryogenic device under test (DUT) is essential for the characterization and operation of superconducting quantum circuits. However, this information is difficult to obtain inside dilution refrigerators because of distributed attenuation, impedance mismatch, switch-path repeatability, and temperature-dependent microwave components. This paper presents an in situ measurement method for RF power at the cryogenic device plane. The method uses a custom variable temperature stage (VTS) as a cryogenic thermal-transfer element. The TVS is alternately heated by a four-wire DC heater and by microwave power dissipated in a 20 dB pass-through attenuator. By fitting the thermal transients and comparing the corresponding steady-state temperatures, the absorbed microwave power is inferred from a directly measured DC electrical power through an AC/DC substitution procedure. The finite reflection and transmission of the attenuator are then accounted for by cryogenic two-port scattering-parameter measurements based on a switch-assisted Short--Open--Load--Reciprocal calibration, so that the result is referred to the DUT reference plane. The system is demonstrated in a dilution refrigerator with powers between -43 and -58 dBm at the DUT input plane. The demonstrated relative standard uncertainty ranges from about 2% at -43.9 dBm to about 40% at -57.6 dBm. The proposed approach combines thermal RF power transfer, cryogenic S-parameter correction, and uncertainty evaluation in a measurement architecture compatible with quantum-device experiments, providing a practical route toward traceable microwave-power calibration at millikelvin stages.
RFMSR: Residual Flow Matching for Image Super-Resolution
arXiv:2607.12753v1 Announce Type: new Abstract: Image super-resolution (ISR) has witnessed remarkable progress with diffusion models and flow matching. The dominant text-to-image (T2I) based approaches leverage large-scale foundation models as generative priors, achieving impressive perceptual quality but at the cost of massive model sizes and prohibitive training expenses. Recent flow-matching-based vision-only approaches have made significant strides; however, they adopt standard flow formulations that transport from a pure Gaussian prior to the data distribution, discarding the rich structural information already present in the low-quality (LQ) input. Furthermore, existing single-step acceleration techniques often forfeit the model's multi-step inference capability. In this paper, we propose Residual Flow Matching for Image Super-Resolution (RFMSR), a vision-only framework that centers the source distribution at the LQ latent, reducing transport distance and preserving structural priors throughout the flow trajectory. We further introduce a two-phase training strategy: Phase I pretrains the velocity field via conditional flow matching, while Phase II applies end-to-end supervision to the single-step prediction while retaining the velocity loss across all timesteps, achieving high-quality single-step generation without sacrificing multi-step refinement. Extensive experiments demonstrate that RFMSR achieves comparable or even superior perceptual quality compared to state-of-the-art (SOTA) methods. The source code is available at https://github.com/Faze-Hsw/RFMSR.
Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement
arXiv:2606.19387v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved remarkable success in software development. However, they are susceptible to hallucinations, meaning that they can introduce subtle semantic and logical errors. Due to the high stakes in chip design and manufacturing, hardware engineers are still reluctant to rely on LLMs for register-transfer level (RTL) generation. In this paper, we propose a hardware generation framework that combines the creativity and broad knowledge of LLMs with the explainability and mathematical rigor of formal methods. Specifically, we devise a set of transformation rules that cover various design decisions and hardware features. By iteratively applying these rules, an LLM agent can convert a design specification into an RTL program with guaranteed correctness. Experimental results demonstrate the effectiveness and efficiency of the framework.
Single-Shot High-Energy Muon and Particle Radiography with a Multi-GeV Laser-Wakefield-Accelerator-Driven Source
arXiv:2607.12984v1 Announce Type: new Abstract: We report the first demonstration of single-shot particle radiography using a 1-10 GeV laser-wakefield-generated beam of muons, pions, and neutrons. The test objects were imaged ~15 m from the beam source, through dense lead shielding followed by the walls of a building and a truck. The muon content of the beam was directly confirmed using large volume scintillator-based detectors, which recorded particle decay events with timing delays consistent with the muon lifetime. Simulations confirm that the high energy component of the beam transmitted through the test object is nearly entirely composed of muons, directly showing their highly penetrative nature, with a single-shot fluence equivalent to >8 hours of integration of cosmic ray muons near the horizon. Our work establishes single-shot high-energy particle radiography with a laser-wakefield-accelerator-driven source.
ChunkFlow: Towards Continuity-Consistent Chunked Policy Learning
arXiv:2607.12992v1 Announce Type: new Abstract: Vision-language action (VLA) models increasingly adopt chunked action heads to satisfy real-time constraints; however, this introduces boundary jitter: overlapping regions between consecutive chunks often yield inconsistent predictions, degrading temporal coherence and the task success rate. Existing methods, such as inference-time blending, merely reweight mismatched proposals without correcting underlying errors, leading to residual accumulation under biased or noisy histories. We propose ChunkFlow, a seam-aware training-and-execution framework for chunked policies that aligns chunk structure with boundary execution. It partitions each chunk into frozen, editable, and future zones, applies deterministic overlap blending at execution, and trains raw predictions with seam and first- and second-order continuity losses. History corruption and scheduled sampling improve robustness to executed-history errors, while an AWAC fine-tuning stage adapts the policy without removing these structural regularizers. Under mild smoothness assumptions, pre-blending seam discrepancies provably decay with increasing overlap. Experiments on CALVIN, LIBERO, and real robots show an improved success-stability trade-off with low-latency inference. Project page: https://cytoderm-ai.github.io/chunkflow.
Real-time fall detection based on vision for low-power edge platforms
arXiv:2607.12909v1 Announce Type: cross Abstract: Falling detection is vital for elderly care and intelligent surveillance; however, prevailing vision-based approaches predominantly frame it as static pose classification or discrete temporal pattern matching, fundamentally overlooking the instability dynamics of the human support system. This paper proposes a physics-informed falling detection framework that recasts falling as a stability-loss event in a coupled dynamical system. We introduce a novel dual-LTC architecture comprising a Center-of-Mass (CoM) subsystem and a Base-of-Support (BoS) subsystem, both instantiated as Liquid Time-Constant (LTC) neural networks to continuously model inertial trajectory evolution and ground-contact adjustment through adaptive time constants, Physical interpretability of falling motion. A learnable coupling module emulates physical interaction between the two subsystems, while a Stability Manifold classifier operates in the joint latent space to detect boundary crossing via Lyapunov-inspired stability metrics. Complementary counterfactual trajectory projection and Time-to-Collision (TTC) estimation further enable irreversibility assessment and early warning. The architecture is designed to support a three-state prediction paradigm (Normal, Falling, Fallen); in this preliminary study, we validate the core stability discrimination capability on a two-class dataset (Normal vs. Falling), leaving the full three-state temporal transition to future work. Unlike conventional CNN--RNN pipelines, the proposed formulation encodes continuous-time mechanical inertia, yielding a sub-50K-parameter network capable of real-time inference on resource-constrained edge devices. Extensive experiments demonstrate competitive accuracy with superior physical interpretability, validating its efficacy for low-compute visual fall detection.
Conditional Enhancement of Radical Pair Dynamics via Chiral State Preparation
arXiv:2605.22130v2 Announce Type: replace Abstract: Chiral-induced spin selectivity (CISS) has been shown to enhance magnetic sensitivity in radical pair mechanism (RPM) models under specific Hamiltonian conditions, yet whether these enhancements persist across a broader parameter space remains untested. We incorporate the CISS effect as a spin-dependent initial state and recombination operator and systematically evaluate the spin dynamics of a model radical pair across a comprehensive parameter sweep of the RPM Hamiltonian. We characterise the orientational response through symmetric and antisymmetric decomposition of the yield distribution under field reversal, providing a direct quantitative signature of CISS-induced symmetry breaking. Our analysis demonstrates that CISS does not function as a generic amplifier of magnetic sensitivity. Claimed enhancements are conditional on the relative alignment of the internal hyperfine and dipolar interaction axes, arising specifically under conditions of non-collinear internal interactions. Extension to a two-nucleus model confirms that these enhancements are sensitive to nuclear spin. CISS-induced effects observed in the single-nucleus model are substantially suppressed when a second collinear nucleus is introduced, with the exception of the hyperfine axis rotation sweep where non-collinear tensor misalignment drives a robust antisymmetric response. These findings indicate that the conditions for CISS-enhanced magnetoreception are more stringent than previously demonstrated, requiring highly ordered and rigid molecular geometries to sustain the effect.
Learning Latent Energy-Based Models via Interacting Particle Langevin Dynamics
arXiv:2510.12311v2 Announce Type: replace-cross Abstract: We develop interacting particle algorithms for learning latent variable models with energy-based priors. To do so, we leverage recent developments in particle-based methods for solving maximum marginal likelihood estimation (MMLE) problems. Specifically, we provide a continuous-time framework for learning latent energy-based models, by defining stochastic differential equations (SDEs) that provably solve the MMLE problem. We obtain a practical algorithm as a discretisation of these SDEs and provide theoretical guarantees for the convergence of the proposed algorithm. Finally, we empirically validate the effectiveness of our method on synthetic and image datasets and demonstrate that using a particle based approach offers significant improvement in computational efficiency.
PRISM Edit: One Vector for All Temporal Answers
arXiv:2607.11327v2 Announce Type: replace Abstract: Model editing keeps large language models (LLMs) up to date without retraining, but temporal facts expose a limitation of the prevailing locate-and-edit paradigm: an update is not always a replacement. When a fact changes, the new answer should become current while the old answer may remain correct in historical time contexts. Building on this insight, we use causal tracing to show that LLMs already support this distinction via a two-stage internal computation: early MLP layers retrieve a time-agnostic subject representation, and later layers modulate it with temporal context to yield the time-correct answer. Motivated by this finding, we introduce PRISM Edit, which optimizes a single polysemous representation across temporal contexts and leverages the model's inherent modulation pathway to route it to temporally correct predictions, without any architectural modification. We evaluate on TimeConflict, a new temporal editing benchmark we introduce, and on temporally augmented CounterFact. PRISM Edit improves over the best baseline by +23.3 Temporal Consistency (TC) and +33.7 Current Relative-time Score (CRS) on average while being more than 2x faster. Code and data are publicly available at https://github.com/AnonymousStudy972/PRISM-Edit.
Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
arXiv:2606.10829v2 Announce Type: replace Abstract: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Existing training-free samplers such as Top-\(k\), Fast-dLLM, and EB-Sampler mainly control how many tokens to reveal, while often ranking candidates by token-wise scores that ignore interactions within the selected set. We propose ADAS, a training-free reranking rule for parallel masked diffusion decoding. ADAS leaves the base sampler's stopping rule unchanged and modifies only subset construction: it greedily discounts a candidate when it attends strongly to already selected positions whose predictions remain uncertain. Unlike graph-constrained methods that turn attention into hard compatibility constraints, ADAS keeps attention continuous and uses it as a soft marginal penalty. Across LLaDA-8B-Base and Dream-7B-Base on GSM8K, MATH500, HumanEval, and MBPP, plugging ADAS into Top-\(k\), Fast-dLLM, and EB-Sampler improves low-NFE performance at matched denoiser evaluations by \(9.11\) and \(10.46\) percentage points on average, respectively, with \(3.1\%\) per-forward runtime overhead. These results show that soft attention-discounted reranking is a simple and modular way to improve quality in highly parallel decoding for masked diffusion language models.
Double negation stable h-propositions in cubical sets
arXiv:2209.15035v2 Announce Type: replace-cross Abstract: We give a construction of classifiers for double negation stable h-propositions in a variety of cubical set models of homotopy type theory and cubical type theory. This is used to give some relative consistency results: classifiers for double negation stable propositions exist in cubical sets whenever they exist in the metatheory; the Dedekind real numbers can be added to homotopy type theory without changing the consistency strength; we construct a model of homotopy type theory with extended Church's thesis, which states that all partial functions with double negation stable domain are computable.
Imputation-free transformer learning enables robust Alzheimer's disease prediction and calibrated uncertainty quantification across heterogeneous clinical cohorts
arXiv:2607.11656v2 Announce Type: replace-cross Abstract: Accurate diagnostic classification and disease-severity prediction for Alzheimer's disease are hampered by the incompleteness and heterogeneity of real-world clinical data. Left unaddressed, these barriers prevent reliable disease modelling and hinder effective clinical evaluation. Conventional imputation strategies introduce systematic bias, distort inter-feature relationships, and yield overconfident predictions, limitations especially consequential in diagnostic settings. Here, we propose NITROGEN, an imputation-free transformer that jointly models within-patient feature dependencies and between-patient relational structure through masked and intersample attention, enabling robust multimodal learning directly from partially observed records. We trained NITROGEN on ADNI (N=7858 scans), and evaluated it on two independent cohorts: OASIS-3 (N=2675 scans) and AIBL (N=1286 scans). Across cohorts and diagnostic and cognitive score prediction tasks, NITROGEN showed robust calibration and uncertainty quantification advantages over tree-based ensemble methods, while maintaining competitive discriminative performance. Cross-cohort and cross-method analyses identified cortical thickness in the temporal pole, age, and APOE genotype as important, though not individually sufficient, features for AD classification. We further introduced a modality-aware uncertainty adjustment that augments predictive uncertainty proportionally to the importance of absent modalities, enabling calibrated confidence when diagnostic information is unavailable. Together, our results show that imputation-free attention learning preserved meaningful discrimination under cohort shift, revealing expected degradation on more distributionally different cohorts, and demonstrate that evaluating models along calibration, interpretability, and cross-cohort reliability, not accuracy alone, is essential for clinical deployment.
Intermodal entanglement in a quantum optical model of HHG due to the back-action on the driving field
arXiv:2603.01315v3 Announce Type: replace-cross Abstract: Preparation of nonclassical light with special quantum properties is essential for quantum technologies. High-harmonic generation (HHG) is a process which not only enables the creation of attosecond pulses but also has the potential to generate light with intricate quantum properties. In a recent experiment [PRX Quantum 5. 040319], nonclassical inter-harmonic correlations have been measured from a HHG source between low-order harmonics. In this work, we theoretically investigate entanglement between different harmonics within an effective, phenomenological quantum optical model. This model implements a significant degree of simplification regarding the processes within the target material, treating the material through susceptibilities, as it is usual in quantum optics. Such an approach yields a general description of HHG in the few-harmonic generation regime, permitting the implications that can be derived within it to hold broadly within the domain of validity. We find that entanglement is produced as a result of the often neglected back-action. We can qualitatively reproduce experimentally measured nonclassicalities, which suggests that intermodal entanglement can, to an extent, be considered a universal phenomenon associated with HHG, rather than a result of using specific material targets.
Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale
arXiv:2607.06233v2 Announce Type: replace Abstract: LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle to generalize to unseen data environments and analytical workflows, especially in heterogeneous enterprise settings. This creates a growing need for synthesizing high-quality data agent trajectories that capture complex analytical workflows for given data environments. Such trajectories support two key downstream uses: they can serve as supervised finetuning (SFT) data that adapts data agent models to the target domain, and as in-context learning (ICL) demonstrations to guide general-purpose LLMs in unfamiliar data environments. Thus, we introduce TOFFEE, a system for synthesizing high-quality data agent trajectories from given data environments via Monte Carlo Tree Search (MCTS) with adaptive model selection and cross-task prefix reuse. We show that TOFFEE can effectively generate scalable trajectory data for complex analytical tasks across heterogeneous environments. In this demonstration, we present the system framework of TOFFEE, including its task pool construction, trajectory explorer, and learned cost model. We also introduce the web interface of TOFFEE and its workflow, and demonstrate two end-to-end scenarios: trajectory synthesis for data agent finetuning, and demonstration-augmented data agent reasoning.
Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models
arXiv:2607.12739v1 Announce Type: new Abstract: A language model may be asked either what experts believe about a contested claim or what it believes about the claim itself. A trustworthy conversational agent should distinguish these two requests and respond in different epistemic registers: neutral attribution in the first case and stance expression in the second. Whether such a shift occurs-and whether it occurs coherently-is not directly assessed by existing benchmarks for accuracy, instruction following, or safety. We introduce ESFP, a behavioral benchmark that treats the contrast between externally attributed and self-attributed prompts as the fundamental unit of measurement. ESFP consists of 104 carefully controlled items spanning six epistemic categories and five phrasing templates, and evaluates model responses along four complementary dimensions: lexical self-attribution, representation-level responsiveness to role framing, sentence-level stance content density assessed by an LLM judge panel, and cross-condition stance consistency. Evaluating eight frontier models from five vendors, we find that epistemic flexibility is largely orthogonal to general model capability: a 27B open-weight model matches the strongest proprietary systems, the flagship model of one family underperforms its lightweight counterpart, and reasoning-optimized models do not consistently exhibit higher flexibility. Stance content density provides the strongest signal, while surface-level lexical markers such as 'I think' can change substantially without corresponding changes in expressed stance. We provide item-level bootstrap confidence intervals, weight-sensitivity analyses, and an explicit discussion of the interpretation limits of the composite score. ESFP measures a model's propensity to adapt its epistemic stance under changing attribution conditions, rather than a general competence measure.
Superimposed Transmission for Cooperative Cellular and Cell-Free Massive MIMO Systems
arXiv:2607.12709v1 Announce Type: new Abstract: This paper proposes a superimposed transmission strategy for cooperative cellular and cell-free massive MIMO systems. By classifying users into near and far, the base station transmits an additional data symbol for each near user, superimposed on the signals from distributed access points. Successive interference cancellation is employed at near-user receivers to decode both symbols. The proposed strategy achieves the highest peak spectral efficiency while maintaining fairness at the cell edge, thereby outperforming all the existing network configurations in system capacity.
High-frequency magnetotransport in LaMnO3 samples synthesized by microwave irradiation versus conventional heating
arXiv:2607.12689v1 Announce Type: new Abstract: Magnetoresistance of hole-doped LaMnO3 at frequencies above a few MHz has been seldom reported compared to numerous studies on magnetoresistance measured at direct current. Here, we contrast the high-frequency magnetoresistance in the frequency range f = 0.9-3 GHz in polycrystalline LaMnO3 (LMO) synthesized by microwave irradiation of oxide precursors in a microwave furnace (MW-LMO) and by conventional heating (CH-LMO) in an electrical furnace. The structure at room temperature changed from orthorhombic in CH-LMO to rhombohedral in MW-LMO. A combination of magnetization and resistivity studies suggest that the CH-LMO sample is a canted antiferromagnetic insulator below 140 K but the MW-LMO is a ferromagnetic metal with T_C = 240 K. While the high-frequency resistance of the CH-LMO sample at 300 K decreases gradually with increasing strength of the applied dc magnetic field for all f, a peak appears at H = H_r larger than 0 gauss when f equal or greater than 1.4 GHz in the MW-LMO sample and H_r increases linearly with f. We attribute the observed features in the high-frequency magnetoresistance of MW-LMO to current-driven resonant excitation of spins in the paramagnetic state. Its absence in the CH-LMO sample highlights the need for an optimum hole density to observe this effect.
Low-Precision Rank Compensation for Matrices and Tensor Trains
arXiv:2607.12969v1 Announce Type: new Abstract: Lower numerical precision reduces storage and memory traffic but raises the perturbation floor. We study rank compensation: reinvesting saved memory in a larger approximation rank. For matrices, the singular-value error identity yields a directly testable sufficient condition requiring the additional singular component to offset the perturbation from storing the rank-augmented approximation in lower precision. On ten SuiteSparse matrices, all 100 truncation-dominated configurations (50 FP32 and 50 FP16) are certified non-increases and strict accuracy wins, with mean error ratio $0.963$ and storage ratios $58.8\%$ and $29.4\%$ relative to the FP64 baseline. FP16 failures occur only in tail-rank stress tests near the perturbation floor. At the largest resident matrix-application batch, compensated FP32 and FP16 achieve geometric-mean A100 speedups of $1.28\times$ and $2.12\times$; neither accelerates the smallest batch. For Tensor-Train (TT) approximation, we give a conditional a posteriori extension based on the measured truncation gain and rounded-core perturbation. Across three-way and six-way synthetic tests, FP32 and FP16 achieve combined accuracy-memory wins in 10 of 20 and 14 of 20 trials. On public hyperspectral tensors and FROSTT top-active subtensors, the corresponding counts are 44 of 60 and 54 of 60; four FP16 Salinas-A tail-stress cases fail. No certified TT case exceeds the FP64 error beyond numerical tolerance. Reconstruction of six public tensors yields geometric-mean compensated speedups of $1.38\times$ (FP32) and $1.94\times$ (FP16). Timings cover resident downstream kernels, not factorization, transfers, or end-to-end acceleration.
Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation
arXiv:2607.12986v1 Announce Type: new Abstract: Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged expected-value scorer for LLM-generated venture routes. Proposition 1 gives the score change from deleting an interior transition while retargeting its predecessor and retaining downstream value: Delta_k = (prod_{i<k} p_i)[c_k + (1 - p_k)R_{k+1}]. On a frozen 26-route cohort, all 57 admissible deletions matched the analytic identity and threshold sign, and every route had at least one score-improving deletion. A score-seeking optimizer, allowed to restructure routes but not told the exploit mechanism, found baseline-beating uncovered structures in 21/26 routes. GATE refused score release for 26/26 silenced routes with 0/26 honest suspensions; after refusal, 47/54 next revisions repaired to a covered structure, and strict covered improvement rose from 1/26 to 13/26. An adaptive compiler-aware co-author exposed the registry-provenance boundary: obligation-channel evasions remained 6/6 across all four v1/v1.5 conditions, while delta-indexed cost floors reduced beat-honest routes from 6/6 to 3/6 and fundability-by-silence from 5/6 to 0/6 without establishing semantic completeness. If a plan scores better only because it omits necessary work, the plan did not improve; the evaluation created an omission incentive. PCSC detects and neutralizes post-hoc omission splices over model-mediated typed-state records. In the cooperative setting tested, GATE acts as a deterministic search-shaping constraint, not merely a post-hoc filter. It does not verify the semantic completeness or real-world quality of arbitrary LLM-generated strategies.
A cryogenic neutral-atom platform with full optical access and 2-hour trap lifetime
arXiv:2607.12988v1 Announce Type: new Abstract: Neutral-atom quantum processors are rapidly scaling toward system sizes of more than ten thousand qubits, allowing for the realization of a new class of quantum computing algorithms and quantum simulation experiments. However, current neutral-atom platforms generally have to find a compromise between the optical accessibility and the storage time of atoms in optical potentials, limiting the available qubit numbers. Here we report on the operation of a novel, cryogenically enhanced, neutral-atom apparatus that overcomes these apparently conflicting requirements. We demonstrate vacuum-limited trapping lifetimes of up to two hours of single $^{88}\mathrm{Sr}$ atoms in an optical tweezer array while preserving full optical access and without the need for complex cryogenic enclosures. Our measurements show that exceptionally long single-atom lifetimes can be achieved with a relatively simple cryostat design. Our architecture can be straightforwardly ported to other atomic species and shows a viable path for scaling up to sorted arrays of tens of thousands of atoms.
Measurement Risk in Supervised Financial NLP: Rubric and Metric Sensitivity on JF-ICR
arXiv:2604.27374v2 Announce Type: replace Abstract: As LLMs become credible readers of earnings calls, investor-relations Q\&A, guidance, and disclosure language, supervised financial NLP benchmarks increasingly function as decision evidence for model selection and deployment. A hidden assumption is that gold labels make such evidence objective. This assumption breaks down when the benchmark ruler itself is sensitive to rubric wording, metric choice, or aggregation policy. We study this measurement risk on Japanese Financial Implicit-Commitment Recognition (JF-ICR; a pinned 253-item test split x 4 frontier LLMs x 5 rubrics x 3 temperatures x 5 ordinal metrics). Three findings follow. First, rubric wording materially changes model-assigned labels: R2--R3 agreement ranges from 70.0% to 83.4%, with the dominant movement near the +1 / 0 implicit-commitment boundary. This pattern is consistent with a pragmatic-boundary interpretation, but is not a validated linguistic-causality claim because the present rubric variants confound semantics, examples, and verbosity. Second, not every metric remains informative under the JF-ICR class distribution. Within-one accuracy is too easy because near misses receive credit and the majority class dominates; worst-class accuracy is too noisy because the rarest class has only two examples. Exact accuracy, macro-F1, and weighted \k{appa} are therefore the identifiable metrics under our operational rule. Third, ranking claims become more defensible only after this metric-identifiability audit: Bradley--Terry, Borda, and Ranked Pairs agree on the identifiable metric subset, while the full five-metric sweep produces disagreement on the closest pair. The contribution is not a new leaderboard, but a reporting discipline for supervised financial benchmarks whose gold labels exist and whose evaluation ruler still requires governance.
Software Supply Chains are Dead: Use-Case-Oriented Regeneration
arXiv:2607.13021v1 Announce Type: new Abstract: Modern software development relies on an increasingly doubtful premise: that the up-front implementation savings from adopting a dependency outweighs the maintenance costs. Two changes are reshaping the build-vs.-reuse calculus: software supply chain attacks have raised the cost of external reliance, while generative AI has lowered the cost of local implementation. We envision use-case-oriented regeneration as a new software sourcing paradigm that shifts the supply chain from external trust to local verification. We evaluate an agentic workflow that synthesizes only the specific slice of dependency functionality that a repository exercises. Our measurements across 180 repository-dependency pairs suggest that this approach is feasible: the replacements preserve 99.8% of repository-observed behavior across baseline validation checks and reduce the exported API surface by 93%. Software sourcing may evolve toward verifiable repository-specific code synthesis, especially when the required functionality is narrow, stable, and well tested.