Forskningsradar

Science Journals

Peer-reviewade publikationer — 55347 artiklar

Metasurface Engineering with Tantalum Pentoxide-Coated Microspheres: Tailoring Optical Resonances and Enhancing Local Density of States
arXiv:2603.25828v2 Announce Type: replace Abstract: Hexagonally-packed polystyrene microsphere monolayers coated with tantalum pentoxide (Ta$_2$O$_5$) form scalable dielectric metasurfaces that support tunable photonic resonances and enhanced local density of optical states (LDOS). Here we combine fabrication, optical and fluorescence spectroscopy, and multiscale electromagnetic simulations to quantify how the thickness of the Ta$_2$O$_5$ shells control far-field resonances and Rhodamine 6G (Rh6G) emission. Experimentally, Ta$_2$O$_5$ shells of 10 - 70 nm deposited on microsphere lattices generate resonances that shift red with the thickness of the shell and systematically enhance the Rh6G fluorescence relative to flat Ta$_2$O$_5$ films. The largest enhancement is obtained for 30 - 50 nm shells, when lattice resonances overlap the Rh6G excitation and emission bands. Finite-cluster finite-difference time-domain simulations reproduce the measured transmittance and reflectance spectra, confirming the assumed geometry of the Ta$_2$O$_5$ shells covering the sphere lattice. Periodic-cell simulations of single electric dipoles yield wavelength-dependent Purcell factors $Fp(\lambda)$ and directional $\beta$-factors $\beta_{top}(\lambda)$, from which we construct emission-weighted figures of merit that link LDOS modulation to the experimentally accessible top-side fluorescence enhancement. As a complementary test of our emitter-environment model, we compare simulated and measured Purcell factors for PS/Ta$_2$O$_5$ microsphere lattices. A physically motivated averaging that accounts for emitter position, orientation and ensemble spectral smoothing yields very good agreement across all shells. Overall, our results establish Ta$_2$O$_5$-coated microsphere lattices as robust dielectric substrates for surface-enhanced fluorescence and clarify how shell thickness and emitter placement jointly control photonic resonances, LDOS and fluorescence response.
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
arXiv:2603.26648v3 Announce Type: replace Abstract: Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end website development remains limited. To address this gap, we introduce Vision2Web, a hierarchical benchmark for visual website development, spanning from static UI-to-code generation, interactive multi-page frontend reproduction, to long-horizon full-stack website development. The benchmark is constructed from real-world websites and comprises a total of 193 tasks across 16 categories, with 918 prototype images and 1,255 test cases. To support flexible, thorough and reliable evaluation, we propose workflow-based agent verification paradigm based on two complementary components: a GUI agent verifier and a VLM-based judge. We evaluate multiple visual language models instantiated under different coding-agent frameworks, revealing substantial performance gaps at all task levels, with state-of-the-art models still struggling on full-stack development.
Retromorphic Testing with Hierarchical Verification for Hallucination Detection in RAG
arXiv:2603.27752v2 Announce Type: replace Abstract: Large language models can still hallucinate in retrieval-augmented generation (RAG), producing claims that are unsupported by or conflict with the retrieved context. Detecting such errors remains challenging when faithfulness is judged solely against the retrieved context: many existing detectors return holistic answer-level scores, while others target open-domain factuality or fail to provide evidence-grounded diagnostics. We present RT4CHART, a retromorphic testing framework for context-faithfulness assessment. RT4CHART decomposes an answer into independently verifiable claims, performs hierarchical local-to-global verification against the retrieved context, and assigns each claim one of three labels: entailed, contradicted, or baseless. It further maps these claim-level decisions back to specific answer spans and returns explicit context-side evidence, enabling fine-grained auditing rather than opaque scoring. We evaluate RT4CHART on RAGTruth++ (408 samples) and our re-annotated RAGTruth-Enhance (2,675 samples). RT4CHART achieves the best answer-level hallucination-detection F1 score among the evaluated baselines. On RAGTruth++, it attains a precision of 0.845, a recall of 0.718, and an F1 score of 0.776, representing an 83% relative improvement over the strongest baseline. It also achieves a span-level F1 score of 47.5% on RAGTruth-Enhance. Ablation studies show that claim-based local processing drives most of the observed improvement, while global verification provides selective benefits across datasets. Finally, our re-annotation identifies 1.68X more hallucination cases than the original labels, suggesting that commonly used benchmarks substantially underestimate the prevalence of hallucination.
Phasor Attention: Mean Root Square Normalization for Phase Manifold Preservation
arXiv:2607.17822v1 Announce Type: new Abstract: While Root Mean Square Normalization has become the de facto standard for accelerating modern sequence models, its reliance on the quadratic accumulation of independent scalars ($\sum x^2$) inherently triggers outlier-induced numerical instability, gradient starvation, and anisotropic phase distortion. We introduce Mean Root Square Normalization (MRSNorm). By structurally pairing channels into 2D phasors, MRSNorm mathematically inverts the traditional scaling paradigm: it computes the localized $L_2$ magnitudes (Root Square) before aggregating them via a global $L_1$ average (Mean). This operational inversion strictly constrains activations to a phasor manifold, preserving conformal invariance. By sharing a single affine weight across phasor components, MRSNorm halves the total number of learnable parameters, proving that unconstrained spatial scaling in standard norms is a harmful redundancy. We analytically demonstrate that this geometric constraint yields a built-in, trigonometric gradient clipper governed by the Pythagorean identity, unconditionally equalizing the local gradient norm to ensure Gradient Homogeneity. Empirical evaluations on a ResNet with CIFAR-100 show that despite halved parameters, MRSNorm provides critical structural stability under rigorous stress tests. Under extreme hyperparameter settings where standard normalizations suffer from gradient divergence, MRSNorm successfully prevents numerical explosion and secures stable optimization trajectories. Our findings propose a fundamental paradigm shift toward phasor-based deep representation learning. The implementation of MRSNorm is available at Appendix C.
Quantum dynamics of a levitated ferromagnetic gyroscope
arXiv:2607.16592v1 Announce Type: cross Abstract: We develop a quantum model for the rotational dynamics of a freely floating levitated ferromagnetic gyroscope (LFG), emphasizing the interplay between intrinsic spin $\boldsymbol{S}$, mechanical angular momentum $\boldsymbol{L}$, and magnetic torque. The conserved total angular momentum projection along the $z$-directed magnetic field $\boldsymbol{B}$, $J_z=S_z+L_z$, is quantized, leading in the small-libration-amplitude limit to discrete precessional states $|m\rangle$ (eigenstates of $J_z$ with eigenvalues $J_z = m\hbar$) and librational harmonic oscillator states $|n\rangle$ ($n=0,1,2,\ldots$). The discreteness of the energies and dynamical variables is governed by the quantum precession scale $\Omega_Q=\hbar/I$, where $I$ is the moment of inertia of the LFG. We find that the phenomenon of LFG precession persists into high-field regimes where the magnitude of the rotational angular momentum associated with precession exceeds the total intrinsic spin. We analyze the complementary quantum limits of localized semiclassical LFG orientation wave packets and exact $J_z$-eigenstates $|m\rangle$, clarifying the relation between classical precession signals and the underlying quantized spin-rotor dynamics. We further show that radio-frequency fields can drive $\Delta m = \pm 1$ and $\Delta n = \pm 1$ transitions, enabling ladder spectroscopy, tilt-angle control, and sideband-like coupling between precession and libration. The coupled dynamics also exhibit branch-point magnetic resonances where precession and librational motion become strongly coupled. These results establish a framework for using LFGs not only as ultrasensitive torque and magnetic-field sensors, but also as controllable mesoscopic quantum systems. The techniques developed here may be applied to searches for exotic, beyond-the-standard model spin-dependent interactions, ultralight dark matter, and spin-gravity couplings.
SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection
arXiv:2511.21325v2 Announce Type: replace Abstract: Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs. A central reason is spectral bias, the tendency of neural networks to learn low-frequency structure before high-frequency (HF) details, which both causes DF generators to leave HF artifacts and leaves those same artifacts under-exploited by common detectors. To address this gap, we propose Spectral-cONtrastive Audio Residuals (SONAR), a frequency-guided framework that explicitly disentangles an audio signal into complementary representations. An XLSR encoder captures the dominant low-frequency content, while the same cloned path, preceded by learnable SRM, value-constrained high-pass filters, distills faint HF residuals. Frequency cross-attention reunites the two views for long- and short-range frequency dependencies, and a frequency-aware Jensen-Shannon contrastive loss pulls real content-noise pairs together while pushing fake embeddings apart, accelerating optimization and sharpening decision boundaries. Evaluated on the ASVspoof 2021 and in-the-wild benchmarks, SONAR attains state-of-the-art performance and converges four times faster than strong baselines. By elevating faint high-frequency residuals to first-class learning signals, SONAR unveils a fully data-driven, frequency-guided contrastive framework that splits the latent space into two disjoint manifolds: natural-HF for genuine audio and distorted-HF for synthetic audio, thereby sharpening decision boundaries. Because the scheme operates purely at the representation level, it is architecture-agnostic and, in future work, can be seamlessly integrated into any model or modality where subtle high-frequency cues are decisive.
Inside Qubic's Selfish Mining Campaign on Monero: Evidence, Tactics, and Limits
arXiv:2512.01437v3 Announce Type: replace Abstract: Qubic's 2025 campaign against Monero provides a rare public case of selfish mining in a privacy-preserving proof-of-work system. We study what can be measured when block ownership, private forks, and release decisions are only partially observable. We combine Monero node data, Qubic pool job observations, coinbase extra-nonce patterns, community-observed Qubic blocks, and view keys later disclosed by Qubic operators to identify withholding intervals and validate attribution. Our analysis finds no evidence that Qubic sustained majority mining power, despite public takeover claims. The campaign did, however, increase orphaning and reorganization depth. Its release behavior is not explained by one fixed selfish-mining policy: most visible releases resemble lead-one behavior, while some periods show more conservative or timing-dependent releases. This variation matters economically. The withholding intervals do not show consistent reward gains over honest mining, mainly because tie-breaking success was low and release execution varied. Difficulty-adjustment spillovers later increased rewards during non-selfish gaps, so withholding intervals and gap phases must be analyzed separately. We conclude that the incident is best understood as a blend of imperfect selfish-mining execution, delayed difficulty effects, and disruption-oriented incentives, with community monitoring and Qubic countermeasures limiting what can be inferred from public artifacts.
Preventing Learning Stagnation in PPO by Scaling to 1 Million Parallel Environments
arXiv:2603.06009v2 Announce Type: replace Abstract: An agent's performance stagnating at a suboptimal level is a common problem in deep on-policy RL. Focusing on PPO, we show that plateaus in certain regimes arise not because of known exploration, capacity, or optimisation challenges, but because sample-based estimates of the loss eventually become poor proxies for the true objective over the course of training. Looking deeper, PPO alternates between sampling rollouts from several parallel environments online using the current policy (which we call the "outer loop") and performing repeated minibatch SGD steps against this offline dataset (the "inner loop"). In our work, we abstract away the inner loop, and conceptually model the outer loop as standard stochastic optimisation. The step size is then controlled by the regularisation strength towards the previous policy and the gradient noise by the number of samples collected between policy update steps. This framing predicts that, much like in SGD, if the outer step size is too large relative to the noise, updates become uninformative and lead to the policy thrashing around a local optimum instead of converging. Recasting PPO in this light makes it clear that there are two ways to address this particular type of learning stagnation: either reduce the step size or increase the number of samples collected between updates. We validate the predictions of our model and conclude that increasing the number of parallel environments is a simple way to avoid these plateaus by simultaneously altering both these factors. Applying our analysis and scaling PPO to more than 1M parallel environments enables monotonic performance improvement up to one trillion transitions and leads to vastly superior performance compared to prior baselines in a complex open-ended domain.
Semilinear single-track vehicle models with distributed tyre friction dynamics
arXiv:2601.06854v3 Announce Type: replace Abstract: This paper introduces a novel family of single-track vehicle models that incorporate a distributed representation of transient tyre dynamics, whilst simultaneously accounting for nonlinear effects induced by friction. The core of the proposed framework is represented by the distributed Friction with Bristle Dynamics (FrBD) model, which unifies and extends classical formulations such as Dahl and LuGre by describing the rolling contact process as a spatially distributed system governed by semilinear partial differential equations (PDEs). This model is systematically integrated into a single-track vehicle framework, where the resulting semilinear ODE-PDE interconnection captures the interaction between lateral vehicle motion and tyre deformation. Two main variants are considered: one with rigid tyre carcass and another with flexible carcass, each admitting a compact state-space representation. Local and global well-posedness properties for the coupled system are established rigorously, highlighting the dissipative and physically consistent properties of the distributed FrBD model. A linearisation procedure is also presented, enabling spectral analysis and transfer function derivation, and potentially facilitating the synthesis of controllers and observers. Numerical simulations demonstrate the model's capability to capture micro-shimmy oscillations and transient lateral responses to advanced steering manoeuvres. The proposed formulation advances the state-of-the-art in vehicle dynamics modelling by providing a physically grounded, mathematically rigorous, and computationally tractable approach to incorporating transient tyre behaviour in lateral vehicle dynamics, when accounting for the effect of limited friction.
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune
arXiv:2607.18213v1 Announce Type: new Abstract: Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code classifier, we find the agent itself encodes internal representations indicating the relevance of code context when reading tool output. Based on this finding, we propose SWE-Pruner Pro, which prunes tool outputs directly inside the agent. Concretely, a small head turns the agent's own internal representations into a keep-or-prune label for each line, with a length-aware embedding keyed to each tool output's line count. Across two open-weight backbones and four multi-turn benchmarks, SWE-Pruner Pro saves up to 39% of prompt and completion tokens while preserving task quality, with bounded inference overhead. Notably, on MiMo-V2-Flash SWE-Pruner Pro additionally raises the SWE-Bench Verified resolve rate by +3.8% and the long-context Oolong accuracy by +2.2 points.
Parallel Decoder Transformer: Planner-Conditioned Latent Coordination for Model-Intrinsic Parallel Generation
arXiv:2512.10054v3 Announce Type: replace Abstract: Autoregressive language models expose one causal token frontier, even when the requested document contains sections that could be developed concurrently. Existing parallel-generation systems arrange external branches around an otherwise unchanged model. We instead formulate model-intrinsic parallel generation: a single trained architecture owns multiple causal frontiers and produces one next-token distribution for each frontier in every synchronized decoding round. The Parallel Decoder Transformer (PDT) retains a frozen shared lower knowledge trunk and replaces the upper trunk with three independently parameterized physical decoder stacks. A prompt-time set planner produces three unordered continuous outlines, each hard-routed to one decoder as persistent Plan-KV memory, while a finite product-quantized notes bus carries block-delayed latent messages among the decoders. Autoregression is preserved within each lane; same-round lane tokens are conditionally independent given the source, plans, private histories, and previously committed messages. We specify source-grounded supervision for long-form historical exposition with single-owner cited facts and token-aligned cross-lane dependencies, a composite objective, a staged curriculum, and preregistered causal evaluations: plan swap and removal, delayed-message ablation, a parameter-matched self-only control, dependency-token likelihood, and blinded human fact audits. The architecture and evaluation pipeline are implemented; scientific training and held-out evaluation are in progress. This paper presents the theory, design, and falsifiable protocol, not a positive empirical result.
CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs
arXiv:2603.27958v2 Announce Type: replace Abstract: Analogical reasoning tests a fundamental aspect of human cognition: mapping the relation from one pair of objects to another. Existing evaluations of this ability in multimodal large language models (MLLMs) overlook the ability to compose rules from multiple sources, a critical component of higher-order intelligence. To close this gap, we introduce CARV (Compositional Analogical Reasoning in Vision), a novel task together with a 5,500-sample dataset as the first diagnostic benchmark. We extend the analogy from a single pair to multiple pairs, which requires MLLMs to extract symbolic rules from each pair and compose new transformations. Evaluation on the state-of-the-art MLLMs reveals a striking performance gap: even Gemini-2.5 Pro achieving only 40.4% accuracy, far below human-level performance of 100%. Diagnostic analysis shows two consistent failure modes: (1) decomposing visual changes into symbolic rules, and (2) maintaining robustness under diverse or complex settings, highlighting the limitations of current MLLMs on this task.
HiCI: Hierarchical Construction-Integration for Long-Context Attention
arXiv:2603.20843v3 Announce Type: replace Abstract: Long-context language modeling is commonly framed as a scalability challenge of token-level attention, yet local-to-global information structuring remains largely implicit in existing approaches. Drawing on cognitive theories of discourse comprehension, we propose HiCI (Hierarchical Construction--Integration), a hierarchical attention module that constructs segment-level representations, integrates them into a shared global context, and broadcasts both to condition segment-level attention. We validate HiCI through parameter-efficient adaptation of LLaMA-2 with only <5.5% additional parameters, extending context from 4K to 100K tokens (7B) and 64K tokens (13B). Across language modeling, retrieval, and instruction-following benchmarks, HiCI yields consistent improvements over strong baselines, including matching proprietary models on topic retrieval and surpassing GPT-3.5-Turbo-16K on code comprehension. These results demonstrate the effectiveness of explicit hierarchical structuring as an inductive bias for long-context modeling.
Teaching AI Interactively: An Experience Report in Higher Education
arXiv:2603.28679v2 Announce Type: replace Abstract: Introductory artificial intelligence (AI) courses present significant learning challenges due to abstract concepts, mathematical complexity, and students' diverse technical backgrounds. This paper presents an experience report examining the redesign of in-class instructional time in a university-level Introduction to Artificial Intelligence course, inspired by CS Unplugged approaches. We redesigned the summer offering, integrating embodied, unplugged simulations, collaborative programming labs, and structured reflection to provide students with a first-person perspective on AI decision-making. We maintained identical assignments, exams, and assessments as the traditional lecture-based offering. We found that students in the redesigned course reported higher attendance, stronger agreement that assessments measured their understanding, and greater overall course effectiveness, despite no significant differences in self-reported learning. Post-course interviews indicate that unplugged simulations and collaboration fostered a safe, supportive learning environment that increased engagement and confidence with AI concepts. These results highlight the importance of in-class instructional design in improving students' learning experiences without compromising rigor.
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
arXiv:2603.29219v2 Announce Type: replace Abstract: Sign language is the primary approach of communication for the Deaf and Hard-of-Hearing (DHH) community. While there are numerous benchmarks for high-resource sign languages, low-resource languages like Arabic remain underrepresented. Currently, there is no publicly available dataset for Syrian Arabic Sign Language (SyArSL). To overcome this gap, we introduce SyriSign, a dataset comprising 1500 video samples across 150 unique lexical signs, designed for text-to-SyArSL translation tasks. This work aims to reduce communication barriers in Syria, as most news are delivered in spoken or written Arabic, which is often inaccessible to the deaf community. We evaluated SyriSign using three deep learning architectures: MotionCLIP for semantic motion generation, T2M-GPT for text-conditioned motion synthesis, and SignCLIP for bilingual embedding alignment. Experimental results indicate that while generative approaches show strong potential for sign representation, the limited dataset size constrains generalization performance. We will release SyriSign publicly, hoping it serves as an initial benchmark.
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
arXiv:2604.00594v2 Announce Type: replace Abstract: As the focus in LLM-based coding shifts from static single-step code generation to multi-step agentic interaction with tools and environments, understanding which tasks will challenge agents and why becomes increasingly difficult. This is compounded by current practice: agent performance is typically measured by aggregate pass rates on benchmarks, but single-number metrics obscure the diversity of tasks within a benchmark. We present a framework for predicting success or failure on individual tasks tailored to the agentic coding regime. Our approach augments Item Response Theory (IRT) with rich features extracted from tasks, including issue statements, repository contexts, solutions, and test cases, and introduces a novel decomposition of agent ability into LLM and scaffold ability components. This parameterization enables us to aggregate evaluation data across heterogeneous leaderboards and accurately predict task-level performance for unseen benchmarks, as well as unseen LLM-scaffold combinations. Our methods have practical utility for benchmark designers, who can better calibrate the difficulty of their new tasks without running computationally expensive agent evaluations.
Is "Knowing It's Malicious Enough?" Evaluating LLMs for Fine-Grained Malware Behavior Auditing
arXiv:2509.14335v2 Announce Type: replace Abstract: Automated malware classifiers achieve strong detection performance, but auditing requires more than flagging a sample: analysts must explain malicious behaviors and justify them with code evidence. Traditional signature-based methods and learning-based XAI often fail to provide such support in a human-interpretable form. Large Language Models (LLMs) appear promising, yet their reliability for malware auditing remains unclear. Evaluation faces three challenges: (1) the lack of human-written behavioral ground truth; (2) real-world codebases that exceed current context limits; and (3) the lack of reliable mechanisms to verify whether generated claims are grounded in code evidence. These obstacles make benchmarking difficult and leave model capabilities and failure modes opaque. We introduce MalEval, a diagnostic framework for measuring the capability boundaries of LLMs in malware auditing. MalEval pairs real-world application codebases with expert-written audit reports to provide fine-grained behavior-level ground truth. It compresses large codebases into behavior-relevant program contexts through a context-driven intermediate representation that preserves call relations. Expert reports and model outputs are mapped, via constrained reasoning, into structured evidence chains linking code-level facts to high-level behaviors in a shared space. MalEval decomposes auditing into four stage-wise tasks, enabling each intermediate judgment to be verified under limited context windows. We evaluate seven LLMs and find that they rely on surface cues rather than verifiable evidence, struggle to compose dispersed facts into coherent attack chains, and are highly sensitive to context formulation. These findings shift attention from isolated outputs to reliable LLM and agentic workflows for malware auditing. MalEval is publicly available at https://github.com/ZhengXR930/MalEval.git
Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
arXiv:2607.17761v1 Announce Type: new Abstract: Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to controlled additive noise scenarios, lacking systematic investigation of the complex distortions introduced by acoustic front-end (AFE) processing pipelines in real-world deployments. In this work, we simulate a unified AFE pipeline comprising acoustic echo cancellation, noise suppression, automatic gain control, and voice activity detection (VAD), and conduct a comprehensive evaluation of current state-of-the-art models. The results show that the nonlinear and time-frequency coupled distortions introduced by AFE significantly degrade detection performance. To address this issue, we propose a Time-Frequency Consistency Learning (TFCL) framework, which aims to learn invariant spoofing representations that remain stable before and after AFE processing. We observe that AFE not only introduces temporal misalignment (e.g., segment-level shifts caused by VAD), but also weakens or distorts critical frequency-domain cues. To this end, TFCL employs an attention-driven soft alignment mechanism to capture cross-temporal dependencies, along with frequency-domain structural consistency constraints to enforce feature invariance. As a result, the model is able to maintain stable representations under both temporal perturbations and spectral distortions. Extensive experimental results demonstrate that the proposed method effectively mitigates the performance degradation caused by AFE processing, significantly improving the robustness of SDD in real-world scenarios. The code is available at https://github.com/JunXue-tech/TFCL.
Attosecond delay metrology beyond the photon coherence time with spectrally resolved Hong-Ou-Mandel interferometry
arXiv:2607.16849v1 Announce Type: cross Abstract: Hong-Ou-Mandel (HOM) interferometry enables delay estimation at the quantum precision limit but is traditionally constrained to path differences within the coherence time of the interfering photons. Here, we demonstrate single-measurement path-delay sensing at the measurement Cramer-Rao bound using spectrally resolved HOM interference, thereby removing the conventional dynamic-range limitation imposed by the photon coherence window, with no scanning required for calibration. By extracting delay information from the spectral interference fringes of spectrally entangled photon pairs, we retain near-optimal sensitivity over an operational range exceeding the photon coherence time by over two orders of magnitude. Using one million detected photon pairs, we achieve a time-delay precision of 20 attosecond (6 nm), while real-time operation (at 1 Hz) yields 330 attosecond (100 nm) precision. Because the estimator relies on fringe periodicity rather than absolute coincidence rates, the method is intrinsically robust to photon losses and variations in interference visibility, eliminating the need for recalibration. As a practical demonstration, we measure the thickness of a 300 um transmissive target with nanometer-scale precision. These results mark a significant step towards deploying quantum-limited measurements in real-world sensing applications using HOM interferometry.
FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches
arXiv:2607.17765v1 Announce Type: new Abstract: We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events. For every one of the 104 matches of the 2026 FIFA World Cup, four frontier models -- Claude Opus 4.8, ChatGPT (GPT-5.5, high reasoning), Gemini 3.1 Pro, and Grok (Expert Mode) -- ran an identical search-act-reflect loop: gather evidence with a web tool, commit to a 1X2 (team-A win / draw / team-B win) distribution and a virtual 100-USD bet, and, after the match, reflect given only the final score. Because every match kicked off after the models' training cutoffs, the benchmark is contamination-free by construction. Crucially, we pair the four agents with a fifth competitor drawn from the same information environment -- the pre-match betting market -- collected as per-match 1X2 odds, giving an economically grounded baseline and letting us score not just what an agent predicts but what it does with money. The release contains 416 forecasts and 414 reflections with verbatim reasoning, ground truth (including penalty shootouts), odds, and a reproducible evaluation suite. A reference evaluation surfaces findings that raw accuracy hides: the four agents issue an identical top pick in 92% of matches and none beats the market's Brier score; indeed, a naive flat stake on the market favorite out-earns all four agents. Yet the agents diverge sharply as decision-makers: betting return-on-investment ranges from -18% to +10%, fading the market is unprofitable for all four, the share of forecasts that cite the market ranges from 12% to 100%, and self-reported error rates on wrong picks range from 36% to 86%. The benchmark thus measures calibration, decision quality, and self-knowledge -- axes on which frontier models differ even when their predictions do not. Data and code: https://github.com/graphuofm/FIFA2026LLM
Receiver-Centered Robot-to-Human Handover with Grasp-Aware Object Orientation
arXiv:2607.17839v1 Announce Type: new Abstract: Collaborative robots are increasingly sharing workspaces with human operators, making tool handover a frequent and safety-critical micro-interaction. However, traditional static handovers often lead to awkward grasps when handling asymmetric industrial tools. This paper presents a receiver-centered voice-driven adaptive handover system for mechanical tools, built on a Franka cobot. Using an LLM for intention recognition and MediaPipe for real-time 3D hand tracking, the framework dynamically adjusts the end-effector's orientation to present tools in an ergonomically optimal, handle-first pose. A within-subjects study compared this adaptive approach with an object-agnostic static baseline. The results demonstrate that the adaptive system reduces the grasp delay for asymmetric tools, improving the fluency of the interaction. Furthermore, the adaptive strategy improved specific trust-related perceptions, particularly motion predictability and perceived task simplicity.
Hippasus: Effective and Efficient Automatic Feature Augmentation for Machine Learning Tasks on Relational Data
arXiv:2602.02025v2 Announce Type: replace Abstract: ML models critically depend on feature quality, yet in real-world settings, useful features are often distributed across multiple relational tables rather than a single dataset. Feature augmentation addresses this problem by automatically discovering and joining additional tables to enrich a base table with predictive features. However, scaling feature augmentation to complex schemas with many tables and multi-hop relationships is challenging. It requires exploring a large space of join paths, executing costly joins, and selecting useful features from noisy results. Existing approaches suffer from either limited effectiveness or efficiency. Restricting exploration to simple joins limits predictive performance, while more expressive methods rely on expensive training data, lack scalability, or fail to fully exploit schema-level semantics. We present Hippasus, a cost-aware, LLM-augmented feature discovery framework over relational schemas that addresses these challenges. Hippasus combines lightweight statistical signals with adaptive semantic reasoning, invoking stronger (LLM-based) analysis only when necessary. It further introduces efficient multi-way join execution with cross-path feature consolidation, and a hybrid feature selection strategy that integrates statistical relevance with semantic refinement. Experiments on real-world datasets show that Hippasus improves feature augmentation accuracy by up to 26.8% over state-of-the-art methods, while achieving a favorable effectiveness-cost tradeoff.
Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs
arXiv:2607.18086v1 Announce Type: new Abstract: Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved. Using a structured evidence-sufficiency prompt as a test case, we asked whether it reduces unsafe overconfident answers, how far that effect depends on the scoring judge, and what it costs in helpfulness. Methods: In a retrospective public-data benchmark (Real-POCQi, HealthBench, MedRBench), four models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3) answered a fully paired common panel (1,200 cells) with a standard prompt and the wrapper. The pre-specified endpoint was the paired reduction in unsafe overconfidence scored by the primary judge (GPT-5.4-nano); secondary analyses added a different-family judge (Claude Sonnet 5), a correctness judge, matched scaffold controls, and a blinded three-clinician review. Results: Unsafe overconfidence fell from 49.3% to 24.7%, a paired reduction of 24.7 points (95% CI 21.8-27.7; p<0.001), robust in direction across models and paraphrases. Magnitude was judge-dependent: Sonnet agreed on direction but nearly halved the effect (+13.1 points), with one-directional disagreement. Blinded clinicians characterized the primary judge as a high-sensitivity (1.00), low-specificity (0.55) screen, not a calibrated rate. The gain carried a model-specific helpfulness cost (correct diagnosis 80.3% to 50.3%): near-free for GPT-5.5, near-total for Gemini (-58 points). Matched scaffold controls showed genuine behavior change, not judge circularity. Conclusions: LLM-judged clinical safety effects should be reported as directional and relative, anchored to human review and evaluated jointly with helpfulness, not as calibrated absolute rates. This does not establish clinical deployment readiness.
Some prospects for semiproducts and products of modal logics
arXiv:2607.17928v1 Announce Type: cross Abstract: We consider products and semiproducts of propositional modal logics L with S5 and present new examples of product and semiproduct logics axiomatized in the minimal way and enjoying the product (or semiproduct) FMP. An essential part of the proof is local tabularity of these (semi)products for L of finite depth; it is obtained by using bisimulation games. These results readily imply decidability for 1-variable fragments of predicate modal logics QL and QL+Barcan formula. We also present new counterexamples, i.e. (semi)products not axiomatizable in the simplest way.
A Centrality Measure Using Magnitude Homology
arXiv:2607.16377v1 Announce Type: cross Abstract: The magnitude of a metric space constitutes an expressive invariant that subsumes numerous different geometrical-topological invariants. Building on recent advances in magnitude homology, i.e., a bigraded homology theory that recovers the magnitude, we develop a novel local measure of the centrality or importance of nodes in a graph. Our measure is inspired by the concept of relative homology as it considers the change in magnitude homology when removing a vertex. We show that our proposed measure satisfies several properties a centrality measure is reasonably expected to respect and demonstrate that we introduce a new perspective on centrality by comparing to several established centrality measures.