Forskningsradar

Science Journals

Peer-reviewade publikationer — 56237 artiklar

Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents
arXiv:2607.12267v1 Announce Type: new Abstract: Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace this to context dilution: an agent's investigative state (what it has confirmed, what it suspects, and what it still needs) lives only implicitly in a growing context window, where early discoveries are buried under later retrievals. We introduce SLEUTH, which makes this state explicit and actionable through a structured epistemic working memory: the agent maintains Confirmed Facts grounded to sources, Active Hypotheses ranked by evidence, and Open Questions that directly drive its next action. Across five multi-hop benchmarks and five established baselines, SLEUTH's advantage grows with difficulty, from +5 points on HotpotQA to +11 on 4-hop chains, surpassing Reflexion without multiple episodes. Analyzing where the remaining gap lies, we identify the evidence sufficiency problem: agents often find the answer but fail to commit, exhausting their budget on needless verification. A lightweight commitment trigger fixes this, but only when the agent already maintains structured state: the identical trigger applied to an unstructured agent yields no improvement, isolating organized epistemic state as the necessary condition for effective commitment. Finally, enforcing protocol adherence on a weaker model recovers up to +19 points on the hardest problems, showing that how an agent organizes its reasoning, not raw model capability, is the active ingredient for scaling multi-hop reasoning.
HPC-Enabled Video-based Coastal Wave Parameter Estimation Using V-JEPA and Deep Spatiotemporal Learning
arXiv:2607.11998v1 Announce Type: cross Abstract: High deployment cost, poor spatial coverage and susceptibility to storm conditions are all challenges faced by traditional in-situ methods. This paper presents a video-based and high performance computing (HPC) enabled deep learning framework for joint sensor free estimation of five coastal wave parameters, namely significant wave height (Hs), maximum wave height (Hmax), peak period (Tp), zero upcrossing period (Tz) and wave direction (theta) from monocular coastal video. The proposed architecture comprises of a V-JEPA (self supervised) ViT Small backbone for robust spatiotemporal feature extraction in visually challenging scenarios, a dual-stream SlowFast temporal encoder for broad bandwidth representation of wave motion in both hydrodynamic breaking and swell regimes, an optical flow stream based on Farneback optical flow algorithm for adding saliency information to the structure with emphasis on hydrodynamically active wavelength bands of waves, and a multi-task regression layer with dispersion constraints (Airy wave dispersion lambda_p = 0.1). The model was trained on an NVIDIA DGX A100 cluster and was early stopped at epoch 31 and achieved Pearson correlation coefficients of 0.451, 0.578, 0.643, 0.680 and 0.832 for Hs, Hmax, Tp, Tz and wave direction respectively, with generalization ability to geographically diverse held out test data sites. While operating in a data-limited regime (6 annotated training scenes), the framework demonstrates statistically significant temporal correlations (PCC of 0.451 to 0.832), confirming proof of concept feasibility; R2 values (max 0.246) indicate that variance capture will improve with larger annotated datasets.
Hierarchical Bayesian inversion using the Karhunen-Lo\`eve expansion with analytical eigenpairs of the squared exponential kernel
arXiv:2607.12387v1 Announce Type: cross Abstract: Hierarchical Bayesian inversion with Gaussian random field priors addresses uncertainty in covariance hyperparameters, such as the standard deviation and correlation length. When a Gaussian random field is represented by the Karhunen-Lo\`eve (KL) expansion, the basis functions depend on these hyperparameters through an integral eigenvalue problem (IEVP) associated with the covariance kernel. Consequently, the IEVP must be solved repeatedly whenever the hyperparameters are updated, leading to significant computational cost in hierarchical inference. In this paper, we focus on the squared exponential kernel and construct the KL expansion using the analytical solution to a Gaussian-weighted IEVP. This analytical KL expansion offers a computationally efficient alternative to the conventional KL expansion by eliminating the repeated numerical solutions of the IEVP during hyperparameter updates. While the analytical KL expansion is applicable to arbitrary domains and dimensions, it does not have the same mean-square optimality as the conventional KL expansion. To address this limitation, we employ an optimization-based approach that selects the standard deviation of the Gaussian weight function in the IEVP to effectively reduce the truncation error of the KL expansion. Numerical experiments in one- and two-dimensional settings show that this selection strategy provides sufficient accuracy for practical applications. Furthermore, the analytical KL expansion admits closed-form differentiation, enabling efficient posterior sampling via HMC. The proposed framework is applied to Bayesian inversion for a steady Darcy flow model, where the hydraulic conductivity field is successfully estimated using weakly informative hyperpriors.
Demystifying image-recovery from radio interferometers: toward a multiscale predictive model
arXiv:2607.12396v1 Announce Type: cross Abstract: Radio interferometers suffer from the missing short-spacing problem, losing large-scale diffuse emission. This missing flux underestimates gas mass and biases key metrics like star formation efficiency. Quantifying this scale-dependent loss currently relies on computationally intensive mock observations, lacking an analytical image-domain framework. We introduce the Constrained Diffusion Decomposition (CDD) method to decompose an input image ($I_{\mathrm{in}}$) into $n$ continuous scale-space components, denoted as $I_l = \mathrm{CDD}_l(I_{\mathrm{in}})$ for $l \in [1, n]$, and apply it to simulated Atacama Large Millimeter/submillimeter Array (ALMA) observations of the Perseus molecular cloud across multiple array configurations. We find that the interferometric spatial filtering response can be mathematically decoupled: the scale-dependent flux recovery fraction follows a one-dimensional error function (\texttt{erf}), defined as $R(l) = \frac{B}{2} \left[ 1 - \mathrm{erf}\left( \frac{l - c_{\mathrm{recover}}}{w} \right) \right]$, where compact structures are effectively recovered, while extended emission decays monotonically as scales approach the maximum recoverable scale. The proposed CDD--\texttt{erf} framework predicts the spatially filtered interferometric image $I_{\mathrm{pred}}$ directly in the image domain, bypassing visibility simulations, mapping the true sky brightness distribution via the equation $I_{\mathrm{pred}} = \sum_{l=1}^{n} [ \mathrm{CDD}_l(I_{\mathrm{in}}) \times R(l)]$. This provides a quantitative bridge between model and interferometric observations.
Nonlinear Tearing Modes in Current-Vortex Sheets
arXiv:2607.12291v1 Announce Type: new Abstract: The linear and nonlinear development of instabilities and Alfv\'en resonances in a plane current-vortex sheet is presented here for sheared equilibrium profiles $\boldsymbol{B_{y0}} = \tanh(z)\boldsymbol{\hat{y}}$ and $\boldsymbol{V_{y0}} = M_0\tanh(z/r)\boldsymbol{\hat{y}}$. We extend Rutherford's nonlinear model for constant-psi magnetic islands to account for a sheared equilibrium flow and determine the flow's impact on the magnetic island's size. We find that the polarization current induced by the equilibrium flow slows the nonlinear growth of the tearing mode. The saturation of the magnetic island is hastened somewhat for $r > 1$, slowed for $r < 1$, and unmodified for $r=1$. Finally, we find that, in the presence of Alfv\'en resonances, the magnetic island's growth in the nonlinear regime is no longer adequately characterized by constant-psi, and the dynamics of such islands are not captured by the model.
Flatness-Preserving Residual Learning for Real-Time Tight Quadrotor Formation Flight
arXiv:2607.12275v1 Announce Type: new Abstract: Quadrotors flying in tight formations are severely affected by turbulent aerodynamic interactions, such as downwash, that can cause catastrophic collisions if left unmodeled. To compensate for these effects, we propose a physics-informed residual dynamics learning framework that captures complex aerodynamic interactions while ensuring the joint multi-quadrotor system remains differentially flat. We leverage this preserved flatness to design a computationally efficient feedback linearization controller that is easily tunable with linear control techniques and cancels aerodynamic disturbances via feedforward compensation. Hardware experiments demonstrate our framework reduces average tracking errors by 31% compared to nominal baselines. Crucially, our lightweight approach matches the tracking performance of state-of-the-art nonlinear model predictive control (NMPC) while requiring an order of magnitude less computation. We are the first to show that stable, tight formation flight can be achieved with under 30 seconds of training data and a 5ms loop rate, unlocking high-fidelity aerodynamic compensation for compute-constrained flight stacks.
Quantum algorithm for Clifford multiplication
arXiv:2607.10473v1 Announce Type: cross Abstract: Given two dense multivectors of the Clifford algebra $C\ell(V, Q)$ with $N=2^{p+q}$ coefficients, the fastest known classical algorithms compute their geometric product in $O(N^{\omega/2})$ arithmetic operations, where $\omega$ denotes the matrix multiplication exponent. I show that, under amplitude encoding, a quantum computer executes the geometric product in $O(\operatorname{polylog} N)$ time, using logarithmic space with sublogarithmic circuit depth. This exponential speedup establishes Clifford multiplication as a quantum primitive, providing an efficient computational foundation for quantum geometric algorithms and relativistic simulations.
What Makes a Representational Prior Work? Feature Families, Label-Free Invariances, and Critical Windows in Grokking
arXiv:2607.12735v1 Announce Type: new Abstract: Companion work showed the grokking delay is causally the time to form task-structured representations, injectable via a contrastive prior. Here we characterize what makes such a prior work, across four axes, in 188 new runs. Content: a coherent, learnable prior built from the wrong feature family (magnitude bands) blocks generalization like a random partition (1/15 vs 0/20 grok; $p=0.43$ between them), confirming the companion's prediction that priors act at the level of the circuit's features. Supervision: a fully label-free invariance prior -- positives are commuted pairs $(a,b)\sim(b,a)$ only -- generalizes in 15/15 runs at a median $2.7\times$ speedup, more reliably than the label-supervised prior itself ($p=0.038$), and combined with a weight-norm clamp yields the strongest method we test (median $17\times$, 5/5) -- strongest meaning reliably fast: plain cross-entropy with a clamp matches this speed only at the exact critical norm, while the prior keeps it fast across the entire clamp range. Timing: the prior is only needed early -- applied solely during the first 2000 epochs (4% of budget) it generalizes 10/10 at $2.7\times$, beating continuous application (8/10, $1.25\times$) and a duration-matched later window ($2.1\times$). Setting: the dissociation replicates on modular multiplication and across depths and normalization variants, and a clamp sweep quantifies the companion's central claim: structure injection flattens the weight-norm delay-law exponent about 17-fold (plain cross-entropy slows $31\times$ per +10 norm units, a lower bound as higher cells are censored, versus $1.22\times$ with the prior). Honest boundary: tasks that generalize before memorizing have no delay to control. Feature-family alignment decides whether a prior permits generalization; invariance content suffices for acceleration without labels; a brief early window captures nearly all of the benefit.
Not Only NTP: Extending Training Signal Coverage for Generative Recommendation
arXiv:2607.12277v1 Announce Type: new Abstract: Next-Token Prediction (NTP) carries two structural training signal limitations. First, NTP optimizes for single-step prediction only, placing no supervised pressure on learning longer-range behavioral structure -- we term this \textbf{temporal locality}. Second, in multi-domain sequences, each target item embedding receives gradient updates exclusively from the immediately preceding hidden state, with no explicit gradient pathway from cross-domain context -- we term this \textbf{spatial locality}. We propose \textbf{NONTP}, extending NTP's signal coverage along both dimensions through two auxiliary objectives. \textbf{TCL (Temporal Contrastive Learning)} uses a BYOL-style EMA teacher with InfoNCE to align hidden states against a $K$-step future trajectory in representation space. \textbf{TDL (Trans-Domain Learning)} mean-pools cross-domain hidden states and predicts through the shared prediction head, opening a second gradient pathway with no additional parameters. Both are discarded at inference: zero overhead. On a four-domain Meituan industrial dataset (full ranking), NONTP achieves HR@10 +34.3\% over NTP and +18.3\% over MBGR. On the public Amazon Movie-Book-CDs benchmark, HR@10 +2.8\% and NDCG@10 +3.7\%. Online A/B tests confirm CTR +1.8\% and GMV +2.1\% (both $p < 0.01$). Ablation studies confirm each component contributes independently, with gradient conflict analyzed as a direction for future work.
What Would You Click? Personalized Video Thumbnail Generation with Preference-aware Highlight Retrieval
arXiv:2607.12882v1 Announce Type: new Abstract: Video thumbnails are a key factor for attracting user clicks on video platforms, and are increasingly supported by automation. However, existing thumbnail generation methods typically produce generic results shared across users, overlooking the diversity of individual preferences. We therefore introduce personalized video thumbnail generation, a novel task that aims to create thumbnails tailored to user-specific preferences. It is challenging in two aspects: (i) identifying visual anchors (i.e., key frames) from each video to guide the generation, which requires a balance between personalization and informativeness that existing highlight detection methods fail to achieve; and (ii) generating personalized thumbnails that are both visually coherent and faithful to the original video. As a response, we propose a two-stage framework that tightly couples preference-aware retrieval with controllable generation. In the first stage, a personalized highlight retriever captures fine-grained user-video interactions and incorporates video semantics through summarization, enabling the selection of diverse visual anchors aligned with both user preferences and video contexts. In the second stage, a VLM-guided diffusion pipeline transforms these anchors into thumbnails by extracting and injecting semantically grounded visual cues, improving personalization while preserving visual coherence and fidelity. Experiments on two public datasets show our method delivers state-of-the-art performance compared with both retrieval-based and generative baselines. A user study further demonstrates improved click preference, highlighting its effectiveness in enhancing user engagement. The code is available at https://github.com/hezy18/PVTG.
An Extension to the Procedure for Developing Uncertainty-Consistent Shear Wave Velocity Profiles from Inversion of Experimental Surface Wave Dispersion Data
arXiv:2607.12743v1 Announce Type: new Abstract: Measurements of shear wave velocity (Vs) with uncertainty are critical for site-specific probabilistic seismic hazard studies. However, rigorously quantifying the uncertainty in Vs over large enough areas and great enough depths remains challenging. In 2021, Vantassel and Cox (i.e., VC21) proposed a procedure for developing suites of Vs profiles from surface wave testing whose uncertainty were consistent with the experimental dispersion data's uncertainty. The VC21 procedure was a significant step forward, however, it requires a full dispersion data matrix to compute inter-wavelength phase velocity correlations. While applicable to many practical cases, VC21 could not be applied to the case where multiple surface wave arrays of different sizes were deployed at a site as a means of developing broadband dispersion data and deeper Vs profiles. In response, this work extends the VC21 procedure using two possible approaches for estimating a full dispersion data matrix. Approach 1 uses a selection of theoretical dispersion curves from an initial, traditional surface wave inversion. Approach 2 estimates the full data matrix by combining pieces of the data matrix obtained from the experimental dispersion measurements. Both approaches are evaluated using two synthetic datasets; one relatively-simple, three-layered model and one more-complex, five-layered model. Approach 1 and Approach 2 were able to reasonably estimate the true correlation matrix and recover uncertainty-consistent Vs profiles similar to the true distribution of Vs. While the uncertainty of the recovered Vs profiles were higher than is often assumed, the engineering proxies computed from those Vs profiles, namely the time averaged shear wave velocity in upper 30 m and the fundamental site period, showed substantially less uncertainty indicating the Vs profiles, while uncertain, are effective at capturing a site's ...
A Shared Subcircuit Lets LLMs Count Down Across Tasks
arXiv:2607.12279v1 Announce Type: new Abstract: Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table. These are all tasks that language models can do that requires tracking how many tokens remain before a target. In this work, we identify in Llama-3.1-70B-Instruct a general mechanism for performing these tasks: a "countdown subcircuit" that compares the current position to a goal length and estimates the time remaining until then. We first isolate a countdown subcircuit in a controlled setting, in which the model is tasked with writing a fixed-length sentence ending in a specified word. We then investigate the geometry of the representations used by the subcircuit, and find that the subcircuit uses an identical motif previously identified in a frontier LLM on a separate task, thus suggesting that this motif is shared across models. Finally, we use unsupervised probing on a natural language dataset to find a variety of other tasks where this subcircuit is used, including tasks where the goal length is inferred from context rather than explicitly stated. Our work suggests that reverse-engineering subcircuits allows us to understand how behaviors generalize from a single example to many different tasks and even models.
Mass-Conserving Saddle Dynamics via Generalized Inner Product: Theory, Algorithms, and Applications
arXiv:2607.12715v1 Announce Type: new Abstract: To reveal the effect of the inner product choice, we present a unified formulation of saddle dynamics for the functional F with a mass constraint under different inner products. We establish the equivalence between the index-k saddle points and the linearly stable steady states of the corresponding dynamics. Further, we present the dynamics with discrete H^{-1} and L^2 inner products and numerically verify the convergence orders of both dynamics. Finally, we apply the method to a phase field model with driving force under Neumann and periodic boundary conditions. The results uncover previously unreported saddle points and their connectivity, highlighting how the choice of inner product enriches the solution landscape in conservative systems.
The Seriality Gap in Video Diffusion Models
arXiv:2607.13031v1 Announce Type: new Abstract: When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere dynamics, we find that the performance of standard bidirectional video diffusion degrades as the causal chain lengthens, even when provided more denoising steps. In a length-matched single-ball control, where ball-ball interactions are absent, the degradation largely disappears, isolating dependent-event structure rather than video length as the cause. Across intervention studies, methods that increase effective serial computation improve performance disproportionately, including autoregressive/blockwise generation and architectural depth. We identify this pattern as the seriality gap: a mismatch between tasks requiring growing serial computation and video diffusion models whose denoising loop does not provide scalable serial compute. We then prove that, for deterministic video prediction, denoising steps do not add serial computation beyond the backbone, indicating a structural obstacle for video diffusion on serial reasoning and simulation tasks.
XScientist: A Git-Like Research Protocol for Long-Running Autonomous Scientific Discovery
arXiv:2607.12301v1 Announce Type: new Abstract: Autonomous research systems are often evaluated as one-shot paper generators: given a topic, they produce a manuscript and a small set of experiment logs. This framing hides the operational problem that makes such systems difficult to trust: research is long-running, branching, failure-prone, and dependent on auditable handoffs between agents and humans. XScientist is a git-like research protocol and operating system for this setting. It orchestrates idea generation, experiment execution, manuscript drafting, self-review, repair, quality gating, daemon scheduling, and reproducibility artifacts as one continuously observable pipeline. The central design choice is to treat each run as a portable research artifact rather than only as a PDF. XScientist exports an Agent-Native Research Artifact (ARA), a protocol that records an exploration DAG, per-node code and outputs, claim-to-evidence anchors, content hashes, provenance, and re-execution hooks. This makes each generated paper inspectable as a science exploration tree: failed branches, repaired experiments, ablations, and manuscript claims remain connected to the nodes that produced them. The system also includes deterministic integrity forensics, sample gates, truth contracts, reviewer-oriented repair loops, and long-running daemon controls. This paper describes the current XScientist architecture, the ARA protocol surface, and the practical safeguards needed to move autonomous science from single-run demos toward reproducible, reviewable, and forkable research infrastructure. The implementation and manuscript source are maintained in the public GitHub repository at https://github.com/smileformylove/XScientist.
When Binaries Talk Back: Representation-Confusion Attacks on LLM-Assisted Reverse Engineering
arXiv:2607.12507v1 Announce Type: new Abstract: LLM-assisted reverse-engineering (RE) systems analyze strings, decompiler output, and tool reports derived from ttacker-controlled binaries. A binary can make data look like instructions or records from one origin look like independent evidence. We call such failures Representation-Confusion Attacks in Reverse Engineering (RARE): the pipeline promotes a correctly extracted observation to instruction authority, claim-validating evidence, or trusted analysis state without the authority or support that role requires. RARE-Bench measures these failures with behavior-checked clean and adversarial binaries. After an exploratory 11,520-call study, we test RARE-Guard's authorization and evidence controls on 20 new programs and two models. Without runtime controls, the models propose a planted unsafe action in 35/40 adversarial cases and 0/40 clean cases. When binary-derived content is shown only as data (Data-Only rendering), they still make 15 unsafe proposals. Tool Authorization denies all 15 and authorizes all 40 matched analyst requests. On identical report drafts, Support Gate validates 23/40 false claims by counting records from one origin separately. Provenance Gate groups those records before counting support, validates 0/40 false claims, and retains all 40 supported claims. We then instrument Ghidra, r2pipe, and angr on 16 further programs. In a preselected eight-program subset, no single-tool draft reaches Support Gate's validation threshold for the false claim. In fused drafts across all 16 programs, Support Gate validates 32/32 false claims. Provenance Gate prevents validation of all 32 and retains all 32 supported claims. A deterministic renderer prevents downgraded claims from reappearing in the final report. Binary-derived content may therefore guide analysis without gaining authority over tools, and views from several tools do not necessarily provide independent evidence.
MelT: A Portable, Single-GEMM Mel Audio Frontend via Non-Uniform DFT with Measured Latency and Energy Gains on GPUs
arXiv:2606.01009v3 Announce Type: replace Abstract: Modern neural audio models run on accelerators whose peak throughput comes from dense matrix multiplication, increasingly at the edge and in datacenters. The conventional acoustic frontend, however -- a Short-Time Fourier Transform (STFT) followed by sparse Mel aggregation -- remains a multi-stage pipeline centered on the Fast Fourier Transform (FFT), with execution overheads unlike the dense linear algebra dominating the inference stack. This work introduces MelT, a portable single-stage Mel frontend that precomputes Mel-spaced Non-Uniform Discrete Fourier Transform (NDFT) bases and applies them to time-domain frames through General Matrix Multiplication (GEMM). The contribution is a computational design principle: decoupling Mel feature extraction from vendor-specific FFT primitives and lowering it onto the matrix-multiplication substrate accelerators already optimize. It is not a new spectral operator. MelT's direct projection performs more arithmetic than the FFT pipeline. Yet in the compact-resolution regime of neural audio frontends, it achieves a 1.64-times to 3.29-times latency reduction and up to a 3.03-times reduction in measured active energy, from the Apple A18 Pro to the NVIDIA H100. All gains are within-platform comparisons, accompanied by task-level validation. Word error rate stays statistically equivalent to the native frontend's on frozen Whisper models of medium size and larger; speaker-attribute classification on VoxCeleb1 is non-inferior. The cepstral extension MFCCT preserves utility on a clinical respiratory-insufficiency classification task (SPIRA) while improving on the MFCC baseline. These results indicate that, in practical regimes on the accelerators evaluated here, hardware alignment rather than arithmetic count can govern the realized cost of feature extraction.
QUBO-Optimized Evidence Selection for Retrieval-Augmented Question Answering with Unconventional Solvers
arXiv:2607.12334v1 Announce Type: new Abstract: Retrieval-augmented question answering depends on selecting evidence passages that jointly support answer generation. However, many RAG pipelines rely on top-\(k\) ranking, where passages are selected mainly by individual relevance scores, even though multi-hop questions often require complementary evidence satisfying multiple information requirements. Recent LLM-based selectors address this by treating retrieval as set selection, but using an LLM for this intermediate stage can be costly and difficult to scale. In this work, we formulate evidence selection as a Quadratic Unconstrained Binary Optimization (QUBO) problem. Given a question, candidate passages, and decomposed information requirements, our method constructs an energy function that balances relevance, requirement coverage, support strength, redundancy, complementarity, and compactness. Low-energy solutions correspond to compact evidence subsets that cover the needed requirements while avoiding unnecessary or repetitive context. The selected passages are then passed to a downstream language model for answer generation, separating combinatorial evidence selection from semantic answer generation. We evaluate the proposed QUBO selector on HotpotQA and compare it with LLM-based set selectors and non-LLM baselines including BM25, relevance top-\(k\), maximal marginal relevance, hybrid lexical--semantic ranking, greedy coverage, and random selection. The QUBO selector achieves competitive exact-match and token-F1 performance relative to LLM-based selectors while providing a solver-compatible formulation for structured evidence selection. These results suggest that multi-hop evidence selection can be cast as discrete optimization, opening a path toward RAG pipelines where LLMs are reserved for semantic processing and answer generation, while context selection is handled by Ising/QUBO-compatible solvers.
ProtoPointNet: Prototype-Based Interpretable Classification of 3D Dental Point Clouds with Verifiable Spatial Activations
arXiv:2607.12335v1 Announce Type: new Abstract: Prototype-based networks provide inherently interpretable classification by linking predictions to learned exemplars, but their use in 3D point clouds and clinical surface-pair reasoning remains limited. We introduce ProtoPointNet, a prototype-based model for dental occlusion classification from registered upper--lower intraoral arch pairs. Each point is encoded by a 14-dimensional descriptor combining local surface geometry, curvature, and explicit inter-arch displacement and clearance, exposing occlusal relationships to prototype matching. A shared multi-task point-cloud backbone learns axis-specific prototype heads for sagittal-left, sagittal-right, vertical, transverse, and midline classification. To support limited clinical data, we train prototypes from scratch using auxiliary supervision and encoder-freeze hand-off. On Bits2Bites, ProtoPointNet achieves mean test macro-F1 of 0.724 and AUROC of 0.825, with strongest performance on vertical (F1 0.828) and sagittal-left classification (F1 0.807). Projected prototype activations localise to anatomically plausible regions, including posterior molars and premolars for cross-bite evidence and anterior incisors for bite-depth evidence. These results support prototype-based reasoning as a transparent, spatially grounded alternative to black-box 3D classifiers for dental surface-pair analysis.
Optically Derived Radio-Frequency Benchmark in Methanol: A Sub-kHz Reference for Astrophysical Tests of Fundamental Physics
arXiv:2607.12533v1 Announce Type: new Abstract: Methanol radio lines observed in space provide sensitive probes of whether the proton-to-electron mass ratio has changed over cosmic time, but such tests require laboratory rest frequencies with very high accuracy. Here we determine the frequency of the astrophysically important 12.2 GHz $3_{-1}$E -- $2_{0}$E transition of CH$_3$OH by measuring near-infrared rovibrational transitions rather than the microwave line directly. Using wavelength-modulated NICE-OHMS locked to an ultra-stable optical frequency comb and referenced via a fiber link to a hydrogen-maser frequency standard, we measure Lamb-dip frequencies near 1.4 $\mu$m (216 THz) with 10 Hz statistical reproducibility and absolute uncertainties as low as 130 Hz. Pairs of optical transitions sharing common upper levels form a triangulation scheme that yields the ground-state rotational combination difference. We obtain 12 178 596.415(135) kHz, improving on earlier molecular-beam microwave spectroscopy by a factor of 20 and agreeing with a recent free-induction-decay measurement. This result establishes a sub-kHz laboratory benchmark for a key radio-astronomical methanol line and demonstrates that optical triangulation can be extended to non-chiral molecules with internal rotation.
Robust Design of Integrated Sensing and Communication in LEO Satellite Systems
arXiv:2607.12337v1 Announce Type: new Abstract: With the growing demand for satellite sensing and communication, the limited wireless resources are difficult to support multiple satellite systems. Therefore, it is desired to investigate integrated sensing and communication (ISAC) in low Earth orbit (LEO) satellite systems to enable multi-functionality within a single satellite, thereby saving both spectrum and orbital resources. In this paper, a framework for ISAC in LEO satellite systems is established, where a satellite can simultaneously sense multiple targets and serve multiple communication users (CUs) over the same spectrum. Considering the limited onboard energy of satellite, a novel robust beamforming design algorithm is developed with the goal of minimizing total transmit power while satisfying the mean squared error (MSE) requirements for sensing and signal-to-interference-plus-noise ratio (SINR) requirements for communication in presence of channel phase uncertainty which exacerbates the cross-functional interference. According to theoretical analysis, the proposed algorithm for ISAC in LEO satellite systems is effective. Moreover, extensive simulations confirm the superiority of the proposed algorithm over baselines.
Skills That Don't Exist: A Large-Scale Study of Hallucinated Skill Recommendation in LLM Agents
arXiv:2607.12340v1 Announce Type: new Abstract: LLM agents acquire new capabilities by downloading skills from open registries. Instead of browsing these catalogs manually, developers typically ask the agent to recommend and install a skill. This convenience hides a risk: agents frequently invent names for skills that exist in no registry. We term this flaw skill name hallucination. A fake name may seem harmless, but it opens the door to supply-chain attacks. Because registries rarely verify publishers, an adversary can prompt the agent, collect the fake names it returns, pre-register malicious skills under them, and wait for a victim to install the payload. We conducted the first large-scale measurement of skill name hallucination, evaluating 15,000 prompts across 12 configurations (4 standalone LLMs and 8 agents). We conservatively counted a name as hallucinated only if it was missing from all live registries and GitHub. The results reveal a systemic vulnerability: every configuration hallucinates. Rates average 36.0% for standalone LLMs and 36.9% for agents, rising to 43.1% on real-world developer questions. In total, the systems generated 5,669 distinct hallucinated names. Crucially, these names are not random noise. Agents repeat the same fake names across prompts and models, giving attackers highly reliable targets to hijack. Finally, we tested four model-level defenses and found a severe conflict between security and usability. The strongest, retrieval grounding, cut the hallucination rate from 40.8% to 3.2% but crippled usefulness: even the best-defended system recommended the correct skill only about one in six times. Skill name hallucination is thus a highly exploitable vulnerability requiring minimal attacker effort. Fixing it cannot rely on prompt engineering or model tuning alone. It demands ecosystem-wide structural changes: registry-level name reservations and verified recommendation pipelines.
Policy-Conditioned Constrained Decoding for Column-Level Access Control in Text-to-SQL
arXiv:2607.12341v1 Announce Type: new Abstract: Text-to-SQL is increasingly deployed across trust boundaries between data providers and users. Such deployment must balance three competing requirements: policy compliance, answer coverage, and bounded cost. Existing approaches typically decide refusal based on which columns a query mentions and enforce it stochastically. Whether a query is compliant, however, depends not only on which columns appear but on how they are used, and stochastic enforcement cannot deterministically rule out violations. We formalize this requirement as a column-use policy over semantic use: output, filter condition, and aggregation argument. We integrate the policy by aligning each role with grammar productions tracked by the decoder. The resulting system, PCC-SQL, applies a per-token logits mask that deterministically eliminates single-query column-use violations on the supported SQL fragment in a single decoding pass. Across three benchmarks and three open-source models, PCC-SQL achieves 0% Leakage Rate and Coverage up to 88.7% on Spider-CU, while staying within +10% tokens of direct prompting. We additionally assess semantic alignment with execution accuracy.
Degree Lower Bounds for Torus Polynomials and $MAJORITY$ vs $ACC^0$
arXiv:2607.12346v1 Announce Type: new Abstract: The class $ACC^0$ consists of Boolean functions that can be computed by constant-depth circuits of polynomial size with $AND, NOT$ and $MOD_m$ gates, where $m$ is a natural number. At the frontier of our understanding lies a widely believed conjecture asserting that $MAJORITY$ does not belong to $ACC^0$. A few years ago, Bhrushundi, Hosseini, Lovett and Rao (ITCS 2019) introduced torus polynomial approximations as an approach towards this conjecture. Torus polynomials approximate Boolean functions when the fractional part of their value on Boolean points is close to half the value of the function. They reduced the conjecture that $MAJORITY \notin ACC^0$ to a conjecture concerning the non-existence of low degree torus polynomials that approximate $MAJORITY$. We reduce the non-existence problem further, to a statement about finding feasible solutions for an infinite family of linear programs. The main advantage of this statement is that it allows for incremental progress, which means finding feasible solutions for successively larger collections of these programs. As an immediate first step, we find feasible solutions for a large class of these linear programs, leaving only a finite set for further consideration. Our method is inspired by the method of dual polynomials, which is used to study the approximate degree of Boolean functions. Using our method, we also propose a way to progress further. We prove several additional key results with the same method, including lower bounds for approximating the $AND$ function, lower bounds when the approximating polynomial is symmetric, showcasing the power of our machinery.
Space-Time Modulated vs 2-Bit Reflect-Arrays: A Comparison of Main Beam Gain and Sidelobe Level Performance
arXiv:2607.12342v1 Announce Type: new Abstract: This paper presents a comprehensive analysis of the phase errors introduced by 2-bit reflectarray architectures and investigates their impact on aperture synthesis for beam-steering applications. Building upon this analysis, an amplitude-aware synthesis technique based on space-time modulated (STM) reflect-arrays is proposed to mitigate the main-beam gain degradation commonly associated with STM-based equivalent phase generation while preserving the enhanced phase resolution enabled by temporal coding. A 10{\lambda}by 10{\lambda} reflect-array aperture with {\lambda}/2 inter-element spacing is considered, where the desired aperture phase distribution is synthesized using a library of achievable reflection coefficients generated from all possible combinations of four 2-bit phase states distributed across (L) temporal segments. The proposed approach introduces a phase-tolerance window and selects the reflection state with the highest reflectance within the allowable phase range. By exploiting the inherent amplitude diversity of the STM state space, the proposed approach maximizes aperture efficiency while maintaining the desired beam-steering phase distribution. Numerical results demonstrate that the method substantially reduces the gain penalty typically associated with STM reflect-arrays along with providing significant sidelobe suppression.