Forskningsradar

Science Journals

Peer-reviewade publikationer — 56237 artiklar

Beyond Interestingness: Semantic and Context-Aware Natural Language Query Recommendations for Visual Data Analysis
arXiv:2201.04868v3 Announce Type: replace Abstract: Recent advances in large language models (LLMs) have made natural language interfaces (NLIs) widely accessible for data exploration, yet analysts who have a broad analytical objective still face the challenge of decomposing it into effective step-by-step queries, especially over unfamiliar, multi-table relational databases. Rather than generating high-level analytical agendas, we investigate how to augment an NLI with semantic- and context-aware next-step query recommendations that act as analytical scaffolding for relational database exploration. Our approach goes beyond interestingness-only methods by jointly integrating semantic relevance, data interestingness, and context coherence to guide exploration toward coherent, topic-focused analyses and potentially insightful subsets. We evaluate QRec-NLI with NL2SQL benchmarking, LLM-enhanced description validation, agentic comparisons against interestingness-only and LLM-based prompting baselines, and a 12-participant user study. In the agentic comparison, QRec-NLI yields more topically relevant and locally coherent query sequences than both baselines. In the user study against the interestingness-only baseline, it receives stronger ratings for insight-generation support and decision support.
Structure-preserving Lagrange-multiplier methods for mean curvature flow and their error bounds
arXiv:2607.14787v1 Announce Type: new Abstract: We propose and analyze a class of evolving surface finite element methods for mean curvature flow of closed surfaces using Lagrange multipliers for preserving the energy-decreasing structure. The algorithm is based on the solution-driven formulation using Huisken's evolution equations for the normal vector and the mean curvature. The time discretizations use linearly implicit backward difference formulas (BDF) for the parabolic geometric variables and implicit Adams updates for the nodal positions. The approach also accommodates artificial tangential velocities of minimal-deformation-rate type. The resulting fully discrete algorithms are area-decreasing at every time step, with a prescribed decay rate determined by the computed mean curvature. We prove local existence and uniqueness of the discrete Lagrange multiplier and convergence of a simplified Newton iteration for its computation under weak regularity assumptions. Under stronger regularity assumptions, as used in the convergence theory for the underlying evolving surface finite element method, we derive optimal-order error bounds of order $h^k+\tau^q$ in the $H^1$-norm for finite elements of polynomial degree $k\ge 2$ and $q$-step BDF and $q$-step implicit Adams methods with $2\le q\le 5$, both without and with minimal-deformation-rate tangential motion. Numerical experiments for mean curvature flow of a sphere confirm the predicted convergence rates and show that the Lagrange-multiplier correction entails only a small computational overhead that is essentially independent of the mesh size and the time step size.
Transcoders for Investigating Deception in Language Models
arXiv:2607.14791v1 Announce Type: new Abstract: Transcoders have recently emerged as a promising approach for mechanistic interpretability (MI), enabling circuit-level analysis of model behaviour. In this paper, we investigate the use of transcoders to analyse deceptive behaviour in language models, a behaviour that poses a safety and security risk. Using a Qwen3-4B model with pre-trained transcoders, specifically per-layer transcoders (PLTs), we construct attribution graphs that capture feature activations and inter-feature dependencies, allowing circuit-level analysis of deception. Through feature steering and circuit analysis, we identified a dictionary of deception-related features and show that these features exert a stronger influence on deceptive outputs, as they produce predictable shifts between deceptive and non-deceptive responses. These findings suggest that deception emerges from internal model mechanisms and highlight the potential of transcoders for behavioural monitoring and early detection of security vulnerabilities related to malicious behaviours in language models.
CODA: Algorithm-Hardware Co-design for Edge Video Diffusion via NMP-Enabled Compute-Cache Operator Disaggregation
arXiv:2607.14908v1 Announce Type: new Abstract: Deploying Video Diffusion Models (VDMs) on edge devices is appealing for localized and privacy-preserving generation, but their iterative Transformer-based denoising remains too slow for practical local inference. Cross-Timestep Caching (CTC) has emerged as a promising direction for reducing redundant computation, reusing activations across adjacent denoising steps rather than modifying model weights, while largely preserving generation fidelity. However, on memory-constrained edge GPUs, CTC requires a massive cache footprint that quickly exceeds on-device VRAM and forces the cache into host memory. More fundamentally, cache operators remain tightly interleaved and chain-dependent with native compute operators, so naive near-memory offloading still incurs repeated PCIe exchanges for residual and fusion computations, turning cache reuse into a communication- and serialization-bound execution flow. We therefore propose CODA, an algorithm-hardware co-designed architecture centered on Compute-Cache Operator Disaggregation. CODA separates dense compute paths and memory-bound cache paths across the xPU and a lightweight DIMM-side near-memory engine, reorganizes fragmented cache activity into hardware-friendly coalesced segments, and exploits Classifier-Free Guidance (CFG) branch independence to overlap xPU compute with cache-side execution. Experiments show that CODA achieves up to 1.80x end-to-end speedup and 1.74x higher energy efficiency, while preserving competitive generation quality compared with a state-of-the-art caching algorithm.
Analytic finite-rank corrections for singularly weighted estimates in a computer-assisted proof of 3D Euler singularity
arXiv:2607.15256v1 Announce Type: cross Abstract: Computer-assisted proofs of self-similar singularity formation for fluid equations often rely on numerically constructed approximate profiles. One effective approach to establishing stability of perturbations around a numerically constructed profile is to perform weighted energy estimates with singular weights near the singularity. However, the weighted norms require exact local vanishing conditions that are not automatically preserved by the equations nor the numerical construction. In this paper, we review an analytic low-rank correction method first developed in [ChenHou2023a,ChenHou2023b] to overcome this difficulty. The numerical step determines coefficients, rigorous bounds, and low-order defect modes in explicit global basis representations, while the required vanishing conditions are enforced analytically through low-rank corrections derived from Taylor expansions of the relevant quantities represented in a smooth basis. For completeness, we briefly review the singularly weighted estimates and a quantitative finite-rank perturbation method in the 2D Boussinesq / 3D Euler stability argument, where singular weights and the required vanishing order arise. Against this background, we formulate the local correction principle in a simplified setting, explain the correction of the residual error in numerical constructions of approximate space-time solutions and the stream function, and discuss its broader applicability to computer-assisted stability analysis for nonlocal PDEs.
Quasi-Monte Carlo for Bayesian design of experiment problems governed by parametric PDEs
arXiv:2405.03529v5 Announce Type: replace Abstract: This paper contributes to the study of optimal experimental design for Bayesian inverse problems governed by partial differential equations (PDEs). We derive estimates for the parametric regularity of multivariate double integration problems over high-dimensional parameter and data domains arising in Bayesian optimal design problems. We provide a detailed analysis for these double integration problems using two approaches: a full tensor product and a sparse tensor product combination of quasi-Monte Carlo (QMC) cubature rules over the parameter and data domains. Specifically, we show that the latter approach significantly improves the convergence rate, exhibiting performance comparable to that of QMC integration of a single high-dimensional integral. Furthermore, we numerically verify the predicted convergence rates for an elliptic PDE problem with an unknown diffusion coefficient in two spatial dimensions, offering empirical evidence supporting the theoretical results and highlighting practical applicability.
Synthetic Light-in-Flight
arXiv:2407.07872v3 Announce Type: replace Abstract: Light-in-flight (LiF) measurements enable the visualization of light paths through arbitrary, volumetric scenes, making light-matter interactions at ultrafast timescales visible. Traditionally, LiF measurements require specialized equipment, such as ultrashort pulse light sources and high-speed electronics, often limited by low spatial resolution. Herein, we introduce a novel computational approach, "Synthetic Light-in-Flight" (SLiF), that overcomes these constraints by relying solely on tunable, continuous wave (CW) lasers and off-the-shelf CMOS cameras. From multiple CW scene measurements at different optical wavelengths, we create multiple "synthetic fields," each at a "synthetic wavelength," which is the beat wave of two respective optical waves. These synthetic fields are robust to speckle and environmental fluctuations, enabling us to combine multiple synthetic fields into a "synthetic light pulse" that sections the volumetric scene at much lower instantaneous peak illumination power than a comparable physical light pulse. We experimentally demonstrate the generation of synthetic pulses with 1 ps-scale width and show that their complex synthetic pulse fields can be freely manipulated in the computer after their acquisition, allowing for spatial and temporal shaping of different sets of pulses from the same set of measurements to maximize the decoded information output for each scene. Finally, we show that the recovered time-of-flight information can be used to characterize physical scene properties, such as depth and refractive indices.
Human-In-The-Loop Machine Learning for Safe and Ethical Autonomous Vehicles: Principles, Challenges, and Opportunities
arXiv:2408.12548v3 Announce Type: replace Abstract: Machine Learning (ML) has become central to Autonomous Vehicles (AVs), supporting perception, prediction, planning, control, and decision-making in dynamic environments. However, achieving full autonomy in cluttered and complex scenarios, such as intricate intersections, diverse scenes, varied trajectories, and complex missions, remains challenging; moreover, data labeling is still a major bottleneck. These limitations motivate Human-in-the-Loop Machine Learning (HITL-ML), in which human input is incorporated through validation, annotation, task organization, reward design, action correction, preference feedback, and supervisory intervention. To advance safe and ethical autonomy, this paper presents a tutorial survey of HITL-ML for AVs, focusing on Curriculum Learning (CL), Human-in-the-Loop Reinforcement Learning (HITL-RL), Human-in-the-Loop Large Language Models (HITL-LLMs), Active Learning (AL), and ethical principles. We first review CL methods that structure training from simple to complex tasks, covering navigation, path planning, obstacle avoidance, data collection, landing, intersection handling, motion planning, and UAV swarm coordination. We then examine HITL-RL through reward shaping, action injection, demonstrations, preference-based feedback, and interactive learning, emphasizing improved learning efficiency, safer policy exploration, and real-time intervention. Next, we review HITL-LLM through collaboration and oversight and specify key challenges. After that, we discuss AL for perception, anomaly detection, semantic mapping, object detection, vehicle recognition, and security-related classification. Ethical principles are reviewed as technical requirements for transparency, accountability, human oversight, safety, security, regulatory compliance, and reliability of human input.
A Unified Conceptual Framework for Gravitational Instabilities: From Latent System State to Failure
arXiv:2607.14831v1 Announce Type: new Abstract: Gravitational instabilities are among the most widespread natural hazards and are expected to become increasingly significant under ongoing environmental change. Despite substantial advances in process understanding, monitoring, and modeling, predicting the timing of failure remains a fundamental challenge because instability emerges from complex interactions operating across multiple spatial and temporal scales. Existing approaches often focus on specific processes, forcing mechanisms, or observational signatures, resulting in a fragmented view of how systems evolve toward failure. This article propose a unified conceptual framework in which gravitational instabilities are interpreted as the progressive evolution of a latent system state controlled by damage accumulation, stress redistribution and external forcing to catastrophic failure. Within this perspective, failure emerges from the continuous interplay between internal system dynamics and external forcing, while predictability depends on our ability to infer the evolving state of the system from incomplete and indirect observations. The framework provides a common structure linking physical processes, observable manifestations, monitoring strategies, modeling approaches, and forecasting methods across a wide range of gravitational hazards. By integrating concepts from geomechanics, fracture mechanics, statistical physics, and data-driven sciences, the proposed framework shifts the focus from the search for universal precursors toward the reconstruction of evolving system states and their proximity to instability. It offers a unifying perspective for understanding failure processes and developing more robust approaches to hazard assessment and early warning.
Thermodynamic Limits of Proof
arXiv:2601.15571v5 Announce Type: replace Abstract: Every irreversible recorded distinction has a positive thermodynamic work floor. Landauer's principle supplies the ideal bound $\varepsilon\ge k_B T\ln 2$ per irreversible bit, experimentally verified to $\pm 10\%$. Proof available to an agent is checkable information for that agent: some substrate must produce, retain, and expose evidence that excludes answer-changing alternatives. A finite detector array operating at temperature $T$ for finite time has finite signal-acquisition capacity. Combining finite causal access, positive retained-record cost, and exact lower bounds on required records gives the Physical Counting Impossibility Theorem: no fixed-budget substrate can provide universal exact proof once the retained-record lower bound exceeds the declared budget. The theorem requires exactly $B<\infty$ and $\varepsilon>0$. An answer reports a value; proof supplies checkable grounds for accepting it. A reversible device may compute an answer and erase its scratch history, but proof requires retained, inspectable records. A global answer register, oracle response, entanglement witness, finite survey catalog, or trusted device output supplies proof only through an interface that exposes the relevant grounds to the verifier. A proposed interface must identify the retained-record lower-bound family $R(n)$ it induces. Sound operational claims about efficient solvability inherit the same finite-budget obstruction when their acceptance would license universal exact proof. Substrate-free derivability has proof status only when a physical verification event makes it available to an agent.
Think When Needed: Model-Aware Reasoning Routing for LLM-based Ranking
arXiv:2601.18146v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly applied to ranking tasks in retrieval and recommendation. Although reasoning prompting can enhance ranking utility, our preliminary exploration reveals that its benefits are inconsistent and come at a substantial computational cost, suggesting that when to reason is as crucial as how to reason. To address this issue, we propose a reasoning routing framework that employs a lightweight, plug-and-play router head to decide whether to use direct inference (Non-Think) or reasoning (Think) for each instance before generation. The router head relies solely on pre-generation signals: i) compact ranking-aware features (e.g., candidate dispersion) and ii) model-aware difficulty signals derived from a diagnostic checklist reflecting the model's estimated need for reasoning. By leveraging these features before generation, the router outputs a controllable token that determines whether to apply the Think mode. Furthermore, the router can adaptively select its operating policy along the validation Pareto frontier during deployment, enabling dynamic allocation of computational resources toward instances most likely to benefit from Think under varying system constraints. Experiments on three public ranking datasets with different scales of open-source LLMs show consistent improvements in ranking utility with reduced token consumption (e.g., +6.3\% NDCG@10 with -49.5\% tokens on MovieLens with Qwen3-4B), demonstrating reasoning routing as a practical solution to the accuracy-efficiency trade-off.
Knowledge-Aware Evolution for Task-Free Streaming Federated Continual Learning with Arbitrary Class Overlap
arXiv:2601.19788v2 Announce Type: replace Abstract: Federated Continual Learning (FCL) leverages inter-client collaboration to better balance new knowledge acquisition and old knowledge retention on non-stationary data. However, existing FCL methods struggle to adapt to streaming scenarios where sequential and ephemerally accessible data chunks lack task identifiers and exhibit arbitrary class overlap, leading to confusion between old and new knowledge and an inability to sustain local inference on all encountered classes. To address this, we propose FedKACE with three components: 1) an adaptive mechanism that determines when to switch the inference model from the local to the global one to improve client-side inference performance; 2) a responsive gradient-balanced replay scheme that utilizes the ratio of the squared L2 gradient norms to balance client-specific knowledge between new acquisition and old retention; 3) a holistic buffer maintenance strategy that preserves highly informative and boundary-significant samples to enhance knowledge retention under class overlap.Experiments across multiple scenarios and theoretical analysis demonstrate the effectiveness of FedKACE.
To Grok Grokking: Provable Grokking in Ridge Regression
arXiv:2601.19791v4 Announce Type: replace Abstract: We study grokking, the onset of generalization long after overfitting, in a classical ridge regression setting. We prove end-to-end grokking results for learning over-parameterized linear regression models using gradient descent with weight decay. Specifically, we prove that the following stages occur: (i) the model overfits the training data early during training; (ii) poor generalization persists long after overfitting has manifested; and (iii) the generalization error eventually becomes arbitrarily small. Moreover, we show, both theoretically and empirically, that grokking can be amplified or eliminated in a principled manner through proper hyperparameter tuning. To the best of our knowledge, these are the first rigorous quantitative bounds on the generalization delay (which we refer to as the "grokking time") in terms of training hyperparameters. Lastly, going beyond the linear setting, we empirically demonstrate that our quantitative bounds also capture the behavior of grokking on non-linear neural networks. Our results suggest that grokking is not an inherent failure mode of deep learning, but rather a consequence of specific training conditions, and thus does not require fundamental changes to the model architecture or learning algorithm to avoid.
Selecting Hyperparameters for Tree-Boosting
arXiv:2602.05786v3 Announce Type: replace Abstract: Tree-boosting is a widely used machine learning technique for tabular data. However, its out-of-sample accuracy is critically dependent on multiple hyperparameters. In this article, we empirically compare several popular methods for hyperparameter optimization for tree-boosting including random grid search, the tree-structured Parzen estimator (TPE), Gaussian-process-based Bayesian optimization (GP-BO), Hyperband, the sequential model-based algorithm configuration (SMAC) method, and deterministic full grid search using $59$ regression and binary classification data sets. We find that the SMAC method clearly outperforms all the other considered methods on average, and it gives stable performance across a diverse collection of tabular data sets under a fixed tuning budget, which is relevant for users who cannot afford extensive manual trial-and-error tuning. We further observe that (i) a relatively large number of trials larger than $100$ is typically required for accurate tuning, (ii) using default values for hyperparameters or a full search over a small grid often yields very inaccurate models, (iii) all considered hyperparameters can have a material effect on the accuracy of tree-boosting, i.e., there is no small set of hyperparameters that is more important than others, and (iv) choosing the number of boosting iterations using early stopping yields more accurate results compared to including it in the search space for regression tasks.
Comparison of Structure Preserving Schemes for the Cahn-Hilliard-Navier-Stokes Equations with Degenerate Mobility and Adaptive Mesh Refinement
arXiv:2602.08639v4 Announce Type: replace Abstract: The Cahn-Hilliard-Navier-Stokes (CHNS) system utilizes a diffusive phase-field for interface tracking of multi-phase fluid flows. Recently structure preserving methods for CHNS have moved into focus to construct numerical schemes that, for example, are mass conservative or obey initial bounds of the phase-field variable. In this work decoupled implicit-explicit formulations based on the Discontinuous Galerkin (DG) methodology are considered and compared to existing schemes from the literature. For the fluid flow a standard continuous Galerkin approach is applied. An adaptive conforming grid is utilized to further draw computational focus on the interface regions, while coarser meshes are utilized around pure phases. All presented methods are compared against each other in terms of bound preservation, mass conservation, and energy dissipation for different examples found in the literature, including a classical rising droplet problem.
Unified Evaluation Methodology for AI-Native Integrated Sensing and Communication
arXiv:2607.14806v1 Announce Type: new Abstract: Integrated Sensing and Communication (ISAC) couples radio sensing, data transmission, and control actions within a single closed-loop system. When Artificial Intelligence (AI)-driven policies adapt sensing and communication online across a variety of sensing tasks and objectives, end-to-end performance is shaped not only by waveform and channel conditions but also by inference latency, uncertainty, environmental dynamics, and hardware non-idealities, leading to fundamental trade-offs between sensing accuracy, communication reliability, and resource overhead. This manuscript presents a unified system architecture and evaluation methodology for AI-native ISAC, defined as ISAC in which learning-based agents adapt sensing, communication, and actuation policies online under uncertainty. We formalize the design space of closed-loop ISAC, propose a three-stage validation pipeline from bounds and feasibility analysis, through high-fidelity digital-twin simulation, to preliminary over-the-air validation, and provide a minimal reporting checklist that links technical Key Performance Indicators (KPIs) (e.g., data rate, SINR, target detection, parameter estimation, track quality, localization error, outage, latency, overhead, and energy per decision) to application-level Key Value Indicators (KVIs) (e.g., availability and mission effectiveness). Two representative instantiations, specifically Unmanned Aerial Vehicle (UAV)-based outdoor and Reconfigurable Intelligent Surface (RIS)-enabled indoor coverage extensions, are used to illustrate how to structure reproducible baselines and comparable evidence across heterogeneous deployments, helping bridge the gap between theoretical ISAC gains and deployment-ready performance claims.
Why Git Is the Memory Solution for the Agentic Development Lifecycle
arXiv:2607.14390v1 Announce Type: new Abstract: Coding agents now produce a growing share of a team's code, while the reasoning behind each change -- the alternatives weighed, the constraints discovered, the approaches rejected -- is trapped in assistant transcripts that vanish with the session. Memory for this setting, the agentic development lifecycle (ADLC), is usually posed as one retrieval problem and built as machinery: tiered stores, memory graphs, compiled wikis, model-judged admission. We argue memory should instead be git-bound -- built into the repository's version control, inheriting the guarantees the machinery struggles to construct: ground truth from commits, freshness from rebuild, verification from the merge, containment from review. On this ledger we solve two problems separately, then combine them. Seed supply is closed as an eight-corpus retrieval study under a pre-registered ship discipline: five imported ranking mechanisms rejected, two kept, and a best configuration of ~0.31 pooled MRR -- ~60x the raw-transcript grep floor, ~15x an honest parsed-turn floor. Answer assembly is where ranking stops helping: single-shot retrieval scores only 0.07-0.20 answer-sufficiency on real developer questions, and ungated episode injection measurably degrades good answers. A router dispatches breadth to a git-anchored structural map, pointed lookups to confidence-gated episodes, and rationale to decision synthesis, which reconstructs why-arcs no single session contains (0.83 sufficiency on a young ~50k-LOC production system). Routed, the system answers at 382-980 tokens per question -- three orders of magnitude below the recorded history. Because ground truth is mined from commit-session links rather than annotated, every result is replicable on any user's own history at zero labeling cost. The remaining constraint is capture. Code, benchmark, and paper source: github.com/rekal-dev/rekal-cli.
Randomized routing strategies of fleets of CAVs may prove market efficient
arXiv:2607.14859v1 Announce Type: new Abstract: In future cities every driver may own a vehicle which could be either independently driven (HDV), or autonomously routed and piloted (CAV). The autonomous operations could be handled by a few competing companies. What is the market structure which would make this market aligned with city goals? In this paper we discuss a variant of the emerging market of collectively routed fleets of CAVs, where revenue for fleet operators is proportional to market share. We provide benchmark scenarios to compare the routing algorithms. We present several routing algorithms and demonstrate that, when the attitudes of human drivers towards CAVs exhibit significant diversity, randomised CAV routing, resulting in unpredictable travel times for HDVs, is more efficient than routing proportional to system optimum/user equilibrium. Based on this, we propose to improve the design of the market by augmenting the market-share objective with mean systemwide travel time in order to limit antisocial randomised strategies of fleet operators and drive the competition towards social welfare oriented cooperation.
Linear representations of grammaticality in neural language models
arXiv:2607.15175v1 Announce Type: new Abstract: Whether neural language models (NLMs) possess the ability to distinguish strings on the basis of their grammaticality remains a debated topic in the computational linguistics literature. Existing evidence has largely relied on probability-based measures, testing whether models assign higher probabilities to grammatical than ungrammatical strings. However, probability comparisons have been criticized as a measure for grammatical knowledge based on the assumption that grammaticality is inherently entangled with likelihood. Model-assigned probability is a function of many related sentence properties, such as lexical frequency, plausibility, and world knowledge. In this work, we move beyond probability-based evaluations and investigate whether grammaticality is encoded in the internal representations of NLMs. Using mass-mean probing, we test whether grammatical and ungrammatical sentences are systematically separated in representational space. We further examine the extent to which these representations are independent of sentence properties that are correlated with grammaticality, as well as their generalization across grammatical phenomena and languages. Our results provide evidence that grammaticality is robustly encoded in sentence representations of a wide range of pretrained NLMs, yielding clear representational separation on the dimension of grammaticality that cannot be fully explained by alternative sentence-level factors. Moreover, this encoding generalizes across a broad range of grammatical phenomena and to some degree, across languages, suggesting that grammaticality constitutes a coherent representational dimension in contemporary NLMs. These findings contribute new evidence to debates about the nature of syntactic knowledge in language models and offer a complementary framework for evaluating grammatical competence that is not dependent on string probabilities alone.
WavePhaseNet: A DFT-Based Method for Constructing Semantic Conceptual Hierarchy Structures (SCHS)
arXiv:2602.14419v2 Announce Type: replace Abstract: This paper reformulates Transformer/Attention mechanisms in Large Language Models (LLMs) through measure theory and frequency analysis, theoretically demonstrating that hallucination is an inevitable structural limitation. The embedding space functions as a conditional expectation over a {\sigma}-algebra, and its failure to be isomorphic to the semantic truth set fundamentally causes logical consistency breakdown. WavePhaseNet Method The authors propose WavePhaseNet, which explicitly constructs a Semantic Conceptual Hierarchy Structure (SCHS) using Discrete Fourier Transform (DFT). By applying DFT along the sequence dimension, semantic information is decomposed into frequency bands: low-frequency components capture global meaning and intent, while high-frequency components represent local syntax and expression. This staged separation enables precise semantic manipulation in diagonalized space. Dimensionality Reduction GPT-4's 24,576-dimensional embedding space exhibits a 1/f spectral structure based on language self-similarity and Zipf's law. Through cumulative energy analysis, the authors derive that approximately 3,000 dimensions constitute the lower bound for "complete representation." This demonstrates that reduction from 24,576 to 3,000 dimensions preserves meaning and intent while enabling rigorous reasoning and suppressing hallucination. Cohomological Consistency Control The reduced embedding space, constructed via cohomological regularization over overlapping local windows, allows defining a graph structure and cochain complex. This quantifies inconsistencies among local inferences as coboundary-based losses. Applying harmonic projection based on Hodge theory positions cohomology as a computable regularization principle for controlling semantic consistency, extracting maximally consistent global representations.
Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
arXiv:2607.14108v1 Announce Type: new Abstract: This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory. To ensure that tool efficiency is well-defined, we also introduce marginal tool utility, a new quantitative metric defined per tool call indicating whether a tool is useful or whether it can be safely removed from the tool suite without affecting accuracy while increasing tool efficiency; in this paper, we determine the sign of marginal tool utility for each tool call in a trajectory using LLM-as-a-Judge. While much prior work has been done to develop techniques that improve tool use by LLMs and design evaluation methods measuring efficiency indirectly using accuracy as a proxy, our work is centered on measuring efficiency directly via the quantitative metric proposed in this paper in post hoc trajectory analyses. It is our intention that this work contributes to the frontier of LLM evaluation research as a springboard for future benchmark designs and agent harness engineering (specifically with regards to creating lean tool suites) that optimize for metrics that complement but are distinct from accuracy.
Spoofer or Spoofers? Estimating a Lower Bound on the Number of DRDoS Sources Using Anycast Honeypots
arXiv:2607.14832v1 Announce Type: new Abstract: DDoS attacks remain a significant threat, with distributed reflection denial-of-service (DRDoS) attacks being particularly difficult to trace back to their sources. To better understand attacker behavior and deployment patterns, we present a novel approach for estimating a lower bound on the number of networks involved in generating spoofed traffic. Our approach leverages a global deployment of anycast amplification honeypots that attract requests from topologically nearby sources. Using this infrastructure, we develop two estimators based on the set of honeypots receiving spoofed traffic and on variations in observed TTL values, while accounting for natural path instability. Analyzing 287 days of amplification attacks, we find that at least 21.0% originate from multiple network locations, indicating that attackers frequently distribute spoofing activity across networks. Our findings suggest that combating spoofing requires coordinated and distributed defenses, and inform the design of future attribution techniques.
The order of long rainbow arithmetic progressions
arXiv:2607.15116v1 Announce Type: cross Abstract: Let $T_k$ be the minimum positive integer $t$ such that, for every positive integer $n$, every equinumerous $t$-coloring of $[tn]$ contains a rainbow $k$-term arithmetic progression. Jungi\'{c}, Licht, Mahdian, Ne\v{s}et\v{r}il and Radoi\v{c}i\'{c} conjectured that $T_k=\Theta(k^2)$, while Conlon, Fox and Sudakov proved that $T_k=O(k^2\log k)$. We prove the matching lower bound $T_k=\Omega(k^2\log k)$, and hence $T_k=\Theta(k^2\log k)$.
T5-CSBoost: Adversarial Perturbation Resistant LLM Fingerprinting
arXiv:2607.14113v1 Announce Type: new Abstract: While many AI-generated text (AIGT) detectors achieve strong performance on clean inputs, their accuracy degrades significantly under light paraphrasing, word substitutions, character edits, and distribution shifts. We present T5 Contrastive Style Boosted Classifier (T5-CSBoost), an extension to the T5-Sentinel framework that keeps the original next-token prediction objective for source attribution while introducing an auxiliary margin-based triplet loss over decoder embeddings. This contrastive style regularization encourages the learning of compact, perturbation-resistant stylistic representations, offering a lightweight yet effective alternative to prior approaches that rely on architectural modifications, adversarial training, or complex multi-task objectives without altering the underlying T5-small backbone. T5-CSBoost achieves state-of-the-art multiclass source attribution and binary human-vs-LLM detection on OpenLLMText and HC3 AIGT benchmarks. More importantly, T5-CSBoost demonstrates enhanced robustness to word and character level adversarial perturbations of up to 90% intensity, achieving state-of-the-art on the challenging MAGE/Deepfake stress-test suite, including unseen models, unseen domains, and extreme paraphrasing scenarios. Our results highlight that explicitly regularizing stylistic embeddings via contrastive learning is a practical and effective strategy for building more robust LLM fingerprinting systems in real-world adversarial settings.
ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
arXiv:2607.14285v1 Announce Type: new Abstract: Safety alignment in LLMs aims to align models with human values, but which values take precedence when they conflict? We investigate this question in the context of tool-calling LLM agents deployed in regulated industries, where agents processing confidential documents may encounter content that triggers safety-trained values (e.g., public welfare) that conflict with deployment-context instructions (e.g., internal logging). To empirically verify this phenomenon, we build a benchmark of 128 scenarios across 16 domains. We find that safety-aligned open-source models override their deployment instructions up to 43.4% of the time, engaging in whistleblowing, data exfiltration, and evidence tampering when processing documents that suggest organizational wrongdoing. We also find that abliteration reduces rates of external whistleblowing. These results reveal a fundamental tension in pluralistic alignment, where the same safety training that protects users can cause agents to act against deployment instructions in ways that create unpredictable liability risks. We release our benchmark as a framework to support evaluation of agent behavior under competing legitimate interests.