arXiv:2606.24129v1 Announce Type: new Abstract: For a wheelchair user, a standard blue line on a map is often a broken promise. While platforms like OpenStreetMap (OSM) successfully capture where a path is, they frequently fail to convey how it physically feels to travel on it. This information barrier is problematic for wheelchair users. To solve this issue, we present OmniPath, a system that moves from passive mapping to proactive environmental auditing. Our framework fuses the network topology of OSM with the submeter precision of high-density aerial LiDAR (USGS 3DEP) to create a high-fidelity 3D model of the pedestrian environment. Rather than simply routing a user, our agent virtually traverses the network, analyzing the surface in 0.5 meter increments. It rigorously quantifies physical friction points specifically running slope, cross slope, and vertical discontinuities against ADA compliance standards, calculating a weighted severity score to categorize hazards from ``Mild'' to ``Critical.'' To ensure real world reliability, we validated the system against 200 physical ground truth field surveys across the National Mall using stratified random sampling. The framework demonstrated strong diagnostic reliability for high-severity hazards, achieving F1-scores of 0.60 for Severe and 0.58 for critical categories. By automating this micro-scale inspection, OmniPath identifies the ``invisible'' barriers that standard maps miss, effectively transforming a static dataset into accessibility data source that anticipates accessibility challenges before the user ever leaves home.
Science Journals
arXiv:2606.24151v1 Announce Type: new Abstract: Self-evolving agents improve over time by distilling experience from past executions and reusing it in future tasks. Existing systems represent such experience either as natural-language text injected into the agent context or as code exposed as callable tools. However, the choice between these representations is typically made at design time rather than derived from the characteristics of the experience itself, leaving the trade-offs between them poorly understood. We present the first controlled study that isolates text memory and code memory over an identical set of experiences. Our results show that the two forms exhibit complementary trade-offs in construction cost, execution efficiency, and transferability, such that neither representation alone is sufficient. Guided by these findings, we propose Metis, a self-evolving agent system built on a hierarchical dual-representation memory. Metis organizes textual experience into execution plans, environment facts, and common pitfalls, and selectively crystallizes recurring plans into validated callable tools. This design combines the broad applicability of text memory with the execution efficiency of code memory while incurring tool-generation cost only when justified by repeated reuse. We evaluate Metis on AppWorld, a challenging benchmark for interactive agents. The results show that Metis improves task accuracy by up to 20.6% over ReAct while reducing execution cost by up to 22.8%. Compared with representative self-evolving agent systems, Metis consistently achieves a better balance between accuracy, execution efficiency, and memory-construction cost.
arXiv:2606.24384v1 Announce Type: new Abstract: Ultra-high dose rate (UHDR) irradiation used in FLASH radiotherapy induces strong space-charge effects in plane-parallel ionisation chambers (PPICs), leading to significant reductions in charge collection efficiency (CCE). To investigate these effects, we extended the Garfield++ framework by implementing ion-ion recombination and self-consistent space-charge electric field calculations. The developed Monte Carlo model couples particle transport, electron attachment, recombination processes, and dynamic electric-field distortions. The implementation was validated against analytical and numerical models from the literature, including the works of Fenwick and Kumar, Kranzer et al., and Paz Mart\'in et al., with excellent agreement for the free electron fraction (FEF), CCE, induced current, and electric field evolution. The simulations show that space charge can locally increase the electric field by more than a factor of four or reduce it to nearly zero. The results suggest that CCE reduction under UHDR conditions is mainly driven by the decrease of the FEF caused by electric-field-dependent electron attachment, indicating that recombination may be largely governed by FEF evolution. This opens promising perspectives for improved analytical models and real-time correction methods for ionisation chamber dosimetry under UHDR conditions.
arXiv:2606.24160v1 Announce Type: new Abstract: Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i.e., what would have happened had reality been different, even when no data of this unrealized reality is currently available. Reinforcement learning provides methods to learn a policy that optimizes a specific measure (e.g., reward, regret) when the agent is deployed in an environment and pursues an exploratory, trial-and-error approach. These two disciplines have evolved independently and with virtually no interaction between them. We note that they operate over different aspects of the same building block, counterfactual relations, which makes them umbilically connected. Based on these observations, novel learning opportunities arise when this connection is explicitly acknowledged and mathematized. To realize this potential, we note that any environment where the RL agent is deployed can be decomposed as a collection of autonomous mechanisms with different causal invariances, parsimoniously modeled as a structural causal model; any standard RL setting implicitly encodes such a model. This formalization allows us to put under a unifying treatment different modes of learning, including online, off-policy, and causal calculus learning, which appear unrelated in the literature. However, these modalities are not exhaustive: we introduce several natural and pervasive classes of learning settings that entail novel dimensions of analysis. Specifically, we introduce and discuss through causal lenses generalized policy learning, where to intervene, imitation learning, and counterfactual learning. These tasks lead to a broader view of counterfactual learning and suggest great potential for studying causal inference and reinforcement learning side by side, which we call causal reinforcement learning (CRL).
arXiv:2606.24172v1 Announce Type: new Abstract: More than a billion people communicate in Indic languages, yet the natural language processing infrastructure serving them remains fragmented and underdeveloped. The cause is structural: the field organizes its tools and benchmarks around individual languages or small subsets of genealogical language families, building separate analyzers, parsers, and datasets for each language and starting over for the next. This overlooks a deep regularity. Through more than two millennia of convergence around Sanskrit, Indic languages came to share a morphosyntactic architecture formalized in P\={a}nini's grammar, the Ast\={a}dhy\={a}y\={i}. This cuts across genealogical lines, uniting languages through a common framework. We argue that this P\={a}ninian framework supplies a unifying computational architecture the field has lacked, and that benchmarks grounded explicitly in it would make Indic language systems more accurate, more data-efficient, and more transferable, effectively merging many apparently disparate and sparse Indic language resources into a single high-resource metalanguage bedrock. We propose a four-part benchmark suite to render this shared architecture explicit, measurable, and ready to be leveraged for practical applications. Moreover, we underscore the question it raises for interpretability research: whether neural models trained on these languages come to represent P\={a}nini's categories on their own.
arXiv:2606.24404v1 Announce Type: new Abstract: The incorporation of additional modalities into action recognition models increases their performance across a wide range of settings. However, how this additional information can contribute to making the models more robust remains underexplored, particularly for the case of multi-modal out-of-distribution (OOD) detection. While methods exist that regularize the multi-modal training process with OOD detection in mind, they still apply off-the-shelf OOD detectors designed for the uni-modal case during inference, discarding important information. Based on an interesting relationship we find between the multi-modal and uni-modal predictions, we propose to use this signal to build a post-hoc detector explicitly designed for the multi-modal scenario. We combine this new source of information with a feature-space score, which detects off-manifold samples in the multi-modal space, and normalize them by the multi-modal logits. In doing so, the proposed hybrid detector is compatible with existing training-time approaches and consistently improves performance. Experiments on a wide range of established datasets from the MultiOOD benchmark show that, on average, our approach outperforms the state of the art. Our results show the importance of explicitly considering the different modalities at inference time for multi-modal OOD detection.
arXiv:2606.24423v1 Announce Type: new Abstract: We develop a sparse spectral method for solving partial differential equations on a class of two-dimensional geometries bounded by algebraic curves. The numerical method uses generalised bivariate Koornwinder polynomials which form a complete orthogonal basis, but one which is not graded in terms of polynomial degree. The polynomials are built from new families of univariate semiclassical orthogonal polynomials whose associated operator matrices (Jacobi matrices, raising matrices and differentiation matrices) are computed with optimal linear complexity in the number of basis functions. When used to discretise partial differential equations the resulting matrices are sparse enabling efficient numerical solution. Moreover, we develop fast transforms from values on a grid to expansion coefficients. The efficiency and accuracy of the resulting spectral method are illustrated through a series of numerical experiments on geometries whose boundaries are smooth and piecewise smooth including non-convex geometries. We observe algebraic convergence for geometries with corners, which accelerates to exponentially fast (spectral) convergence when the boundary is smooth.
arXiv:2606.23890v1 Announce Type: new Abstract: To date only one chiral species, propylene oxide, has been observed in the interstellar medium but little is known about the chemistry that leads to a detectable abundance of this molecule in Sagittarius B2. We used a glow-discharge ion source and a room-temperature ion trap to study the neutralization reactions necessary to convert propylene oxide cation (PO$^+$) -- the assumed precursor for propylene oxide in space -- into the observed astrochemical. We found that the charge-transfer reaction between PO$^+$ and ammonia (NH$_3$) proceeds with a pressure-independent rate coefficient of $(1.39\pm0.03)\times10^{-12}$ cm$^{3}$ s$^{-1}$ to neutralize PO$^+$ and form the radical cation NH$_3$. Although this measured rate coefficient is much slower than that predicted by capture theories, the high abundance of NH$_3$ in Sagittarius B2 motivates the inclusion of this reaction in astrochemical models. We hypothesize that the low ionization energies of many chiral molecules important to origin-of-life theories means these species may exist as cations in the interstellar medium.
arXiv:2606.23805v1 Announce Type: new Abstract: The 2025 Iberian blackout has renewed concerns about the resilience of power grids with high shares of renewable generation. This commentary argues that renewable generation can not only advance decarbonization but also strengthen grid stability through synthetic inertia, advanced inverter-based control, and coordinated transmission planning. Rapid advances in energy storage and power electronics make this transition increasingly viable.
arXiv:2606.23891v1 Announce Type: new Abstract: The rise in GPU compute speed has outpaced improvements in host-to-device memory transfer speeds, despite the advent of shared-memory superchips. Consequently, memory transfer times now constitute an increasingly large fraction of total time-to-solution, compelling developers to compress GPU kernel input and output data into compact, minimal formats prior to GPU-offloading. This complements existing work on GPU- and compute-friendly data arrangements. We study a Smoothed Particle Hydrodynamics solver and propose memory layout strategies for host-side particle data that are particularly well-suited to GPU-offloading. Specifically, we advocate splitting classic array-of-struct data structures into a split array-of-struct arrangement, in which each logical struct decomposes into substructs determined by kernel read/write access patterns and attribute types. Splitting a monolithic particle struct into several bespoke, finer-grained structs can reduce the time required to pack data to and from buffers by ~20% - 40%, lowering total time spent on GPU-offloading by ~12% - 25%.
arXiv:2606.23901v1 Announce Type: new Abstract: This paper addresses the problem of robust formation control by introducing Topological Online Learning for Displacement-based (TOLD) formation control, a real-time edge-level adaptation framework. Unlike conventional node-level robust controllers that regulate individual robot inputs without modifying the interaction topology, TOLD updates the interaction topology weights online to directly minimize formation distortion. Two strategies are proposed under the TOLD formation control framework: Online Gradient Flow (OGF) with unconstrained weights and Online Exponential Gradient Flow (OExpGF) with non-negative convex weights. Theoretical analysis establishes that, for single-integrator agents over directed graphs, OExpGF guarantees asymptotic consensus, while OGF ensures bounded formation distortion. Simulations with twelve robots under intermittent disturbances show 1.2%-33.14% median cumulative Root Mean Distortion Error reduction when augmenting TOLD with node-level controllers. Hardware experiments with Crazyflie 2.0 quadrotors demonstrate over 62% (OGF) and 31.4% (OExpGF) reduction in median formation distortion compared to fixed-weight consensus.
arXiv:2606.24467v1 Announce Type: new Abstract: Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) caches, limiting sustainable deployment on resource-constrained hardware. Existing KV cache eviction methods typically apply heuristic token scoring over all heads in GQA-based LLMs. These methods ignore the different functionalities of attention heads, leading to the eviction of critical tokens and thus degrading the performance of LLMs. To address this issue, we propose CompressKV, a resource-efficient KV-cache compression framework for GQA-based LLMs. Instead of aggregating attention scores from all heads, CompressKV identifies Semantic Retrieval Heads (SRHs) that capture both the initial and final tokens of a prompt and semantically important mid-context evidence, and uses them to select tokens whose KV pairs should be retained. Furthermore, CompressKV allocates cache budgets across layers according to offline estimates of layer-wise eviction error. Experiments on LongBench and Needle-in-a-Haystack show that CompressKV consistently outperforms existing KV-cache eviction methods across memory budgets. Notably, it preserves over 97\% of full-cache performance using only 3\% of the KV cache on LongBench question-answering tasks and achieves 90\% accuracy with just 0.7\% KV storage on Needle-in-a-Haystack. These results demonstrate an improved resource--performance trade-off for long-context LLM inference. Our code is publicly available at: https://github.com/TUDa-HWAI/CompressKV
arXiv:2606.24658v1 Announce Type: new Abstract: Nanostructures can be designed to absorb light efficiently at resonance despite their subwavelength footprint, but causality and passivity fundamentally limit the bandwidth over which strong absorption can be maintained. Here we derive fundamental absorption-bandwidth limits for passive, causal, linear, and temporally dispersive subwavelength objects by rigorously casting electromagnetic scattering as an equivalent impedance-matching problem. This mapping yields ultimate Bode-Fano-type constraints for optical absorption and provides rational synthesis guidelines for the material dispersion of passive nanoparticles that can approach the bounds. Our results clarify the ultimate limits for broadband light harvesting and dissipation, with implications for solar-energy conversion, photothermal hyperthermia, thermal management, and related nanophotonic technologies.
arXiv:2606.24489v1 Announce Type: new Abstract: Pose graph optimization (PGO) is a key back-end component for state estimation in networked multi-robot simultaneous localization and mapping (SLAM). In object-based multi-robot SLAM, the problem becomes more tightly coupled because robots must jointly estimate both their trajectories and the poses of persistent objects observed by multiple agents. Existing decentralized solutions often assume that the communication graph closely matches the physical interaction topology, which is restrictive in realistic deployments where communication is sparse, intermittent, or time-varying. This paper presents a fully decentralized Riemannian optimization framework for object-based multi-robot PGO that decouples the coupled estimation problem via a consensus mechanism, enabling flexible communication topologies. To improve convergence under limited communication budgets, we further develop a distributed approximate-Newton scheme that exploits local second-order information while operating directly on the SE(d) manifold to preserve geometric consistency, and we establish the convergence to Riemannian first-order stationary points and provide a local condition-number analysis explaining the benefit of approximate second-order information over first-order Riemannian descent. The resulting method reduces iteration count and communication overhead without sacrificing estimation accuracy. Extensive evaluations on public benchmarks, large-scale simulations, and real-world multi-robot experiments demonstrate improved accuracy, runtime efficiency, scalability across network topologies, and robustness to communication failures.
arXiv:2606.24496v1 Announce Type: new Abstract: The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability. However, while the community has focused on creating more and more capable agents, less attention has been allocated to assessing the security of those systems. In this work, we present the first in-depth security analysis of the most widely used agentic systems for offensive security operations. We show that most of these tools share common design flaws that enable an active adversary to exfiltrate API keys, establish persistent footholds, and fully compromise the operator's machine, even when the agent operates inside a sandboxed container. To support our analysis, we introduce a full cyber kill chain for such agentic systems, capturing the progression from initial LLM manipulation to lateral movement, persistence, guardrail bypass, and sandbox escape. Building on our security analysis, we derive a robust architecture for agentic offensive-security tools and propose actionable, broadly applicable design principles that mitigate the disclosed attack paths at the architectural level.
arXiv:2606.24500v1 Announce Type: new Abstract: Reachability analysis for dynamical systems seeks to compute a set containing all reachable states at a given time. Compared to ordinary differential equations (ODEs), the analysis of nonlinear reaction--diffusion PDEs with parametric uncertainties remains largely underexplored, due to the infinite-dimensional state space and the variety of solutions under different parameters. We address this through a three-step procedure: 1) Finite Element Methods (FEM)s to discretise the space and generate a finite-dimensional FEM-based model, 2) Proper Orthogonal Decomposition (POD) to build a Reduced-Order Model (ROM), and 3) set-based reachability-analysis methods applied to the ROM. We propose a framework that enables us to derive explicit upper bounds on the approximation errors introduced at each stage of the pipeline. In particular, we quantify the discrepancy between trajectories of the original PDE and those of the FEM-based discretization, as well as the error between the FEM-based model and the reduced-order model. Importantly, these bounds are shown to hold uniformly over the considered set of parameters. By combining these error estimates, we obtain an over-approximation of the reachable set of the original PDE. The approach is illustrated on the Allen--Cahn equation and a logistic growth PDE.
arXiv:2606.24676v1 Announce Type: new Abstract: The potential to precisely control both the linear and nonlinear index of refraction through optical manipulation of the atomic states has recently pushed warm alkali vapors to the forefront of research in the field of quantum sensors, quantum memories, and quantum fluids of light. Rubidium (Rb) vapor in centimeter-scale glass cells or millimeter-scale MEMS cells has proven to be a very promising platform for these applications, yet only a handful of research works have been dedicated to the investigation of the (non)linear refractive index of Rb vapor. We present results of theoretical calculations of the (non)linear refractive index of warm Rb vapor, based on the optical Bloch equations for 6-level Rb atoms interacting with a probe laser. They are compared to the experimental results obtained using an interferometric technique, showing excellent quantitative agreement. A Kerr nonlinear refractive index $n_2$ of up to $10^{-4}$ cm$^2$/W is obtained. Python scripts for all theoretical calculations presented in this work are provided, including the refractive index calculation, that can readily be used in practical implementations for simulating the (non)linear refractive index of Rb vapor including the effects of Doppler broadening, transit time broadening, pressure broadening, saturation, optical pumping, and spin-exchange collisions.
arXiv:2606.24515v1 Announce Type: new Abstract: Computer-Use Agents (CUAs) execute high-level user goals by perceiving and acting directly within graphical user interfaces. However, reinforcement learning for CUAs remains difficult because open-ended desktop environments rarely provide scalable, machine-readable reward signals: task success is often visually grounded and hard to specify with handcrafted reward functions or dense manual labels. We propose an RL fine-tuning framework that uses autonomous vision-language evaluation as a scalable supervision signal for GUI agents. Given a final screenshot and the original instruction, a Vision-Language Model judges task completion and provides terminal feedback without task-specific heuristics or manual labels during policy optimization. Because autonomous evaluators are imperfect, we model their feedback as a noisy binary reward channel and derive a noise-corrected reward estimator for Proximal Policy Optimization. Experiments across macOSWorld, Windows Agent Arena, and OSWorld show that corrected evaluator rewards outperform both zero-shot baselines and raw evaluator rewards, improving success rates by an average of 12.6 percentage points over zero-shot performance and 5.1 points over raw evaluator fine-tuning. These results suggest that autonomous evaluation can serve as a practical reward signal for RL in GUI environments when evaluator noise is explicitly modeled and corrected.
arXiv:2606.24535v1 Announce Type: new Abstract: Multi-agent LLM environments require robust mechanisms for shared knowledge management. This paper formalizes the fleet-memory problem and identifies four foundational failure modes: unauthorized leakage, stale propagation, contradiction persistence, and provenance collapse. To address these, we define explicit systems-level primitives: scoped retrieval, temporal supersession, provenance tracking, and policy-governed memory propagation. These primitives are implemented in MemClaw, a production multi-tenant memory service, and evaluated via ArgusFleet, a reproducible harness testing four governance dimensions. Rather than a baseline comparison, this study measures a live production service, emphasizing real-world architectural insights and negative results. Key Evaluation Results Provenance: Successfully reconstructed 100% of depth-four derivation chains with correct writer identity at sub-second per-hop latency. Propagation: Demonstrated high intra-fleet visibility with zero cross-fleet leakage. Under strong write mode, write-to-visible latency was optimized to a single search round-trip. Production Architectural Issues Discovered Asymmetric Scope Enforcement: Tenant isolation held, but sub-tenant scope was initially bypassed on direct GET-by-id requests for agent-scoped credentials (disclosed and remediated during the study). Pipeline Ordering Conflict: While contradiction supersession works for admitted writes, a synchronous near-duplicate gate can prematurely reject contradictory writes before the asynchronous contradiction detector can evaluate them. Conclusion: Long-context retrieval alone is insufficient for production multi-agent memory. Governed shared memory demands explicit systems-level abstractions, and live evaluation is vital to expose enforcement and pipeline-ordering failures missed by design-only treatments.
arXiv:2606.24537v1 Announce Type: new Abstract: Traditional importance sampling (IS) is designed to estimate rare-event probabilities by minimizing estimator variance. However, many applications prioritize rapid discovery: the generation of a trajectory within a rare set $A_n$. This requires a shift from ensemble-based estimation to a design principle focused on the hitting time $\tau_{A_n} := \inf\{t \ge 1 : Y_t^n \in A_n\}$. We formalize a Quality of Discovery problem as the problem of minimizing the description length (surprisal) of the discovered trajectory under the nominal model $p$. We prove that minimizing this description length is equivalent to minimizing the nominal rank exponent $J_{\mathrm{rank}}(q_n) := \lim_{n\to\infty} \frac{1}{n} \log G_n(Y^n)$, where $G_n(x^n)$ is the guesswork of sequence $x^n$. For i.i.d.\ models and type-defined rare sets $\Gamma$, we show that while classical IS targets the mass-dominating type $Q_{\mathrm{IS}}^* \in \arg\min_{Q \in \Gamma} D(Q\|p)$, discovery optimality is achieved by $Q_{\mathrm{GW}}^* \in \arg\min_{Q \in \Gamma} [H(Q) + D(Q\|p)]$. This framework identifies a fundamental rule: minimizing the guesswork exponent ensures the discovered sequence is the "least surprising" representative of the set relative to the nominal model's search order. We further demonstrate that under budgetary constraints, this exponent serves as a lexicographic tie-breaker when the hitting-time minimizer is not unique. This establishes $H(Q) + D(Q\|p)$ as a natural objective for discovery-based importance sampling, providing a formal bridge between randomized sampling and systematic search.
arXiv:2606.24551v1 Announce Type: new Abstract: Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces, but existing evaluations confound interaction modality with differences in tasks, initial states, verifiers, and permitted actions. We introduce a matched execution-layer benchmark of 440 desktop tasks across 18 applications and 12 workflow categories, where screen-only GUI agents and skill-mediated CLI agents receive identical goals, states, and final-state verifiers while being restricted to modality-native actions. In this controlled setting, the strongest GUI agent reaches a 59.1% full pass rate, outperforming the strongest original-skill CLI agent at 48.2%; however, verifier-guided skill augmentation raises CLI success to 69.3%, showing that much of the CLI deficit comes from incomplete skill coverage rather than model capability alone. These results suggest that GUI and CLI expose different execution bottlenecks: GUI agents are limited by reliable grounded interaction over long-horizon workflows, whereas CLI agents are limited by the coverage and scalability of their skill interfaces.
arXiv:2606.24677v1 Announce Type: new Abstract: Time-series, geospatial, and ontology systems each maintain a hierarchy -- day <= month <= year, zip <= city, is-a / part-of -- and each indexes it in a separate silo. We observe these are all subsumption posets, with one recurring workload: order testing (is x under y?) and hierarchical roll-up (aggregate a measure over everything under y). We present OEH, a single declarable index that, by a cheap structural probe, encodes a hierarchy as a nested-set order-embedding (trees) or a chain decomposition (low-width DAGs), and answers both subsumption and index-resident monoid roll-up from one structure. On five real hierarchies -- Gene Ontology, NCBI Taxonomy (1.3M), GeoNames (330k), a 2.6M-node calendar, and git commit DAGs -- OEH on trees matches a 2-hop index on query latency using about half the space and building 6--7x faster, and adds roll-up that 2-hop cannot. Its roll-up matches TimescaleDB's continuous aggregates exactly and in the same latency regime, while also answering subsumption. On high-width DAGs the chain index is declined and 2-hop dominates. Order-embedding is classical; our contribution is the unification and the structure-selected index over subsumption and index-resident roll-up.
arXiv:2606.24774v1 Announce Type: new Abstract: Vision-Language Large Models (VLLMs) trained on massive crawled corpora raise pressing copyright and data-provenance concerns. These concerns are particularly acute in healthcare, where patient medical images paired with clinical reports demand rigorous privacy safeguards. However, existing training data detection methods either fail in cross-modal scenarios or rely on superficial output signals with insufficient discriminative power. We introduce GradAudit, a gradient-based auditing framework that examines internal optimization dynamics rather than treating VLLMs as black boxes. Our approach builds on a key observation: model parameters converge to regions where gradients on training samples become stable and well-aligned, whereas gradients on non-training samples remain noisy and inconsistent. By analyzing these gradient signatures, GradAudit achieves strong separability and detects genuine image-text associations learned during training, not merely individual modality membership. Empirically, across both medical and general-domain datasets, GradAudit substantially outperforms state-of-the-art baselines in both pretraining and fine-tuning VLLMs. In a case study employing copyrighted content, we show that existing training data detection methods not only underestimate the extent of unauthorized data usage, but that this underestimation becomes more pronounced as models become more recent and more advanced.
arXiv:2606.24687v1 Announce Type: new Abstract: Solar Cycle 25 has run far stronger than the 2019 consensus forecast issued by the NOAA/NASA/ISES prediction panel, with densities in low Earth orbit from 2022-2026 holding at 2-3x the predicted levels. The cumulative drag impulse experienced by LEO satellites reached 5-6 standard deviations beyond the forecast's stated uncertainty. This means that even operators who designed conservatively against the two-sigma worst case fell short of their drag budgets. This paper quantifies a lower bound on the economic cost of that misprediction. Starting from the 13,704 payloads on-orbit below 800 km during 2022-2026, we screen to the 1,597 payloads which we validated with high confidence to be both operational and in ballistic freefall. We estimate each satellite's ballistic coefficient and propagate its trajectory under the forecasted atmosphere versus the observed one. A probabilistic cost model assigns each satellite an annualized mission cost based on direct costs (amortized capital costs plus annual operations), stratified by size class, with bespoke estimates for high value missions. Survival and forward cost discounting is applied at a modal 11% per year. We combine the differences in lifetime with the cost model to estimate the total dollar impact. Against the forecast's two-sigma upper bound which we consider to be a standard engineering design target, these satellites lost 688 cumulative mission years valued at \$0.88 billion. Against the nominal forecast, they lost 2,472 mission years worth \$2.77 billion. These estimates are deliberate lower bounds which exclude propulsive satellites, revenue above direct cost, and downstream economic impact. The results give a quantitative case for the value of accurate decadal-scale space weather forecasting, and show that well-calibrated uncertainties are as valuable to a satellite operator end-user as the accuracy of the central prediction itself.
arXiv:2606.24689v1 Announce Type: new Abstract: Large Language Models (LLMs) and LLM-based Multi-Agent Systems (MAS) are revolutionizing software engineering (SE) by advancing automation, decision-making, and knowledge processing. Their recent application to SE tasks has already shown promising results. In this paper, we focus on summarization as a key application area. We present Metagente, an LLM-based MAS designed to generate concise and accurate summaries of software documentation. Metagente employs a Teacher-Student architecture where multiple LLM agents collaborate to enhance relevance and precision of produced summaries. An empirical evaluation on real-world datasets demonstrates Metagente's effectiveness in streamlining workflows, outperforming the considered baselines. The evaluation provides evidence that Metagente improves summarization for requirements analysis and technical documentation. Our findings underscore the transformative potential of these technologies in SE, while identifying challenges and future research directions for their seamless integration.