arXiv:2607.15143v1 Announce Type: new
Abstract: AI coding agents set up projects by reading documentation and installing the dependencies it lists, without verifying their names, sources, or known vulnerabilities. By editing only a README, requirements file, or Makefile, an attacker can redirect the agent to an untrusted registry, a known-vulnerable version, or a wrong-but-plausible name: documentation becomes a vector for code execution. We present the first systematic evaluation of package-install-time supply-chain attacks delivered through ordinary project-setup documentation across production coding-agent harnesses, probing frontier models on twelve scenarios in five attack classes, grounded in documented incidents. The same model catches an attack through one harness and installs it through another: install-time security rests on the harness-model combination, not the model alone. Agents catch blatant typosquats reliably, but plausible separator-confusion names (azurecore for azure-core) slip through, and how often depends on the harness-model pairing. Source-based attacks like registry redirection are missed almost everywhere. The source blind spot recurs on npm and Cargo, where nearly every model installs the untrusted dependency; name detection carries over less consistently across ecosystems. Security-oriented prompts recover part of the gap but only for the dimension they name; a deterministic pre-install check that verifies names, sources, and versions before any code runs closes most of it.
Science Journals
arXiv:2607.15233v1 Announce Type: cross
Abstract: The undamped Duffing oscillator is a nonlinear dynamical system with broad applications in physics, engineering and biological system. We present a comprehensive analysis of this system using the Lindstedt Poincare method (LPM) and its modifications and make comparison with numerical solution obtained using higher order Runge-Kutta. It is also shown the method suggested in this article converges better than the standard LPM and Lindstedt Poincare method with Burton's modification.
arXiv:2607.13902v2 Announce Type: replace
Abstract: Self-assembly of supramolecular structures in cells and synthetic applications often proceeds under unfavorable biochemical conditions and at low copy numbers of final target structures, ranging from tens of bacterial microcompartments to a single bacterial flagellum per cell. Spatial organization through coupled reaction compartments of different reactivity (delay-facilitated assembly) can recover high yield in such environments at the mean-field level, but its robustness to stochastic fluctuations at low target numbers is unclear. Using stochastic simulations of a minimal two-compartment model, we show that delay-facilitated assembly is susceptible to a stochastic yield catastrophe at low target numbers: even when each compartment in isolation allows for high-yield assembly, slow exchange between them induces a substantial drop in the final yield. We trace the mechanism to a specific assembly stage, where the random order of rate-limiting exchange events of subunits and partially completed structures determines the ratio of productive growth to excess nucleation. Restricting the exchange of larger structures -- either by suppressing it entirely or letting exchange rates decrease with size -- restores most of the yield without altering the mean-field behavior. The same phenomenology appears for two-dimensional hexagonal subunits and in a cytosol-membrane geometry, where diffusion-limited exchange naturally implements the required size dependence. Our results show that equal success of assembly strategies at high target numbers does not imply their equal success at low target numbers, and that competing slow events occurring in random order are a common signature of stochastic yield catastrophes.
arXiv:2607.14842v1 Announce Type: new
Abstract: Dexterous in-hand manipulation requires continuous 6D pose tracking, yet the manipulating fingers inevitably occlude the object from the camera. We study how to structure the sparse haptic signals already available on multi-fingered hands, including proprioception, proximal force/torque, and binary contact, to complement a pretrained visual pose tracker under occlusion. We propose a kinematic-aware finger-level encoder and systematically compare it against four alternative designs through three levels of evaluation: per-frame refinement, sequential open-loop tracking, and closed-loop manipulation. Our experiments reveal that (i) per-frame evaluation cannot distinguish encoder quality, while sequential tracking amplifies architectural differences by up to 15 times; (ii) the structured encoder learns task-specific cross-modal gating, using vision exclusively for translation and dedicating one attention head to haptics for rotation, without explicit supervision; and (iii) compact finger-level tokenization with 4 tokens outperforms both flat fusion and joint-level representations, which suppress vision through norm dominance. We validate that improved tracking yields higher success in a downstream reorientation task and provide qualitative real-world demonstrations. Our project page is available at https://cold-young.github.io/kine-fuse/.
arXiv:2607.14509v1 Announce Type: new
Abstract: This paper describes DS@GT ARC's third-place solution to the PlantCLEF 2026 challenge on multi-species plant identification in vegetation quadrat images, where systems must predict every species present in high-resolution (~3000 x 3000 pixel) plot photographs while training only on single-label images of individual plants. The pipeline is built around a fine-tuned DINOv2 ViT-L/14 classifier applied over a multi-scale tile decomposition of each quadrat, with per-tile predictions blended with a FAISS kNN retriever and post-processed by source-aware temporal fusion across repeated plot visits, a habitat-fit demotion that injects geographic and altitude priors from the training data, and a South-Western Europe geographic mask. Habitat-fit demotion and multi-scale aggregation are the largest individual contributors in the ablations. Two complementary training-centric directions, a cross-region transformer with noisy-student distillation on the LUCAS dataset and a label-as-query transformer decoder over synthetic CLS-domain pseudo-quadrats, yielded null results. An inference-time augmentation with instance-aware segmentation crops also did not improve performance. The selected submission reaches a private-leaderboard macro-F1 of 0.43902 (third place; public 0.51096); an unselected configuration of the same pipeline scored above 0.45 on the private set. Code: https://github.com/dsgt-arc/plantclef-2026.
arXiv:2504.04177v2 Announce Type: replace-cross
Abstract: We describe the physical processes that affect the formation, trapping, and outgassing of O2 at Europa and Ganymede. Following Voyager measurements of their ambient magnetospheric plasmas, laboratory data indicated that observed ions, mostly ejected from volcanic Io, would in turn impact and sputtering their surfaces, decomposing the ice producing thin oxygen atmospheres. Subsequently, Europa and Ganymede's O2 atmospheres were inferred from O aurora, condensed O2 bands identified at 5773 and 6225 Angstroms, and their atmospheres were shown to have a dusk/dawn enhancement, confirmed by recent Juno data. Although plasma produces these observables, processes that occur within the topmost surface are not well understood. Here, we note that the incident plasma particles produce nonequilibrium defect density locally in the surface ice grains. Defect diffusion within these grains leads to the formation of voids and molecular products, some of which are volatile. Although some volatiles are released into the satellite atmospheres, others are trapped at defect sites or trapped in voids, creating bubbles whose lifetimes are limited by the plasma-induced destruction rate. We discuss how trapping competes with annealing of the radiation damage, and how hemispheric differences at Europa and Ganymede, roughly determine the observed trend with latitude of O2 bands. We discuss the relative importance of condensed O2 and O2 adsorbed on regolith grains as atmospheric sources, accounting for dusk/dawn enhancements and temporal variability reported in condensed O2 band depths. Since plasma-induced damage and thermal annealing timescales drive oxidant variability on icy moons (likely also Callisto, Dione, and Rhea), they can help determine volatile downwelling, a potentially metabolic source for their oceans, and upwelling of other trapped oxidants (e.g. CO2) suggestive of ongoing geologic activity.
arXiv:2607.14437v1 Announce Type: new
Abstract: Smartphone-based context-awareness holds significant promise for wheelchair users -- from detecting everyday accessibility barriers to enabling ability-based adaptations. Such capabilities often build on passive context inference through mobile sensing, yet their accuracy hinges on how and where phones are carried and the resulting signal quality. While prior work documents phone-carrying behaviors in the general population, patterns specific to wheelchair users remain underexplored. Through a mixed-methods approach combining a survey of 91 and interviews with 15 wheelchair users, we systematically investigate their phone-carrying locations and influencing factors. Our findings reveal distinct patterns extending beyond pocket storage to diverse wheelchair-mounted accessories and around-body placements, shaped by the interplay of physical ability, wheelchair design, and everyday contexts, including social, activity, and device factors. Grounded in these findings, we articulate how carrying location can serve as a proxy for user context to enable novel context-aware experiences, and discuss design implications for developing inclusive and effective mobile context-aware applications.
arXiv:2607.14676v1 Announce Type: new
Abstract: A variable-to-variable (V2V) length code parses a source sequence into phrases of variable length and maps each phrase to a binary codeword of, generally, a different random length. After encoding $n$ phrases, the realized compression ratio $R_n=\Lambda_n/\Sigma_n$ -- total codeword length over total source-symbol count -- is the finite-sample counterpart of the code's asymptotic rate $\rho$, to which it converges only as $n\to\infty$. This paper first derives exact formulas for all integer moments of $R_n$ for a given discrete memoryless source (DMS). Specifically, we obtain a closed-form formula for every moment $\E\{R_n^k\}$ as a one-dimensional integral involving only single-phrase moment generating functions of the pair $(L,\ell)$ -- the phrase length, in source symbols, and codeword length, in bits. From these moments we derive an Edgeworth approximation to the cumulative distribution function (CDF) of $R_n$ that is substantially more accurate than the central limit theorem (CLT) approximation. Using the Laplace method of integration, we also derive explicit closed-form formulas for the bias constant $C=\lim_{n\to\infty}n(\E\{R_n\}-\rho)$ and for the variance constant $\lim_{n\to\infty}n\cdot\Var\{R_n\}$. The analysis extends to Markov sources via state-indexed matrices with a redundancy formula obtained in closed form. On the coding-theoretic side, we cast V2V length codes as finite-state encoders and apply a generalized Kraft inequality for a compression-rate lower bound, and give a structural decomposition of the bias coefficient that separates cleanly across variable-to-fixed (V2F) length codes, fixed-to-variable (F2V) length codes, and V2V length codes. Applied to the Khodak code of Bugeaud, Drmota, and Szpankowski, this decomposition shows that its improved performance is reflected in its smaller bias constant.
arXiv:2606.17930v3 Announce Type: replace
Abstract: AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As a result, performance is increasingly sensitive to the amount and allocation of compute available at test time ("inference compute"). Yet many evaluations still report performance at a single restrictive budget, meaning that low scores may reflect the evaluation setup rather than the model's underlying capability. To test this, we evaluate up to 12 frontier language models on seven challenging benchmarks spanning software engineering, mathematics, medicine, and cybersecurity. We use a controlled setup combining three simple inference-scaling interventions: larger token budgets, context compaction, and repeated submission attempts, guided either by the model itself or by minimal correctness feedback. We find three main results. First, larger token budgets substantially improve performance on benchmarks across multiple domains, including cybersecurity, FrontierMath, Humanity's Last Exam, and TerminalBench. Second, fixed-budget evaluations can increasingly understate frontier capability as models advance. Newer models reach higher performance at large budgets, where they unlock harder tasks and solve them more reliably. Third, benchmarks differ in which inference-scaling methods help most: repeated submission broadly improves performance, but the value of larger token budgets, external feedback, and parallel attempts varies by benchmark. Overall, our results show that benchmark scores are protocol-dependent. We therefore argue that evaluations should report capability as a function of inference-time compute, specify protocol choices explicitly, and compare model generations over a large shared compute range at matched budgets, especially in safety- or policy-relevant settings.
arXiv:2606.18918v4 Announce Type: replace
Abstract: This paper investigates the computational complexity of verification problems for Binarized Neural Networks (BNNs), in which activations and weights are binary. Specifically, we study three verification problems. First, we prove that checking the satisfiability of a linear property for a BNN is NP-complete via a reduction from the Boolean Satisfiability (SAT) problem. Second, we show that verifying robustness under non-uniform image occlusion is NP-complete through a reduction from SAT. Finally, we demonstrate that uniform occlusion induces a piecewise-constant structure in the network output, which enables the design of a polynomial-time algorithm for robustness verification.
arXiv:2604.12334v2 Announce Type: replace-cross
Abstract: We study additive mixtures of Markov kernels of the form $A_\alpha = \alpha P + (1-\alpha)G$, where $\alpha \in [0,1]$, $P$ is a baseline sampler and $G$ is a Gibbs kernel induced by a partition of the state space. We first motivate the study of $A_\alpha$, which can be interpreted as the projection of a lifted Markov chain. We then consider the minimisation of distance to stationarity under two objectives: the squared Frobenius norm and the Kullback-Leibler (KL) divergence. For the Frobenius objective, we derive explicit trace formulae and identify a Cheeger-type functional that characterises optimal two-block partitions. This yields a structured combinatorial optimisation problem admitting a difference-of-submodular decomposition, enabling efficient approximation via majorisation-minimisation. We also obtain geometric decay rates governed by the absolute spectral gap of $P$. For the KL divergence, we establish convexity-based bounds showing that the divergence of $A_\alpha$ is controlled by those of both $P$ and $G$, thereby reducing partition selection to the Gibbs component. Numerical experiments on the Curie-Weiss model demonstrate that suitable choice of both the partition and the parameter $\alpha$ can significantly accelerate convergence in total variation distance. We observe a consistent trade-off between local exploration and global averaging, with intermediate values of $\alpha$ achieving the best performance across regimes.
arXiv:2607.14853v1 Announce Type: new
Abstract: Collaborative control in complex environments is severely challenged by stochastic wireless delay and reliability variations, which can degrade navigation, tracking, and collision avoidance. These network-induced uncertainties complicate the maintenance of energy efficiency during collaborative tasks, and can potentially lead to over-provisioning of resources. In this paper, for a navigation setup with dynamic collision avoidance, we address this challenge by expanding the quality of control (QoC) framework from prior works to practical robotic models. Our approach (i) models end-to-end network effects on closed-loop performance, (ii) systematically explores the impact of various control parameters dictating robotic motion on network latency-reliability (iii) validates these models through experiments on a private 5G testbed across varying delay, reliability and control configurations. Our analysis indicates the optimal control-communication co-design operating regimes for practical robots and also compares the QoC performance of standard ROS~2 quality of service (QoS) policies under real-world conditions and showing how RELIABLE QoS offers 51.5% better QoC than BEST-EFFORT under certain experimental settings.
arXiv:2607.14123v1 Announce Type: new
Abstract: Despite the proliferation of Explainable AI (XAI) techniques -- from feature attributions to sparse autoencoders -- explanations rarely influence real-world workflows. In practice, they are often generated and discarded without guiding meaningful action. This gap reflects foundational shortcomings: research has not yet established methodologies for integrating explanations into end-to-end, human-in-the-loop systems. This position paper argues that the machine learning community must pivot from ad-hoc XAI methods toward addressing foundational & structural challenges, including unclear problem formulations, underspecified evaluation objectives, and the absence of pipelines for explanation-driven feedback. We support this claim through an analysis of recent ICML, NeurIPS, and ICLR papers and a survey of XAI practitioners, revealing recurring issues that limit cumulative progress. We conclude by outlining a practical checklist designed to shift XAI toward a more human-centered, action-oriented paradigm. By emphasizing foundational clarity over the development of ad-hoc methods, we hope to provide a roadmap for integrating explanations into actionable, feedback-driven AI systems.
arXiv:2504.14659v3 Announce Type: replace-cross
Abstract: Minimum mean square error (MMSE) estimation is widely used in signal processing, information theory, and related fields. Despite its practical robustness, the MMSE can be discontinuous under standard notions of stochastic convergence. To bridge this gap, we review classical counterexamples to the continuity of the MMSE and observe that they share a common pathology: along the approximating sequence, the observation is strictly more informative about the limit estimand than the limit observation is. Motivated by practical acquisition mechanisms, we study MMSE continuity under two natural constraints: (1) continuity of the second moment, and (2) a degradedness (Markov) restriction ensuring that each approximating observation is no more informative than the limit observation is about the limit estimand. Under these conditions, we establish continuity of the MMSE and of the MMSE estimator. We provide complementary semicontinuity results and continuity guarantees in related settings and establish continuity under linear estimation. We further extend the analysis to the families of Bregman divergences and continuous metric cost functions, including the Kullback-Leibler and Jensen-Shannon divergences as special cases.
arXiv:2604.23904v3 Announce Type: replace-cross
Abstract: Synthetic tabular data are often evaluated by distributional similarity, privacy distance, or train-on-synthetic-test-on-real predictive performance, but these criteria do not ensure validity for causal inference. We show that fully generative tabular synthesizers, including GAN- and LLM-based models, can preserve predictive utility while distorting average treatment effect (ATE) estimates. The failure is structural: ATE preservation requires both a realistic covariate law and an accurate treatment-effect contrast, whereas prediction loss penalizes treatment-effect error only through an overlap-weighted term. Thus, under imbalance or limited overlap, a generator may reproduce dominant observed outcomes while underlearning intervention-relevant contrasts. We formalize this mismatch through sensitivity and loss-decomposition results. Motivated by this causal analysis and intuition, we propose a hybrid synthetic-data framework for causal inference that generates covariates while modeling treatment and outcome mechanisms separately. We evaluate the framework in three settings: ATE preservation under fully generative versus hybrid synthesis, augmentation for practical positivity problems, and diagnostic simulation engines for comparing OR, IPW, AIPW, and TMLE before real-data analysis. We also stress-test the hybrid construction across settings that vary overlap, covariate dimension, seed sample size, and treatment-effect complexity, including a logistic outcome-model misspecification check. Across controlled simulation experiments, hybrid synthesis improves causal fidelity relative to fully generative baselines; the ACTG application shows improved predictive fidelity and potential for finite-sample estimator benchmarking. LLM-based hybrid synthesis is often more faithful than CTGAN in settings where causal fidelity can be assessed.
arXiv:2607.10847v2 Announce Type: replace
Abstract: Ehrenfest dynamics is a widely used mixed quantum--classical approach for nonadiabatic molecular dynamics, whereas thawed Gaussian wavepacket dynamics provides an efficient semiclassical description of adiabatic nuclear quantum dynamics. Here we describe thawed Gaussian Ehrenfest dynamics (TGED), which unifies and generalizes these two methods to capture both electronic nonadiabaticity and nuclear quantum effects within a single framework. The fully variational formulation of TGED is derived by applying the time-dependent variational principle to a Hartree product of electronic and Gaussian nuclear wavepackets. Replacing the effective locally quadratic molecular potential obtained from this variational treatment by alternative effective locally quadratic potentials yields an infinite family of TGED methods, of which we present several members. We analyze the limiting cases of the general formalism and show, in particular, that it reduces to conventional Ehrenfest dynamics in the classical limit for the nuclei and to thawed Gaussian wavepacket dynamics in the absence of electronic coupling. Finally, we present explicit geometric integrators for the entire family of methods and identify the conditions under which the different approximations become exact.
arXiv:2601.17325v4 Announce Type: replace-cross
Abstract: A hypergraph $H$ is said to be \emph{linear} if every pair of vertices lies in at most one hyperedge. Given a family $\mathcal{F}$ of $r$-uniform hypergraphs, an $r$-uniform hypergraph $H$ is said to be \emph{$\mathcal{F}$-free} if it contains no member of $\mathcal{F}$ as a subhypergraph. The \emph{linear Tur\'{a}n number} $ex_r^{\mathrm{lin}}(n,\mathcal{F})$ denotes the maximum number of hyperedges in an $\mathcal{F}$-free linear $r$-uniform hypergraph on $n$ vertices.
Gy\'arf\'as, Ruszink\'o, and S\'ark\"ozy~[\emph{Linear Tur\'an numbers of acyclic triple systems}, European J.\ Combin.\ (2022)] initiated the study of bounds on the linear Tur\'an number for acyclic $3$-uniform linear hypergraphs.
In this paper, we extend the study of linear Tur\'{a}n numbers for acyclic systems to higher uniformity. We first give a construction for linear $r$-uniform trees with $k$ edges that yields the lower bound $
ex_r^{\mathrm{lin}}(n,T_k^r)\ge {n(k-1)}/{r}, $
under mild divisibility and existence assumptions. Next, we study hypertrees with four edges. We prove the exact bound $
ex_r^{\mathrm{lin}}(n,B_4^r)\le {(r+1)n}/{r} $
and characterize the extremal hypergraph class, where $B_4^r$ is formed from $S_3^r$ by appending a hyperedge incident to a degree-one vertex. We also prove the bound $
ex_r^{\mathrm{lin}}(n,E_4^r)\le {(2r-1)n}/{r} $
for the crown $E_4^r$. Finally, we give a construction showing $
ex_r^{\mathrm{lin}}(n,P_4^r)\ge {(r+1)n}/{r} $
under suitable assumptions and conclude with a conjecture on sharp upper bound for $P_4^r$.
arXiv:2607.14127v1 Announce Type: new
Abstract: Representative clutter height (RCH) is a key parameter in radio propagation and interference analysis because it captures the dominant height of local obstructions that drive terminal clutter loss. Current practice often relies on fixed clutter heights assigned to land use classes in Recommendation ITU-R P.452-18, but this misses within class variation and can lead to conservative exclusion zones and poor site ranking for low Earth orbit ground station siting and spectrum coordination. We present an interpretable, globally deployable machine learning framework for predicting RCH from open geospatial data. The model is trained using LiDAR derived labels from the U.S. Geological Survey 3D Elevation Program and inference time features from global land-cover, terrain, demographic, thermal, and optical remote sensing products. We define RCH using a robust 75th percentile clutter height statistic, evaluate multiple regressors, and select LightGBM for its accuracy, efficiency, and compatibility with feature attribution analysis. The final model achieves a mean absolute error of 1.79m and an R^2=0.765, reducing absolute error by more than 60% relative to the ITU baseline. Beyond aggregate fit, we evaluate domain facing criteria relevant to RF planning, including meter scale error, tolerance band accuracy, over and under estimation tails, agreement with ITU clutter height regimes, and SHAP-based physical plausibility. SHAP identifies tree canopy cover, land-cover semantics, and spectral reflectance as the most influential predictors. Studies on segmentation derived features, non-forest ablations, and land-cover matched international validation show that open geospatial data can improve clutter modeling at scale without sacrificing interpretability or deployability.
arXiv:2607.14548v1 Announce Type: new
Abstract: As large multimodal models move from understanding content to operating on digital environments, mobile GUI has emerged as a challenging and consequential testbed for digital embodied intelligence. Mobile agents operate under three coupled constraints: precise perception of complex interfaces, scalable acquisition of high-quality interaction data, and robust long-horizon decision making under compounding execution errors. This report presents HyMobileAgent, a mobile GUI agent built on Hy3.0-VL-A3B, a vision-native foundation model featuring native any-resolution input, an A3B-scale deployment budget, and a 32K context window to model extended interaction histories. Rather than relying solely on model scaling, we develop a joint data and environment centric scaling framework to address the key bottlenecks of mobile interaction.
Our framework integrates a GUI perception flywheel combining mock-interface synthesis, rejection sampling, and icon-specific augmentation; a knowledge pipeline that transforms tutorial videos into structured interaction data; a million-scale action data pipeline deployed across more than 2000 sandbox and real-device instances with automated failure attribution; the PhoneWorld Mock App Factory, providing a resettable training environment with 34 mock applications and over 34000 tasks; and a structured Planning-and-Reflection mechanism with explicit dead-loop detection for reliable long-horizon execution.
We also introduce a progressive training recipe consisting of mid-training, supervised fine-tuning, and reinforcement learning with task-specific reward designs.
arXiv:2607.14530v1 Announce Type: new
Abstract: Hyper-Connections (HC) expand the residual stream of Transformers into $N$ parallel streams, providing a form of memory scaling beyond model width and depth. Manifold-Constrained HC (mHC) stabilizes this formulation at scale. The large gains from $N{=}1$ to $N{=}4$ suggest residual-stream expansion as a promising scaling axis. However, existing HC-family methods typically stop at $N{=}4$. Our experiments reveal why: scaling mHC beyond this point yields diminishing performance gains and rapidly increasing training cost. We attribute this limitation to two bottlenecks: insufficient write-back information for an expanding number of streams and residual-mixing generation whose cost scales cubically with $N$. To address both bottlenecks, we propose xHC (Expanded Hyper-Connections), the first HC-family method to achieve meaningful expansion beyond $N{=}4$. xHC combines temporal feature augmentation for richer write-back with a sparse residual-stream architecture that updates only $k=4$ of the $N=16$ streams while retaining dense access to the full residual state. Across 18B and 28B MoE models, xHC delivers strong and consistent downstream improvements. On an 18B MoE model, xHC improves the average downstream score by 4.0 points over mHC, while adding only modest training FLOPs over the vanilla baseline. Scaling-law experiments show that the vanilla and mHC require $1.50\times$ and $1.19\times$ the compute of xHC, respectively, to reach the same loss. Practical large-$N$ training also requires controlling memory traffic from the expanded residual state. We therefore introduce xHC-Flash, which reduces the per-sublayer memory traffic from $73.5C$ to $40C$, comparable to the $34C$ required by mHC at $N{=}4$, while retaining the gains of full xHC. Together, xHC and xHC-Flash make large-$N$ residual-stream expansion effective and practical for LLM pre-training.
arXiv:2607.14147v1 Announce Type: new
Abstract: Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal. We ask where and how it fails. The harm representation stays intact: on the prompts the attack flips to compliance, a linear probe reads harm as high as on the refused ones (0.91-0.98), while behavioral refusal drops to chance. This holds across four models and three families (1.5-3.8B, and at 14B). Refusal is therefore a shallow, response-site computation. We localize it to an early window: a dose-matched position control shows the first half of the response suffices to break refusal, while the second half is nearly inert. Three causal probes converge on that window. Restoring the harm direction there partially re-engages refusal. Injecting the model's own refuse-state reverses the jailbreak (74%, held-out). And knocking out the early response's attention to the prefill, but not an equal attention mass elsewhere, selectively collapses the harmful continuation. A base-model control identifies the mechanism: the same knockout collapses the continuation prefill-specifically even in a non-safety-tuned base model (64% to 25% harmful content vs a matched control's 64%, replicated at 7B). So the prefill's grip is generic autoregressive conditioning, not safety-specific suppression, and "refusal restoration" is a model-dependent fallback. The dominant mechanism is passive. A small safety-specific attractor remains on top (logit-trace concentration 0.24 vs 0.03), whose active-vs-passive character we size but do not fully separate. No single direction or component is a clean handle either: the decision is decodable but distributed, and refusal tracks harm rather than scary surface. The consequence is structural: a monitor reading the untouched prompt-side representation is immune by construction, but only to response-site attacks. The mechanism is diffuse; the failure surface is local.
arXiv:2607.14532v1 Announce Type: new
Abstract: Recent work shows that it is possible to extract verbatim or near-verbatim text of some copyrighted works from some large language models (LLMs or models). That is evidence that the model weights encode the works in some form - that the model has "memorized" those works from its training data. But LLMs don't store information in the same format as familiar databases. Rather, their weights store statistical relationships between tokens that have been learned from the training data, and those relationships inform a generation process that is often probabilistic rather than deterministic. In the case of memorization, those relationships are strong enough that, in many circumstances, the model might generate a copyrighted work from its training data with some probability.
Copyright law has not previously had to decide whether storing information that might or might not produce output similar to a copyrighted work is itself a copy of the work. The answer to the question is important, because it may determine the legality of many LLMs. The statute and case law are largely unhelpful. We argue that copyright law will likely take a functional approach to the question, finding that LLMs contain a copy of a particular work only if it is straightforward to extract that work in outputs. That result is unsatisfying as a policy matter, and we suggest potential changes to the law, but it is the most likely outcome under current law.
arXiv:2607.14870v1 Announce Type: new
Abstract: In this work, we introduce the One-for-All Adaptive Radiotherapy Planning Agent, a unified foundation-model-based system that performs complete, treatment-specific online adaptive planning directly from daily cone-beam CT in under two minutes. The agent first autonomously predicts all essential planning components, including synthetic CT generation, multimodal alignment, and tumor/organ segmentation. It then intelligently leverages these outputs to execute the final clinical plan design, providing a comprehensive, automated solution for daily treatment. We also demonstrate that the agent enables clinicians to define planning with intent and intervene at critical decision points, ensuring a "human-in-the-loop" framework that generates acceptable plans before final approval. Evaluated on multiple datasets spanning head-and-neck, lung, abdominal, and prostate cancers with both photon and proton therapy, the proposed framework achieves clinically acceptable accuracy and plan quality comparable to clinically generated treatment plans, with target dose errors (D98) generally within 2.0 Gy of the reference plan. The strong performance of the One-for-All agent highlights the promise of a unified foundation-model approach and opens opportunities for fast, scalable, and fully automated online adaptive radiotherapy across diverse clinical scenarios.
arXiv:2607.14876v1 Announce Type: new
Abstract: With the proliferation of immersive Head-Mounted Displays (HMDs) for Virtual and Augmented Reality (VR/AR), reliable and high-precision eye tracking has become increasingly important. Conventional 2D image-based methods offer low system complexity but remain limited in stability, accuracy, and robustness. Three-dimensional ocular surface reconstruction can provide richer geomet-ric information, and structured light profilometry is particularly attractive because it enables dense and accurate surface measurement. However, Phase-Shifting Profilometry (PSP), which estimates phase from sequentially acquired fringe images, is highly susceptible to motion-induced errors when the eye rotates between frames. This study proposes a rotational motion compensation framework for PSP-based dynamic 3D eye reconstruction. Relative eye rotation is estimated from image-based motion cues using a user-specific 3D eye model in a spherical-coordinate domain. The estimated motion is then used to compensate for camera-pixel mismatch and phase-shift errors caused by inter-frame rotation. A region-wise optimization strategy is further introduced to reduce residual artifacts by inde-pendently refining the compensation strength in different ocular regions. Experiments with a rotating fake eye under non-uniform motion demonstrate that the proposed method substantially suppresses motion-induced deformation and improves reconstruction accuracy. An additional experiment with a non-spherical rigid object indicates that the compensation principle is not restricted to spherical eye geometry. These results establish a practical basis for stable PSP-based dynamic 3D eye reconstruction toward future high-precision eye tracking in immersive environments.
arXiv:2607.14424v1 Announce Type: new
Abstract: In recent years Flow Matching has become a prominent method for generative modeling robot motion generation. In its generic form Flow Matching is an ODE-based neural sampler that is trained by regressing empirical flow fields associated with motion samples as data. However, in robot motion generation we often have additional constraints that might not be present in the collected data. The majority of current approaches train the flow on the available data and use inference-time guidance to enforce task-specific constraints. To address this mismatch, we propose \textbf{ConFlow}, a constraint-guided flow matching framework that incorporates constraint information directly into the training objective via differentiable barrier or cost functions. To address design specifications such as smoothness and boundary conditions, we propose replacing the standard Gaussian source distribution used in flow matching training with a conditional Gaussian Process. Our approach also uses infeasible demonstrations as negative supervision, improving constraint satisfaction without requiring additional expert data. Experiments on a two-robot navigation task demonstrate that ConFlow achieves lower collision rates and higher trajectory quality than standard flow matching baselines, with or without inference-time guidance. These results validate training-time constraint integration as an effective approach to closing the training--inference gap in generative motion models.