arXiv:2607.17898v1 Announce Type: new
Abstract: Motion planning for rigid-body assembly poses a fundamental challenge in robotics due to tight geometric constraints. In such scenarios, feasible motions often require passing through (near-)zero clearance configurations in which the parts are tightly constrained by contact. In this work, we introduce Critical-Manifold Guided RRT (CMG-RRT), a sampling-based planner designed specifically for tight assembly problems. Our key observation is that in tight assemblies, valid solution paths lie on or near a critical manifold: the subset of configuration space consisting of poses with at least one contact point between parts. CMG-RRT guides exploration by adaptively biasing sampling toward neighborhoods of the critical manifold using a hierarchical subdivision of the configuration space. We prove that CMG-RRT is probabilistically complete under standard clearance assumptions. Empirical evaluation on challenging rotational assembly benchmarks demonstrates a 100% success rate across all tested instances, including, to the best of our knowledge, the first fully automatic solution of the Elk disentanglement puzzle. Our open source software is available through our project page: https://www.cgl.cs.tau.ac.il/projects/tight-assembly-planning.
Science Journals
arXiv:2607.16936v1 Announce Type: new
Abstract: Pediatric bone age prediction is a crucial task in clinical practice that can help diagnose endocrine disorders and provide insight into a child's growth and development. However, conventional bone age prediction methods are often labor-intensive and require specialized radiological expertise. This paper presents a Deep Learning (DL)-based approach to pediatric bone age prediction using EfficientNet with Additive Attention, a state-of-the-art neural network architecture for image classification and regression tasks. The method utilizes over 12,000 X-ray images from the RSNA bone age dataset. It involves image preprocessing, transforming them into three-channel images, and training a Convolutional Neural Network (CNN) to automatically learn the features of hand bone images. This approach provides a more effective and accurate solution for predicting bone age, which is critical in diagnosing pediatric endocrine diseases. This work uses two variations of the EfficientNet model (B0 and B4), where EfficientNetB4 is also finetuned with the Additive Attention mechanism. These three models predict the age for the original age, and their comparison is shown in curves. The predicted ages depict that in most cases, EfficientNetB4 and EfficientNetB4 with Additive Attention (EN-AA) successfully predicted the bone ages more accurately regarding the original age, and their performance was better than the EfficientNetB0. Specific performance metrics are provided to underscore this improvement. Learning curves for training and validation loss confirm effective learning without overfitting or underfitting, further validating our approach's efficacy in pediatric endocrine disease diagnosis.
arXiv:2607.17913v1 Announce Type: new
Abstract: Distributed Fine-Tuning (DFT) of large-scale Foundation Models (FMs) on resource-constrained edge devices is limited by local compute constraints and communication overhead. Parallel Split Learning (PSL) reduces client-side computation by keeping few model layers on each client and offloading the remaining computation to the server; however, clients must exchange intermediate activations and gradients with the server at every training step. Existing SL communication-compression methods mainly rely on task-agnostic heuristics, such as sparsification and quantization. While learnable SL compressors can better adapt to intermediate representations, they require co-training with the target model. Therefore, directly inserting them into off-the-shelf FMs introduces feature-distribution misalignment and degrades DFT performance. To address this, we propose AE-PSL, a communication-efficient PSL framework that compresses intermediate activations and gradients using a lightweight AutoEncoder (AE) placed at the split layer. To ensure compatibility of AE compression with pre-trained FMs, AE-PSL introduces a novel two-stage alignment mechanism, which adapts the AE to the pre-trained model's feature manifold and client-specific feature distributions before DFT.
arXiv:2607.16517v1 Announce Type: cross
Abstract: Neural activity is widely held to organize on low-dimensional structure embedded in a high-dimensional state space. Persistent homology reads such structure directly from the pattern of pairwise correlations, without assuming in advance which variables are relevant. We apply persistent homology to microelectrode-array (MEA) recordings of spontaneous activity from human (Lancaster) and mouse (Pa\c{s}ca) cortical organoids, spanning $26$--$234$ simultaneously sorted units, and ask whether topological data analysis resolves structure at the node counts that neural recordings actually deliver. Building weighted networks in correlation space and characterizing them by Vietoris--Rips filtration, we find that the first homology ($H_1$, loops) rises significantly above a rate- and population-preserving null in $14$ of $18$ datasets. This loop structure occupies a non-redundant core: it is robust to random removal of units yet disrupted by targeted removal of the units that carry it. Topological richness grows with network size, and second homology ($H_2$) emerges significantly above the null only in the larger networks. These results show that persistent homology resolves structured topology in neural recordings at the scale experiments actually deliver.
arXiv:2607.16828v1 Announce Type: new
Abstract: Despite the impressive generative capabilities of text-to-image diffusion models, they remain vulnerable to implicit sexual prompts, where subtle cues disguised as benign terms or adversarial tokens unexpectedly generate the inappropriate content due to model biases or latent correlations in training data. Existing safety mechanisms face fundamental limitations: detection methods primarily identify explicit content and fail to capture implicit malicious intent, while mitigation approaches rely on static negative prompts inadequate for diverse implicit scenarios. To address these challenges, we propose UniNDM, a unified noise-driven framework that rethinks safety mechanisms through the lens of noise dynamics in diffusion processes. Our key insight is that early-stage predicted noise exhibits inherent separability between normal and sexually explicit content, which we theoretically demonstrates quadratically increasing semantic concentration with timestep. Leveraging this property, we develop a lightweight noise-based detector achieving superior accuracy with virtually no computational overhead. For mitigation, we introduce noise-enhanced adaptive negative guidance: dynamically generating context-specific negative prompts via large language models to handle diverse implicit content, while optimizing initial noise by suppressing attention concentration on explicit tokens to provide comprehensive protection. Besides the U-Net-based diffusion models, we further extend our framework to emerging Diffusion Transformer architectures through region-constrained semantic guidance tailored for their unified multimodal attention. Comprehensive experiments across U-Net models and DiT models on both natural and adversarial datasets demonstrate substantial improvements over state-of-the-art methods, including SLD, UCE, Safree, etc. Our code is publicly available at https://github.com/Aries-iai/UniNDM.
arXiv:2607.16529v1 Announce Type: cross
Abstract: We introduce a positive-allocation companion construction for Koopman-inspired finite-dimensional prediction of nonlinear dynamical systems. The method determines recurrence coefficients by representing a target observable snapshot as a nonnegative, normalized combination of earlier training snapshots. These coefficients define a companion matrix whose spectral structure is induced by the allocation constraints at the construction stage. We prove that normalized positive allocation places the companion spectrum in the closed unit disk and, because the coefficients sum to one, includes $1$ as an eigenvalue. Additionally, we develop modal and non-modal diagnostics for the resulting model trajectory. When the companion matrix is diagonalizable, the modal representation shows that first and second finite differences act as spectral filters through factors of $\lambda_\ell-1$ and $(\lambda_\ell-1)^2$. We also derive $C$-based finite-difference bounds that avoid diagonalization and can be evaluated directly from the training data and companion matrix. Numerical experiments on the FitzHugh--Nagumo and susceptible--infectious--recovered (SIR) models illustrate the behavior of the construction in oscillatory and transient dissipative settings. The examples demonstrate both the interpretability of the companion recurrence and its limitations, particularly when pointwise trajectory agreement degrades while finite-difference and modal diagnostics remain informative.
arXiv:2607.16541v1 Announce Type: cross
Abstract: Optimal Polynomial Intersection (OPI) is a structured optimization problem for which Decoded Quantum Interferometry (DQI) attains a satisfaction guarantee governed by the semicircle law. Sun and Wootters recently showed that, for balanced OPI over prime fields, a Fourier-defined distribution $P_u$ gives a strict worst-case improvement from limiting rate $0.6225$ onward and asymptotically perfect solutions from rate $0.7496$ onward, and asked whether $P_u$ can be sampled efficiently. We answer this question for Reed--Solomon OPI parameters satisfying their exponent condition strictly below the dual Johnson radius. Under coherent membership-oracle access, we give a bounded-error polynomial-time quantum sampler for $P_u$. The ideal circuit samples $P_u$ exactly conditioned on success, while a finite-precision implementation achieves any prescribed inverse-polynomial total-variation error. Consequently, every fixed limiting rate $0.6225\le r<1$ admits a strict worst-case improvement over the DQI semicircle value, and every limiting rate $r\ge 3/4$ admits solutions of satisfaction $1-o(1)$ with high probability. The algorithm coherently sums the amplitudes of all low-weight errors in each syndrome class using deterministic complete list decoding. Complete Reed--Solomon list decoding and the Sun--Wootters denominator estimate make the list size and postselection overhead polynomial. In concurrent and independent work, Horinaga and Yamakawa obtain worst-case OPI algorithms over prime-power fields and exact satisfaction at every fixed rate strictly above $3/4$.
arXiv:2607.17051v1 Announce Type: new
Abstract: The convergence of a finite element discretization for the two-dimensional Ricci flow is proved. In this method, the Ricci flow on a two-dimensional surface is formulated into solution-driven metric evolution, with the metric evolution driven by the Gauss curvature. The Gauss curvature satisfies a parabolic equation that in turn depends on the metric, thereby enhancing the parabolic structure of the problem. The solution-driven metric evolution formulation is discretized by the finite element method, and the convergence of finite element approximations is proved by adapting the matrix-vector formulation developed in the literature initially for studying solution-driven surface evolution in extrinsic curvature flow. In addition to its convergence, the proposed method also preserves important geometric structures of the Ricci flow at the discrete level, such as area conservation and the Gauss-Bonnet theorem. Extensive numerical experiments are presented to demonstrate the convergence of the proposed method as well as the simulation of Ricci flow.
arXiv:2607.17417v1 Announce Type: new
Abstract: Large language models confabulate chemical objects (molecular formulas, space groups, formation energies) in fluent reasoning traces, concentrated on long-tail entities where confidence is least trustworthy. Deterministic, database-grounded verification can catch and repair such errors without the coverage cost of blanket retrieval; the binding constraint, we find, is detection, not repair. Our tiered verifier extracts each checkable claim, checks it against authoritative databases and physics, and feeds the reference into a gated correction loop. Across four models and 528 condition-pinned prompts, gated correction cuts committed-formula error from 22% to 4% at $3.2\times$ fewer retrievals than blanket augmentation, beating a conversational oracle. Repair succeeds wherever a flag fires (80--97%); the bottleneck is in-loop detection recall. Grounding improves the final answer only when the verifier's scope reaches the deliverable (83% to 90%), and the lift appears only where extractable long-tail error exists: absent on near-ceiling physical constants, large on isotope half-lives (11% to 0%).
arXiv:2607.16570v1 Announce Type: cross
Abstract: Organic mixed ionic electronic conductors (OMIECs) are a promising class of polymer materials for applications spanning neuromorphic computation to energy efficient electronics and bioelectronics. Despite being highly tunable, the relationship between structural features and key performance properties such as charge carrier mobility is poorly understood. Scanning nanodiffraction in the transmission electron microscope (TEM) is a powerful probe for elucidating this structure-property relationship, but produces large, noisy datasets that are difficult to interpret because polymer reflections exhibit several distinct morphologies. To address the complexity, we trained a machine learning (ML) model to detect these polymer diffraction peaks and their intensities from synthetic data. Compared to correlative peak detection algorithms, the conventional method for analyzing nanobeam 4D scanning transmission electron microscopy (4DSTEM) data, we show that the ML model is significantly faster and outperforms correlative algorithms in almost all cases, opening up the possibility of near-live visualization of 4DSTEM experiments.
arXiv:2607.17937v1 Announce Type: new
Abstract: Agent Skills package procedural instructions and checks for use by general-purpose agents, but loading a skill does not guarantee that every requirement remains active throughout a long tool-using trajectory. We study this problem in a production-derived, white-box code-audit workflow. Holding the task and 24 artifact checks fixed, we vary the surrounding context and classify where failures first become visible: lost requirements, editing drift, failed checking, or non-agent evaluator/runtime failures. Codex with gpt-5.4-mini passes 8/10 runs in a 10,991-character clean context but only 3/10 in both a 299,140-character relevant context and an equal-length irrelevant context. This 50-percentage-point difference is large but remains trend-level under two-sided Fisher tests (p = 0.0698). Requirement coverage nevertheless stays above 92% in both long conditions, showing that a few omissions can invalidate an otherwise complete artifact. A second task passes all clean and long runs, so the evidence does not support a universal context-length threshold. A detailed external checklist passes 10/10 runs, compared with 5/10 for a generic self-check (p = 0.0325). Coding-agent scaffolds may help by selecting a smaller working set, but they do not eliminate failures. We do not introduce context rot or a new general monitoring method; we provide a bounded failure classification and empirical case study for white-box code auditing.
arXiv:2607.17057v1 Announce Type: new
Abstract: This study investigates the coupled effects of an insoluble surfactant and a radial electric field on the stability of a viscous liquid film flowing down a vertical fiber. Starting from the governing equations in two dimensions, a reduced model in one dimension is derived using the long wave approximation to describe the coupled evolution of the interface and surfactant transport. Linear stability analysis identifies two distinct unstable modes: the Rayleigh-Plateau mode, which dominates at lower values of the Marangoni number $Ma$, and the Marangoni mode, which becomes dominant at higher values of $Ma$. The influence of the radial electric field is determined by the position of the outer electrode $\beta$. When $\beta<\mathrm{e}$, the electric field enhances both instabilities and narrows the stable interval in $Ma$ between the two modes. When $\beta>\mathrm{e}$, the electric field suppresses both modes and can completely eliminate the unstable region associated with the Marangoni mode even at a relatively small electric Weber number $E_b$. Continuation of the traveling wave solutions further shows that, when $\beta<\mathrm{e}$, the magnitude of the relative interfacial motion $I_{RP}$, generally increases with $E_b$. By contrast, the intensity of Marangoni convection $I_{M}$ varies only weakly at smaller values of $E_b$ and increases appreciably only when the electric field becomes sufficiently strong. Analysis of the stream function and the relative interfacial velocity reveals that the electric stress primarily intensifies the recirculation beneath the wave crest and reshapes the spatial distribution of the relative interfacial velocity.
arXiv:2607.16580v1 Announce Type: cross
Abstract: Structural reliability analysis supports the design and safety assessment of buildings and civil infrastructure but requires specialized expertise throughout the workflow. This study presents a multi-agent large language model framework that automates component-level reliability analysis from a natural-language problem statement to interpreted estimates of the reliability index and failure probability. Specialized agents handle problem formulation, method planning, code generation, execution, and result interpretation, with human confirmation at key decision points. The Method Planner is fine-tuned using QLoRA for a priori reliability-method category selection. Analysis results are not generated directly by an LLM; instead, validated deterministic solvers compute the reliability estimates, improving reproducibility and reducing hallucination risk. The framework uses open-weight models and supports local execution without closed APIs. Results show that it lowers the expertise barrier to structural reliability assessment while preserving computational trustworthiness.
arXiv:2607.17651v1 Announce Type: new
Abstract: Flow policies can represent multimodal action distributions for robot manipulation, yet a robot must execute one action at each control step. When several proposals are sampled, critic-based ranking makes data collection depend on value estimates over candidate actions that may be weakly represented in replay. We introduce HCPG-Flow, an analytic rollout-time selector that augments SAC-Flow with hierarchical, object-centric contact-progress guidance while preserving its actor and critic objectives. HCPG switches from end-effector approach to task progress after contact, scores each proposal by the first-order reduction of a task-relevant distance, standardizes scores within the candidate set, and executes a temperature-controlled action embedding. Across ten simulated tasks, HCPG improves mean success over SAC-Flow on both benchmarks, including a 9.5 percentage-point gain on Maniskill. Four physical tasks further show high success with a 17.4% reduction in successful completion steps.Project page: https://hitxraz.github.io/HCPG-Flow/
arXiv:2607.18237v1 Announce Type: new
Abstract: Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects of similarity. Benchmarking a broad range of frontier vision-language models (VLMs) reveals a considerable performance gap compared to human annotators' consensus. Leveraging our data, we fine-tune a VLM to produce our Text-Prompted Image Perceptual Similarity (TPIPS) metric, capturing multiple senses of visual similarity depending on the specified text prompt. We demonstrate that TPIPS aligns more closely with human perception and generalizes reliably beyond the training distribution. Finally, we show that TPIPS unlocks new capabilities in text-guided retrieval, compositional search, and the fine-grained evaluation of generative models. Our code, data, and trained models are at https://peterwang512.github.io/TPIPS
arXiv:2607.16606v1 Announce Type: cross
Abstract: Chirality in a discrete $Z_q$ spin interaction distinguishes clockwise from counterclockwise phase differences, but its manifestation in continuous nonlinear dynamics is unclear. We show that any pairwise $Z_q$ Hamiltonian admits a unique equilibrium-preserving embedding into a continuous phase-energy landscape that matches the discrete energy on the $q$-state phase grid, where every grid point is stationary. This embedding reveals that a $Z_q$ kernel is nonchiral if and only if the sine components of the relaxation vanish. Chirality of the discrete $Z_q$ model is therefore exactly the odd part of the continuous phase interaction. In the induced nonlinear phase dynamics, this odd part becomes an orientation-dependent phase shift in the multi-harmonic coupling, and chiral reversal flips this shift while preserving the coupling magnitudes. In self-sustaining oscillator networks, the shift is further realized as a direction-dependent delay. Transistor-level ring-oscillator simulations validate the predicted phase locking and reversal of directed phase bias. These results show that algebraic handedness in a discrete spin Hamiltonian can be represented as tunable time-domain asymmetry in continuous nonlinear dynamics.
arXiv:2607.17946v1 Announce Type: new
Abstract: Large Language Models (LLMs) often struggle to navigate value conflicts when trained with the compressed scalar rewards of Reinforcement Learning from Human Feedback (RLHF). To address this challenge, we investigate how chain-of-thought (CoT) reasoning can help improve performance in this domain. Geometrically, we show that CoT correlates with further smoothing the model's loss landscape in its sharpest direction, helping resolve the optimization instability of traditional scalar rewards. We also demonstrate via relevant downstream benchmarks that value conflict-focused CoT may generalize to different kinds of moral reasoning, demonstrating that this CoT has the potential to be an effective mechanism for better moral reasoning. To capitalize on this potential, we create a new value conflict-focused CoT design that further smooths the sharpest direction of the loss landscape and increases moral reasoning performance. This finding shows that explicitly modifying and improving the design of reasoning dynamics offers a promising avenue for improving model performance on user requests with complex value conflicts, advancing pluralistic alignment in LLMs.
arXiv:2607.18003v1 Announce Type: new
Abstract: Susceptible-infected-susceptible (SIS) epidemic models on networks are governed by hierarchical moment equations where the dynamics of smaller subsystems depend on the state of larger ones. Moment closure approximations, which truncate this hierarchy by expressing higher-order state probabilities in terms of lower-order ones, are essential for obtaining tractable reduced systems. Higher-order networks, which extend the pairwise structure to include group interactions, introduce a combinatorial explosion of closure configurations, making systematic derivation harder. Consequently, existing higher-order SIS models are derived heuristically, where structural and dynamical assumptions underpinning their closures are not always apparent from the formulation alone. We develop a bottom-up derivation of higher-order SIS dynamics, building systematically from node-level equations to pairs, triplets, and three-body interactions. Central to our approach is a network-dependent closure operator that generates topologically appropriate approximations from local pairwise and triadic structure. Using this framework, we recover three existing higher-order SIS models--Burgio et al.'s maximal clique, Malizia et al.'s pair-based and inter-order models--as special cases, each arising under specific topological and dynamical assumptions. Our derivation reveals assumptions that are invisible from heuristic approaches: for instance, Malizia et al.'s inter-order overlap parameter is insufficient alone to express the model within our framework despite performing well against simulations, with the original derivation implicitly invoking additional structural assumptions. Our framework offers both a foundation for higher-order epidemic modeling and a constructive pathway for understanding the assumptions implicit in heuristically derived mean-field closures and provides a principled method of generating new models.
arXiv:2607.17621v1 Announce Type: new
Abstract: Existing self-evolving memory systems mainly improve agent memory based on textual outputs, such as task trajectories and reflections. However, this text-based paradigm rarely incorporates internal mechanistic signals, leaving how retrieved memory is actually utilized during task execution underexplored. This limitation can lead to unreliable error attribution and hallucinated memory modifications. In this work, we show that retrieval-head attention provides a mechanistic signal for revealing segment-level memory utilization. By aggregating attention over memory segments and decision steps, we construct a context utilization matrix that exposes recurring memory-use patterns and indicates corresponding refinement strategies. Building on this observation, we propose Attention-Guided Memory Refinement (AGMR), a framework that uses utilization patterns revealed by attention to guide targeted segment-level memory updates. AGMR corrects or enhances memory for failed executions, simplifies memory for successful executions, and verifies each update through re-execution. Experiments on interactive decision-making benchmarks show that AGMR improves both task performance and memory efficiency over text-only memory refinement baselines. Code is available at https://anonymous.4open.science/r/AGMR_code-3262/
arXiv:2607.17458v1 Announce Type: new
Abstract: In conventional optical microscopy, accessing subwavelength information is classically constrained by the Abbe diffraction limit which remains a significant challenge in nanophotonics. To overcome this limitation, we introduce a scheme that generates a tightly confined near-field focal spot under non-diffracting Bessel beam illumination, employing a series of high refractive-index nanoscale structures. Furthermore, the utilization of higher-order Bessel beams enables additional resolution enhancement beyond the diffraction limit. Our results demonstrate a subwavelength resolution of (lambda/53) with a pyramidal nanostructure. To elucidate the factors governing the focusing behavior, we systematically investigate the influence of different geometrical parameters related to the incident Bessel beam on the focusing performance and shows that these factors critically determine the stability and spatial profile of the focal spot. These findings highlight the potential of the proposed approach for various applications in super-resolution imaging and quantum technology.
arXiv:2607.16890v1 Announce Type: new
Abstract: Spread codes are a well-known family of constant-dimension subspace-metric codes. For constant dimension $k$ and ambient space dimension $n$ being a multiple of $k$, these codes have minimum distance $2k$ and a rich geometric structure. In this paper, we study the decoding capabilities of the Nearest Neighbor Decoder for Desarguesian spread codes, establishing that unique decoding is still achievable beyond half the minimum distance. Motivated by this, we develop a new decoding algorithm to uniquely decode Desarguesian spread codes in the presence of both insertions and deletions, which increase and decrease, respectively, the dimension of the transmitted codeword. Even when the sum of the dimensions of insertions and deletions exceeds half the minimum distance, provided that deletions are of dimension at most $k-2$, the algorithm succeeds with a small decoding failure. We also propose two refinements to this algorithm that, empirically, can handle nearly as many insertions as the Nearest Neighbor Decoder.
arXiv:2607.17059v1 Announce Type: new
Abstract: Laser-plasma accelerators have been the subject of extensive research in recent years. The electron beams they generate exhibit a broad energy spread. To conveniently characterize beams from laser wakefield acceleration (LWFA), electron spectrometers employing scintillating screens coupled with CCD cameras are typically used. In this work, we calibrate a series of DRZ phosphor screens and measure the spectra of the light they emit. The calibration was performed using the radio-frequency linear electron accelerator at Tsinghua University, which provided monoenergetic electron beams with peak energy of approximately 30 MeV.
arXiv:2607.16410v1 Announce Type: cross
Abstract: The EWF-(FCI,SQD) method, a wave-function-based embedding approach combining full configuration interaction (FCI) and sample-based quantum diagonalization (SQD), is a promising new tool for the simulation of molecular systems. However, applications of EWF-(FCI,SQD) have so far been limited to single-point calculations, whereas the study of complex chemical processes requires the ability to explore potential energy surfaces. In this work, we demonstrate geometry optimization with EWF-(FCI,SQD), scaling our simulations to molecules as large as menthone and benzidine within the STO-3G basis set. Without fragmentation, these systems comprise 73 and 82 molecular orbitals respectively, presenting an intractable Hilbert space for conventional exact or high-level subspace solvers and establishing a clear necessity for fragmentation-based methodologies. The underlying fragment SQD simulations in the EWF-(FCI,SQD) geometry optimizations use up to 70 qubits. The resulting geometries show exceptional accuracy relative to the classical reference, with deviations below 4 picometers.
arXiv:2607.17062v1 Announce Type: new
Abstract: In this paper, we introduce the General Lotto game with a regulator (R-Lotto), a leader-follower extension of the classical two-player General Lotto game. The model captures regulatory interventions in competitive resource allocation, where a regulator first chooses an intervention parameter to influence the subsequent competition between two resource-constrained followers. The intervention parameter represents favoritism toward one of the followers, and the followers then play a general Lotto subgame with favoritism. We derive the followers' equilibrium payoff and characterize the Nash--Stackelberg equilibrium (NSE) intervention of the regulator. We further develop a multi-battlefield R-Lotto model with a regulator budget constraint. In this setting, the follower subgames on different battlefield is decoupled, while the regulator's intervention decisions are coupled through a common budget. Numerical simulations demonstrate the proposed equilibrium characterizations and provide practical decision-making guidance for the regulator.
arXiv:2607.16612v1 Announce Type: cross
Abstract: Backpropagation makes training deep networks memory intensive because it must store intermediate activations. Forward-mode methods avoid this cost, but their gradient estimates become increasingly noisy as the number of trained parameters grows. We introduce Split Forward Gradient (Split-FG), which splits a network at an intermediate representation: it computes the output head gradient exactly and estimates only the trunk gradient with a Jacobian--vector product. This reduces estimator variance and requires no backward pass through the trunk, while retaining an Adam-style convergence guarantee. Our experiments reveal an important practical failure mode. On WikiText-103, naive forward-gradient training of the trunk performs worse than leaving a randomly initialized trunk frozen, likely because Adam updates every noisy, under-determined trunk coordinate too aggressively. Simply using a much smaller learning rate for the trunk reverses this result: a $16$M-parameter GPT-2-style model reaches validation perplexity $387$, compared with $668$ for the frozen-trunk control and $2{,}885$ for a matched pure forward-gradient baseline (backpropagation reaches $150$). Split-FG also produces the strongest backprop-free results on our tabular benchmarks and reaches $60.5\%$ on CIFAR-10 and $35.2\%$ on CIFAR-100 with a heavy-head design. It reduces peak memory by up to $35\%$ relative to matched backpropagation, although the performance gap widens as the forward-mode trunk grows.