arXiv:2607.12255v1 Announce Type: new
Abstract: We study interaction-aware mixture-of-experts for post-stroke rigidity prediction using multi-level views of structured health records. Despite minimal performance gains, routing attribution reveals systematic importance differences across views, underscoring view construction as key to interpretability.
Science Journals
arXiv:2607.12256v1 Announce Type: new
Abstract: We develop a contour-count indicator method for visible pole clusters in outward meromorphic continuation from circular boundary data. The method starts from determinant characteristics built from positive Fourier coefficients. In the pure finite-pole model, the correct determinant characteristic factors into a polynomial whose zeros are the reciprocals of the exterior poles. In the presence of a holomorphic background, finite sampling, and noise, roots of individual determinants are unstable and are used only as local evidence. We aggregate this evidence into a scalar indicator field on the reciprocal pole plane: at each sampling point, the field records the fraction of determinant orders and shifts for which a small contour centered at that point encloses exactly one empirical determinant zero. The resulting field plays the role of a sampling-type imaging functional for pole visibility. Fixed superlevel sets give visible-pole clusters, while zero-dimensional persistent homology is used only as a threshold-robust post-processing step. We prove deterministic results linking pure-pole contour counts, Rouch\'e stability, indicator-field contrast, fixed-threshold component recovery, and persistence-gap stability. These results explain why isolated poles with sufficient residue and separation generate stable high-value components, whereas weak, close, boundary-near, or noise-dominated poles may give low, short-lived, or merged components. The framework is a cluster-certification and imaging method, not an unconditional all-pole recovery procedure.
arXiv:2603.18191v2 Announce Type: replace
Abstract: Optical beam steering is a key technology for free-space optical communication, sensing, and imaging. Mechanical beam steering systems suffer from limited scanning speed and bulky form factors, while existing solid-state solutions rely on pixelated synthetic aperture that requires complex fabrication and control architectures. Integrated acousto-optic beam steering (AOBS) is an emerging technology that enables continuous one-dimensional beam steering using integrated acoustic transducers and fixed-wavelength laser sources. Here, we integrate AOBS with an optical ring resonator on the same thin-film lithium niobate (TFLN) platform to significantly enhance beam steering efficiency and system functionality. The resulting device achieves a resonance-enhanced beam steering efficiency of up to $26\%$ and a field of view of $18^\circ$. Moreover, by leveraging integrated electro-optic control, we dynamically lock the ring-resonator's resonance to a chirped laser frequency, enabling frequency-modulated continuous-wave (FMCW) LiDAR operation. By combining lithium niobate's piezoelectric and electro-optic properties, this work establishes a compact, efficient, and scalable beam-steering platform with co-integrated acousto-optic modulation and electro-optic control for multifunctional applications.
arXiv:2607.11019v2 Announce Type: replace
Abstract: Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineering, it operates in an open, ambiguous, and continuously evolving environment. These characteristics call for a data-agent architecture that treats semantics, methodology, execution, and evolution as first-class system concerns. To this end, we introduce QwenPaw-Data, an agentic data system designed for enterprise intelligent data analysis. QwenPaw-Data consolidates heterogeneous assets from warehouses, dashboards, documents, interaction logs, and historical tasks into reusable, governable, and evolvable analysis assets, then turns natural-language requests into end-to-end analytical workflows spanning data understanding, retrieval, analysis, report generation, and decision support. Its architecture decomposes the problem into three collaborative subsystems: DataBridge provides trustworthy semantic grounding through interconnected metadata, knowledge, and trace graphs; Skill-Hub codifies expert analytical methodology into reusable and verifiable skills; and Host materializes these evidence and method assets into controllable, artifact-centric runtime execution. Across these subsystems, semantics, methods, traces, and feedback are continuously deposited back into the system, forming a self-evolving asset flywheel. Experiments on public benchmarks and real-world industrial BI workloads show that QwenPaw-Data improves both verifiable data access capability and higher-level analytical quality, offering a practical foundation for reliable, traceable, and continuously improving enterprise data agents.
arXiv:2604.16944v2 Announce Type: replace
Abstract: For an extensive-form game, logistic quantal response equilibrium (QRE) is defined with respect to its associated normal form and provides a natural equilibrium-selection mechanism as the rationality parameter tends to infinity. However, direct computation of logistic QRE in the normal form is generally impractical because the strategy space grows exponentially in the number of information sets. To address this difficulty, we construct a dilated-entropy-barrier artificial game in the sequence form and prove that its Nash equilibria characterize the corresponding logistic QREs. Building on this characterization, we further develop a sequence-form formulation of logistic QRE relative to a totally mixed strategy profile. This formulation gives rise to a differentiable path-following method for tracing the associated logit-QRE path, and we establish the existence of the corresponding smooth path. By recasting the dilated-entropy terms as the standard entropy terms, we additionally derive an equivalent smooth path. Numerical experiments illustrate the equilibrium-selection process of the proposed methods and evaluate their computational performance.
arXiv:2607.12939v1 Announce Type: new
Abstract: Point tracking in surgery is crucial to enable applications in downstream tasks such as segmentation, 3D reconstruction, virtual tissue landmarking, autonomous probe-based scanning, and subtask autonomy. This paper introduces the 2025 iteration of a point tracking challenge to address this, wherein participants submit their algorithms for quantification. Their algorithms are evaluated using a dataset named surgical tattoos in infrared (STIR), with the challenge named the STIR Challenge 2025 (STIRC2025). The STIR Challenge 2025 comprises two quantitative components: accuracy and efficiency. The accuracy component tests the accuracy of algorithms on in vivo and ex vivo sequences. The efficiency component tests algorithm inference latency. The challenge was conducted as a part of MICCAI EndoVis 2025, and seven teams participated in this challenge. In this paper we summarize the challenge results and participant methods. The challenge dataset is available at: https://zenodo.org/records/20191078, and the code for baseline models and metrics calculation is available here: https://github.com/athaddius/STIRMetrics
arXiv:2607.12624v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly deployed as purpose-specific agents to handle domain-specific tasks such as customer service and code generation. These agents are expected to comply with not only generic safety guardrails but also purpose-specific restrictions tailored to their designated roles. Such additional restrictions enlarge the attack surface, particularly to prompt injection (PI) attacks. To defend against such attacks, existing detection methods primarily rely on analyzing input-output patterns, yet yield limited effectiveness. To address this limitation, we turn to analyzing the hidden activation space and discover that LLMs inherently retain latent policy-violation (PV) concepts when prompted with requests beyond their designated purpose. Particularly, PV concepts capture the semantics of conflicts between user queries and predefined restrictions, implicitly reflecting LLMs' intrinsic awareness of recognizing policy violations. Building on this insight, we propose PVDetector, a training-free framework that detects PI attacks during LLM inference by measuring hidden-state alignment with PV concepts, which are derived offline from the contrastive pairs of policy-violating and policy-compliant prompts. Experiments across multiple LLMs and datasets show that PVDetector achieves <1\% false negative rate with minimal auxiliary overhead, consistently outperforming state-of-the-art methods. Our code is available at https://github.com/Claresigle/PVDetector .
arXiv:2607.12627v1 Announce Type: new
Abstract: We propose an architecture for learning the dynamics of mechanical systems based on discrete forced Euler-Lagrange equations on Lie groups using only position data. By formulating the dynamics directly on manifold-valued configuration spaces, the method naturally respects the geometric structure of the systems and preserves geometric invariants and conservation laws. The reliance on position measurements alone makes the framework applicable in settings where velocity data are unavailable or noisy. The approach extends naturally to multibody systems, accommodates external control inputs, and demonstrates strong performance on both synthetic and real-world datasets.
arXiv:2607.12549v1 Announce Type: new
Abstract: We investigate the capacity of an integrated sensing and communication system operating with orthogonal frequency division multiplexing, where the integrated sidelobe level of the transmit signal is adopted as an input cost. The problem is reduced to the capacity of a complex additive white Gaussian noise channel under power and kurtosis constraints. The sensing-optimal and communication-optimal operating points are characterized by circularly symmetric constant-modulus and complex Gaussian inputs, respectively, while the capacity-achieving input in between has discrete amplitude support with a concentric multi-ring geometry. We propose the superposition of the two extremal inputs as a simple and tunable signaling strategy, whose rate admits an exact expression via successive decoding. Numerical results show that the proposed strategy attains a worst-case gap of approximately 0.043 bits to the capacity.
arXiv:2607.12632v1 Announce Type: new
Abstract: We propose a data-driven procedure, based on convolutional variational autoencoders, to identify the presence of signal pulses in long time-series. The dataset consists of synthetic waveforms, each composed of non-gaussian noise and a log-normal shaped signal of variable intensity, with a length of 10,000 samples. The model heavily compresses the input waveforms, allowing a direct study of such a reduced representation. After training for 150 epochs on 7,500 waveforms, a region in the latent space where the network encodes time-series presenting only background noise emerges, allowing in turn to tag as candidates for containing a signal those falling outside. When applied on a test dataset of freshly generated waveforms, 100% of the events with a large pulses are correctly labelled, and this fraction only decreases for signal amplitudes comparable with accidental noise pulses. This approach was designed to fully exploit the measurements in dual-phase Liquid Argon Time Projection Chambers, as the one of the Recoil Directionality experiment, built in the context of the Darkside project. The goal is the identification of delayed electroluminescence signals, produced by low energy (~ a few keV) nuclear recoils, with a sensitivity at least comparable to the conventional reconstructions.
arXiv:2604.16668v2 Announce Type: replace-cross
Abstract: We derive distance relay characteristics in terms of incremental phasors. We use a circuit model of the network to estimate the incremental remote current. If we assume that all sources are stationary, i.e., remain periodic shortly after a fault, then the incremental remote current, and thus the characteristics, do not depend on the real-time voltages or current injections of the sources.
arXiv:2607.12578v1 Announce Type: new
Abstract: Post-click conversion rate (CVR) is a crucial element in online recommendation systems, which addresses significant challenges such as data sparsity (DS), sample selection bias (SSB), and delayed feedback. However, the impact of item discount rate-a key factor influencing both pricing and user purchasing behavior, has received limited attention. In this paper, we introduce the Discount-Aware Network (DANet) to model the relationship between item discount rates and CVR. DANet comprises three main components: 1) a time-frequency transformation module that utilizes Fourier transform to derive the frequency spectrum and capture the long-term discount rate trends of items; 2) a distribution de-bias module designed to mitigate the biases in user-specific discount rates caused by various purchase combinations and promotional activities, as well as periodic deviations linked to different promotion periods on e-commerce platforms; and 3) a supervised regression auxiliary task that establishes the explicit item discount labels to enhance the model's performance in terms of value accuracy, facilitating an effective representation of item discount rates. Experimental results on real datasets demonstrate the superiority of DANet, with offline AUC improving by 1.61%, and online A/B test also shows that DANet achieves impressive gains of 3.63% on pCVR and 2.23% on GMV. DANet has been successfully deployed on Alibaba Tmall APP. The code is available at https://github.com/tangrc/DANet.
arXiv:2607.12582v1 Announce Type: new
Abstract: For moving-load problems whose geometry and material properties are approximately invariant along the traveling direction, 2.5D analysis retains three displacement components at lower cost than full three-dimensional discretization. We present a 2.5D Non-Uniform Rational B-spline (NURBS)-trace infinite-element method (NBIEM), formulated as a coupled finite/infinite-element scheme, for wave propagation in linear viscoelastic semi-infinite geotechnical media. The bounded near field is discretized by isogeometric analysis, while the exterior is represented by tensor products of the boundary NURBS basis and admissible outgoing or evanescent exponential radial functions. Both subdomains share the same NURBS trace space and control-point degrees of freedom, enforcing displacement continuity without projection or mortar variables. For the selected radial functions, far-field stiffness and mass contributions are evaluated through closed-form radial moments, eliminating finite radial cutoff and radial quadrature. Closed-form half-space solutions verify displacement and stress frequency-response functions in sub-Rayleigh, super-shear but sub-compressional, and super-compressional moving-load regimes. Low-frequency studies assess sensitivity to radial parameters and artificial-boundary placement. Additional tests examine complex-valued response accuracy, phase fidelity, computational cost, and the frequency-dependent working range of the default S-wave-informed exterior realization. Applications to layered media, track--subgrade systems, and buried structures demonstrate the ability to handle heterogeneous materials, multi-patch configurations, curved interfaces, and cover-depth-dependent geotechnical responses. The framework provides a geometrically consistent and computationally efficient treatment of moving-load wave propagation and soil--structure interaction in semi-infinite domains.
arXiv:2607.12663v1 Announce Type: new
Abstract: Accurate differentiation between gastric adenoma and carcinoma during endoscopy is critical for clinical decision-making. Yet, this task is highly challenging due to high inter-class similarity and ambiguous boundaries between the two classes. Existing ROI-based classification methods often suffer from detection/segmentation error propagation and loss of surrounding global context. In contrast, full-image classification lacks the necessary spatial focus. Furthermore, we observe that deep neural networks gravitate towards domain-specific texture biases(e.g. bleeding, lighting artifacts), often causing models to predict based on spurious correlations instead of intrinsic morphological features. To address these limitations, we propose a novel framework, Masked Achromatic Guidance Expert (MAGE). During training, we introduce an auxiliary local expert branch trained on masked achromatic views of the neoplasm. By suppressing background context and color, this branch is forced to learn highly discriminative, purely structural features. We then employ a dual-objective distillation strategy, transferring both classification logits and spatial attention maps to provide implicit spatial supervision to the main branch that receives full WLI as input. This dual-objective distillation forces the model to ground its predictions in morphology rather than relying on shortcuts, while still retaining clinically relevant color cues. At inference time, our deployable model operates on images without annotated masks, ensuring real-time deployability . Extensive experiments on a clinical gastric endoscopy dataset show that our method significantly outperforms existing detection-based methodologies (e.g. YOLO) and classification-based methodologies (e.g. Swin-Transformer), providing not only superior classification performance but also interpretable attention maps for clinical reliability.
arXiv:2607.12905v1 Announce Type: new
Abstract: The transport of runaway electrons (RE) in ergodic magnetic geometries is an area of active study. Computing the transport from the direct simulation of particle trajectories is computationally expensive. Instead, diffusion models, such as the one by Rechester and Rosenbluth, are often employed to incorporate transport effects into reduced simulations. However, the comparison of diffusion-based to direct simulations reveals that the transport is typically not purely diffusive. In this paper, we introduce a simple transport model, based on chaos theory, which goes beyond the Rechester-Rosenbluth approximation. Besides chaotic diffusion, our model takes into account the effect of so-called sticky regions, a trapping layer around magnetic islands, where particle escape slows down to a power-law decay rather than an exponential decay. We demonstrate the applicability of the model both in the Ullmann-Caldas map with parameters corresponding to the TBR-1 tokamak, and in a JOREK simulation of a JET disruption scenario, with remarkably good fits achieved in both cases.
arXiv:2607.12682v1 Announce Type: new
Abstract: This work investigates the nonlinear transition to large heat fluxes observed in local gyrokinetic simulations of electromagnetic turbulence in STEP. Using the stress-balance framework of Zhang et al. (arXiv:2606.04616, arXiv:2607.11789), we confirm that the onset of extreme transport correlates with a critical value of $q^{2}\beta_{e}$, where $q$ is the safety factor and $\beta_{e}$ is the ratio of electron thermal pressure to magnetic pressure, and relate this to a limit on the poloidal beta $\beta_{\mathrm{pol}}$. Crucially, this critical value lies below any relevant linear stability limit in the ($q$, $\beta_{e}$) space (e.g., the onset of ideal or kinetic ballooning modes). Using an extensive set of nonlinear gyrokinetic simulations, we demonstrate that the transition to large fluxes in STEP is governed by a balance between the electrostatic and magnetic-flutter stresses. We argue, and also show numerically, that larger-major-radius tokamaks reach the electromagnetic non-zonal regime at lower $\beta_{e}$, making this MHD-controlled saturation limit more accessible in reactor-scale devices than in small spherical tokamaks. We also demonstrate that access to a second-stable regime enables re-saturation at larger values of $\beta^{\prime}$. We further show that the ideal ballooning mode (IBM) threshold serves as a useful proxy for delineating this second-stable region and also as a qualitative guide for the onset of large fluxes. These results provide a predictive framework for identifying no-go zone predictions from local gyrokinetics and offer new insight into the electromagnetic saturation physics relevant to STEP and other high-$\beta_{e}$ devices.
arXiv:2607.12126v1 Announce Type: new
Abstract: Medical imaging systems require reliable front-end electronics that can acquire sensor data, process image information, identify errors, and communicate results to other parts of the system. In applications such as X-ray imaging, CT, PET, ultrasound, and other diagnostic imaging systems, the electronics must often handle large amounts of data while maintaining predictable timing and low-latency operation. Because of these requirements, FPGAs (Field Programmable Gate Arrays) are commonly useful for medical imaging and signal-processing applications. In this project, the medical imaging concept is simplified into a small FPGA-based frontend demonstration.
arXiv:2607.12489v1 Announce Type: new
Abstract: Distributed swarms typically rely on either active wireless communication or passive vision, and they are frequently hindered by bandwidth constraints or environmental sensitivity. This paper proposes Infra-Swarm, a robust vision-based swarm. Each robot is equipped with a near-infrared light source and four ordinary gray-scale cameras. The Infra-Swarm system directly measures the centimeter-level 3D position of neighbors based on the position (bearing) and intensity (strength) of optical flares in the captured images. By utilizing 940 nm narrow-band filters to physically reject 99.2% of ambient light interference, the perception front-end achieves hardware-level robustness against illumination variations. Furthermore, its minimal computational overhead provides a resilient foundation for the massive scalability of robotic collectives on resource-constrained hardware.
arXiv:2607.12127v1 Announce Type: new
Abstract: Learning-based methods for the traveling salesman problem (TSP) are often evaluated through the tours produced after decoding or search, but the learned object itself frequently lives in a surrogate space such as heatmaps, assignments, construction policies, or search-guidance scores. This hides the fundamental question: what Hamiltonian structure has actually been learned before decoding? In this study, we directly answer this question by learning TSP through a structurally meaningful latent object, rather than leaving most of the Hamiltonian structure to the final decoding stage. Based on a connected-by-construction rooted $1$-tree Gibbs family, we propose an end-to-end unsupervised learning pipeline called \emph{C2TSP}. The pipeline learns residual edge perturbations from unbiased TSP cost through implicit differentiation. For structural correction, a smoothed Held--Karp layer restores expected degree balance, while certificate-guided sharpening further pushes the connected distribution toward more tour-like structures. Experiments show that C2TSP yields strong decoding performance while preserving interpretable structural information. Ablations further verify that edge perturbation and certificate-guided sharpening jointly improve both tour cost and tour-like structure.
arXiv:2607.12517v1 Announce Type: new
Abstract: A common intuition holds that a region's music mirrors the temperament of its people, so that melancholic melodies mark melancholic populations. We test the measurable half of that intuition and reject the inferential half. Using the Essen Folksong Collection, a corpus of thousands of notated folk melodies, we extract real melodic and affect-related features from 2393 deduplicated melodies spanning 16 countries and 7 geographic regions, with the analysis performed on symbolic scores rather than audio. The mode of each melody is computed with a key-finding algorithm rather than read from the file, because the collection's own documentation warns its major and minor labels are unreliable. Cross-country differences in melodic structure are large and highly significant. All 8 tested features differ across countries at p<0.001, with the leap-related features reaching p<10^-90, and China carries a distinctive wide-leap, high-activity signature (arousal composite +1.24 standard deviations, mean absolute interval 2.77 semitones against Germany's 2.17). We then test the inferential half. We correlate the regional musical-affect measures with two published, validated national indices, the World Happiness Report ladder score and the Hofstede individualism index. None of the 6 correlations is significant (0 of 6). The geography of musical affect is real and measurable, but it does not predict how happy or how individualist a population is, and any claim that it does is an ecological fallacy. We release the full extraction and analysis pipeline, and a fail-closed checker re-derives every number in this paper from the data.
arXiv:2607.12177v1 Announce Type: new
Abstract: The analysis of satellite and aerial imagery has entered a new era with the advent of foundation models. This paper describes the concept of Geospatial Foundation Models (GeoFMs), which are artificial intelligence/machine learning (AI/ML) models pre-trained on massive geospatial datasets through varied methodologies. We first articulate the core paradigm shift that GeoFMs enable: a separation of duties, where large-scale model providers perform the computationally intensive pretraining, allowing domain experts to rapidly fine-tune or prompt these models for specific, mission-critical tasks. This approach democratizes access to state-of-the-art AI/ML while maintaining the security and confidentiality of the downstream task. We then explore the novel capabilities unlocked by different types of GeoFMs, distinguishing between the finetunable vision models produced by self-supervised techniques like masked auto-encoding, and the vision-language models produced by contrastive learning which enable zero-shot tasks like open-vocabulary image analysis. Next, we discuss the practical considerations for operationalizing GeoFMs, from performance-cost analysis to the broader MLOps ecosystem. To that end, we introduce a taxonomy of model adaptation strategies and propose a framework for domain experts to select the most cost-effective adaptation approach for their particular mission set. Finally, we present a forward-looking vision of Agentic Geospatial Reasoning, where Large Language Models act as intelligent orchestrators, leveraging GeoFMs as tools to answer high-level user queries in natural language and automate complex analytical workflows, moving the field from perception to cognition.
arXiv:2607.12524v1 Announce Type: new
Abstract: Developers now draw code from two very different sources, the accumulated human answers on sites such as Stack Overflow and the output of large language models. We ask two questions about that split. First, can the provenance of a code snippet be recovered from the code itself, and second, do the two sources differ in the security patterns they adopt for the same task. Using only open sources, a public gateway of open-weight language models and the public Stack Overflow API, we build a fully reproducible pipeline that collects real implementations of 31 security-sensitive programming tasks, among them OAuth with PKCE, JWT verification, password hashing, and SQL access, from 9 language models and from human answers, and scores every sample with deterministic security and style detectors. On 528 real samples we train a cross-validated classifier that recovers human versus model provenance with 93 percent accuracy against a 78 percent baseline, and a 7-way classifier that attributes a sample to the specific model that wrote it at 48 percent. We then report where the sources diverge on security, which patterns models adopt more often than the human corpus and which they inherit from it. Running the same tasks in Python, JavaScript, Go, and Java, we find the security divergence holds in every language while the provenance boundary is partly language-specific and does not transfer symmetrically between them. A vulnerability repair case study, in which models are handed insecure code and asked to fix it, finds a 77 percent repair rate across 21 seeds and 12 weakness classes, but a recurring partial-fix failure in which the model removes the insecure pattern without adding the correct defense. The pipeline is data driven, so any new task or language is added as a single specification entry, and a fail-closed checker re-derives every number in this paper from the stored data.
arXiv:2607.12179v1 Announce Type: new
Abstract: We explore how the response of a conceptual model of the marine carbon cycle depends on the way in which carbon is injected from the atmosphere. We find that, for single-injection pulses, the threshold amount required for a large response of the excitable system depends on pulse duration but not on its specific form. We do, however, see differences in the number of large transient responses in carbon and, correspondingly, the duration of the response for different pulse shapes. These differences are magnified as the system is pushed towards increased excitability and can be understood in terms of the geometry of an increasingly winding heteroclinic orbit. Inspired by Large Igneous Provinces (LIPs), we also consider random sequences of injection pulses. We find a wide range of possible responses for a given overall amount injected and duration, depending on the mean characteristics of the individual pulses. We also identify a resonance-like "Goldilocks" zone, in which intermediate pulse durations or arrival frequencies produce the largest number of repeated transients, and we test the framework with illustrative scenarios motivated by the Siberian Traps and Columbia River Basalt Group.
arXiv:2607.12181v1 Announce Type: new
Abstract: Robot gaze is a major component of human-robot dialogue coordination. Most studies of gaze in human-robot dialogue focus on face-to-face social conversations, but little is known about gaze in demanding task-focused interactions. In this paper, we investigate how the gaze of a robot game partner affects human visual attention and if humans tend to direct confirmation-seeking gazes towards the robot. In our study, we let participants play a collaborative word association game with a NAO robot acting as an embodied, LLM-driven conversational partner. Our experiments are conducted under two conditions, which implement mutual and referential gazes of the robot respectively. We record participants' gaze using eye tracking glasses and analyze the interactions using gaze coordinates, speech segments, key events and areas of interests. We find that robot gaze orientation does not affect the time to first fixation on words the robot proposed. We also find that participants gaze more often at the robot when their dialogue line contains confirmation requests, compared to when it does not. Our results indicate (likely also due to the cognitively demanding nature of the game) that the verbal aspect of this task overshadows the effects of referential robot gaze. These findings offer valuable insights for designing and validating robot gaze and turn-taking behavior in collaborative tasks which require coordination and efficient communication.
arXiv:2607.12590v1 Announce Type: new
Abstract: Reinforcement learning (RL) is traditionally concerned with learning a control policy for a fixed environment. In many engineering systems, however, the environment itself is alterable: physical or operational parameters can be tuned to shape the transition dynamics and costs experienced by the agent. This motivates jointly optimizing both the policy and the environment design parameters. To this end, we establish an Environment Parameter Gradient Theorem -- a formal expression for the gradient of the value function with respect to environment parameters. The key theoretical device is a generalized action-value function $Q_{\pi,\xi}(s,a,\zeta)$, which comprises two copies of the environment parameters: $\zeta$ governs the cost and transition dynamics at the current state--action pair, while $\xi$ governs the future rollouts. This decoupling yields a tractable closed-form gradient expression and is essential to the theorem's derivation. Building on this result, we develop a model-free algorithm that simultaneously learns the optimal policy and the environment parameters. We demonstrate the efficacy of our framework on a UAV network design problem, where the optimal UAV placement (environment parameters) and communication routes (governed by the policy) are learned jointly to minimize the total communication cost in the network.