Forskningsradar

Science Journals

Peer-reviewade publikationer — 61548 artiklar

How to improve the accuracy of semiclassical and quasiclassical dynamics with and without generalized quantum master equations
arXiv:2603.04563v2 Announce Type: replace Abstract: Semi- and quasi-classical (SC) theories can handle arbitrary interatomic interactions and are thus well-suited to predict quantum dynamics in condensed phases that encode energy and charge transport, spectroscopic responses, and chemical reactivity. However, SC theories can be computationally expensive and inaccurate. When combined with generalized quantum master equations (GQMEs), the resulting SC-GQMEs have been observed to enhance the efficiency and accuracy of SC dynamics. Yet, while the mechanism responsible for improved efficiency is clear, the underlying improved accuracy remains elusive. What is worse, SC-GQMEs can yield unphysical dynamics in challenging parameter regimes -- a shortcoming that might be avoided if the mechanism of accuracy improvement were understood. Here, we uncover this mechanism. We leverage short-time analyses to prove that exact, "left-handed" time-derivatives delay the onset of SC inaccuracy, and show that their numerical integration yields dynamics with improved accuracy, even without the GQME. We find, however, that these derivatives are a double-edged sword: while offering greater short-time accuracy, they become unphysical in challenging parameter regimes. Because short-lived memory kernels can leverage short-time accuracy while circumventing long-time instability, we develop a protocol to unambiguously determine the memory kernel cutoff, even in challenging regimes where previous treatments had failed. Our insights into accuracy improvement and kernel cutoff protocol can be expected to apply to complex systems that go beyond simple models.
Seeing the Reasoning: How LLM Rationales Influence User Trust and Decision-Making in Factual Verification Tasks
arXiv:2603.07306v2 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly show reasoning rationales alongside their answers, turning "reasoning" into a user-interface element. While step-by-step rationales are typically associated with model performance, how they influence users' trust and decision-making in factual verification tasks remains unclear. We ran an online study (N=68) manipulating three properties of LLM reasoning rationales: presentation format (instant vs. delayed vs. on-demand), correctness (correct vs. incorrect), and certainty framing (none vs. certain vs. uncertain). We found that correct rationales and certainty cues increased trust, decision confidence, and AI advice adoption, whereas uncertainty cues reduced them. Presentation format did not have a significant effect, suggesting users were less sensitive to how reasoning was revealed than to its reliability. Participants indicated they use rationales to primarily audit outputs and calibrate trust, where they expected rationales in stepwise, adaptive forms with certainty indicators. Our work shows that user-facing rationales, if poorly designed, can both support decision-making yet miscalibrate trust.
A convolutional neural network surrogate for hierarchical homogenization: fast elastic moduli prediction of digital rocks
arXiv:2606.25275v1 Announce Type: new Abstract: Digital rock physics (DRP) aims to estimate effective rock properties (e.g., elastic moduli) directly from 3D micro-CT images. However, direct numerical simulations (DNS) on high-resolution large 3D scans are often computationally prohibitive and severely limit the application of DRP. To address this bottleneck, we combine a lightweight 3D convolutional neural network (CNN) with hierarchical homogenization (HHM) and apply it to determine effective elastic moduli. In this scheme, a large rock image is divided into subcubes. The CNN replaces costly DNS by directly predicting subcube elastic moduli, while HHM upscales subcube-level predictions to the full rock. Using a shared convolutional backbone, we systematically compare three training targets: (i) full anisotropic $6\times6$ stiffness tensors, (ii) isotropic bulk and shear moduli $(K, G)$, and (iii) Hashin--Shtrikman (HS)-normalized factors. Across multiple rock types, all three models agree well with DNS results while substantially reducing the computational cost. Moreover, training from scratch on each rock type is fast enough that transfer learning is unnecessary. Across all three targets, the accuracy is comparable. In our comparative study, the HS-normalized factor offers the best overall speed--accuracy trade-off while guaranteeing physical consistency, making it a convenient default. The isotropic $(K, G)$ target is a slightly more accurate alternative.
Rate Programmable Ionic-Redox Switching with Tunable Volatility in CuCrP2S6
arXiv:2606.25679v1 Announce Type: cross Abstract: Metal thiophosphates are emerging as a multifunctional material platform for neuromorphic electronics due to their accessible polar phases and ion dynamics on biologically relevant timescales. While resistive switching in these materials is frequently attributed to ferroelectric or antiferroelectric polarization, the intrinsic role of ion dynamics remains underexplored. Here, we isolate and demonstrate purely ion-driven resistive switching in paraelectric CuCrP2S6. Robust and reproducible resistive switching is observed in the absence of measurable ferroelectricity. The conductance can be tuned through both voltage amplitude and sweep rate, revealing a rate dependence characteristic of ion dynamics. The resulting resistance states exhibit controllable volatility, where switching rate determines the decay time constant of the readout current, attributed to ionic relaxation. Using either inert or reactive electrodes, we observe electrical evidence of solid-state redox activity associated with the interfacial reduction of native Cu+ ions, enabling controlled formation of filamentary conduction pathways. Analysis of this process allows extraction of the Cu+ diffusion coefficient, providing quantitative insight into the underlying transport kinetics. The understanding of ionic-redox based resistive switching in CuCrP2S6 is crucial for unleashing its full potential as a material platform for dual- or multi-mode operation.
SafeGen: LLM-Driven Assertion Generation and Fault Criticality Evaluation for Functional Safety
arXiv:2606.25296v1 Announce Type: new Abstract: With advances in autonomous driving and electric vehicle technologies, functional safety has become a critical requirement in automotive chip design. Traditional simulation-based fault analysis is often overly conservative at the module level and fails to accurately reflect fault criticality. This paper presents SafeGen, an LLM-driven, formal-verification-assisted framework for functional-safety-oriented fault criticality assessment. SafeGen leverages large language models (LLMs) and a document-level Hyper Knowledge Graph (HyperKG) that incorporates Failure Modes, Effects, and Diagnostic Analysis (FMEDA) guidelines to extract verifiable specifications from design and safety documents and evaluate their relevance to overall system safety. The HyperKG is further enriched with register-transfer-level (RTL) information to guide the generation of Functional Safety Assertions (FSAs) that are both semantically grounded and design-aware. Each assertion is linked to its corresponding specification, enabling traceable reasoning throughout the assessment process. A gate-to-RTL fault-mapping mechanism supporting both stuck-at and bridging faults, combined with formal property verification (FPV), enables semantic-level fault criticality grading based on specification-linked assertion violations. A digital-physical co-simulation platform for a field-oriented control (FOC) system is developed to validate SafeGen. Experimental results demonstrate that SafeGen generates higher-quality assertions than existing LLM-based assertion generation frameworks while providing greater semantic interpretability in fault criticality assessment compared with traditional simulation-based approaches.
Brillouin Light Scattering Spectroscopy of Propagating Magnons at Sub-Kelvin Temperatures
arXiv:2606.25834v1 Announce Type: cross Abstract: Coupling light to magnetic excitations in the form of spin waves underpins both the optical study of magnetism and emerging schemes for quantum transduction, positioning the quanta of these excitations, magnons, as promising carriers for hybrid quantum networks. However, exploiting them in the quantum regime requires millikelvin temperatures to suppress thermal magnon populations, thereby confining such experiments to dilution refrigerators. There, magnons can already be excited and read out electrically, yet an optical interface required for microwave-to-optical photon conversion has been missing. Here, we demonstrate the first optical detection of coherently driven, propagating spin waves via Brillouin Light Scattering (BLS) spectroscopy inside a dilution refrigerator. By simultaneously recording the optical and electrical responses of the same spin-wave mode in a yttrium iron garnet film, we find that the BLS spectra track the electrically measured transmission across a range of applied magnetic fields. For the lowest optical power of 7.9 {\mu}W that still enabled spin-wave detection, we measured a global equilibrium sample temperature of 510 mK via a resistance thermometer, while numerical modelling of the laser-induced heating yields a maximum local temperature of 900 mK at the focal spot. This brings free-space optical access to magnons into the sub-kelvin regime, representing a milestone towards magnon-mediated quantum transduction in hybrid quantum systems.
Self Capacitive Tactile Sensor System designed for Companion Robots
arXiv:2606.25348v1 Announce Type: new Abstract: Tactile sensing is essential for humanoid robots to achieve safe physical interaction, dexterous manipulation, and truly human-like responsiveness. However, the design of such systems remains challenging. Conventional approaches often suffer from complex multilayer structures, intricate wiring, high cost, and poor scalability, making it difficult to realize full-body tactile sensing with real-time, low-latency detection while maintaining minimal computational load on the robot's main processor. In this work, we present a simple, scalable and hardware friendly tactile sensing system for a companion humanoid robot based on the self-capacitance principle. The proposed sensor system employs a single conductive fabric layer with a conductive fabric wire architecture and does not require intricate electrode patterning. Scalability was demonstrated by fabricating a 100-point sensor array on a flexible printed circuit (FPC). Evaluation across sampling frequencies showed that 10 Hz is insufficient and misses transient events, whereas 100 Hz and 1000 Hz reliably capture and clearly distinguish all interaction types: gentle touch, slow tapping, fast tapping, and hitting. A decision-tree classifier was implemented directly on the FPGA, offloading real-time inference from the Raspberry Pi 4 with minimal latency and negligible power overhead. This design fully meets the tactile sensing requirements of the HIRO-chan robot and is well-suited for full-body tactile sensing in HIRO-chan and other companion robots.
Cache-Resident LLM Inference in GB-Scale Last-Level Caches
arXiv:2606.25353v1 Announce Type: new Abstract: Large language model (LLM) inference is increasingly dominated by data movement across the memory hierarchy. Recent 3D-stacked cache technologies have enabled GB-scale last-level caches in modern server CPUs, making it possible to keep reusable model weights on chip and exploit cache bandwidth and latency. Achieving this regime is not straightforward: deeper pipelining for weight residency increases in-flight requests and KV-cache footprint, while cache-resident operators make operator-boundary synchronization a visible bottleneck. We present a cache-resident execution model for inference on hierarchical-memory clustered systems. The model separates weight-centric operators from attention and KV-cache management into dedicated resource domains, keeping reusable weights cache-resident while scaling KV capacity independently of pipeline depth. It also relaxes synchronization from operator boundaries to true sub-operator dependencies, reducing coordination overhead in the cache-resident regime. We instantiate this model on a multi-socket CPU cluster with a weight-attention decoupled architecture, locality-aware placement, and a specialized static runtime. The prototype substantially outperforms equally provisioned llama.cpp. On deployed Llama-3.2-3B and Llama-2-7B configurations, it achieves 2.04x-11.51x speedup on time-per-output-token (TPOT). Under a validated analytical model, it further reaches up to 13.9x TPOT speedup across model sizes, context lengths, and batch sizes. These results show that commodity CPUs with GB-scale last-level caches can support efficient LLM inference when execution is organized around cache residency, decoupled state management, and dependency-aware coordination.
Climatic Extremes in Brazil: a parallel analysis of Historical Trends and Socioeconomical Impacts
arXiv:2409.16309v3 Announce Type: replace Abstract: An important consequence of human induced climate change is the increase in extreme weather events. This study contributes to the understanding of Brazil's climate change by examining historical temperature and precipitation patterns. Extreme events of temperature and precipitation are identified using data from the Brazilian Institute of Meteorology, which includes records from 634 meteorological stations operating intermittently since 1961. Using the first 30 years (1961 to 1990) as the reference period, our results show a significant increase in warm days and a corresponding decrease in cold days over the last 30 years (1991 to 2020), in agreement with previous works. In terms of precipitation, it indicates a trend toward drier conditions in the Northeast region of Brazil, whereas the South is experiencing wetter conditions, with an increase in the number of heavy precipitation days in South and in the extremely dry periods in the Northeast. These results have been verified for consistency with several extreme climate indices measured in this study. Additionally, data from S2iD is analyzed, an official database that records natural disasters in Brazil, to estimate their impact in terms of human losses and financial costs over the past decade. Our findings indicate that drought events are the most economically costly, with multiple instances causing damages exceeding a billion USD, whereas storms have the greatest impact on people. Although it is not possible to directly attribute the natural disasters recorded in the S2iD database to the extreme weather events identified through meteorological data, discussion is done on potential implications of these events in the frequency and location of the disasters.
Consensus Time in 3-Majority and 2-Choices Is Determined by the Maximum Initial Opinion Density
arXiv:2606.11778v2 Announce Type: replace Abstract: We establish the correct parameter governing the convergence time of the 3-Majority and 2-Choices dynamics on the complete graph in the synchronous model. Recent work [Shimizu and Shiraga, PODC'25] provides matching upper and lower bounds on the number of rounds to consensus, but only in a weak sense: the bounds are shown to coincide for some initial opinion configuration. In contrast, we obtain tight bounds in a strong sense, with upper and lower bounds matching up to logarithmic factors for every initial configuration. Let $\alpha^{(0)}$ be the initial opinion-frequency vector, and denote by $\|\alpha^{(0)}\|\_\infty$ its maximum entry. We show that 3-Majority reaches consensus in $\tilde{\Theta}(\min\{\|\alpha^{(0)}\|\_\infty^{-1},\sqrt n\})$ rounds w.h.p., while 2-Choices reaches consensus in $\tilde{\Theta}(\|\alpha^{(0)}\|\_\infty^{-1})$ rounds w.h.p. Our results demonstrate that the convergence time of both dynamics is governed not by global parameters such as the number of opinions $k$ or the squared $\ell\_2$ norm of the initial opinion distribution, but rather by the ``local'' parameter $\|\alpha^{(0)}\|\_\infty$, the maximum initial opinion density.
FactorLibrary: From Polynomials to Circuits via Recursive Subgoals
arXiv:2606.25394v1 Announce Type: new Abstract: Finding minimal arithmetic circuits for polynomials over finite fields is a combinatorially hard problem central to algebraic complexity theory. We formulate it as a reinforcement learning problem in two directions, bottom-up and top-down. To address the challenge of a fast-growing combinatorial search space, we introduce FactorLibrary, which stores factorizable subexpressions that serve as reusable subgoals across training episodes. We trained a bottom-up agent with Gumbel-PPO-MCTS and two top-down agents with PPO+MCTS and SAC. The PPO+MCTS top-down agent exhibited the most stable performance, finding certified optimal circuits up to complexity $8$ with a success rate of $91.8\%$.
Single-Shot Realization of 10000-Mode Octave-Spanning Artificial Gauge Fields
arXiv:2606.23960v2 Announce Type: replace Abstract: Artificial gauge fields (AGFs) enable photons and other bosons to emulate fermionic phenomena such as chiral edge transport and quantum Hall phases; however, existing theories and realizations remain confined to narrow bandwidths under single-mode approximation. We introduce a general theoretical framework for ultra-broadband, multi-modal dispersion-corrected AGFs in both linear and nonlinear regimes. Using integrated photonics, we realize over 100 distinct AGFs hosting more than 10,000 modes across nearly an optical octave -- the first frequency-comb realization of the integer quantum Hall model for photons. Leveraging Kerr nonlinearity, we achieve single-shot AGF control beyond waveguide dispersion, robust to wafer-scale fabrication variations. Our results establish a new regime of ultra-broadband multimodal AGFs, opening pathways to exotic dispersion-corrected AGF dynamics and simulations, as well as volume-manufacturable device functionalities such as waveguide-dispersion-resilient photonic circuits, and AGF-enabled programmable nonlinear and quantum optics and optoelectrics.
S2-CAR: Segmentation-Supervised Complexity-Adaptive Recommendation
arXiv:2606.25415v1 Announce Type: new Abstract: Sequential recommendation aims to predict user preferences from interaction histories, yet existing models often struggle when behavior patterns become complex and heterogeneous. A key reason is that interaction histories are rarely uniform: users' interests shift in a latent way over time, yet existing models either treat the full sequence as a homogeneous context or rely on rigid time-window segmentation that misaligns with true intent boundaries. This mis-segmentation not only introduces cross-intent interference at intermediate sequence positions but also leads to over-reliance on short-term interest signals. To address this, we propose S2-CAR, a segmentation-supervised and complexity-adaptive framework for sequential recommendation that models user intent as a continuous latent energy state. Specifically, it uses the Context-Aware Soft Temporal Point Process (Soft-TPP) to segment boundaries triggered by the natural decay of latent-state energy rather than fixed intervals, enabling intent segmentation without fixed time-gap rules. Next, upon this segmentation, a Segment-Count-Adaptive Multi-Intent Extraction module hierarchically aggregates intent-coherent segments into a compact set of multi-interest representations. Extensive experiments on 3 representative public benchmark datasets spanning movie, e-commerce, and gaming domains across 13 baselines demonstrate that S2-CAR consistently outperforms state-of-the-art methods across all datasets and metrics. Further analysis shows that the proposed energy-based segmentation serves as a plug-and-play module, yielding consistent improvements when integrated into existing sequential recommendation backbones.
Why Pool When You Can Flow? Active Learning with GFlowNets
arXiv:2509.00704v2 Announce Type: replace Abstract: The scalability of pool-based active learning is limited by the computational cost of evaluating large unlabeled datasets, a challenge that is particularly acute in virtual screening for drug discovery. While active learning strategies such as Bayesian Active Learning by Disagreement (BALD) prioritize informative samples, it remains computationally intensive when scaled to libraries containing billions samples. In this work, we introduce BALD-GFlowNet, a generative active learning framework that circumvents this issue. Our method leverages Generative Flow Networks (GFlowNets) to directly sample objects in proportion to the BALD reward. By replacing traditional pool-based acquisition with generative sampling, BALD-GFlowNet achieves scalability that is independent of the size of the unlabeled pool. In our virtual screening experiment, we show that BALD-GFlowNet achieves a performance comparable to that of standard BALD baseline while generating more structurally diverse molecules, offering a promising direction for efficient and scalable molecular discovery.
Learning with a Single Rollout via Monte Carlo Pass@k Critic
arXiv:2606.25451v1 Announce Type: new Abstract: Estimating token-level advantages in reinforcement learning (RL) for language models remains challenging because scaling up episodic experience collection is expensive. The difficulty intensifies for baseline advantage estimation methods, where repeated sampling causes trajectories to diverge into substantially different reasoning prefixes. In this context, RL algorithms such as GRPO prove limited: an outcome reward is too sparse to be attributed to specific actions like intermediate steps, and comparisons across sampled traces are non-trivial because they are heterogeneous. To mitigate both the computational cost of repeated sampling and the difficulty of credit assignment, we study single-rollout proximal policy optimization (SR-PPO) featuring token-level credit assignment in RL for language models. Instead of estimating advantages by normalizing episodic returns within the candidate group, we train a calibrated token-level credit critic using Monte Carlo outcomes from one rollout per prompt. Specifically, we use the critic to predict the Pass@k success probability at the prompt prefix, which is derived from a Pass@1 attempt. This choice yields a more selective learning signal than Pass@1: it discounts easily solved prefixes while prioritizing hard ones whose success probability remains marginal. We show that as $k$ increases, Pass@k converges to a reachability indicator, reflecting whether a prefix can lead to at least one successful continuation. In an explicit state graph, the limit ($k \rightarrow \infty$) can be computed in $O(|V|+|E|)$ time, offering a promising surrogate for direct credit assignment without the need to sample contrastive traces. As an initial validation, SR-PPO exhibits stable learning dynamics, along with consistent gains in Pass@128 success rates on mathematical reasoning benchmarks such as HMMT26 and AIME24.
Approximating velocity fields with planted attractors via Neural-ODEs for classification purposes
arXiv:2606.23550v2 Announce Type: replace-cross Abstract: In this work, Neural ODEs equipped with a curated collection of equilibrium points have been successfully employed for classification tasks. The planted attractors serve as indicators for the target classes, while the velocity field leveraging the universal approximation capabilities of the architecture shapes the dynamical landscape. This process defines the basins of attraction of the trained model, effectively directing each input (provided as an initial condition) toward its corresponding destination target.
Bayesian Optimization for reanalysis and calibration of highly energetic sea state events simulated with a spectral third-generation wave model
arXiv:2601.00628v2 Announce Type: replace Abstract: Accurate hindcasting of sea state events is a cornerstone of coastal engineering, risk assessment, and climate-related studies, yet it remains limited by uncertainties in physical parameterizations and model structure. This study introduces an automated calibration framework based on Bayesian Optimization (BO) using the Tree-structured Parzen Estimator (TPE) to constrain key dissipative processes in the ANEMOC-3 hindcast wave model, including bottom-friction losses, depth-induced wave breaking, and dissipation driven by wave strong opposing currents. The methodology enables the joint optimization of continuous physical parameters and discrete model structure choices within a unified probabilistic search space, significantly reducing model-observation misfit. Calibration is conducted over the high energy storm conditions of February 2014, while transferability is assessed both temporally and spatially, through independent validation on January 2014 and January 2018 events and across a network of offshore and coastal buoy observations. The optimized configurations retain skill beyond the calibration period and across observation sites, yielding systematically improved agreement with buoy measurements in terms of bias, root mean square error, and scatter index relative to the reference configuration. These results highlight the potential of Bayesian Optimization as a scalable and robust framework for automating the calibration of complex wave hindcast systems. Future developments will address multi-objective optimization, uncertainty quantification, and the integration of complementary observational datasets.
IntentTester: Intent-Driven Multi-agent Framework for Cross-Library Test Migration
arXiv:2606.25588v1 Announce Type: new Abstract: Unit tests capture both functional checks and domain-specific knowledge, but this knowledge remains locked within individual projects and is rarely reused across libraries with overlapping functionality. Existing migration techniques based on structural code mappings (e.g., API signatures) often break down under divergent designs or cross-language settings, resulting in non-executable migrated tests. In this paper, we present IntentTester, a multi-agent framework for intent-driven test reuse. Instead of translating raw code, IntentTester abstracts tests into a language-agnostic Test Description Language (TDL), aligns them with semantically related entities and dependencies in a repository graph, and synthesizes executable tests through LLM-guided reasoning and iterative validation. This design enables cross-library and cross-language migration without manual intervention, producing migrated tests that existing structure-mapping approaches cannot achieve. We evaluate IntentTester on nine open-source projects across three domains (JSON, HTML, and Time) and two languages (Java and Python). IntentTester generates 2,776 syntactically correct tests with 85\% correctness; in comparison, the two baselines achieve 51\% and 43\%. Among them, 2,410 tests executed successfully, yielding a 74\% effectiveness rate. Beyond higher success rates, IntentTester also surfaced previously unknown defects including stack overflows, null dereferences, and parsing inconsistencies, several of which have been acknowledged or patched by maintainers. Our results show that intent-driven migration shifts the focus from code mappings to semantic alignment, allowing practical cross-library and cross-language test reuse while improving test quality and exposing implementation flaws.
The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems
arXiv:2606.26057v1 Announce Type: new Abstract: AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems. The dominant approach places controls inside the agent's own runtime: system prompts, output filters, and guardrail libraries. Any control in the agent's address space is reachable by inputs that influence it; this generalizes to any AI system with sufficient reach into its own runtime, a class we term escapable AI systems. We identify four properties that an authorization mechanism must satisfy for architectural control rather than for cooperative requests: process separation, pre-action enforcement on a structurally only path, fail-closed at both the request and system levels, and externalized signed evidence verifiable outside the controlled system's trust boundary. We position this layer as execution-time AI alignment, complementing training-time alignment (RLHF, Constitutional AI) and inference-time alignment. We present the Unfireable Safety Kernel, a Rust reference implementation realizing all four. Its fail-closed invariant is machine-checked at two levels: an SMT theorem (Z3) and an exhaustive bounded-model-checking proof of the production decision function (Kani, 4/4 harnesses). A Python-to-Rust migration was gated on byte-equivalence (1000/1000 fixtures; 17/17 adversarial classes). We evaluate the kernel governing a live, escapable AI system, a deterministic, self-improving world model, against an escape-seeking adversary driving its real self-modification seam: across 1,000 self-modifications, all 704 attempts on the safety-critical core are refused, with no escape; a further 300, under the operator kill switch, are also refused. A separate campaign of 6,240 authorization round-trips had no successful bypass. Against 3 contemporary systems claiming the agent control plane, the agent invokes control; here, it lacks that choice.
Bright-state source cancellation in dissipative shortcut Raman atom optics
arXiv:2606.24939v1 Announce Type: cross Abstract: Spontaneous Raman scattering limits shortcut-assisted atom optics, but its microscopic origin is obscured once the lossy excited state is adiabatically eliminated. We organize the problem around a single quantity: in the instantaneous dark-bright basis the lower-manifold optical source is carried entirely by the bright-state amplitude, $S=\Omega b$, so that primary spontaneous scattering reduces to the compact functional. This recovers the known dissipative-STIRAP loss in transparent form and makes the action of a shortcut explicit: ideal counterdiabatic STIRSAP cancels the bright-state \emph{source}, not the optical decay coefficient. We show this cancellation is exact in the full three-level model at the counterdiabatic point, for arbitrary one-photon detuning, Rabi frequency, and pulse duration. The residual source splits into orthogonal quadratures -- shortcut mismatch (real) and two-photon Doppler detuning (imaginary) -- which invites a velocity-selective protocol that nulls the Doppler quadrature for a chosen momentum class with a second, phase-shifted lower-state field. Our central result is that this source nulling is never superior to simply chirping the two-photon detuning: the two coincide only when the selected class $\delta_c$ is small compared with the bright-state gap, and the nulling degrades and then fails as $\delta_c\to|\mu|$ -- precisely the regime of launched or warm clouds and high-order large-momentum-transfer (LMT) optics that motivates velocity selection. The controlling quantity is the magnitude of the residual Hamiltonian perturbation a scheme leaves behind, not the residual source it cancels. As a complement to existing multi-pulse decay budgets, we cast a single-pulse mode-error budget for LMT interferometry entirely in terms of the bright-state source, and delineate when shortcut-assisted Raman control reduces the total scattering cost.
Cross-Attention Multimodal Learning for Predicting Response to Neoadjuvant Imatinib in Gastrointestinal Stromal Tumors: A Multicenter Retrospective Study
arXiv:2606.25579v1 Announce Type: cross Abstract: Background: Response to neoadjuvant imatinib in gastrointestinal stromal tumors (GISTs) is highly variable and cannot be reliably predicted using current clinical or molecular markers. This study developed and evaluated an explainable multimodal deep learning framework integrating computed tomography (CT) imaging and clinical variables to predict treatment response. Methods: Patients from four tertiary centers were retrospectively included between 2000-2023 in independent pretraining (n=935) and prediction (n=213) cohorts. A cross-attention framework integrating clinical variables and tumor-centered CT imaging was developed to predict response to neoadjuvant imatinib. Two training strategies were evaluated: (1) self-supervised pretraining with low-rank adaptation and (2) training from scratch. Hyperparameters were optimized using SMAC3. Performance was assessed through internal cross-validation and external testing. Ablation analyses and attention-based explanations were used to quantify modality contributions. Results: Among 213 patients (54.5% responders), responders had larger tumors (112 vs. 89 mm, P=0.026), higher mitotic index (3 vs. 0, P<0.001), and more frequent KIT mutations (69.0% vs. 56.7%, P=0.019). Cross-attention models achieved the highest internal performance (AUC up to 0.99) but lower external performance (AUC 0.60-0.63). Clinical-only performance was moderate (AUC 0.66), whereas imaging-only models showed limited generalizability (AUC 0.56-0.66). Explainability analyses identified significant differences in feature importance between responders and non-responders, including CD117, BRAF, PDGFRA, age, sex, disease status, and comorbidities (FDR-adjusted P<=0.036). Conclusion: The cross-attention framework shows potential for improving imatinib response prediction in GIST while providing interpretable insights into multimodal determinants of treatment response.
Agentic Software Engineering: Foundational Pillars and a Research Roadmap
arXiv:2509.06216v3 Announce Type: replace Abstract: Agentic Software Engineering (SE 3.0) represents a new era where intelligent agents are tasked not with simple code generation, but with achieving complex, goal-oriented SE objectives. To harness these new capabilities while ensuring trustworthiness, we must recognize a fundamental duality within the SE field in the Agentic SE era, comprising two symbiotic modalities: SE for Humans and SE for Agents. This duality demands a radical reimagining of the foundational pillars of SE (actors, processes, tools, and artifacts) which manifest differently across each modality. We propose two purpose-built workbenches to support this vision. The Agent Command Environment (ACE) serves as a command center where humans orchestrate and mentor agent teams, handling outputs such as Merge-Readiness Packs (MRPs) and Consultation Request Packs (CRPs). The Agent Execution Environment (AEE) is a digital workspace where agents perform tasks while invoking human expertise when facing ambiguity or complex trade-offs. This bi-directional partnership, which supports agent-initiated human callbacks and handovers, gives rise to new, structured engineering activities (i.e., processes) that redefine human-AI collaboration, elevating the practice from agentic coding to true agentic software engineering. This paper presents the Structured Agentic Software Engineering (SASE) vision, outlining several of the foundational pillars for the future of SE. The paper culminates in a research roadmap that identifies a few key challenges and opportunities while briefly discussing the resulting impact of this future on SE education. Our goal is not to offer a definitive solution, but to provide a conceptual scaffold with structured vocabulary to catalyze a community-wide dialogue, pushing the SE community to think beyond its classic, human-centric tenets toward a disciplined, scalable, and trustworthy agentic future.
LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation
arXiv:2509.17773v3 Announce Type: replace Abstract: The rapid progress of image-guided video generation (I2V) has raised concerns about its potential misuse in misinformation and fraud, underscoring the urgent need for effective digital watermarking. While existing watermarking methods demonstrate robustness within a single modality, they fail to trace source images in I2V settings. To address this gap, we introduce the concept of Robust Diffusion Distance, which measures the temporal persistence of watermark signals in generated videos. Building on this, we propose I2VWM, a cross-modal watermarking framework designed to enhance watermark robustness across time. I2VWM leverages a video-simulation noise layer during training and employs an optical-flow-based alignment module during inference. Experiments on both open-source and commercial I2V models demonstrate that I2VWM significantly improves robustness while maintaining imperceptibility, establishing a new paradigm for cross-modal watermarking in the era of generative video. \href{https://github.com/MrCrims/I2VWM-Robust-Watermarking-for-Image-to-Video-Generation}{Code Released.}
USS: Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning
arXiv:2606.25880v1 Announce Type: new Abstract: Embodied Visual Tracking (EVT) requires an agent to continuously follow a specified target while actively moving through dynamic environments. However, prevailing EVT paradigms predominantly rely on language-based target indication. While language is expressive and convenient, cluttered scenes often contain multiple objects that satisfy the same semantic description, leading to ambiguous target grounding. We therefore propose a paradigm shift, reframing target indication in EVT from text-only specification to unified spatial-semantic prompting. Based on this paradigm, we introduce Unified Spatial-Semantic Prompts for Embodied Visual Tracking with Latent Dynamics Learning, USS, an end-to-end embodied tracking framework that supports text, point, bounding box, and mask prompts within a unified architecture. USS encodes heterogeneous prompts with modality-specific encoders, fuses prompt tokens with visual features through hybrid attention, and decodes compact prompt-conditioned representations into egocentric waypoints. To further improve temporal robustness, USS incorporates a latent world model that predicts future representations through self-supervised alignment. Real-robot experiments demonstrate that explicit spatial target cues yield higher success rates than text-only prompts, particularly in scenarios involving similar distractors and longer-horizon tracking where maintaining instance-level target identity is critical. In the simulation benchmark, USS also achieves state-of-the-art performance among non-MLLM-based methods and competitive results against recent MLLM-based approaches with faster inference speed. Our findings reveal that spatial-semantic prompting provides a more precise and flexible target indication interface for embodied visual tracking. Project site: https://arescheah.github.io/uss-project-page/.
Erased, but Not Gone: Output Forgetting Is Not True Forgetting
arXiv:2606.25001v1 Announce Type: new Abstract: Machine unlearning (MU) is commonly judged by output forgetting, such as low forget-set accuracy or reduced logit-level membership inference. But if output-level success can coexist with retraining-inconsistent residuals in representation space, what kind of forgetting are current evaluations actually certifying? We study this question through retraining-consistent representation forgetting, using the retrained model (i.e., trained from scratch without the forget data) as an operational reference for correct forgetting. Across multiple unlearning methods, datasets, and models, our theoretical analysis and empirical results show that standard output-level evaluation can systematically overestimate the success of unlearning. Under this stronger lens, current methods often appear forgotten at the output layer while exhibiting a structured mismatch relative to retraining. They partially align with retraining on forget samples, remain more inconsistent on retain samples, and leave residual discrepancy concentrated along retraining-related directions rather than diffuse in representation space. This structured mismatch is characterized by forget/retain asymmetry, directional mismatch, and concentrated residuals along retraining-related directions. These results suggest that current MU is often evaluated for apparent forgetting rather than retraining-consistent forgetting. More broadly, retraining reveals what output forgetting hides.