arXiv:2607.13027v1 Announce Type: new
Abstract: Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent systems run on desktops or servers, which support tool use and task automation. Mobile devices are also important agent environments because they are widely accessible and contain users' data, sensors, and daily-use applications. Existing mobile agents mainly operate smartphones through graphical user interface (GUI) actions such as tapping, swiping, and typing, which often form long, interface-dependent sequences, cannot directly access device capabilities, and make execution boundaries difficult to define. We present \textbf{PalmClaw}, an open-source agent framework that runs natively on mobile phones and manages the sessions, memory, skills, tools, and agent loop directly on the device. PalmClaw exposes device capabilities as device tools with explicit arguments, structured results, and clearly defined execution boundaries. This design enables agents to use mobile capabilities directly while keeping each action explicit and controlled. Experiments show an 11.5\% relative improvement in task success and a 94.9\% reduction in completion time over the strongest baseline, with lower setup burden and traces illustrating how execution boundaries are applied. Code is available at https://github.com/ModalityDance/PalmClaw.
Science Journals
arXiv:2510.13698v4 Announce Type: replace
Abstract: Even modern AI models often remain vulnerable to multimodal queries in which harmful intent is embedded in images. A widely used approach for safety alignment is training with extensive multimodal safety datasets, but the costs of data curation and training are often prohibitive. To mitigate these costs, inference-time alignment has recently been explored, but they often lack generalizability across diverse multimodal jailbreaks and still incur notable overhead due to extra forward passes for response refinement or heavy pre-deployment calibration procedures. Here, we identify insufficient visual attention to safety-critical image regions as one of the key causes of multimodal safety failures. Building on this insight, we propose Multimodal Risk-Adaptive Steering (MoRAS), which enhances safety-critical visual attention via concise visual contexts for accurate multimodal risk assessment. This risk signal enables risk-adaptive steering for direct refusals, reducing inference overhead while remaining generalizable across diverse multimodal jailbreaks. Notably, MoRAS requires only a small calibration set to estimate multimodal risk, substantially reducing pre-deployment overhead. We conduct various empirical validations across multiple benchmarks and MLLM backbones, and observe that the proposed MoRAS consistently mitigates jailbreaks, preserves utility, and reduces computational overhead compared to state-of-the-art inference-time defenses.
arXiv:2510.18939v2 Announce Type: replace
Abstract: Long-horizon agentic search requires iteratively exploring the web over long trajectories and synthesizing information across many sources, enabling powerful applications like deep research systems. In this work, we show that popular agentic search frameworks struggle to scale to long trajectories primarily due to context limitations--they accumulate long, noisy content, hit context window and tool budgets, or stop early. We therefore introduce SLIM (Simple Lightweight Information Management), a simple framework that separates retrieval into distinct search and browse tools, and periodically summarizes the trajectory, keeping context concise while enabling longer, more focused searches. Across a wide range of long-horizon tasks, SLIM achieves comparable performance at substantially lower cost and far fewer tool calls than strong open-source frameworks with both proprietary and open-weight models, including RL-trained models for deep research. Specifically, with o3 as the base model, SLIM achieves 56% on BrowseComp and 33% on HLE, outperforming all open-source frameworks by 8 and 6 absolute points, respectively, while incurring 4-6x fewer tool calls. With GLM-4.7 Flash, SLIM achieves 10 points improvement over the next best open-source framework, Search-o1, on BrowseComp using a third of the cost. To systematically understand failure modes in long-horizon agentic search, we develop an automated fine-grained trajectory analysis pipeline and error taxonomy, and find that SLIM exhibits significantly fewer hallucinations than prior systems. We hope our analysis framework and simple tool design inform future long-horizon agents.
arXiv:2506.21326v2 Announce Type: replace
Abstract: We study transport phenomena involving chemically reactive species, modeled by advection--diffusion--reaction systems coupled with flow fields governed by Darcy's law. Both the velocity field and the species concentrations are discretized using the Virtual Element Method, while time integration is performed through a discontinuous Galerkin scheme. This work represents a preliminary study, in which we introduce some simplifications of the full model. In particular, we assume a concentration-independent viscosity in the Darcy problem, constant diffusion tensors in the advection--diffusion--reaction systems, and first-order reaction networks with liquid-phase degradation. We derive an abstract error estimate by means of a technique that combines Gauss--Radau interpolation with numerical integration. The theoretical results are supported by numerical experiments that exhibit arbitrary-order accuracy in both space and time.
arXiv:2507.02513v5 Announce Type: replace
Abstract: Pedestrian detection in RGB images is a key task in pedestrian safety, as the most common sensor in autonomous vehicles and advanced driver assistance systems is the RGB camera. Low-light pedestrian detection lacks large public datasets and autolabelling pipelines. This research proposes a solution in the form of an automated infrared-RGB pipeline. The pipeline consists of 1) Infrared detection, where a fine-tuned model for infrared pedestrian detection is used 2) Label transfer process from the infrared detections to their RGB counterparts 3) Training object detection models using the generated labels for low-light RGB pedestrian detection. The research was performed using the KAIST dataset. For evaluation, three object detection models, DETR, YOLO, and RCNN, were trained on generated and ground truth labels. When compared on previously unseen images, the results showed that the models trained on generated labels out-performed the ones trained on ground-truth in 5 out of 6 cases for the mAP@50 and LAMR metrics, and outperformed ground-truth on mAP@50-95 in all cases. Acquired results indicate that the proposed auto-labelling pipeline could be used for scalable annotation of low-light datasets for pedestrian detection. The source code for this research is available on GitHub: https://github.com/BouzoulasDimitrios/IR-RGB-autoamed-low-light-pedestrian-labelling
arXiv:2607.12362v1 Announce Type: new
Abstract: Recent 4D Gaussian Splatting (4DGS) methods often fail under fast motion with large inter-frame displacements, where Gaussian attributes are poorly learned during training, and fast-moving objects are often lost from the reconstruction. In this work, we introduce Spatiotemporal Position Implicit Network for 4DGS, coined SPIN-4DGS, which learns Gaussian attributes from explicitly collected spatiotemporal positions rather than modeling temporal displacements, thereby enabling more faithful splatting under fast motions with large inter-frame displacements. To avoid the heavy memory overhead of explicitly optimizing attributes across all spatiotemporal positions, we instead predict them with a lightweight feed-forward network trained under a rasterization-based reconstruction loss. Consequently, SPIN-4DGS learns shared representations across Gaussians, effectively capturing spatiotemporal consistency and enabling stable high-quality Gaussian splatting even under challenging motions. Across extensive experiments, SPIN-4DGS consistently achieves higher fidelity under large displacements, with clear improvements in PSNR and SSIM on challenging sports scenes from the CMU Panoptic dataset. For example, SPIN-4DGS notably outperforms the strongest baseline, D3DGS, by achieving +1.83 higher PSNR on the Basketball scene.
arXiv:2607.12695v1 Announce Type: new
Abstract: Pinching antenna systems (PASS) employing dielectric waveguides have recently emerged as a promising flexible antenna architecture for high-frequency wireless communications. While prior work has focused primarily on millimeter-wave regimes, extending PASS to the terahertz (THz) band introduces distinct electromagnetic phenomena that invalidate conventional modeling assumptions. This paper develops the first analytical framework for THz-PASS that integrates in-waveguide propagation attenuation, evanescent coupling via coupled-mode theory, and THz-specific free-space effects including molecular absorption and its re-radiation noise. Using this model, we benchmark THz-PASS against conventional phased arrays under identical propagation scenarios. Our comparative evaluation reveals that THz-PASS achieves effective gains in spectral efficiency through proximity exploitation, making it particularly well-suited for confined and linear deployment topologies.
arXiv:2607.12638v1 Announce Type: new
Abstract: X-ray detection underpins a wide range of applications in medicine, security, industrial inspection, scientific research for non-destructive imaging and material analysis. The rapid development of Ga2O3-based X-ray detectors offers a promising pathway toward next-generation detectors with high sensitivity, low noise, and harsh environment applications, benefiting from its intrinsic material properties such as high density, wide band gap energy, and high thermal-chemical stability. However, the underlying device operating mechanisms, including both carrier excitation and transport processes, have not yet been adequately studied, largely due to the misuse of X-ray sources in previous studies. Besides, benchmarking of device characteristics has been problematic due to experimental or data analysis issues, as well as misunderstandings of the applied equations associated with parameter definitions. In this work, we have designed and performed an instructive research work based on epitaxial beta-Ga2O3:Si and its planar Schottky detectors, measured with energy-tuneable monochromatic X-ray beams on a synchrotron beamline, clarifying the device excitation and carrier transport mechanisms with properly benchmarked device performance. In the end, we propose a set of protocols for correctly measuring and analysing the device performance. The proposed protocols are broadly applicable and can be readily extended to other semiconductor X-ray detectors.
arXiv:2607.12642v1 Announce Type: new
Abstract: One strategy for reasoning about programs that have certain kinds of effects is to use program logics that provide specialized rules for reasoning about these effects. However, developing program logics requires skills that are distinct from those needed for using program logics, making the development of new logics challenging and less accessible. Moreover, when developing new logics, it can be difficult to reuse components from prior logics or combine support for different effects.
In this paper, we propose an approach for operationally building extensible program logics based on effect handlers. Our starting point is an expressive program logic for reasoning about programs written in a pure, sequential language with support for effect handlers. Within this language, we implement handlers that model concurrency, distributed execution, and crash-recovery behavior. Then, by proving properties about these handlers, we extend the program logic and derive expressive rules for reasoning about these effects. In some cases, this approach leads to stronger reasoning rules than those found in prior program logics targeting these features.
In addition, we develop a relational logic for proving contextual refinements between programs using effects. As with unary reasoning, handlers enable this relational logic to be developed in an extensible way.
arXiv:2607.11951v1 Announce Type: new
Abstract: Large language models can write SQL, but enterprise deployment demands more than plausible text: outputs must be syntactically valid, must respect per-role and per-schema policy, must carry provable (not best-effort) guarantees, must not slow down as generations grow, and must leave a compliance-grade record of every decision. We present GRID (Grammar-Railed Decoding), a grammar-constrained decoding engine that keys exact next-token masks on parser configurations (lexer scan state x LALR(1) stack) rather than on token sequences, and uses the incrementally advanced LALR(1) parser itself as a viable-prefix oracle. LLM tokens are bridged to grammar terminals by a byte-level trie walk with a context-independent/context-dependent split that makes cache-key soundness hold by construction. Role-based access control is compiled into the language: role projections subset the grammar's productions and schema lexicons restrict identifier terminals, so forbidden verbs and identifiers are unreachable at mask level. Four guarantees (soundness, completeness, termination, and near-constant per-token cost) are stated with explicit preconditions and each paired with a test or benchmark. Rust kernels bring the per-token mask to a 3.6-6.7 us median, ahead of llguidance at p50 and p90 on two tokenizers with zero false rejects; per-token guard cost is position-flat at n=16,000. On Spider, constrained decoding is worth +13 execution-accuracy points at 0.5B, and one checker-guided repair pass over the provably mask-unenforceable residue (column-level policy) lifts a 7B model to 94.5% executable. A hash-chained per-token audit trail replays bit-identically with 100% tamper detection. We state plainly what the mask cannot do (distribution faithfulness, column-level RBAC, non-LALR(1) languages) and where measured cost remains.
arXiv:2607.12643v1 Announce Type: new
Abstract: Electronic resonances play an important role in electron attachment-induced processes in biomolecules, and their properties can be significantly influenced by the local molecular environment. Here, we investigate the effect of amino acid micro-solvation on the uracil resonances by employing uracil-glycine as a model system. The resonance spectrum of the uracil-glycine complex consists of four {\pi}-type shape resonances and three core-excited resonances, including an additional glycine-centered resonance, as characterized using the CASSCF/Resonance via Pad\'e (RVP) methodology. Comparison with isolated uracil and the uracil(ghostGly) model shows that explicit interaction with glycine stabilizes both the shape and core-excited resonances by lowering their energies and increasing their lifetimes, while the ghost calculations demonstrate that basis-set extension alone cannot account for the observed stabilization. The core-excited resonances exhibit states that retain non-negligible lifetimes despite their much higher energy, suggesting that they may play an important role in electron-induced dissociation pathways. Overall, the present results demonstrate that amino acid micro-solvation significantly modifies the resonance landscape of uracil, highlighting the importance of explicitly accounting for local biomolecular interactions in theoretical studies of electron attachment.
arXiv:2607.11980v1 Announce Type: new
Abstract: Like many optimization-driven domains, railway rescheduling relies on Mixed-Integer Linear Programming (MILP), yet the field's modeling knowledge is scattered across hundreds of papers in incompatible notations, and narrative surveys organize it subjectively: they classify models by vocabulary rather than by structure, and reproduce neither. We present LP Mining with LP2Graph, a method that mines the structure of published LP and MILP formulations into a reproducible dataset and an induced taxonomy. Its core, LP2Graph, represents each formulation admitted by its canonical grammar as a typed variable--equation graph derived from a single canonical model; once a source is extracted into that model, everything downstream is deterministic. Each source is parsed into this model, homologized, and clustered bottom-up (over variables, then constraints and the objective, then whole-model structure) and, separately, by application domain and solution approach; the resulting groups are labeled by a rule-seeded, self-updating classifier. We validate the representation rather than assume it: per-cluster representatives are regenerated as independent LaTeX and re-solved across CBC, HiGHS and Gurobi against the optimum reported in the source paper. The outcome is an objective, repeatable taxonomy of variables, constraints and model types: the principled foundation on which our raiLPminer line of automated railway-rescheduling model development builds.
arXiv:2607.12135v1 Announce Type: cross
Abstract: This paper investigates the control of discrete-time linear time-invariant (LTI) systems subject to incomplete and corrupted measurements. Specifically, we focus on designing a Linear Quadratic Gaussian (LQG) controller without relying on explicit state estimation. By leveraging minimum variance duality, our approach allows the current control input to be represented as a linear function of available measurements and previously applied inputs, successfully reducing the task to a tractable deterministic optimization problem. We provide theoretical justification for this framework and demonstrate its practical effectiveness through numerical experiments.
arXiv:2607.12652v1 Announce Type: new
Abstract: We formalize strong barbed similarity for the pi-calculus in the Beluga proof assistant, completing a line of work addressing the Concurrent Calculi Formalization Benchmark. By extending previous developments to include replication, we give a coinductive encoding of behavioral equivalence based on barbs and internal actions. Using Beluga's copattern-based coinduction, we obtain concise and compositional proofs, including compatibility properties and a context lemma characterizing barbed precongruence. The case study demonstrates the effectiveness of combining HOAS and coinductive reasoning for mechanizing concurrent calculi.
arXiv:2607.12429v1 Announce Type: new
Abstract: Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same scene across extreme viewpoint gaps. Cross-view localization naturally provides a promising pathway toward this ability, as it requires a model to align ground-view imagery with geo-referenced satellite-view imagery despite drastic appearance changes to estimate camera poses. Recent visual foundation models have made this long-standing localization problem increasingly feasible by providing rich 2D representations for cross-view matching. However, we argue that cross-view localization should not be viewed merely as 2D matching or pose estimation. In this work, we revisit cross-view localization as more than pose estimation and investigate how it can help the model develop consistent cross-view understanding under extreme viewpoint changes, including stable semantics, reliable structure, and transferable geometry. We identify three key limitations of existing methods that prevent them from achieving this. They usually lack explicit 3D grounding, rely on strict point-wise matching that can weaken semantic consistency, and learn from an absolute objective that provides limited guidance for geometric reasoning. To address these limitations, we propose CROSS, a unified cross-view localization framework built upon 3D-grounded alignment, structure-aware matching, and hypothesis ranking. This formulation makes structure learning an intrinsic requirement, encourages semantic representations to remain stable, and enables the model to acquire transferable geometry. Extensive experiments on the KITTI and VIGOR datasets show that CROSS achieves state-of-the-art performance in cross-view localization. More importantly, CROSS effectively learns stable semantics, reliable structure, and transferable geometry across extremely different viewpoints.
arXiv:2604.13853v2 Announce Type: replace
Abstract: Safe and explainable motion planning remains a central challenge in autonomous driving. While rule-based planners offer predictable and explainable behavior, they often fail to grasp the complexity and uncertainty of real-world traffic. Conversely, learned planners exhibit strong adaptability but suffer from reduced transparency and occasional safety violations. We introduce Mosaic, a framework for structured decision-making that integrates both paradigms through arbitration graphs. By decoupling trajectory verification and selection from the generation of trajectories by individual planners, every decision becomes transparent and traceable. This separation lets verification and trajectory selection contribute independently: centralized verification acts as a safety floor, reducing at-fault collisions from 25 for each standalone planner to 16. In contrast, per-step trajectory selection acts as a performance ceiling, combining the complementary strengths of a rule-based and a learned planner. In experimental evaluation on nuPlan, Mosaic achieves 95.56 CLS-NR and 94.18 CLS-R on the Val14 closed-loop benchmark, setting a new state of the art. On the interPlan benchmark, focused on highly interactive and out-of-distribution scenarios, Mosaic scores 54.10 CLS-R, outperforming its best constituent planner by 22.8% -- all without retraining or requiring additional data. The code is available at github.com/KIT-MRT/mosaic.
arXiv:2607.12450v1 Announce Type: new
Abstract: This work introduces a unified formulation for vision models, where diverse forms of visual information beyond natural images, such as masks, depth maps, and other structured visual signals, are all represented as RGB images, while general visual tasks can be converted into a common RGB-to-RGB image editing problem. In this paradigm, different types of visual information internally share the same encoding and decoding architecture and parameters as natural images, enabling a single model to transfer across tasks through a unified visual interface, in a way analogous to how language models operate over text. We refer to this formulation as RGB In and RGB Out (RINO). Built upon a generic image editing backbone without task-specific fine-tuning, RINO demonstrates robust and competitive zero-shot performance on both dense understanding tasks such as segmentation and depth estimation (where we unify outputs as RGB), and dense-conditioned generation tasks such as pose-to-image generation (where we unify inputs as RGB). We hope this study provides useful insights toward general unified vision-language systems, where diverse visual tasks can be expressed, interpreted, and solved through a shared visual language. Code is available at https://github.com/yangtiming/RINO.
arXiv:2607.12903v1 Announce Type: new
Abstract: In operational 1:N face identification, a crucial question arises for each probe: is this person enrolled in the gallery or not? The stakes are high and asymmetric. Rejecting a mate-present (MP) probe loses a valid lead; accepting a mate-absent (MA) probe makes every returned candidate a false identification, at worst a wrongful arrest. Most approaches threshold match scores, but scores shift substantially with image quality and gallery size and composition, making thresholds fixed before deployment brittle under realistic conditions. Our prior work introduced 1-consistency, the only method based on rank consensus across multiple independently trained matchers: a probe is labeled MP if all matchers return the same rank-1 identity. This work stress-tests 1-consistency across 36 (gallery, probe quality) scenarios spanning four quality levels and two structural axes: images per identity and total enrolled identities. We benchmark against two score-thresholding methods that bracket what any deployed threshold could achieve. Fixed Score-Thresholding (FST), calibrated once on baseline conditions, collapses asymmetrically as quality degrades: MP recall falls below 2% while MA recall holds near 100%. Oracle Score-Thresholding (OST), re-tuned per scenario, is the best any threshold could theoretically do, yet for degraded probes 1-consistency matches it with zero tuning. The two differ mainly in error type (OST favors MP recall, 1-consistency favors MA recall), but on one axis 1-consistency does not merely match the oracle: when it labels a probe MP, it returns the correct mate 97-100% of the time versus OST's 66-84% under severe degradation. In short, 1-consistency delivers oracle-level accuracy without the impossible requirement: it sets no threshold, so it needs no advance knowledge of the conditions a probe will arrive in, which is what makes it usable.
arXiv:2607.12048v1 Announce Type: new
Abstract: Deploying medical visual question answering (MedVQA) systems in real-world clinical settings requires models that adapt to new clinical tasks without forgetting previously acquired knowledge. Continual learning (CL) provides a practical framework for this setting. Despite rapid progress in medical vision-language models, the behavior of CL methods when training these models across heterogeneous MedVQA tasks remains underexplored. This work presents a systematic evaluation of CL for MedVQA across diverse clinical objectives, including classification, multi-label classification, detection, cell counting, and report generation. Specifically, we explore (1) the ability of existing CL methods to mitigate catastrophic forgetting; (2) their sensitivity to task ordering, analyzing how different task sequences influence performance retention and forgetting; and (3) the evolution of low-rank adaptation parameters as new tasks are learned, revealing patterns of weight drift under different CL methods. Our findings suggest that existing CL methods struggle to maintain stability-plasticity balance when tasks with different objectives and supervision formats are interleaved. Code and full experimental setup will be publicly available.
arXiv:2607.12665v1 Announce Type: new
Abstract: The Swing-UP of quantum EmitteR population (SUPER) scheme has recently been proposed as a deterministic method for the preparation of collective radiative states in two strongly dipole-coupled quantum emitters (Phys. Rev. Res. \textbf{8}, 013179 (2026)). Here, we extend this approach to an equilateral subwavelength triangular trimer of dipole-coupled two-level quantum emitters (QEs), loosely inspired by biological light-harvesting ring geometries. Using tailored, time-overlapping, red-detuned ultrashort SUPER pulses, we numerically investigate the selective preparation of collective target states. We find that both the state selectivity and the preparation efficiency depend strongly on the inter-emitter spacing. In particular, at deep-subwavelength separations, the symmetric collective state can be deterministically prepared with near-unity efficiency, whereas the inversion efficiency and state selectivity are significantly lower at larger inter-emitter separations. Furthermore, this state preparation technique inherits a certain degree of robustness against reasonable static position imperfections and on-site frequency inhomogeneities of the individual QEs. Our results demonstrate that deep-subwavelength triangular trimers and, more broadly, highly compact ring geometries are excellent candidates for the deterministic preparation of collective radiative states via SUPER excitation. These predictions could be realized with solid-state emitters and molecules. Our findings offer a route toward the direct probing of the `pure' electromagnetic layer of interaction in biological and bio-inspired synthetic nanophotonic ring configurations, with possible relevance in photonics, quantum information processing, and metrology.
arXiv:2604.14060v2 Announce Type: replace
Abstract: Wavelength shifting (WLS) fibers are widely used in particle physics for light collection from scintillators. Light production by charged particles directly in WLS fibers is traditionally ignored. In this study, light produced by charged particles in WLS fibers is clearly observed. The light yield of different batches of Y11(200) 1 mm diameter WLS fibers expressed in the number of detected photo electrons is as large as $23\pm2~\%$ with respect to the light yield of the Bicron BCF-12 1 mm diameter scintillating fiber. This corresponds to about 1200 photons per MeV light production after taking into account different fiber light trapping efficiencies and emission spectra. In clear fibers of the same diameter, no scintillation light is produced, while Cherenkov light is clearly seen at the 45-degree crossing angle. The observed amount of light produced by charged particles in the WLS fibers is not small and should be taken into account in advanced detector simulations.
arXiv:2607.12082v1 Announce Type: new
Abstract: We introduce a tangent-space multiscale manifold method for nonlinear heterogeneous elliptic problems. The method represents the fine-scale solution by a nonlinear reconstruction of a coarse state. The ideal reconstruction eliminates fine scales through a constrained variational problem, and the computable reconstruction approximates this map by localized nonlinear patch solves blended with a partition of unity. Because the approximation set is a nonlinear manifold, the coarse equation is posed with tangent multiscale test functions. We also formulate a network-interpolated variant in which only the restricted patch outputs used by the partition-of-unity blend, together with their tangent actions, are approximated by local learned maps. For heterogeneous monotone nonlinear diffusion, we record the structural monotonicity, differentiability, patch-map regularity, and conditional perturbation estimates that separate the geometric stability mechanism from localization, residual, and optional learning defects. A rigorous a priori theory for the decay of the localization defect, and the resulting convergence rates in the coarse mesh size, is deferred to a separate analysis; here these defects are controlled conditionally and their decay is demonstrated numerically.
arXiv:2607.12680v1 Announce Type: new
Abstract: Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an explicit closed-loop mechanism for tracking semantic progress, diagnosing execution failures, and recovering from error accumulation in long-horizon navigation. To address this gap, we propose ReflectVLN, an agentic VLN framework that organizes decision-making through bidirectionally interactive intention and execution agents. The intention agent performs subtask decomposition and reflection, generating executable subtask descriptions as corrective plans. Conditioned on these descriptions, the execution agent grounds them into short-horizon actions under current observations while monitoring sub-goal progress and detecting off-track behavior. Crucially, ReflectVLN enables closed-loop bidirectional communication: the execution agent emits progress and deviation signals to trigger reflection and subtask updates on demand, and the intention agent returns structured guidance that reconditions subsequent actions for recovery. To encourage temporally coherent decisions with interpretable intermediate rationales, we introduce Action Chain-of-Thought (Action-CoT), a path-conditioned dual-query training scheme for action generation. Experiments on standard VLN benchmarks show that ReflectVLN improves success rates and path efficiency under a constrained data budget, with favorable training cost and fewer high-level intention calls at inference time, while providing interpretable intermediate decisions for analysis and collaboration. Code is available at: https://github.com/AIprogrammer/ReflectVLN
arXiv:2607.12497v1 Announce Type: new
Abstract: Beyond perception, reasoning is essential in remote sensing for advanced interpretation, inference, and decision-making. Recent advances in large language models (LLMs) have enabled tool-augmented agents that leverage external tools to perform complex analytical tasks. However, existing studies in remote sensing primarily focus on perception-oriented tasks, leaving cognitive geospatial reasoning largely underexplored. To address this gap, we introduce TerraLogic, a benchmark for geospatial reasoning. TerraLogic comprises 545 scenario-driven, hierarchy-aware tasks, such as hazard vulnerability assessment, urban heat island analysis, and forest fragmentation dynamics, spanning optical, Synthetic Aperture Radar (SAR), and infrared (IR) imagery. It advances evaluation beyond recognition and monitoring toward cognitive-level geospatial analysis. To facilitate evaluation on TerraLogic, we further propose HieraPlan, a tool-augmented agent that organizes toolkits into functional hierarchies and performs fault-tolerant reasoning. HieraPlan enables structured abstraction, robust recovery from tool failures, and stable long-horizon planning. Extensive experiments demonstrate that current approaches struggle with hierarchical geospatial reasoning, while HieraPlan provides a strong baseline with improved reasoning, cross-modal generalization, and error handling. The dataset and agent code are publicly available at https://github.com/Ireliya/TerraLogic.
arXiv:2607.12512v1 Announce Type: new
Abstract: Fast predictive modelling of radio-frequency heating and current drive is important for integrated tokamak scenario design, yet kinetic calculations of helicon-wave absorption remain too computationally expensive for large-scale parameter scans. We present a reduced model for helicon-wave heating and current drive that retains the dominant parallel electron Landau-damping channel. The wave response is evaluated on the cold-plasma dispersion root, and a single-Landau-pole correction is introduced to obtain compact expressions for the local damping rate and current-drive efficiency. The model is benchmarked against the Chiu-Chan heating model using approximately 1.6 million samples covering representative conditions of EAST, HL-3, DIII-D and KSTAR. The reduction error is found to be governed primarily by the electron Landau parameter and electron beta. Within an identified sub-lower-hybrid-frequency validity window, results from different devices collapse onto a common error curve, which enables an empirical correction that is further tested using ITER-like and BEST-like extrapolation cases. Near and above the lower-hybrid frequency, the agreement deteriorates rapidly owing to changes in the cold-dispersion root structure and the breakdown of the single-branch WKB description. When coupled to a reduced current-drive source, the corrected heating model gives a median deviation of 10.8 percent from the Landau-channel Ehst-Karney reference and reproduces published CFETR current-density profiles. The resulting model provides a computationally efficient reduced closure for helicon-wave heating and current-drive calculations, together with physically interpretable limits on its range of validity.