arXiv:2601.22443v2 Announce Type: replace
Abstract: Can a diffusion model trained on bedrooms recover human faces? Diffusion models are widely used as priors for inverse problems, but standard approaches usually assume a high-fidelity model trained on data that closely match the unknown signal. In practice, one often must use a mismatched or low-fidelity diffusion prior. Surprisingly, these weak priors often perform nearly as well as full-strength, in-domain baselines. We study when and why inverse solvers are robust to weak diffusion priors. Through extensive experiments, we find that weak priors succeed when measurements are highly informative (e.g., many observed pixels), and we identify regimes where they fail. To explain this behavior, we combine Bayesian-consistency theory with local-correlation analysis: the theory gives conditions under which high-dimensional measurements make the posterior concentrate near the true signal, while the correlation analysis shows that weak and stronger natural-image priors can share similar local spatial structure. These results provide a principled justification on when weak diffusion priors can be used reliably. Code is available at https://github.com/jjia131/weak-diffusion-priors-inverse-problem.
Science Journals
arXiv:2606.03362v1 Announce Type: new
Abstract: This study examined emerging and established topics in drone research, focusing on citation impact and knowledge flows across China, the United States, the EU, Ukraine, and Russia between 2020 and 2025 using OpenAlex bibliographic data. The findings revealed that drone-related science is characterised by growing geopolitical asymmetries in scientific production, citation concentration, and international knowledge exchange. In particular, China increasingly dominated scientific production, fractional authorship contribution, and domestic citation circulation. In contrast, the United States and EU countries maintained comparatively more internationally distributed citation structures. However, China-affiliated publications became increasingly integrated into global citation networks, particularly through growing citation exchange with the United States and European countries.
Notably, the interpretation of authorship and citation patterns was complicated by the high proportion of publications with unidentified affiliations, which reached 50% in 2025 within weak-signal topics. These findings underscore the importance of developing comprehensive national Research Organisation Registries (RORs).
Although China demonstrated a citation advantage, this was partly driven by high internal domestic citation concentration rather than exclusively by global integration. Moreover, China still imported proportionally more knowledge from the EU-14 and the United States than it exported, with this asymmetry increasing over time. EU-14 countries maintained the strongest citation impact in weak-signal topics, suggesting a more prominent role in shaping emerging research directions. At the same time, China-affiliated publications cited the United States more frequently than the EU-14 in both strong- and weak-signal topics, with this pattern being particularly pronounced in weak-signal areas.
arXiv:2606.03845v1 Announce Type: new
Abstract: We present and analyze an embedded Trefftz discontinuous Galerkin method for reaction-diffusion problems on anisotropic meshes. The method is constructed by imposing a relaxed local Trefftz condition via an embedding into a tensor-product DG space, yielding a reduced global system while preserving the approximation properties of the underlying high-order discretization. We prove stability and quasi-optimality on anisotropic, possibly curved, quadrilateral elements, and derive anisotropic a priori error estimates. Numerical experiments for $h$- and $hp$-refinement, including curved-domain examples, validate the theoretical results.
arXiv:2606.03268v1 Announce Type: new
Abstract: Dexterous manipulation learning has long been hindered by the high costs of data and training, as pure reinforcement learning typically requires large-scale interactive exploration and imitation learning depends on high-quality demonstrations that are expensive to collect. To address this problem, we propose EaDex, a multi-embodiment dexterous manipulation learning framework under low-cost demonstration conditions, which enables rapid generation of demonstration data and consequently reduces training time for efficient dexterous manipulation. At the data level, EaDex captures human hand motions using only a single RGB-D camera and constructs structured demonstration data through MANO-based hand modeling, data normalization, and motion retargeting. At the learning level, we introduce a contact-reward-based dynamic demonstration annealing mechanism, which guides early-stage exploration under demonstration and gradually transitions to autonomous optimization with accumulating contact rewards. Using our custom dataset, we evaluate EaDex on three dexterous hands and three articulated object-opening tasks, covering nine cross-embodiment manipulation settings, achieving a 55.3% relative improvement over the baseline without demonstration annealing. These results validate the effectiveness of the proposed low-cost demonstration pipeline and the dynamic demonstration annealing strategy for dexterous manipulation learning.
arXiv:2606.03783v1 Announce Type: new
Abstract: Reliable and affordable electricity supply remains a challenge for remote and regional communities, motivating the deployment of renewable-based microgrids supported by flexible storage and advanced planning methods. This paper proposes an integrated techno-economic framework for optimal microgrid design and robustness assessment, and applies it to a 1000-household residential community in Rockhampton, Queensland (Australia). The framework links time-series simulation, dispatch-based operation, and lifecycle costing to evaluate hybrid configurations comprising photovoltaic and wind generation, battery storage, diesel backup, grid exchange, and an optional hydrogen subsystem (electrolyzer--hydrogen storage--fuel cell). Key indicators include net present cost (NPC), cost of energy (COE), renewable penetration, energy purchased/sold, and emissions-related outcomes. To avoid conclusions that depend on a single set of assumptions, the study performs systematic sensitivity analysis across financial, technical and policy drivers: discount rate, technology capital costs, fuel price, load uncertainty, renewable resource variability, carbon pricing/emissions cost, and grid outage duration, supplemented by a no-hydrogen attribution case. The results demonstrate that several sensitivity dimensions induce nonlinear shifts in the optimal design, including breakpoints where capital-intensive renewable--storage expansion becomes economically preferable. The proposed framework enables transparent comparison of hydrogen-enabled and battery-centric solutions and provides planning guidance for resilient, low-emission community microgrids under Australian operating conditions.
arXiv:2606.03890v1 Announce Type: new
Abstract: Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence outside the current view. Existing benchmarks either evaluate offline over full videos or target events rather than spatial structure. We introduce OVO-S-Bench, a fully human-annotated benchmark for streaming spatial intelligence, comprising 1,680 questions over 348 source videos. Annotation involves 12 trained annotators, each also serving as a blind cross-reviewer, across roughly 804 person-hours of multi-round quality assurance. Each question carries a query timestamp and an evidence interval, and at evaluation, the model sees only the prefix preceding the query. Questions span four levels of increasing abstraction: instantaneous egocentric perception, spatiotemporal context tracking, spatial simulation and reasoning, and allocentric mapping. Across 38 proprietary and open-source MLLMs, Gemini-3.1-Pro trails human experts by 27 points, 59.2 vs. 86.6, with allocentric mapping as the dominant bottleneck. Notably, streaming and spatially fine-tuned MLLMs underperform their own backbones. We further find that chain-of-thought reasoning amplifies spatial errors when ungrounded in the stream. By exposing these limitations, OVO-S-Bench establishes a demanding testbed for next-generation streaming spatial MLLMs.
arXiv:2606.03893v1 Announce Type: new
Abstract: Accurate execution of preoperative plans in corrective femoral osteotomies remains challenging. Current techniques are limited by variable accuracy, invasiveness, and radiation exposure, with free-hand methods and patient-specific instrumentation (PSI) often requiring >30 and >6 fluoroscopic images, respectively. We present an integrated, electromagnetic tracking (EMT)-based navigation system for femoral osteotomies that minimizes dissection and intraoperative fluoroscopy. The system couples CT-based preoperative planning with one-time intraoperative C-arm calibration and accurate X-ray-to-CT registration from two fluoroscopic images acquired at initialization. This enables real-time, fluoroscopy-free EMT navigation of the saw blade and bone fragments relative to the preoperative plan, and is compatible with uniplanar and biplanar osteotomies. In a feasibility study using 18 synthetic femora, EMT guidance significantly outperformed free-hand execution in total angular error ($(3.05 \pm 0.75)^\circ$ vs.\ $(6.32 \pm 2.36)^\circ$, $p=0.031$), assuming the same minimal surgical exposure for both. No EMT-guided trials exceeded the >5{\deg} clinical threshold, whereas free-hand produced 4 outliers of 6 trials. The system achieved statistical equivalence ($\pm 2^\circ$, $\pm 2,\text{mm}$) to PSI for total angular ($p \le 0.02$) and total translational ($p=0.048$) errors, with no significant differences in user questionnaire scores. By transferring preoperative plans using only two fluoroscopic images while matching PSI accuracy without additional surgical exposure, the proposed system motivates subsequent cadaveric and clinical validation.
arXiv:2606.02617v1 Announce Type: new
Abstract: In decaying magnetohydrodynamic turbulence, energy can be transported from small to large scales, known as inverse transfer. We explore the mechanism behind this phenomenon using shell-to-shell transfer functions. Independent of magnetic net-helicity, large magnetic scales receive energy directly from the integral scale in both the magnetic and kinetic reservoirs, leading to increasingly non-local transfer for larger receiving scales. The resulting rate of energy increase in each receiving scale is proportional to its energy, resulting in self-similar, multiplicative growth. Even though the system is magnetically dominated, contributions from kinetic-magnetic and magnetic-magnetic energy-exchange are similar in magnitude. In the case of vanishing net-helicity, transfer functions between the positively and negatively helical parts of the field are computed. We find that inverse transfer only occurs within each helical sector, not across them. Our findings are consistent with the theory underlying the conservation of the Hosking integral, which explains inverse transfer as merging of local magnetic islands with equal-signed helicity.
arXiv:2510.15780v2 Announce Type: replace-cross
Abstract: Artificial intelligence (AI) is increasingly used to support renewable energy forecasting and grid operations. As renewable penetration grows, reliable probabilistic forecasting is becoming essential for managing uncertainty and supporting risk-aware operational decision-making. However, these forecasts often suffer from miscalibration due to temporal variability, changing weather conditions, and heterogeneous operating regimes. In many real-world settings, renewable energy forecasts are provided by external sources, vendors, or independently trained systems, making retraining infeasible because of limited model access or computational constraints. This creates a need for efficient and model-agnostic methods that can improve forecast reliability after they are produced. This paper presents Context-Aware Conformal Prediction (CACP), a framework for calibrating renewable energy forecasts. The proposed method relies on a weighting mechanism during the calibration procedure which assigns higher weights to historical observations that are more similar to the target forecasting condition. This enables adaptive prediction intervals that reflect local uncertainty regimes without requiring access to, or retraining of, the underlying forecasting model. Experiments are performed on a large-scale dataset from National Renewable Energy Laboratory (NREL) day-ahead solar forecasting, covering multiple systems including MISO, ERCTO, and SPP. The results show that CACP improves the reliability-efficiency tradeoff at both site and system levels compared to NREL's base forecasting model and the other conformal prediction baselines. These results suggest that CACP can serve as a practical reliability-enhancement layer for trustworthy AI-enabled renewable energy forecasting and operational decision support.
arXiv:2606.03928v1 Announce Type: new
Abstract: Reasoning models improve accuracy through extended chains of thought, but their long outputs create a memory and compute bottleneck. KV cache eviction methods reduce this cost by evicting unimportant key-value pairs from the cache, yet they often yield worse accuracy than selection-based sparse attention alternatives, which keep the full KV cache. We identify key factors crucial to KV cache eviction accuracy. First, a small fraction of value states have abnormally large magnitudes, and evicting them causes catastrophic failure where models enter repetitive reasoning loops. Second, introducing stochasticity during eviction improves accuracy by increasing cache diversity. Based on these findings, we propose Value-aware Stochastic KV Cache Eviction (VaSE), a training-free recipe that protects large-magnitude value states and promotes diverse eviction decisions. Across six reasoning tasks, Qwen3 models using VaSE with 4x KV cache compression yield higher average accuracies than SOTA selection method at the same sparsity, while outperforming the strongest eviction method by more than 4%. Overall, VaSE bridges the gap between efficiency and accuracy, supporting FlashAttention2 and enabling a static memory footprint for reasoning models.
R2IF: Aligning Reasoning with Decisions via Composite Rewards for Interpretable LLM Function Calling
arXiv:2604.20316v2 Announce Type: replace
Abstract: Function calling empowers large language models (LLMs) to interface with external tools, yet existing RL-based approaches suffer from misalignment between reasoning processes and tool-call decisions. We propose R2IF, a reasoning-aware RL framework for interpretable function calling, adopting a composite reward integrating format/correctness constraints, Chain-of-Thought Effectiveness Reward (CER), and Specification-Modification-Value (SMV) reward, optimized via GRPO. Experiments on BFCL/ACEBench show R2IF outperforms baselines by up to 34.62% (Llama3.2-3B on BFCL) with positive Average CoT Effectiveness (0.05 for Llama3.2-3B), enhancing both function-calling accuracy and interpretability for reliable tool-augmented LLM deployment.
arXiv:2606.02618v1 Announce Type: new
Abstract: We present Cognitive Loop via In-Situ Optimization (CLIO), an agent that couples a continuously-updated belief-state graph with a recursive plan-then-act loop. The result is a reasoning agent that can contribute something qualitatively different, which we term \emph{calibrated deference}: the capacity to recognize when its own tools or assumptions are failing, to adapt its strategy in response, and to generate mechanistic hypotheses that guide experimental revision. We tested CLIO in a closed-loop human-AI campaign to design an aqueous organic redox flow battery (AORFB) negolyte, with CLIO leading proposal and interpretation in close partnership with chemists who synthesized, characterized, and weighed in on design choices. Across 17 candidates over three rounds, CLIO converged on a top phosphonate candidate; characterization confirmed a 130~mV improvement in redox potential over the literature baseline. Characterization then revealed unexpectedly poor electrochemical reversibility -- a regression no property predictor had flagged. CLIO generated competing mechanistic hypotheses, prioritized discriminating diagnostics, traced the failure to phosphonate-potassium ion pairing, and prescribed a sulfonate replacement. The resulting compound showed substantially improved electrochemical reversibility and maintained a 90~mV improvement in redox potential, closing the design-make-test-redesign loop.
arXiv:2606.03182v1 Announce Type: cross
Abstract: The stochastic incompressible Navier-Stokes equations on $\TT^3$, completed by the fluctuation-dissipation noise, have a Fokker-Planck generator that decomposes into a self-adjoint Ornstein-Uhlenbeck (dissipative) part and an antisymmetric (convective) part. We prove two results about this generator. First, the logarithmic Sobolev inequality holds with the same optimal constant as the pure Ornstein-Uhlenbeck operator, $c_\mathrm{LSI} = \nu\lambda_1$ (where $\nu$ is the viscosity and $\lambda_1$ is the smallest nonzero eigenvalue of the Laplacian on $\TT^3$), independent of the number of retained Fourier modes. Second, the full semigroup is hypercontractive with the same rate as the Ornstein-Uhlenbeck semigroup. Both results follow from a single structural property: the convective generator is antisymmetric in $L^2(P_\mathrm{eq})$ (where $P_\mathrm{eq}$ is the Gibbs measure), and therefore contributes nothing to the Dirichlet form or the $L^q$ norm evolution. The antisymmetry is a consequence of two properties of the incompressible Navier-Stokes nonlinearity: energy conservation and phase-space volume preservation (the Liouville property). These are the same properties that underpin the fluctuation-dissipation theorem for the nonlinear Navier-Stokes equations.
arXiv:2606.00096v2 Announce Type: replace
Abstract: Visual agents employ external visual tools within visual chains of thought to incorporate fine-grained evidence. While prior work has mainly studied these tools in visual search tasks, their role in more complex visual reasoning remains underexplored. In this paper, we move beyond simple visual search tasks to investigate more challenging tasks, including 3D spatial reasoning and medical visual question answering, where agents must integrate tool-acquired local evidence with the global context. We identify a {tool-use collapse phenomenon: models progressively stop using tools while still achieving higher task accuracy. Moreover, we observe a clear asymmetry: (i) completely eliminating tool use degrades performance, whereas (ii) incentivizing tool use yields only marginal gains despite substantially increasing usage. We find that vanilla training and tool-use encouragement both reduce rollout diversity, explaining why higher tool use does not yield stronger reasoning performance. Motivated by these findings, we add an entropy regularization term to encourage diverse rollout exploration, achieving the best performance despite gradually declining tool usage. Overall, our findings suggest a training-time view of tools as scaffolding, where broader exploration over language generation and visual tool invocation improves reasoning despite tool-use collapse. Project page: https://scaffolded-exploration.github.io
arXiv:2606.03097v1 Announce Type: new
Abstract: Incorporating news into time series forecasting is appealing because news can reveal abrupt exogenous events that historical values alone cannot recover. However, existing LLM-based news-forecasting pipelines face two practical limitations: relevant news articles often exceed the model's context window, and iterative retrieval of supplementary news is typically unguided, leading to redundant updates and slow convergence. We address these issues with a novel framework that combines importance-aware news compression and process-level retrieval supervision. First, we train an importance reward model that estimates the forecasting utility of each article and uses this signal to allocate compression budgets during sequential pairwise fusion, preserving informative content within a fixed context limit. Second, we introduce a process reward model (PRM) that ranks multiple supplementary-news candidates conditioned on the current error profile and the history of previously selected articles, replacing one-shot blind retrieval with quality-controlled selection. Both components are trained offline using historical data with ground truth; inference uses the frozen filtering logic and compression modules without any reflection loop. Experiments on finance, energy, traffic, and bitcoin forecasting benchmarks show that our method improves prediction accuracy over strong baselines, significantly reduces the number of refinement iterations compared to the iterative baseline, and remains effective when relevant articles span thousands of tokens.
arXiv:2606.03617v1 Announce Type: new
Abstract: Digital Twins (DTs) are emerging as a cornerstone of the 6G vision, enabling real-time cyber-physical mirroring for smart manufacturing, autonomous vehicles, and remote healthcare. However, maintaining high-fidelity synchronization at scale demands an enormous and sustained uplink bandwidth, threatening both the feasibility and the energy efficiency of large deployments. We propose a Semantic-Aware DT Synchronization (SA-DTS) framework that radically redefines the synchronization pipeline: instead of streaming raw sensor or video data, a lightweight neural semantic encoder at the physical-world source extracts only task-relevant features and transmits compact semantic descriptors over the 6G air interface. At the DT replica, a paired decoder coupled with a dynamic Knowledge Graph (KG) reconstructs the full contextual state. A hierarchical KG partitioning strategy with an adaptive partition count $G = \lceil N / \log_2 N \rceil$ ensures that aggregate update overhead scales as $O(N \log N)$ rather than $O(N^2)$, making the framework viable for deployments with hundreds of simultaneously twinned entities. Extensive simulations on three canonical DT workloads -- industrial robot control, patient-monitoring, and vehicular platooning -- demonstrate bandwidth savings of up to 94%, end-to-end synchronization latency reductions of 87%, and KG-assisted state-reconstruction accuracy exceeding 97%, all under realistic 6G channel conditions. Empirical correlation confirms that the proposed Semantic Fidelity Score tracks standard task metrics (collision accuracy, alarm F1, spacing deviation) with Pearson $r > 0.97$ (95% CI: [0.961, 0.982]). Our results reveal that semantic communication is not merely a compression tool but a fundamental enabler for truly real-time, scalable DT ecosystems.
arXiv:2606.03966v1 Announce Type: cross
Abstract: The International Astronomical Union's Office of Astronomy for Development (IAU OAD) uses astronomy as a tool to address societal challenges and contribute to sustainable development. Building on more than a decade of project funding and implementation, the OAD has developed a portfolio of flagship projects that represent tested and scalable applications of Astronomy for Development across thematic areas including socio-economic development, science diplomacy, skills development, inequality reduction, and technology transfer. To support the growth and long-term sustainability of these initiatives, the OAD has established the Flagship Ecosystem, a framework built around four interconnected pillars: Resources, Training, Community, and Implementation.
This paper presents an overview of the OAD Flagship projects, the structure and components of the Flagship Ecosystem, and explores how it supports the translation of astronomy-based interventions into sustainable development outcomes. The ecosystem provides open-access resources, capacity-building opportunities, communities of practice, funding mechanisms, and evidence-generation activities that enable individuals and organizations to implement and scale astronomy-for-development initiatives. Grounded in principles of inclusivity, openness, sustainability, participation, and evidence-informed practice, the ecosystem aims to strengthen the global impact of Astronomy for Development while fostering collaboration across diverse sectors and regions.
arXiv:2606.02747v1 Announce Type: new
Abstract: Planning records define restrictions over geographic areas, but their source documents often provide only indirect spatial evidence rather than machine-readable boundaries. We introduce Plan2Map, a 208-case multimodal benchmark for document-grounded geospatial boundary reconstruction from UK planning records. Given only a source planning document, systems must reconstruct a valid geospatial boundary from notice text, schedules, map plates, map labels, and boundary annotations; the reference GeoJSON is held out for scoring. We propose GeoPlanAgent, a document-grounded, geospatial-tool-in-the-loop system that decomposes the task into evidence extraction, localisation, map registration, boundary segmentation, projection, and verification. On Plan2Map, GeoPlanAgent achieves 0.736 mean IoU and 0.904 median IoU, with 67.8\% of predictions at or above 0.8 IoU, substantially outperforming direct VLM-to-GeoJSON baselines. Diagnostic analysis shows that direct VLM prediction remains unreliable, while remaining errors are concentrated in localisation and map registration, and supervised boundary segmentation substantially improves pixel-level mask quality. Plan2Map provides a concrete testbed for multimodal geospatial reconstruction from public planning records. Project page: https://odeb1.github.io/Plan2Map_Project_Page/.
arXiv:2606.03310v1 Announce Type: new
Abstract: Understanding complex interactions between brain regions is critical for early neurodegenerative disease classification such as Alzheimer's Disease (AD) and Parkinson's Disease (PD). While graph-based models are widely used to analyze brain networks, most existing approaches primarily focus on pairwise interactions between directly connected nodes, limiting their ability to capture higher-order dependencies across multiple regions. Although hypergraph-based methods have been proposed to model higher-order relations, many rely on predefined hyperedges or restrict learning to hyperedge weights, reducing flexibility and limiting their capacity to capture multi-resolution structural patterns. In this regard, we introduce an adaptive multi-scale hyperedge learning framework, i.e., MuHL, which constructs hierarchical node features and dynamically learns high-order interactions through continuous hyperedge construction over multi-resolution graph signals. Extensive experiments on multiple brain network benchmarks demonstrate that MuHL consistently improves disease classification performance across different stages, and further identifies key regions of interest (ROIs) and their group-wise interactions from the learned hyperedges that are associated with disease progression, highlighting its potential as a powerful tool for brain network analysis in neurodegenerative disorders.
arXiv:2511.02304v2 Announce Type: replace
Abstract: We study learning multi-task, multi-agent policies for cooperative, temporal objectives, under centralized training, decentralized execution. In this setting, using automata to represent tasks assigned to agents enables breaking down a team-level objective into simpler, smaller sub-tasks. However, existing approaches remain sample-inefficient and are limited to the single-task case, requiring retraining policies for each new task. In this work, we present Automata-Conditioned Cooperative Multi-Agent Reinforcement Learning (ACC-MARL), a framework for learning task-conditioned, decentralized team policies. We identify challenges to the feasibility of ACC-MARL, propose solutions, and prove that our approach is optimal. We further show that learned value functions can be used to assign tasks optimally at test time. Experiments demonstrate emergent task-aware, multi-step coordination among agents, such as pressing a button to unlock a door, holding the door, and short-circuiting tasks.
arXiv:2606.03239v1 Announce Type: new
Abstract: LLM-based search agents are trained predominantly with outcome-only reward, leaving the search process itself unsupervised. This signal degenerates on outcome-homogeneous groups where all sampled trajectories share the same correctness, yielding zero within-group advantage and no gradient. Existing process supervision either trains a costly verifier or generates per-query rubrics that are inconsistent across queries and discarded after one use. We propose ARBOR (Adaptive Rubric Buffer for Online Reward), a reusable process-reward framework that maintains a rubric memory shared across queries. Query-local drafts induced from contrastive trajectories are admitted, consolidated into cross-query common rubrics, and retired as the policy evolves. A small active subset of common rubrics scores trajectories via sparse pairwise judging, and the resulting scores are added to the base reward, providing process-level gradient even when outcome reward is uniform. ARBOR consistently outperforms GRPO and DAPO baselines on four multi-hop QA benchmarks, raising average LLM-judge accuracy by up to 4.2 points and converting up to 42% of otherwise-zero-gradient training groups into informative ones.
arXiv:2606.03603v1 Announce Type: new
Abstract: World models and multimodal large language models (MLLMs) provide complementary capabilities for predicting future outcomes from static visual observations. World models can generate concrete visual rollouts of possible futures, while MLLMs can reason abstractly over questions, goals, and rules. However, generated rollouts are stochastic and may be visually plausible but task-incorrect, making it necessary to determine when visual simulation is useful, whether a rollout is credible, and how it should influence the final answer. We formulate this problem as controlled concrete reasoning, where a model learns to invoke, verify, and integrate visual future simulation alongside abstract reasoning. To study this setting, we construct two human-verified benchmarks, VRQABench for controllable spatial lookahead and OpenWorldQA for open-domain physical prediction, and propose Privileged-Future On-Policy Self-Distillation (PF-OPSD). During training, PF-OPSD uses ground-truth future videos and answers only as teacher-side privileged context to evaluate on-policy concrete-reasoning trajectories, while the deployable student never observes true futures at test time. Experimental results show that PF-OPSD outperforms baseline by 10.6% and 10.9% on VRQABench and OpenWorldQA, respectively, while increasing robustness to noisy or conflicting rollouts. Our code and dataset are available at https://github.com/yczhou001/PF-OPSD.
arXiv:2606.03059v1 Announce Type: new
Abstract: Microchannel flow boiling is a promising technique for micro-device thermal management, and understanding the bubble dynamics in microchannel flow boiling is important for the applications. Previous studies only focused on single, isolated bubbles, but the bubbles in microchannel flow boiling applications often exist as bubble trains, in which the bubbles interact with each other. Here, we investigate numerically vapor bubble trains in microchannel flow boiling by adopting the flow-focusing technique to form monodispersed bubbles in the upstream of the microchannel. With increasing the initial vapor-liquid volume ratio, the bubble frequency increases while the growth rate of the bubbles decreases because of the reduced bubble size. With increasing the heat flux on the wall or reducing the latent heat of the working fluid, the bubble train growth rate increases because of the increased vaporization rate. The vaporization of the fluid in the upstream causes the bubble expansion and accelerates the bubble movement in the downstream. The wall temperature and the Nusselt number fluctuate because of the periodic pass-through of bubbles.
arXiv:2606.03410v1 Announce Type: new
Abstract: Engineering diagrams pose a distinct challenge for vision-language models: unlike natural images or general documents, they encode information through dense spatial layouts, domain-specific symbols, and cross-references between visual callouts and structured parts tables. Despite their centrality to service, repair, and design workflows, there is no public benchmark for measuring VLM capabilities in this domain; existing datasets primarily focus on flowcharts, scientific figures, or business documents. To address this gap, we introduce Enginuity, the first open dataset and benchmark for evaluating VLMs on complex engineering diagrams. We define two tasks over a corpus of U.S. military service and repair manuals: structured parts-table extraction (Task 1) and free-form visual diagram question answering (VQA)(Task 2) for benchmarking. We evaluate four frontier VLMs (GPT-5.2 Chat, Claude Opus 4.7, Gemma 4, Qwen3-VL-32B-Instruct) under zero-shot and chain-of-thought prompting. On Task 1, models reach Recall@all of 0.61-0.87 but Token F1pen of only 0.03-0.18, exposing a systematic gap between part identification and description fidelity. Task 2 reveals a consistent factual-reasoning gap across all models. A supporting analysis shows that token-overlap metrics under-report model capability on technical descriptions by 2-6x relative to semantic similarity, motivating LLM-as-judge calibration for domain-specific evaluation. We release the dataset, annotations, evaluation harness, and per-sample model outputs to support a reproducible study of VLM capability on engineering content.
arXiv:2606.03374v1 Announce Type: new
Abstract: We present eMEM (Embodied Memory), a hybrid graph-based memory system for embodied agents operating in physical environments. Current agent memory architectures, such as Generative Agents, MemGPT, and A-MEM, treat memory as text streams or knowledge graphs, but embodied agents require memory that is simultaneously searchable by meaning, space, and time. eMEM fills this gap with a multi-index architecture (SQL ITE for structured storage, hnswlib for approximate nearest neighbour semantic search, and an R-tree for spatial queries) unified behind a single graph model. A tiered consolidation pipeline transforms raw perceptual observations into compressed summaries, mirroring hippocampal-neocortical consolidation in biological systems. Ten agent-facing recall tools expose memory retrieval primitives, including concept-to-location resolution and cross layer recall, as first-class operations for LLM tool calling. The system is fully embedded and runs in-process alongside the agent. In addition we introduce eMEM-Bench v1, a benchmark we construct over ProcTHOR-10K scenes for embodied memory evaluation. The benchmark is organised explicitly around eight cognitive-psychology paradigms (DRM lures, pattern separation, pattern completion, source monitoring, context-dependent retrieval, long-horizon interference, serial position, and a foil augmented retention curve), each chosen so that the result is interpretable against the broader memory-systems literature in humans and prior agent-memory systems; a level of diagnostic that surface-task benchmarks like LoCoMo or OpenEQA cannot provide. eMEM scores 80.8 weighted mean over 988 probes, with a flat retention curve at ceiling from 1 h to 1 yr of simulated delay on room-unique items. We show that a pure RAG baseline (the flat_rag ablation) loses 30 pt on context dependent retrieval and 29 pt on DRM lure rejection, isolating the contribution of multi-layer storage and consolidation…