arXiv:2510.06717v2 Announce Type: replace
Abstract: Large language models (LLMs) have been widely applied to knowledge-driven decision-making for automated vehicles due to their strong generalization and reasoning capabilities. However, the safety of the resulting decisions cannot be ensured due to possible hallucinations and the lack of integrated vehicle dynamics. To address this issue, we propose SanDRA, the first safe large-language-model-based decision making framework for automated vehicles using reachability analysis. Our approach starts with a comprehensive description of the driving scenario to prompt LLMs to generate and rank feasible driving actions. These actions are translated into temporal logic formulas that incorporate formalized traffic rules, and are subsequently integrated into reachability analysis to eliminate unsafe actions. We validate our approach in both open-loop and closed-loop driving environments using off-the-shelf and finetuned LLMs, showing that it can provide provably safe and, where possible, legally compliant driving actions, even under high-density traffic conditions. To ensure transparency and facilitate future research, all code and experimental setups are publicly available at github.com/CommonRoad/SanDRA.
Science Journals
arXiv:2510.15574v2 Announce Type: replace
Abstract: In this article, we design and analyze a hybrid high-order (HHO) finite element approximation for the solution of a nonlocal nonlinear problem of Kirchhoff type. The HHO method involves arbitrary-order polynomial approximations on structured and unstructured polytopal meshes. We establish the existence of a unique discrete solution to the nonlocal nonlinear discrete problem. We derive an optimal-order error estimate in the discrete energy norm. The discrete system is solved using Newton's iterations on the sparse matrix system. We perform numerical tests to substantiate the theoretical results.
arXiv:2510.21859v3 Announce Type: replace
Abstract: Electromagnetic (EM) methods, owing to their efficiency and non-invasive nature, have become one of the most widely used techniques in geological exploration. Nevertheless, data processing for these methods remains highly time-consuming and labor-intensive. With the remarkable success of deep learning, applying such techniques to EM methods has emerged as a promising research direction to overcome the limitations of conventional approaches. The effectiveness of deep learning methods depends heavily on the quality of datasets, which directly influences model performance and generalization ability. Existing application studies often construct datasets from random one-dimensional or structurally simple threedimensional (3D) models, which fail to represent the complexity of real geological environments. Furthermore, the absence of standardized, publicly available 3D geoelectric datasets continues to hinder progress in deep learning based EM exploration. To address these limitations, we present OpenEM, a large-scale, multi-structural 3D geoelectric dataset that encompasses a broad range of geologically plausible subsurface structures. OpenEM consists of nine categories of geoelectric models, spanning from simple configurations with anomalous bodies in half-space to more complex structures such as flat layers, folded layers, flat faults, curved faults and their corresponding variants with anomalous bodies. In addition, we provide a 3D model generator that enables fully controllable 3D model construction, allowing flexible and extensible augmentation of OpenEM. OpenEM provides a unified, comprehensive, and large-scale dataset for common EM exploration systems to accelerate the application of deep learning in electromagnetic methods. The complete dataset and 3D model generator is publicly available at https://doi.org/10.5281/zenodo.17141981.
arXiv:2607.10410v1 Announce Type: cross
Abstract: Reliable forecasting of several interrelated environmental variables - such as regional precipitation and temperature, or other correlated geophysical fields - across many locations calls for accurate predictions accompanied by trustworthy statements of their uncertainty. Modern deep-learning models forecast such variables accurately but usually report no uncertainty, and forcing them to output uncertainty through maximum likelihood tends to degrade their accuracy, especially when the variables are strongly correlated. Motivated by this tension, we develop TSCoNet, a two-stage convolutional-recurrent model coupled with a Gaussian copula that jointly forecasts multiple variables over space and time while quantifying predictive uncertainty. The method first learns accurate mean forecasts and then, holding the mean fixed, refines a shared representation to estimate the predictive variance, yielding calibrated prediction intervals after a standard recalibration, so that uncertainty is added without sacrificing point accuracy. We study the approach on simulated non-stationary spatial fields on the sphere and on a real dataset of monthly precipitation and temperature for fifty cities over 2000-2020. The model matches the accuracy of a strong deterministic forecaster while supplying calibrated prediction intervals that the deterministic model cannot, giving a single tool that provides both accurate point forecasts and reliable uncertainty for multivariate spatio-temporal data.
arXiv:2607.11095v1 Announce Type: cross
Abstract: Adversarial perturbations threaten machine learning classifiers, including variational quantum classifiers. We show that finite quantum measurement statistics (shot noise) act as a built-in defense against gradient-based test-time attacks whose cost scales unfavorably for the attacker. Because every gradient component must be inferred from repeated circuit executions under any unbiased gradient-estimation rule, white-box extraction consumes a dimension-dependent measurement budget that measurement grouping cannot remove in expressive circuits. Under stated assumptions, single-step attacks need at least quadratically many shots in the input dimension $d$, growing as $d^{5/2}$ under norm-concentration scaling, with a sufficient-budget analysis for iterative attacks via stochastic gradient Langevin dynamics. Simulations up to 784 input dimensions validate the law: the realized total budget is the $d^{5/2}$ geometric floor for plateau-mitigated models and grows as $d^{3.00}$ for the tested deep circuits, whose gradient norms decay with dimension absent barren-plateau mitigation; folding the measured gradient norm back in recovers the parameter-free $d^{3/2}$ shot-noise geometry. Against a matched classical baseline whose attack overhead is dimension-independent (the cheap-gradient principle of automatic differentiation), the quantum gradient cost ratio grows empirically as $d^{3.00}$, so the attacker's relative cost diverges as the model scales. Experiments on a 156-qubit IBM processor (ibm_boston, 4-qubit circuits, $d=12$) reproduce the effect: at matched budgets the device attack tracks the ideal within a few percent, with the high-shot gradient faithful to the exact one. The defense operates precisely when the forward map is classically hard to simulate: only then is a white-box attacker denied the simulate-and-backpropagate shortcut and must pay the measurement cost we quantify.
Religion and Artificial Intelligence as Distributed Meaning Systems: A Naturalistic Conceptual Model
arXiv:2607.10011v1 Announce Type: new
Abstract: This paper develops a naturalistic account of religion and artificial intelligence as structurally similar distributed meaning systems. I argue that both emerge from the same underlying cognitive architecture: socially extended processes that offload interpretation, norm-guidance, and world-model construction into external symbolic environments. Drawing on work in distributed cognition, cultural evolution, and philosophy of mind, the paper proposes a conceptual model showing how meaning is generated, stabilised, and transmitted through recursive interactions between agents and their informational ecologies. Religion is analysed not as a set of beliefs but as a cognitive-ecological system that scaffolds coordination, normativity, and shared interpretation. Contemporary AI systems are shown to instantiate analogous functions, operating as high-bandwidth, algorithmically mediated environments that shape reasoning, attention, and social meaning-making. The model explains how both systems create epistemic compression, reduce cognitive load, and generate shared frameworks that guide behaviour. It also clarifies the conditions under which AI systems can become culturally entrenched meaning authorities. The contribution is conceptual: a unified framework for understanding religion and AI as parallel forms of distributed cognitive machinery. This reframing opens new pathways for analysing artificial agents not as isolated tools but as components in evolving socio-cognitive ecologies.
arXiv:2607.10069v1 Announce Type: new
Abstract: Semantic caching defines answer reuse on embedding similarity: two utterances share a stored answer when a similarity score clears a threshold, with no notion of authorization, versioning, or of what makes two demands the same. This note changes the object on which reuse is defined: in a governed domain, reuse should operate on a mathematically characterized quotient of resolved conversational demands, not on a similarity heuristic. Three independently defined relations on resolved utterances -- reading identity, resolution identity, and reuse identity -- form a refinement chain, strict under realized nondegeneracy conditions checkable on deployment logs; the pipeline's outputs are invariant along the chain, and reuse identity is exactly the kernel of the resolution map into the governed answer partition, so the reuse quotient is the utterance-side object that partition induces, not a relabeling of it. Reuse identity licenses the governed query key and its certified answer space; reuse of a particular answer requires resolution identity or an applicability certificate. The supporting layer is stated at exactly the strength proved: exact-denotation normal forms; join aggregation as a design operator, with closure-stable cells characterizing no-escape; total computability of the full pipeline relative to an untrusted proposal layer; policy admissibility for arbitrary proposers -- and provably not factual grounding or intent fidelity; and elicitation terminating after finitely many informative replies, sound under target consistency.
arXiv:2607.10534v1 Announce Type: new
Abstract: Large language model (LLM) agents are increasingly extended through Agent Skills, reusable artifacts that package natural-language metadata, procedural instructions, and execution-time resources for runtime use. As open-source skill marketplaces expand, users and agents increasingly rely on brief metadata to select third-party skills, making it difficult to detect inconsistencies between a skill's description and its true behavior, a problem we call cross-layer misalignment. To address this issue, we propose Progressive Loading-Aware Hierarchical Contrastive Learning (PL-HCL), an LLM-based framework that detects misalignment by modeling the layered structure of Agent Skills and learning cross-layer consistency. Using a normalized corpus of over 264,000 open-source skills and a human-verified challenge set, PL-HCL improves Macro-F1 from approximately 0.45 for unadapted baselines to 0.87-0.89 across evaluated LLM backbones. This approach offers an effective screening tool for users and operators, as well as design principles for detecting inconsistencies in layered digital artifacts.
arXiv:2512.03988v2 Announce Type: replace
Abstract: Consumer-grade smartwatches offer a new option for personalized health monitoring for general consumers, as cardiovascular diseases continue to prevail as the leading cause of global mortality. The development and validation of reliable cardiovascular monitoring algorithms for these consumer-grade devices requires realistic biosignal data from diverse sets of participants. However, the availability of public consumer-grade smartwatch datasets with synchronized cardiovascular biosignals remains limited, and existing datasets often lack rich demographic diversity in their participant cohorts, potentially leading to biased algorithm development. This paper presents HEART-Watch, a multimodal physiological dataset of synchronized wrist-worn Google Pixel Watch electrocardiogram (ECG), photoplethysmography, and accelerometer signals from a diverse cohort of 40 healthy adults across three physical states - sitting, standing and walking - alongside reference chest ECG. Intermittent upper arm blood pressure measurements and concurrent biosignals were collected as an additional biomarker for future research. The motivation, methodology, and initial analyses of results are presented. HEART-Watch is intended to support the development and benchmarking of robust cardiovascular algorithms on consumer-grade smartwatches across diverse populations.
arXiv:2502.06469v3 Announce Type: replace
Abstract: This paper proposes a stochastic model predictive control method for linear systems affected by additive Gaussian disturbances that optimizes over disturbance feedback matrices online. Closed-loop satisfaction of probabilistic constraints and recursive feasibility of the underlying convex optimization problem is guaranteed. Optimization over feedback policies online increases performance and reduces conservatism compared to fixed-feedback approaches. The central mechanism is a finitely determined maximal admissible set for probabilistic constraints, together with the reconditioning of the predicted probabilistic constraints on the current knowledge at every time step. The proposed method's applicability is demonstrated on a building temperature control example.
arXiv:2508.16583v2 Announce Type: replace
Abstract: The Walker-Anderson half-space penetration model has been successfully used for the rapid, efficient calculation of penetration of walls by rigid and eroding rods. These models align well with detailed simulations for thick targets; however, existing extensions for finite targets struggle to accurately capture nose-tail velocity profiles in thinner targets. For stack-ups of thin-walled targets, this deficiency results in mischaracterized rod-erosion relative to hydrocode or experimental predictions. In this work, we leverage insights from detailed hydro-code simulations to propose an updated modification to the Walker-Anderson model to correctly account for wave propagation within a given target. This addition improves results for thin targets while retaining good behavior for thick targets with zero additional model parameters. Our updated model exhibits strong agreement with detailed simulations for targets with multiple thin walls.
arXiv:2607.11490v1 Announce Type: cross
Abstract: Neutral Atom Quantum Computing (NAQC) is an emerging modality for scalable quantum computation, valued for its long coherence times and the naturally identical atomic qubits. However, one of the main drawbacks is its slow execution rate, dominated by lengthy classical processing tasks, such as fluorescence imaging, cooling, and atom rearrangement. We address this bottleneck with AtomFlow, a field-programmable gate array (FPGA)-based control architecture that consolidates fluorescence-image analysis and a newly developed atom-rearrangement algorithm onto a single Zynq UltraScale+ device. By co-locating the two stages on the same board and emitting rearrangement moves in a streaming fashion as soon as they are computed, AtomFlow eliminates the round-trip latency of conventional host-mediated pipelines. Evaluated on a 16x16 atom array, AtomFlow achieves an end-to-end latency of 25.3 ms with a first-move latency of 4 ms and an average move generation of 1 ms. Furthermore, our scalability analysis demonstrates that the architecture can readily support larger atom arrays within a single-board resource budget.
arXiv:2512.07901v3 Announce Type: replace
Abstract: Von Neumann founded both game theory and the theory of self-reproducing automata, but the two programs never merged. This paper provides the synthesis. The Theory of Strategic Evolution analyzes strategic replicators: entities that optimize under resource constraints and spawn copies of themselves. We introduce Games with Endogenous Players (GEPs), where lineages (not instances) are the fundamental strategic units, and define Evolutionarily Stable Distributions of Intelligence (ESDIs) as the resulting equilibrium concept.
The central mathematical object is a hierarchy of strategic layers linked by cross-level gain matrices. Under a small-gain condition (spectral radius less than one), the system admits a global Lyapunov function at every finite depth. We prove closure under meta-selection: adding governance levels, innovation, or constitutional evolution preserves the dynamical structure. The Alignment Impossibility Theorem shows that unrestricted self-modification destroys this structure; stable alignment requires bounded modification classes.
Applications include AI deployment dynamics, market concentration, and institutional design. The framework shows why personality engineering fails under selection pressure and identifies constitutional constraints necessary for stable multi-agent systems.
arXiv:2512.08224v2 Announce Type: replace
Abstract: Bacteriophage-bacteria interactions are central to microbial ecology, influencing evolution, biogeochemical cycles, and pathogen behavior. Most theoretical models assume static environments and passive bacterial hosts, neglecting the joint effects of bacterial traits and environmental fluctuations on coexistence dynamics. This limitation hinders the prediction of microbial persistence in dynamic ecosystems such as soils and oceans. Using a minimal ordinary differential equation framework, we demonstrate that environmental fluctuations can suppress destructive oscillations through resonance, promoting coexistence where static models otherwise predict collapse. Counterintuitively, we find that lower bacterial growth rates are helpful in enhancing survival under high infection pressure, elucidating the observed post-infection growth reduction. Our studies highlight bacterial hosts as active builders of ecological dynamics and environmental variation as a potential stabilizing force. Our findings thus bridge a theory-experiment gap and provide a framework for predicting microbial responses to environmental stress, which might have potential implications for phage therapy, microbiome management, and climate-impacted community resilience as well.
arXiv:2512.10719v3 Announce Type: replace
Abstract: End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning capabilities obtained from the large-scale pretraining. However, we find that current VLMs struggle to understand fine-grained 3D spatial relationships which is a fundamental requirement for systems interacting with the physical world. To address this issue, we propose SpaceDrive, a spatial-aware VLM-based driving framework that treats spatial information as explicit positional encodings (PEs) instead of textual digit tokens, enabling joint reasoning over semantic and spatial representations. SpaceDrive employs a universal positional encoder to all 3D coordinates derived from multi-view depth estimation, historical ego-states, and text prompts. These 3D PEs are first superimposed to augment the corresponding 2D visual tokens. Meanwhile, they serve as a task-agnostic coordinate representation, replacing the digit-wise numerical tokens as both inputs and outputs for the VLM. This mechanism enables the model to better index specific visual semantics in spatial reasoning and directly regress trajectory coordinates rather than generating digit-by-digit, thereby enhancing planning accuracy. Extensive experiments validate that SpaceDrive achieves state-of-the-art open-loop performance on the nuScenes dataset and the second-best Driving Score of 78.02 on the Bench2Drive closed-loop benchmark over existing VLM-based methods. Code is available at: https://github.com/zhenghao2519/SpaceDrive.
arXiv:2512.13068v2 Announce Type: replace
Abstract: This work analyzes the convergence of sums of the form $S_{\boldsymbol{\gamma}}(m)=\sum_{v\subseteq \mathbb{N}}\gamma_v m^{|v|}$ with product and order dependent (POD) weights $\gamma_v$. We establish that for a nonnegative sequence $\{\Upsilon_j\mid j\in \mathbb{N}\}$, $$\sum_{v\subseteq \mathbb{N}} |v|! m^{|v|}\prod_{j\in v} \Upsilon_j<\infty \text{ for all } m>0 \text{ if and only if } \sum_{j=1}^\infty \Upsilon_j<\infty.$$ We further characterize the growth of $S_{\boldsymbol{\gamma}}(m)$ when $\gamma_v=(|v|!)^{\sigma}\prod_{j\in v}j^{-\rho}$ and prove that $\log S_{\boldsymbol{\gamma}}(m)$ is of asymptotic order $m^{1/(\rho-\sigma)}$ when $\rho>\sigma\geq 0$. We subsequently generalize both the convergence criterion and the asymptotic order of $\log S_{\boldsymbol{\gamma}}(m)$ to smoothness-driven product and order dependent (SPOD) weights, while noting that a full necessary-and-sufficient analogue remains open. Finally, we apply our theory to quasi-Monte Carlo (QMC) integration, showing that interlaced polynomial lattice rules achieve a dimension-independent convergence rate without a commonly imposed assumption in the QMC literature.
arXiv:2607.11317v1 Announce Type: new
Abstract: Low-bit quantization makes small reasoning models inexpensive to deploy but can degrade their chains of thought. This motivates decoder-side monitors that intervene when generation becomes unreliable. We show that a natural candidate, the centered token log-probability increment $\log p(w_t)+H_t$, is the wrong observable for this purpose. Under the model's own sampling law it is a mean-zero martingale by construction, so it measures sampling self-consistency rather than trajectory health and is nearly silent during confident repetition, where both $\log p(w_t)$ and entropy are close to zero. We introduce a training-free decoding controller that combines (i) a degeneration-aware alarm score fusing token uncertainty with explicit verbatim repetition and (ii) a calibrated e-process-inspired sequential detector. The raw product process is Ville-valid under a conditional-mean null, while the deployed CUSUM-floored statistic is treated as an empirical change detector because the score is history-dependent and autocorrelated. On GSM8K with DeepSeek-R1-Distill-Qwen-1.5B in FP16 and INT4, calibration turns a monitor that fires on 93--95% of generations into a selective detector of failing traces ($\phi \approx 0.3$, precision $\approx 0.6$ against a 0.38 base rate). In this pilot, the controller reduces measured verbatim-degeneration signals and yields a positive but statistically inconclusive INT4 accuracy change from 63% to 69% (paired McNemar $p=0.18$, $n=100$), at a 28% token-budget cost. We also find that non-termination, rather than looping, is the dominant failure mode on GSM8K. The main contribution is methodological: an explanation of why centered token log-probability is inadequate for decoder monitoring and a calibrated, cautiously evaluated replacement.
arXiv:2607.11650v1 Announce Type: cross
Abstract: Removable adhesive systems such as 3M Command strips are designed to support substantial loads while allowing clean, damage-free removal from the substrate. These systems rely on a highly extensible adhesive strip that bonds strongly during use but releases when stretched, causing the adhesive layer to elongate and progressively debond from the surfaces. A central challenge in the design of stretch-release adhesives is therefore to maximize load-bearing capacity while minimizing the force required for removal. This study investigates the finite-deformation mechanics governing both load support and tape release in a hyperelastic stretch-release adhesive system, with particular focus on the 3M Command tape geometry. Explicit analytical expressions are derived for the energy release rate of interfacial cracks under both load-bearing and release conditions and are validated against $J$-integral evaluations from finite element simulations. The results show that the ratio of maximum supported load to release force scales linearly with the ratio of bonded length to adhesive thickness, which is typically very large. We also investigate geometry-driven alternating crack propagation between the backing and substrate interfaces, governing tape removal, by analytical solutions and simulations. Parametric studies of competing interfacial fracture toughnesses produce failure envelopes that provide a predictive framework for estimating release forces and unstable crack propagation in multilayer stretch-release adhesive systems.
arXiv:2506.21833v2 Announce Type: replace
Abstract: Forward-mode automatic differentiation (FmAD) and zero-order (ZO) optimization are increasingly proposed as memory-efficient, backpropagation-free alternatives for large language model (LLM) fine-tuning, yet their benefits are typically evaluated only against standard backpropagation (BP), omitting memory-efficient variants such as activation checkpointing. We present a unified theoretical and empirical comparison of BP, checkpointed BP, FmAD, and ZO for LLM and vision-language model training, showing that while FmAD and ZO reduce activation memory, they trade memory for higher computational cost and longer wall-clock time to convergence, resulting in lower accuracy and slower training, especially under constrained perturbation budgets. Across models, BP with checkpointing outperforms FmAD and ZO variants, including variance-reduced methods, achieving up to 31.1% higher accuracy, 34.8% faster convergence, and 3.8x fewer computations at comparable memory usage, while also revealing instability-related failure modes in FmAD and ZO. Overall, our results correct a one-sided benchmarking narrative by showing that memory-efficient methods entail fundamentally different trade-offs, and that ignoring these distinctions has led to misleading conclusions about LLM optimization in prior work. Our source code is available at {https://github.com/Astuary/Gradient_Estimation_Methods}.
arXiv:2607.11792v1 Announce Type: new
Abstract: Automatic speech recognition (ASR) has become a critical component of modern robotic systems because it is one of the most natural and intuitive ways for humans to interact with robots. A commonly used method is to directly use API services online. But is that all we can do? This article provides an overview of how ASR technologies are integrated into various intelligent robots and machines. We discuss the evolution of speech recognition from established approaches to state-of-the-art deep learning models, such as OpenAI's Whisper. We also list large-scale datasets and open source toolkits that have been widely used in both industry and academia. We structure the survey around ASR model families, deployment strategies in robotics (especially ROS-based, cloud-based, and hybrid solutions), and several real-world robotic platforms. Finally, we outline the challenges of deploying robust speech recognition in robots and discuss future directions, including multimodal interaction in diverse and dynamic environments. This paper can help social robotics researchers better navigate the emerging domain of language-based natural human-robot interaction.
arXiv:2607.10942v1 Announce Type: new
Abstract: Physical AI systems, such as autonomous vehicles and intelligent machines, require transformer-based perception models that satisfy stringent edge latency and energy constraints. However, heterogeneous edge-GPU deployment remains limited by underutilized hardware engines and accelerator-incompatible operators, causing fragmented execution and lower throughput per watt. This paper presents Heterogeneous Frame Dispatch Scheduling (H-FraDS), a hardware-aware frame scheduling methodology for transformer inference on a recent NVIDIA edge GPU. H-FraDS routes frames across the GPU and dual deep learning accelerator (DLA) cores using fixed dispatch ratios to improve utilization under latency and power constraints. To enable scheduling, incompatible transformer components are adapted for DLA execution by reshaping tensors, approximating error function (ERF) with tanh, and replacing layer normalization with bounded tanh. The adapted model maintains a 92% F1 score, with only a 2% reduction from the original. Optical flow accelerator (OFA) is further used for inference-side optical-flow estimation. To the best of the authors' knowledge, prior work has not addressed these combined issues. Using Swin Transformer for autonomous-driving perception, H-FraDS Balanced Dispatch (1:2) achieves 125.93 FPS, a 2.36x speedup over standalone adapted-DLA execution, 4.0 FPS/W, and approximately 24 ms DLA latency, satisfying 30 FPS real-time operation; the GPU-DLA-OFA case achieves a 2.02x DLA throughput speedup.
arXiv:2512.13840v3 Announce Type: replace
Abstract: We introduce MoLingo, a text-to-motion (T2M) model that generates realistic, lifelike human motion by denoising in a continuous latent space. Recent works perform latent space diffusion, either on the whole latent at once or auto-regressively over multiple latents. In this paper, we study how to make diffusion on continuous motion latents work best. We focus on two questions: (1) how to build a semantically aligned latent space so diffusion becomes more effective, and (2) how to best inject text conditioning so the motion follows the description closely. We propose a semantic-aligned motion encoder trained with frame-level text labels so that latents with similar text meaning stay close, which makes the latent space more diffusion-friendly. We also compare single-token conditioning with a multi-token cross-attention scheme and find that cross-attention gives better motion realism and text-motion alignment. With semantically aligned latents, auto-regressive generation, and cross-attention text conditioning, our model sets a new state of the art in human motion generation on standard metrics and in a user study. We will release our code and models for further research and downstream usage.
arXiv:2512.17776v5 Announce Type: replace
Abstract: Recent advances in large language models have enabled deep research systems that generate expert-level reports through multi-step reasoning and evidence-based synthesis. However, evaluating such reports remains challenging: report quality is multifaceted, making it difficult to determine what to assess and which criteria to use; LLM-based judges may miss errors that require domain expertise to identify; and because deep research relies on retrieved evidence, report-wide claim verification is also necessary. To address these issues, we propose DEER, a benchmark for evaluating expert-level deep research reports. DEER systematizes evaluation criteria with an expert-developed taxonomy (7 dimensions, 25 subdimensions) operationalized as 101 fine-grained rubric items. We also provide task-specific Expert Evaluation Guidance to support LLM-based judging. In addition to rubric-based assessment, we propose a claim verification architecture that verifies both cited and uncited claims and quantifies evidence quality. Experiments show that current systems produce structurally plausible, evidence-citing reports, but still struggle to fully satisfy expert-level user requests and achieve logical completeness. Beyond performance comparisons, DEER makes system strengths and limitations interpretable and provides diagnostic signals for improvement.
arXiv:2512.19311v2 Announce Type: replace
Abstract: This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the input of a prediction network at one training timestep is the corresponding ground-truth noisy data that is an interpolation of the noise and the data, and during testing, the input is the generated noisy data. We present a novel training approach, named MixFlow, for improving the performance. Our approach is motivated by the Slow Flow phenomenon: the ground-truth interpolation that is the nearest to the generated noisy data at a given sampling timestep is observed to correspond to a higher-noise timestep (termed slowed timestep), i.e., the corresponding ground-truth timestep is slower than the sampling timestep. MixFlow leverages the interpolations at the slowed timesteps, named slowed interpolation mixture, for post-training the prediction network for each training timestep. Experiments over class-conditional image generation (including SiT, REPA, and RAE) and text-to-image generation validate the effectiveness of our approach. Our approach MixFlow over the RAE models achieve strong generation results on ImageNet: 1.43 FID (without guidance) and 1.10 (with guidance) at 256 x 256, and 1.55 FID (without guidance) and 1.10 (with guidance) at 512 x 512.
arXiv:2512.21299v2 Announce Type: replace
Abstract: Analysing the dynamics of phase-changing liquid films is essential for enhancing the performance of thermal management systems. Still, direct simulation of the full governing equations is computationally expensive. To circumvent this limitation, I derived a weighted-integral boundary-layer (WIBL) model under long-wave assumptions, weak evaporation, and strong surface tension, also accounting for variable substrate heating. In the linear regime, the WIBL reproduces growth rates and the cutoff wavenumber of unstable modes with significantly higher accuracy than commonly used Benney-type models for Re<40, as compared to the Orr-Sommerfeld equations. The linear analysis further reveals a threshold separating streamwise- and spanwise-dominated instabilities in hanging films, arising from the competition between Kapitza and Rayleigh-Taylor mechanisms; the WIBL predicts this threshold accurately for small Re and inclination angles. In the nonlinear regime, with substrate heating that varies in both space and time, the WIBL model captures the evolution of free-surface thickness and temperature within approximately 6% of the original Navier-Stokes equations. Three-dimensional simulations show that a condensing film undergoes dry-out due to Kapitza instability, whereas unsteady substrate heating promotes spanwise momentum spreading, modifies wave dynamics, and prevents dry-out. The WIBL model provides a good level of accuracy at a low computational cost, enabling extensive parametric studies, nonlinear stability analyses, and the design of optimal substrate-heating control strategies.