Forskningsradar

Science Journals

Peer-reviewade publikationer — 56237 artiklar

MosaicLeaks:Privacy Risks in Querying-in-the-Open for Deep Research Agents
arXiv:2605.30727v1 Announce Type: new Abstract: Deep research agents increasingly combine private local documents with external tools like web retrieval, creating a privacy risk: an agent's external queries may leak sensitive information from its local context. This risk is amplified by the mosaic effect, where individual queries may appear harmless but become revealing in aggregate. We introduce MosaicLeaks, a benchmark of 1,001 multi-hop deep research tasks that chain private enterprise documents and a public web corpus, forcing agents to make external queries that depend on local information. We evaluate leakage with an adversary LLM that observes only the agent's external queries and attempts to infer private information at three levels: the agent's research intent, answers to specific private questions and verifiable claims about the enterprise documents. We find that models across families and sizes frequently leak at all three levels, that zero-shot privacy prompting reduces but does not eliminate leakage and that reinforcement learning for task performance alone worsens leakage. To address this, we propose Privacy-Aware Deep Research (PA-DR), an RL framework that combines situational rewards for task success with a learned privacy classifier to provide dense credit assignment over both per-query and mosaic-level leakage. Training Qwen3-4B-Instruct with PA-DR improves accuracy from 48.7% to 58.7% and reduces answer and full-information leakage from 34.0% to 9.9%.
Stochastic bifurcation analysis via polynomial chaos: consistency and convergence of branch-approximating solutions
arXiv:2605.31288v1 Announce Type: new Abstract: Parameter-dependent dynamical systems that exhibit bifurcations pose significant computational challenges, as traditional continuation methods require repeated, costly simulations across large ranges of parameter values to capture sudden qualitative changes in the solution. In this work, we propose a systematic approach to reconstruct the branches of the entire bifurcation diagram in a single numerical solver leveraging generalized Polynomial Chaos (PC) expansion. By treating the parameter as a random variable, we cast the deterministic parameter-dependent model in a weak stochastic form, and then use a Galerkin projection to recover bifurcation branches globally across the parameter domain without iterative pointwise continuation. We show that the resulting Galerkin system, in the non-uniqueness regime, produces many discrete algebraic roots that naturally split into two classes: highly oscillatory solutions and branch-approximating ones. We develop a rigorous theoretical framework that establishes consistency, proves convergence of the branch-approximating solutions to the true steady states, and guarantees uniqueness of the Galerkin solution under suitable assumptions. Finally, we confirm these theoretical results with numerical experiments on several parameter-dependent ordinary differential equations (ODEs), demonstrating the accuracy and computational efficiency of our single-run framework in capturing complex bifurcation diagrams for both scalar and vector-valued systems.
Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World Trustworthiness
arXiv:2605.30745v1 Announce Type: new Abstract: Large Vision-Language Models have achieved unprecedented success in zero-shot recognition by aligning visual features with broad semantic concepts. However, this semantic abstraction creates a critical vulnerability in open-world deployment: the ``Hubris of Semantics'', where models force-fit unknown anomalies into known categories with high confidence due to the lack of explicit negative knowledge. To address this \textit{Open-World Trustworthiness Paradox}, we propose \textbf{Immuno-VLM}, a bio-inspired framework that adapts the biological principle of \textbf{Immunological Negative Selection} to high-dimensional latent spaces. Departing from traditional Open-Set Recognition methods that rely on passive density estimation or inefficient pixel-space outlier generation, Immuno-VLM leverages the generative reasoning of Large Language Models to actively hallucinate ``Semantic Antibodies'', textual descriptions of near-distribution outliers (e.g., look-alikes, contextual anomalies) that effectively bound the decision space of known classes.Extensive experiments on ImageNet-1K and four challenging OOD benchmarks reveal that Immuno-VLM establishes a new state-of-the-art.
Apple-Peel Unfolding in Three and Four Dimensions: Spiral and Zonal Selection Rules
arXiv:2605.30373v1 Announce Type: new Abstract: Apple-Peel Unfolding is a greedy algorithm that selects the faces (or cells) of a polyhedron (or polytope) one at a time in a spiral order, producing a net analogous to peeling an apple in a single continuous strip. We define two face-selection rules -- RS (Spiral rule: minimum signed determinant, i.e.\ sharpest clockwise turn) and RZ (Zonal rule: maximum coordinate along the peeling axis) -- and systematically evaluate their unfolding success rates on (i)~the five Platonic solids, (ii)~the thirteen Archimedean solids, and (iii)~the six regular convex 4-polytopes. A principal contribution is a three-way classification of each solid as \emph{Perfect} (every starting pair yields a complete net), \emph{Possible} (at least one pair succeeds), or \emph{Impossible} (no pair succeeds), together with an equivariance argument showing that face-transitive solids are confined to the $0/100\%$ dichotomy. RZ achieves the highest success rates in most cases; for the regular 4-polytopes it is the only rule yielding non-zero results for the 120-cell, where it achieves a Perfect result (1,440/1,440 pairs). We note that \emph{ordering success} (completing the greedy traversal) and \emph{geometric validity} (no self-intersection in the 3D realization) are distinct: every 120-cell ordering produces a self-intersecting 3D net, so the 120-cell has zero valid 3D nets despite its Perfect ordering result. The 600-cell is Impossible under all rules tested.
Optimization of the light detection system of the ICARUS detector
arXiv:2605.31327v1 Announce Type: new Abstract: The ICARUS detector, a key component of the Short Baseline Neutrino (SBN) Program at Fermi National Acelerator Laboratory (FNAL), is a 600-ton Liquid Argon Time Projection Chamber (LArTPC) equipped with a Light Detection System (LDS) that uses 360 Hamamatsu R5912-MOD 8-inch photomultiplier tubes (PMTs), specifically designed to operate under cryogenic conditions ($\sim 87 \ K$). These PMTs feed the trigger signal to the readout, improve the spatial and timing resolution of the events, and contribute to cosmic rays mitigation. During operation at FNAL, a progressive degradation in the PMT gain was observed. We developed an experimental setup to investigate the temperature dependence of PMT performance. Gain measurements were carried out from room temperature to $-70 ^\circ C$ using an environmental chamber. The results show that, while the PMTs exhibit stable performance at room temperature, a significant and irreversible reduction in gain emerges at lower temperatures. Al though $-70 ^\circ C$ remains above the liquid argon temperatures, the trend clearly reveals a gain-sensitive degradation mechanism. A simplified physical model was developed to reproduce and interpret the observed behavior. Based on these findings, a series of mitigation strategies were implemented in the ICARUS detector to preserve PMT performance and ensure reliable operation under cryogenic conditions.
Quantum Walks for Chemical Reaction Networks
arXiv:2509.07890v2 Announce Type: replace-cross Abstract: Near a detailed-balance equilibrium, the perturbed mass-action dynamics of a chemical reaction network (CRN) map exactly onto an electrical-flow problem on the bipartite species-reaction graph: chemical potentials become electrical potentials, Onsager coefficients become conductances, and the instantaneous Gibbs free-energy consumption equals the dissipated electrical energy. We exploit this map to design quantum walk algorithms that decide species reachability, sample reachable species, approximate any individual steady-state reaction flux, and estimate the total Gibbs dissipation. The first three follow from standard electrical-flow quantum walks; the last is non-trivial because the chemical flow is not the minimum-energy electrical flow on the same graph. We resolve this via a new use of alternative neighbourhoods in multidimensional quantum walks, which forces the walker onto the mass-action flow whenever the network is $\sigma-M$ rigid. In an adjacency-matrix QRAM access model the algorithms achieve up to a quadratic speedup over classical methods -- for example $\Omega(n^{3/2})$ vs $\Omega(n^2)$ for reachability -- and dissipation-aware bounds tighten this further when the perturbation is concentrated.
Controllable Lung Nodule Synthesis via Histogram-Regularized Latent Diffusion Models
arXiv:2605.30631v1 Announce Type: new Abstract: While automated diagnosis systems have achieved remarkable success in computed tomography (CT)-based lung cancer screening, their development remains limited by the scarcity of diverse, annotated pulmonary nodule datasets. Diffusion-based generative models offer a promising strategy for data synthesis; however, many existing conditional approaches primarily optimize spatial reconstruction losses, which encourage voxel-wise similarity but may inadequately constrain lesion-level intensity distributions. As a result, these methods may produce over-smoothed texture profiles and underrepresent the distinct attenuation characteristics of different nodule subtypes, including solid, part-solid, and ground-glass nodules. To address this challenge, we propose a controllable latent diffusion model that synthesizes pulmonary nodules within full 3D CT volumes while accurately modeling nodule-specific intensity distributions. Specifically, rather than relying solely on spatial losses, we introduce a histogram-based regularization term that constrains voxel intensity distributions during the generative process. The model combines subtype, spatial mask, and Hounsfield unit (HU) histogram conditioning with the differentiable feature-space histogram regularization term to better align lesion-level intensity distributions, improving the visual plausibility and subtype consistency of synthesized nodules. Extensive experiments on lung CT data demonstrate that our framework achieves strong visual realism, validated through both quantitative metrics and a visual Turing test. Furthermore, when used for data augmentation, the generated nodules improve performance in downstream clinical tasks, particularly for underrepresented nodule subtypes, and show a potential benefit for subtype-informed malignancy classification.
When Entropy Is Not Enough: Multi-Modal Classification of Encrypted and Compressed Data Fragments
arXiv:2605.31337v1 Announce Type: new Abstract: Reliable identification of encrypted data fragments is essential in cybersecurity, with applications to ransomware detection, digital forensics, and large-scale data analysis. Distinguishing encrypted from compressed fragments is particularly challenging, as short fragments lack structural data and exhibit low statistical redundancy. Traditional statistical methods based on byte-level distributions show limited effectiveness on this task. Recent machine learning approaches improve performance by learning subtle patterns from raw bytes, but predominantly rely on single-modal representations, implicitly assuming that a single view of the data is sufficient for accurate classification. This paper shows that this assumption becomes a fundamental limitation in low-information settings, when only small fragments of data are available (512--2048 Bytes). We propose Triumvir, a multi-modal, uncertainty-aware ensemble architecture that integrates statistical, sequential, and spatial representations of raw byte fragments. Extensive experimental analysis demonstrates that Triumvir consistently outperforms state-of-the-art methods with gains of up to +4.5pp in binary and +6.4pp in multiclass classification. Ablation studies confirm that combining modalities is critical, yielding improvements of up to +5pp over partial configurations.
FLAG: Flow Policy MaxEnt-RL by Latent Augmented Guidance
arXiv:2605.30749v1 Announce Type: new Abstract: Maximum entropy reinforcement learning (MaxEnt-RL) enables robust exploration, yet practical implementations often restrict policies to simple Gaussians. While recent approaches incorporate expressive generative policies via importance-weighted supervised learning, they are prone to importance weight collapse, which limits their scalability in high-dimensional action spaces. Our key insight is to mitigate this limitation by localizing the sampling region, avoiding the weight degeneracy induced by importance sampling over the entire action space. To instantiate this insight, we introduce \textbf{FLAG} (\textbf{F}low policy with \textbf{L}atent-\textbf{A}ugmented \textbf{G}uidance). FLAG augments the state space with a flow latent variable and optimizes a provably consistent proxy MaxEnt-RL objective. We empirically demonstrate that FLAG enables expressive policy optimization with limited importance samples and scales to high-dimensional control tasks. Furthermore, FLAG achieves state-of-the-art performance across challenging benchmarks. Our project webpage: https://flag-rl.github.io/
SQEEZ: Energy-efficient Location Sharing for Mobile Ad Hoc Networks
arXiv:2605.31339v1 Announce Type: new Abstract: Periodic network-wide dissemination of node location data is crucial for shared situational awareness and collaborative mapping in mobile ad hoc and mesh networks for public safety, disaster relief, and military. A key challenge is to provide maximally accurate location information with minimal energy expenditure on part of the nodes. We present SQEEZ: a mechanism for reducing the Position Location Information (PLI) load that combines two orthogonal techniques: (1) adaptive suppression of location updates; and (2) temporal and inline compression of update packets. We describe the SQEEZ suppression and compression algorithms, analyze the tradeoff between location error and energy consumption, and introduce a new metric called \textit{Error-Penalized-Energy (EPE)} that normalizes the energy metric using the error incurred. Our simulation results show that, in the range of parameters studied, SQEEZ improves the EPE-efficiency and scalability in a 30-node random waypoint scenario by up to 4.4x and 2.3x respectively; and increases the EPE-efficiency by 7.5x in a 9-node real-world network trace. Compression provides larger improvements than suppression at high mobilities and vice-versa at low mobilities.
Haptic Sorter: A Unified Planning Framework for Online Shape Estimation and Real-Time Pose Inference
arXiv:2605.31352v1 Announce Type: new Abstract: Robotics manipulation usually assumes that the shape and pose of the object are known to the robot prior to motion planning. However, precise geometric information is not always available in practice, and pose inference suffers from sensor uncertainties and view occlusion. In this work, we propose a unified model-based geometric framework integrating robotic haptic perception, modeling, and manipulation planning. Our novelties involve: \textit{i)} Introducing Bayesian Optimization (BO) to guide the haptic exploration for object shape inference, where superellipses are used to approximate geometric boundary; \textit{ii)} Adaptive formulation of manipulation potential encoding object geometry for quasi-static robot-object interaction; \textit{iii)} Proposing an online Ordinary Differential Equation (ODE) for real-time pose inference based on model prediction and tactile feedback. We deploy our system on a 2D robotic sorting task, and vary object geometries to validate the robustness and generalizability of our framework in both simulation and a real-world multi-arm setup.
dashi: A Python library for Dataset Shift Characterization to Support Trustworthy AI Development and Deployment
arXiv:2605.31360v1 Announce Type: new Abstract: The Artificial Intelligence (AI) life cycle requires a thorough understanding of the underlying data dynamics for robust, safe and cost-effective AI development and use. Dataset shifts are defined as changes between train and test data distributions. Whether occurring over time (temporal) or across different sites (multi-source), they can severely degrade model performance and compromise data quality. This is particularly important in health AI, where the safety and fundamental rights of patients can be severely affected by uncontrolled shifts both at training and operational stages. While the theoretical foundations of covariate, prior, and concept shifts are well established, there is a lack of accessible and comprehensive software tools to perform their analysis. We introduce dashi, an open-source Python library designed for the exploration, quantification, and characterization of dataset shifts. dashi provides a dual approach: an unsupervised approach that leverages information geometry and non-parametric statistical manifolds to data variability characterization and analysis (e.g., Information Geometric Temporal plots and Multi-Source Variability metrics like Global Probabilistic Deviation and Source Probabilistic Outlyingness), and a supervised approach that quantifies and characterizes model performance degradation. Both unsupervised and supervised approaches work across user-defined temporal and domain/source batches. We demonstrate the utility of dashi on three simulated and real-world health AI case studies on gestational diabetes mellitus, COVID-19 and emergency medical dispatch. By providing interactive visual analytics and variability metrics, dashi supports trustworthiness of AI life cycle stages enabling robust and safe machine learning pipelines through the assessment of data coherence and AI performance.
Collaborative Navigation and Exploration with $\beta$-Sparse Gaussian Processes
arXiv:2605.26304v2 Announce Type: replace Abstract: Collaborative navigation of heterogeneous robots in unknown environments poses significant challenges due to sensing, communication, and computational limitations. In this work, a lead robot navigates toward a target while a mobile sensor robot (e.g., a drone) assists by transmitting information about its locally observed map under bandwidth constraints. We propose a framework that enables the sensor to jointly select its transmitted map points and navigation actions online, while also predicting unexplored regions of the environment. To this end, we present $\beta$-Sparse Gaussian Processes, a robust variational sparse Gaussian Process model for task-aware inducing point selection under cardinality constraints. Furthermore, we develop an action-selection strategy that balances task relevance with exploration. Simulations on Mars and Earth maps show that the framework can reduce path cost by 18% relative to no communication and decrease transmitted information by 76% compared to raw-data transmission baselines.
AIM: A practical approach to automated index management for SQL databases
arXiv:2605.31406v1 Announce Type: new Abstract: This paper describes AIM (Automatic Index Manager), a configurable index management system, which identifies impactful secondary indexes for SQL databases to efficiently use available resources such as CPU, I/O and storage. It has been validated on thousands of databases which support production systems. With AIM, the physical design of the database adapts itself to the changes in the workload.We lay out the end to end design of AIM while calling out the guarantees and tradeoffs associated with our design choices. Some of the salient features of AIM include fast convergence even while recommending wide composite indexes, reduced reliance on the query optimizer and a "no regression" guarantee for production workloads. Each index recommendation from AIM is accompanied with a metrics driven explanation, making it easier to verify machine driven changes.AIM is one of the few industrial strength index recommendation engines that is deployed on production databases at a large scale. The experimental results show that AIM is quick in identifying the most effective indexes and the resulting physical design is close to optimal.
Skill Availability and Presentation Granularity in Large-Language-Model Agents: A Controlled SkillsBench Study
arXiv:2605.31408v1 Announce Type: new Abstract: Skill documents provide procedural knowledge to large-language-model agents at inference time. This article studies whether the presentation granularity of controlled skill knowledge changes downstream task success. The experiment uses a pinned SkillsBench version, a 30-task domain-balanced subset validated by official oracle runs, two reasoning-enabled model configurations, six skill conditions, and five trials per task-condition-model cell. Skill availability is the clearest empirical signal. Relative to no skill, skill conditions increase task-mean pass rate by 26.7 to 36.0 percentage points for GPT-5.5 and by 18.0 to 26.0 percentage points for DeepSeek V4-Flash. The final data contain 1,800 rows, with 900 rows for each model. The task is the inference unit. Five trials are aggregated within each task-condition-model cell before paired contrasts are estimated over 30 tasks. The primary presentation contrasts are smaller and uncertain. Low-abstraction guidance differs from high-abstraction guidance by +0.7 percentage points for GPT-5.5 and -6.7 percentage points for DeepSeek V4-Flash, with both 95% bootstrap confidence intervals crossing zero. Adding one worked example to medium-abstraction guidance differs from the no-example variant by +0.7 and +1.3 percentage points. Mean-reward robustness checks preserve the same substantive conclusion. In this controlled subset, skill availability is associated with higher success than no skill, while the tested presentation-granularity changes yield small, uncertain, and model-dependent effects.
Chain-of-Thought and Compressed Looped Transformers: A Memory-Budget Separation
arXiv:2605.30757v1 Announce Type: new Abstract: Chain-of-thought prompting and looped Transformers both give a fixed model more test-time computation, but they differ in what they remember. Chain-of-thought stores intermediate state in generated tokens that remain in the context, whereas a looped Transformer carries state through recurrent hidden activations. We argue that this persistent mutable memory is a central resource for test-time reasoning. We compare three memory regimes, the compressed latent loop, the full sequence-state loop, and the chain-of-thought scratchpad. Our main result shows that a compressed loop is limited by the size of its recurrent state. Running the loop longer adds computation but does not by itself create a growing scratchpad, so a loop with a small recurrent state remains a small-space reasoner even when run for many steps. Under a standard complexity assumption, such loops cannot decide problems that are P-complete under logspace reductions, whereas polynomial-length chain-of-thought can. The separation is specific to compressed loops, as full sequence-state loops carry state at every input position and live in a memory-rich regime closer to explicit scratchpads. Controlled pointer-chasing and associative-recall sweeps illustrate this memory-budget view, with performance sensitive to whether the persistent-state budget matches the task's working-memory demand.
Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing
arXiv:2605.31367v1 Announce Type: new Abstract: Token mixing layers play a key role in how language models can learn and generate long-range dependencies. Their efficiency relies on the necessary trade-off between decoding speed and the memory requirements, along with the cache size. Considering causal generation, this paper explores new trade-offs thanks to a unified framework which separates two crucial features: (i) the direct influence of inputs on outputs in one generation step; (ii) the recurrent propagation of information through past outputs. This framework encompasses major architectures such as attention and state-space models, but also generalizes the recurrence equations by allowing each state to depend on multiple past states rather than only the immediate predecessor. By introducing structure, we design new recurrence patterns that provably achieve the desired complexity, while providing theoretical insights on their expressivity -- trading runtime for expressivity in a principled way. Empirical validation is performed on synthetic tasks, along with language modeling. Together, these results provide a unified toolkit for the understanding and design of efficient and expressive token mixers across model families.
Physics Enhanced Deep Surrogates for the Phonon Boltzmann Transport Equation
arXiv:2512.05976v3 Announce Type: replace Abstract: Designing materials with controlled heat flow at the nano-scale is central to advances in microelectronics, thermoelectrics, and energy-conversion technologies. At these scales, phonon transport follows the Boltzmann Transport Equation (BTE), which captures non-diffusive (ballistic) effects but is too costly to solve repeatedly in inverse-design loops. Existing surrogate approaches trade speed for accuracy: fast macroscopic solvers can overestimate conductivities by hundreds of percent, while recent data-driven operator learners often require thousands of high-fidelity simulations. This creates a need for a fast, data-efficient surrogate that remains reliable across ballistic and diffusive regimes. We introduce a Physics-Enhanced Deep Surrogate (PEDS) that combines a differentiable Fourier solver with a neural generator and couples it with uncertainty-driven active learning. The Fourier solver acts as a physical inductive bias, while the network learns geometry-dependent corrections and a mixing coefficient that interpolates between macroscopic and nano-scale behavior. PEDS reduces training-data requirements by up to 70% compared with purely data-driven baselines, achieves roughly 5% fractional error with only 300 high-fidelity BTE simulations, and enables efficient design of porous geometries spanning 12-85 W m$^{-1}$ K$^{-1}$ with average design errors of 4%. The learned mixing parameter recovers the ballistic-diffusive transition and improves out of distribution robustness. These results show that embedding simple, differentiable low-fidelity physics can dramatically increase surrogate data-efficiency and interpretability, making repeated PDE-constrained optimization practical for nano-scale thermal-materials design.
A scalable Ewald-free BIE framework for periodic Stokes flow via hierarchical proxy sums
arXiv:2605.30805v1 Announce Type: new Abstract: Particulate Stokes flow in confined, periodic geometries underlies a broad class of problems in biophysics, microfluidics, and the rheology of complex fluids. Boundary integral equation (BIE) methods are a natural tool for such problems, but existing periodization schemes rely either on periodic Green's functions, which are restrictive for complex confining geometries, or on free-space schemes that solve auxiliary proxy strengths alongside the surface densities in an extended linear system whose cost scales unfavorably in three dimensions. We present a BIE framework for three-dimensional particulate Stokes flow in periodic pipes with circular cross-sections, wall-bounded doubly-periodic, and triply-periodic geometries that uses only the free-space Green's function and avoids both Ewald summation and the extended linear system. Proxy sources placed on equivalent surfaces of the kernel-independent FMM (KIFMM) form the auxiliary basis, and contributions from far image boxes are captured by a hierarchical proxy sum made absolutely convergent by a net-force-zero compatibility condition. The resulting periodization precomputation depends only on the periodic-box geometry, independent of the kernel and of the surfaces inside the box, and is reused verbatim across the Stokeslet, stresslet, and rotlet. Combined with high-order adaptive surface discretizations, the method achieves high-order accuracy at $\mathcal{O}(N)$ cost with a single layer of image boxes in the near field. Numerical examples on dense polydisperse suspensions with thousands of particles and on flow through complex periodic channels, together with strong and weak scaling studies, demonstrate efficient performance on systems with millions of degrees of freedom on distributed-memory architectures.
Non-Equilibrium Thermodynamics of Black-Hole Coronae: QPOs, Turbulence, and Jets
arXiv:2512.09026v2 Announce Type: replace-cross Abstract: The variability of X-rays observed from accreting black hole systems, including quasi-periodic oscillations (QPOs), suggests a complex nonlinear dynamics in the corona. Here, we propose a new theoretical framework for this variability, based on non-equilibrium thermodynamics. In this model, coronal variability arises from feedback between a macroscopic oscillation of the plasma and the rate at which it is cooled by the inverse Compton scattering of soft photons from the disc. The "pair thermostat'' mechanism then allows the corona to act as a heat engine that extracts work cyclically from the underlying thermal disequilibrium between the low-entropy heating from the black hole and the high-entropy cooling by soft photons from the disk, in close analogy to the well-known $\kappa$-mechanism for pulsating stars. This coronal self-oscillation may explain QPOs without invoking an external periodic driving. Moreover, we argue that this mechanism can generate coronal turbulence and jets.
Subcritical transition to turbulence in buoyancy-driven flows with multiple hysteresis loops under quasi-one-dimensional confinement
arXiv:2605.31380v1 Announce Type: new Abstract: We present both static and quasi-static direct numerical simulations of Rayleigh-B\'enard convection in a quasi-one-dimensional domain, revealing for the first time a clear subcritical transition to turbulence in a buoyancy-driven flow. Within a narrow range of Rayleigh number (Ra), three coexisting flow states are identified: steady convection, oscillatory chaos, and intermittent turbulence. The transitions between these states are accompanied by abrupt jumps in both the Nusselt number (Nu) and Reynolds number (Re), the key global transport quantities in buoyancy-driven flows. Additionally, they exhibit pronounced hysteresis, forming three distinct hysteresis loops in the Nu-Ra plane: normal, reverse, and anomalous loops. More importantly, we show that the steady convection state is linearly stable against infinitesimal perturbations but can transition to intermittent turbulence when subjected to finite-amplitude disturbances, which is a defining hallmark of subcriticality. Thus, contrary to the prevailing view that the transition from convection to turbulence is supercritical, our results demonstrate that buoyancy-driven turbulence can emerge via a subcritical route, paving the way for a unified framework that describes instability mechanisms in both buoyancy-driven and shear-driven flows.
Multi-Turn Multi-Agent Dialogue for Collaborative Reconstruction Improves VLM Performance on Spatial Reasoning, But Only Barely
arXiv:2605.31387v1 Announce Type: new Abstract: Robots operating in diverse environments rely on visual input to interpret objects and spatial layouts. In human-collaborative tasks, they are expected to communicate this understanding through language. Vision-language models (VLMs) support robotic tasks involving visual interpretation, question answering, and instruction following, but their capabilities in collaborative dialogue tasks requiring spatial reasoning remain underexplored. We study this gap through a collaborative structure-building task that combines visual interpretation, grounding, language-guided interaction, and action generation. We develop a framework in which VLMs use dialogue to reconstruct a target structure from visual and textual inputs. We evaluate open-weight and closed VLMs across interaction settings, input modalities, and image representations. Results show that spatial reasoning over visual representations remains difficult for the evaluated VLMs. Detailed text representations of the target yield higher reconstruction success across modality conditions, while decomposed image representations improve performance. These findings reveal limits in visual spatial grounding and grounded instruction generation for collaborative VLM agents.
Is Memorization Helpful or Harmful? Prior Information Sets the Threshold
arXiv:2602.09405v2 Announce Type: replace-cross Abstract: We examine the connection between training error and generalization error for arbitrary estimating procedures, working in an overparameterized linear model under general priors in a Bayesian setup. We find determining factors inherent to the prior distribution $\pi$, giving explicit conditions under which optimal generalization necessitates that the training error be (i) near interpolating relative to the noise size (i.e., memorization is necessary), or (ii) close to the noise level (i.e., overfitting is harmful). Remarkably, these phenomena occur when the noise reaches thresholds determined by the Fisher information and the variance parameters of the prior $\pi$.
Transforming and Encoding FTS for SAT Solving: What Helps, What Hurts (Extended Version)
arXiv:2605.30563v1 Announce Type: new Abstract: Factored tasks are a classical planning representation that extends SAS+ with limited forms of disjunctive preconditions, conditional effects, and angelic nondeterminism. This allows for a more compact representation of tasks than traditional formalisms such as STRIPS or SAS+, and supports a wide range of task transformations. However, existing planning approaches for factored tasks have been limited to heuristic search methods. In this work, we investigate how to encode factored tasks in SAT. We propose several ways to encode the tasks, focusing on different strategies for translating the factored transition relation into propositional logic. We also analyze how to exploit parallelism at various levels in this setting and study the impact of common task transformations on the performance of SAT-based planners.
FSM-Net: An Efficient Frequency-Spatial Network for Real-World Deblurring
arXiv:2605.31400v1 Announce Type: new Abstract: Real-world image deblurring demands both high-fidelity restoration and computational efficiency, a balance existing methods often struggle to achieve. In this paper, we propose FSM-Net (Frequency-Spatial Multi-branch Network), a highly efficient solution that secured 2nd place in the NTIRE 2026 Challenge on Efficient Real-World Deblurring. FSM-Net pioneers a dual-domain approach: a novel Frequency Attention module explicitly recovers high-frequency structural details via FFT, while a Cross-Gated Vision E-Branchformer at the bottleneck captures global dependencies with linear complexity. To ensure robust convergence, we employ a progressive curriculum training strategy guided by a composite loss function (Multi-Scale Charbonnier, Structural Edge, and Frequency). Evaluated on the RSBlur benchmark, FSM-Net achieves an outstanding 33.144 dB PSNR with only 4.94M parameters and 159.35 GMACs (at 1920x1200 resolution). By effectively pushing the Pareto frontier of efficiency and quality, FSM-Net establishes a strong baseline for resource-constrained image restoration.