Forskningsradar

Science Journals

Peer-reviewade publikationer — 60005 artiklar

Exposing the Illusion of Erasure in Knowledge Editing for LLMs
arXiv:2606.23276v2 Announce Type: replace Abstract: Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying mechanisms remain poorly understood. In this work, we examine KE from an adversarial elicitation perspective, revealing that edited knowledge is often not fully erased and continues to surface, with consistent failures observed across diverse model architectures. To explain this behavior, we conduct a mechanistic analysis of popular KE methods. We show that low-rank updates do not overwrite existing knowledge but instead redistribute it within the model's representation space. Furthermore, we find that these methods act as targeted suppression mechanisms that reduce the likelihood of expressing original facts, rather than removing them from the model. Analysis of the loss landscape reveals that edited knowledge lies in narrow, anisotropic regions that are highly sensitive to perturbations, making them highly vulnerable to indirect prompting and adversarial attacks. By exposing these profound architectural vulnerabilities, our work proves that KE algorithms are inherently bypassable and motivates a fundamental reevaluation of how we deploy post-hoc updates in several LLM applications.
Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation
arXiv:2606.23743v2 Announce Type: replace Abstract: Modern video diffusion models achieve higher generation quality through scaling, but this also increases inference cost. Although many acceleration methods have been proposed, a central challenge is that the most effective acceleration strategy is highly instance-specific: a recipe that works well for one combination of model, hardware, and inference configuration often does not transfer to another. Different models vary in architecture, numerical sensitivity, and attention concentration patterns. Inference settings differ in spatial and temporal resolution and video duration, while hardware platforms differ in memory hierarchy, supported numerical formats, and kernel throughput. These factors create a large tuning space, making manual performance engineering costly. We present Sol Video Inference Engine, an agentic, native, training-free acceleration framework for video diffusion models. It organizes five broadly applicable techniques, cache, sparse attention, token pruning, quantization, and kernel fusion, into an agentic acceleration stack for instance-specific optimization. For a concrete deployment target defined by a model, hardware platform, and serving configuration, parallel skill agents optimize the implementation of each technique, an agent integrator composes them into a global acceleration stack, and a human validator provides feedback on generation quality. We instantiate this workflow on three video models with different sizes and architectures: 64B Cosmos3-Super, 22B LTX-2.3, and 2B SANA-Video. With little human effort, the full stack achieves more than 2x end-to-end acceleration while maintaining near-lossless VBench quality, demonstrating the effectiveness of the agent framework for video diffusion acceleration.
Ingredient-Level Food Image Segmentation for Nutrition Awareness
arXiv:2606.24059v2 Announce Type: replace Abstract: Food images often contain several visible ingredients, so assigning one dish label to an entire image hides important visual structure. This work studies ingredient-level semantic segmentation on FoodSeg103, where the model predicts an ingredient class for each pixel. Two SegFormer variants were fine-tuned and evaluated under a controlled setup: SegFormer-B0 as the smaller baseline model and SegFormer-B1 as the larger final model. Both models use ImageNet-pretrained MiT backbones with newly initialized 104-class output layers. On the held-out FoodSeg103 test split of 2,135 images, B0 achieved 0.7709 pixel accuracy and 0.2521 mean IoU, while B1 achieved 0.7929 pixel accuracy and 0.3204 mean IoU. B1 improved every saved test metric, including a +0.0683 absolute gain in mean IoU. The system also converts predicted masks into visible ingredient-area percentages, giving a simple visual composition summary of the predicted meal. This summary can serve as a first-pass nutrition-awareness cue by providing a visual alternative to detailed food tracking similar to plate-based meal guidance, but it is not a direct estimate of calories, macronutrients, food mass, volume, density, or true portion size.
FedUP: One-Shot Federated Unlearning via Centroid-Guided Plug-in Filters
arXiv:2606.24113v2 Announce Type: replace Abstract: Federated unlearning (FU) is critical for complying with legal mandates like the right to be forgotten in decentralized systems, yet current methods face a persistent dilemma between non-target knowledge loss and high request latency. To resolve these issues, we propose FedUP, a one-shot federated unlearning framework utilizing lightweight pluggable filters that act as a "knowledge funnel" to screen out target data while preserving original model performance. By freezing original model parameters and training filters at the server side using differentially private (DP)-protected class centroid samples, FedUP bypasses the need for multi-round client-server communication and complex retraining, reducing unlearning latency from minutes to mere seconds. Additionally, the framework's pluggable architecture ensures inherent reversibility, enabling the seamless restoration of forgotten knowledge by simply removing the filters. Extensive experiments on diverse image and text tasks demonstrate that FedUP effectively reduces non-target knowledge loss and achieves superior unlearning precision and efficiency across various scenarios. Code is available at: https://github.com/suows/FedUP-code.
Paleomagnetic signatures of core-mantle interactions inferred from top-heavy thermochemical geodynamo simulations
arXiv:2606.26042v1 Announce Type: new Abstract: The time-averaged geomagnetic field provides crucial insights into deep Earth dynamics and thermal core-mantle interactions. Paleomagnetic observations and numerical dynamo simulations are equivocal regarding the longitudinal structure of the time-averaged field, though the latter have often considered a generic buoyancy source, which may obscure distinct signatures of thermal and chemical buoyancy that arise near the equator and poles, respectively. In this study, we present a new suite of top-heavy geodynamo simulations, varying the relative strengths of thermal and chemical driving and comparing the resultant magnetic signatures to observational field models spanning centuries to tens of thousands of years. None of the spatially-averaged measures of field morphology and variability we tested could robustly distinguish between different levels of chemical driving or the presence of heterogeneous outer boundary heat flux. On the other hand, observational constraints requiring longitudinal variations in time-averaged inclination anomaly are readily matched by simulations with heterogeneous outer boundary thermal forcing, in contrast to those with homogeneous mantle heat flux. Longitudinal field structures are reduced, but not erased, by elevated chemical driving, which also promotes the formation and deepening of polar minima in the radial magnetic field. Our simulations indicate that both the strong heat flux heterogeneity and chemical driving in Earth's core are likely to result in small but persistent departures from the geocentric axial dipole approximation.
Navigating User Behavior toward Personalized Multimodal Generation
arXiv:2606.24196v2 Announce Type: replace Abstract: Modern AIGC pipelines deliver high-fidelity images and videos but presuppose a well-formed creation instruction, while end users rarely articulate visual details, leaving generators misaligned with user demand. We study personalized content generation, which turns a user's interaction history into an executable instruction for downstream synthesis, and identify two obstacles: behavior must be encoded in a form legible to language reasoning, and the model must acquire instruction-writing skill absent from both pretraining and behavior data. We propose NaviGen, which represents each item with a dual identifier coupling a collaborative code and a textual code as a behavioral substrate and a semantic bridge in one token stream. On this representation, a two-stage SFT+RL pipeline first distills preference reasoning and instruction writing from evolutionarily searched supervision, then aligns generation with user intent through hierarchical and self-consistent rewards. Experiments across product, game, and short-video domains show that NaviGen improves personalized image and video generation, strengthens next-item prediction, and yields more specific, relevant, and visually generatable instructions. Our code is released at: https://github.com/iLearn-Lab/NaviGen.
Color-Center-Compatible Freestanding Diamond Directional Couplers for Quantum Photonics
arXiv:2606.24576v2 Announce Type: replace Abstract: Freestanding all-diamond color-center photonics is a promising platform for optical integration of spin-based quantum defects. Within this geometry, we realize a key building block for quantum-network interconnects: a directional coupler that acts as an on-chip beam splitter. We design and simulate directional couplers with triangular cross sections using eigenmode and finite-difference time-domain simulations and target near-50:50 splitting at visible wavelengths. We fabricate the devices directly from bulk single-crystal diamond by angled oxygen reactive-ion-beam etching followed by a dry post-release hard-mask removal process. Room-temperature measurements at $\lambda_0\approx 637 \mathrm{nm}$ yield a mean coupling ratio of $C^\mathrm{meas}=46(16) \%$. Finally, we integrate SnV$^{-}$ centers into the nanophotonic structures and observe near-lifetime-limited optical linewidths and coherent optical Rabi oscillations without post-fabrication annealing, identifying the platform as a viable route towards integrated diamond quantum photonics.
Interfacial Spectral Memory as a State Variable for Finite-Depth Salt-Finger Exchange
arXiv:2606.26045v1 Announce Type: new Abstract: Thermohaline interfaces in the ocean are often treated through local double-diffusive favorability, yet finite interfaces can also inherit roughness from prior waves, stirring, intrusions, and earlier mixing events. Such inherited geometry can matter because salt fingering does not develop from a flat abstract surface in many geophysical settings. We use controlled three-dimensional direct simulations to test whether the spectral state of a finite rough interface changes the pathway by which salt-finger activity develops between adjacent layers. The density ratio, diffusivity ratio, Prandtl number, interface thickness, roughness amplitude, domain, resolution, and analysis window are held fixed; only the imposed roughness spectrum and, for one pair, the realization are changed. Broad low-mode memory produces the largest cumulative salt exchange and the earliest finite-depth contact. High-annulus memory remains localized and intermediate-scale dominated. Mixed memory produces delayed scale transfer and scalar-rich structure that is robust in integrated exchange and broad-memory measures across a second realization, while local plume timing and probe amplitudes remain realization-sensitive. The simulations therefore support treating interfacial spectral memory as an additional state variable for finite-depth double-diffusive exchange, complementary to local thermodynamic descriptors.
Particle Filtering for Non-Deterministic Electrocardiographic Imaging
arXiv:2509.19404v2 Announce Type: replace-cross Abstract: Electrocardiographic imaging (ECGI) aims to non-invasively reconstruct activation maps of the heart from temporal body surface potentials. While most existing approaches rely on inverse and optimization techniques that may yield satisfactory reconstructions, they typically provide a single deterministic solution, overlooking the inherent uncertainty of the problem stemming from its very ill-posed nature, the poor knowledge of biophysical features and the unavoidable presence of noise in the measurements. The Bayesian framework, which naturally incorporates uncertainty while also accounting for temporal correlations across time steps, can be used to address this limitation. In this work, we propose a low-dimensional representation of the activation sequence that enables the use of particle filtering, a Bayesian filtering method that does not rely on predefined assumptions regarding the shape of the posterior distribution, in contrast to approaches like the Kalman filter. This allows to produce not only activation maps but also probabilistic maps indicating the likelihood of activation at each point on the heart over time, as well as pseudo-probability maps reflecting the likelihood of a point being part of an earliest activation site. Additionally, we introduce a method to estimate the probability of the presence of a conduction lines of block on the heart surface. Combined with classical reconstruction techniques, this could help discriminate artificial from true lines of block in activation maps. We support our approach with a numerical study based on simulated data, demonstrating the potential of our method.
Derivation of the fourth-order DLSS equation with nonlinear mobility via chemical reactions
arXiv:2510.07149v2 Announce Type: replace-cross Abstract: We provide a derivation of the one-dimensional fourth-order DLSS equation based on an interpretation as a chemical reaction network. We consider the rate equation on the discretized circle for a process in which pairs of particles occupying the same site simultaneously jump to the two neighboring sites; the reverse process involves pairs of particles at adjacent sites simultaneously jumping back to the site located between them. Depending on the rates, in the vanishing-mesh-size limit we obtain either the classical DLSS equation or a variant with nonlinear mobility of power type. Via EDP convergence, we identify the limiting gradient structure to be driven by entropy with respect to a generalization of diffusive transport with nonlinear mobility. Interestingly, the DLSS equation with power-type mobility shares qualitative similarities with the fast diffusion and porous medium equation, since we find traveling wave solutions with algebraic tails or compactly supported polynomials, respectively.
Laser-intensity-spike-dominated hot electron generation from two-plasmon decay instability driven by moderate-bandwidth pulses
arXiv:2606.26054v1 Announce Type: new Abstract: Our direct-drive-relevant experiments on the low-coherence Kunwu laser facility identify two-plasmon decay (TPD) as the primary source of hot electrons, and demonstrate for the first time that broadband laser pulses enhance TPD. Using particle-in-cell simulations, we attribute this TPD enhancement and the consequent hot electron production to stochastic intensity spikes inherent in broadband laser fields, robust in both weakly- and strongly-driven regimes. These findings suggest that mitigating hot electron generation requires suppressing these intensity spikes.
Hotelling-Downs with Facility Synergy: The Mall Effect
arXiv:2606.26055v1 Announce Type: new Abstract: We consider a variation of the classic Hotelling-Downs model with the addition of facility synergies. Unlike in the classic model, where clients always use the facility closest to them, we study clients who prefer locations with many facilities to those with few facilities while simultaneously attempting to minimize their distance as well. We show that, in contrast with the classic model, Nash equilibria for our setting always exist, and, in fact, there always exists a Nash equilibrium such that the sum of client costs equals the cost of the optimal solution. Our main result is a bound of $\frac{225}{64}\approx 3.516$ on the Price of Anarchy for our model, showing that, although the client behavior is more complex in our model (and often more realistic depending on the application), the cost of Nash equilibrium solutions still cannot be much worse than the cost of the optimal facility placement.
DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation
arXiv:2606.26058v1 Announce Type: new Abstract: Open domain subject-driven text-to-video (S2V) generation has drawn significant interest in academia and industry. Open domain S2V mainly involves two scenarios: in-domain, which requires retaining the reference subject features as much as possible, and cross-domain, which preserves the intrinsic features of the subject while allowing subject-irrelevant properties to vary flexibly according to the text prompt. Existing methods primarily focus on maximizing subject fidelity in in-domain scenarios, which limits their editability and adaptability in cross-domain scenarios, such as novel styles, semantic combinations, or domain attributes. In this study, we propose that an ideal S2V method should flexibly shuttle between different domains, achieving strong performance in both in-domain and cross-domain scenarios. To this end, we propose DomainShuttle, which could achieve high fidelity and generative flexibility for open domain video personalization. Specifically, we introduce Domain-MoT, which decouples videos and reference features and introduces the domain-aware AdaLN for domain-specific modeling of reference images. We then introduce the Video-Reference DualRoPE scheme, which places reference image tokens and video tokens in separate RoPE spaces to enable precise subject-level spatial modeling, and Cross-Pair Consistent Loss, which aims to extract intrinsic subject features unaffected by irrelevant features. Extensive experiments demonstrate that DomainShuttle achieves significant performance improvements over existing methods, exhibiting high subject fidelity and generative flexibility across diverse open domain application scenarios.
Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models
arXiv:2606.26079v1 Announce Type: new Abstract: Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling changes the answer, a baseline reliability property called for by emerging AI evaluation guidelines. We introduce Facet-Probe, a five-facet audit (option, evidence-chunk, document-rank, image-set, and mixed-modality ordering) of 18 frontier and open-weight MLLMs. A Bayesian item-response model separates ordering noise from per-facet bias, and a same-ordering control estimates the decoder-stochastic floor for observed flips. We find that none of the 18 MLLMs we audit are order-invariant: screened per-facet panel-mean flip rates span 24-50%. A Gemini same-ordering control at temperature 0 estimates a substantial ordering excess over a same-input decoder-noise floor in verified cells. Capability predicts but does not eliminate flips; the best model still flips on 13.4% of trials. In our Gemini mitigation tests, training-free prompt changes are modality-conditional and do not transfer from text to visual reasoning. These results suggest that prompt-level mitigation alone is unlikely to provide general order robustness, motivating future work on training-time and architectural approaches. We propose cross-ordering flip rate as a standard reporting axis for MLLMs.
Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
arXiv:2606.26080v1 Announce Type: new Abstract: Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and stochastic environment feedback make both human annotation and Monte Carlo estimation infeasible at scale. In this work, we show that reinforcement learning (RL) post-training already provides the ingredients for effective step-level scoring, eliminating the need for dedicated reward model training altogether. Concretely, we derive an implicit advantage under a general stochastic Markov decision process, which we term progress advantage -- log-probability ratio between the RL-trained policy and its reference policy exactly recovers the optimal advantage function. This formulation makes the resulting signal annotation-free, domain-agnostic, and available as a byproduct of the standard RL post-training pipeline. We validate the effectiveness of the progress advantage across three different applications: test-time scaling, uncertainty quantification, and failure attribution on five benchmarks and four model families. Across all settings, it consistently outperforms confidence-based baselines and, despite requiring no task-specific training, surpasses dedicated trained reward models. We complement these results with deeper analyses on characteristics of progress advantage, offering practical guidance for adoption in real-world agentic systems.
IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
arXiv:2606.19157v2 Announce Type: replace-cross Abstract: AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear whether these models genuinely utilise such context or rely on parametric knowledge learned during pretraining. Existing benchmarks cannot answer this question because they evaluate transcription under fixed prompting conditions and rarely include explicit contextual inputs. We introduce IndicContextEval, a 56-hour multilingual benchmark of natural speech from 555 speakers across 8 Indian languages and 23 professional domains. We design a 7-level prompting framework that progressively introduces contextual signals, including metadata, natural-language descriptions, entity lists in English and native script, and adversarial prompts with incorrect entities. Evaluating five models reveals substantial differences in context utilisation behaviour, highlighting the need for explicit evaluation of contextual grounding in AudioLLMs.
Exact conservation as selection principle: discrete exterior calculus for the incompressible Navier-Stokes and Euler equations
arXiv:2605.13048v4 Announce Type: replace-cross Abstract: We formulate a new discrete-exterior calculus based discretisation of the incompressible Euler and Navier-Stokes equations that preserves the geometric structure of the continuum, and establish a rigorous convergence and structure theory for a new discretisation. The discretisation operates on prismatic Delaunay-Voronoi meshes over closed Riemannian manifolds. The geometry of Euler and Navier-Stokes equations is maintained via a discrete Lie derivative that is built from an extrusion-based contraction for the nonlinear term in vector-invariant form. Conservation of energy and Kelvin circulation links the discrete scheme to the continuum: at the discrete level, energy conservation is a stability property, and in the vanishing-resolution limit it becomes both a constructive route into the conservative weak-solution theory of the continuum equations and a selection principle on the limits the scheme can reach. This correspondence appears in four regimes. \emph{Smooth solutions}: convergence at rate $\mathcal{O}(h^{\min(r_{\rm rec},\,r_\star)}\,|\log h|)$ in dimensions $d=2,3$, uniformly in viscosity $\nu \ge 0$; first order on general meshes, second order under centroid proximity and reconstruction symmetry. \emph{Leray-Hopf weak regime}: subsequential $L^2$ limits of the discrete Navier-Stokes system are weak solutions of the viscous equations. \emph{Inviscid measure-valued regime}: limits are conservative measure-valued Euler solutions, with concentration defect vanishing above the Onsager threshold $\alpha > 1/3$ provided the discrete solutions admit a uniform $C^{0,\alpha}$ bound; the scheme reaches the energy-conserving side of the Onsager landscape but not the dissipative side. \emph{Dissipative regime}: no subsequence converges to an energy-dissipating Euler solution at any H\"older regularity, an exclusion that follows from discrete energy conservation.
Rounding Almost Commuting Hamiltonians
arXiv:2605.26096v2 Announce Type: replace-cross Abstract: Commuting Hamiltonians lie at the boundary between classical constraint satisfaction and quantum many-body physics, exhibiting rich quantum structure while remaining more tractable than general noncommuting models. In contrast, physical Hamiltonians are rarely exactly commuting, which naturally motivates the study of almost commuting Hamiltonians. Despite their relevance, the implications of approximate commutation are only poorly understood. In this work, we show how to efficiently approximate any almost commuting $2$-local qubit Hamiltonian by a commuting one: we give a new locality-preserving algorithmic rounding technique that maps any $2$-local Hamiltonian $H=\sum_{i=1}^m h_i$ with $\|[h_i,h_j]\| \leq \epsilon$ to a nearby Hamiltonian $\hat{H}$ whose terms pair-wise commute, and which is within overall distance $\|H-\hat{H}\| = O(m\,\epsilon^{1/6})$. As a consequence, we show that $\delta$-approximations to the ground energy for $\epsilon$-almost commuting $2$-local qubit Hamiltonians lie in $\mathsf{NP}$ when $\delta \gg m\epsilon^{1/6}$, extending the classical containment well beyond the commuting setting. Finally, we present two applications of our rounding framework: Gibbs sampling and fast Hamiltonian simulation for almost commuting systems.
SC-TauPath: A Structural Connectivity Attribution Framework for Mapping Tau Propagation Pathways in Alzheimer's Disease
arXiv:2606.04066v2 Announce Type: replace-cross Abstract: Understanding how structural connections are associated with tau propagation in Alzheimer's disease (AD) remains a central open question, yet existing computational models either rely heavily on biophysical assumptions or lack neurobiologically interpretable pathway maps. We present SC-TauPath, a structural connectivity (SC) attribution framework that maps tau propagation pathways from in vivo neuroimaging data. SC-TauPath combines a Network Diffusion Model (NDM)-augmented multilayer perceptron with gradient $\times$ input attribution to score each SC edge's contribution to tau prediction, then translates these attribution scores into multi-scale pathway maps (backbone edges, high-traffic routes, and hub ROIs), which validates established Braak staging anatomy. Applied to 234 ADNI participants with paired DTI SC and 18F-Flortaucipir PET, SC-TauPath achieves strong cross-validated tau prediction and yields attribution-based pathway maps consistent with established Braak staging anatomy, demonstrating that SC encode spatially specific information about regional tau distribution in AD.
A Bregman Perspective on Classification and Regression Trees
arXiv:2606.13984v2 Announce Type: replace-cross Abstract: Classification and Regression Trees (CART) constitute one of the most influential paradigms in statistical learning. Although a variety of impurity measures have been proposed for different statistical models, these criteria are typically introduced on a case-by-case basis and analyzed separately. In this paper, we study CART through the lens of Bregman divergences. This perspective places the classical least-squares criterion, Poisson deviance, Kullback-Leibler-type losses, and other impurity measures associated with exponential-family models within a common framework. As a result, key ingredients of the CART methodology -- including node representatives, impurity measures, and split selection rules -- can be expressed and analyzed through general properties of convex functions rather than through separate model-specific constructions. Beyond the algorithmic formulation, we investigate theoretical properties of Bregman-based CART procedures. In particular, we analyze how geometric properties of the generating convex function influence impurity reductions and stability of recursive partitions. We also establish consistency results within the proposed framework, providing a unified theoretical treatment for a broad family of CART type procedures. Our results provide a geometric interpretation of impurity-based tree construction and show that many classical CART impurity criteria admit a common interpretation within a Bregman framework.
Towards Fast and Effective Long Video Understanding of Multimodal Large Language Models via Adaptive Quasi-Gaussian Sampling
arXiv:2606.24187v2 Announce Type: replace Abstract: Long video understanding remains a daunting challenge for Multimodal Large Language Models (MLLMs) due to the excessive computation and memory footprint. Thus, keyframe selection is often adopted to mitigate this shortcoming, which however still suffers from low flexibility and high noise due to its hard sampling principle. In this paper, we define video frame selection as a problem of Quasi-Gaussian Sampling, and propose an adaptive and training-free approach termed AdaQ. Inspired by the 3-$\sigma$ rule of Gaussian distribution, the objective of AdaQ is to achieve the optimal 3-$\sigma$ interval for different examples, i.e., a smaller 3-$\sigma$ interval for the local query and a larger one for the global query, thereby facilitating robust and adaptive frame sampling. To validate AdaQ, we apply it to four MLLMs with three embedding models. The extensive experimental results not only show its obvious performance gains over the default MLLMs and the SOTA keyframe selection methods, e.g., helping Qwen3-VL-8B outperform GPT4o by 15.8% on average by using only 64 frames, but also confirm its superior robustness and high efficiency for long-video understanding, e.g., only 1 hyper-parameter needs to be set.
Obstructions for Minor-Closed Classes of limiting Densities Below 3/2
arXiv:2606.24326v2 Announce Type: replace-cross Abstract: Given a graph class $\mathcal{G}$, the limiting density of $\mathcal{G}$ is defined as $\delta(\mathcal{G})=\lim_{n\to\infty} \mathsf{ex}(\mathcal{G},n)/n$ where $\mathsf{ex}(\mathcal{G},n)$ is the maximum number of edges of a graph in $\mathcal{G}$ on $n$ vertices. The limiting density $\delta(\mathcal{G})$ is known to be a rational number when $\mathcal{G}$ is a minor-closed graph class. For every $\delta\in[0,\frac{3}{2})$, we prove that the set of $\subseteq$-minimal minor-closed graph classes with densities $>\delta$ is finite and we identify it completely. A consequence of our results is an algorithm that, given a finite set of graphs $\mathcal{Z}$, of total size $n$, either outputs the value of $\delta(\mathsf{excl}(\mathcal{Z}))$ or reports that $\delta(\mathsf{excl}(\mathcal{Z}))\geq \frac{3}{2}$, where $\mathsf{excl}(\mathcal{Z})$ is the class of graphs excluding the graphs in $\mathcal{Z}$ as minors. The algorithm runs in $2^{\mathsf{poly}(n)}$ time.
Recursive QLSTM with Dynamic Variational Quantum Circuit Adaptation
arXiv:2606.24932v1 Announce Type: cross Abstract: Recent advances in quantum computing and machine learning have motivated the development of quantum models for sequential data processing. In this paper, we propose a Recursive Quantum Long Short-Term Memory model, or Recursive QLSTM, which extends QLSTM through metacore-based recursive constructions. We numerically test the model under different input sequence lengths, metacore designs, and recursive rules, and identify the best-performing architecture among these variants. For this selected model, we further provide theoretical arguments explaining why its recursive structure improves temporal information propagation and enhances learning performance. Our results suggest that Recursive QLSTM offers a flexible and effective framework for quantum recurrent learning over input time series of various lengths.
Statistically Valid Hyperparameter Selection: From Tuning to Guarantees
arXiv:2606.25601v1 Announce Type: cross Abstract: Hyperparameter selection is a critical step in the deployment of modern artificial intelligence systems, given the need to tune degrees of freedom such as inference-time parameters, implementation-level settings, and thresholds driving decision rules. Despite its practical importance, hyperparameter selection is typically performed using best-effort empirical methods such as grid search or Bayesian optimization, which provide no formal statistical guarantees on reliability or safety. This monograph presents a unified statistical framework for reliable hyperparameter selection, centered on the learn-then-test (LTT) paradigm, which formulates the problem as multiple hypothesis testing over a candidate set of hyperparameters. The framework enables the selection of hyperparameters that provably satisfy application-specific reliability requirements -- such as bounds on average risk, quantile risk, or information-theoretic constraints -- with explicit, finite-sample control of error probabilities. The supporting statistical machinery, namely p-values, e-values, and concentration inequalities, is developed from first principles in a dedicated appendix.
Engineering a Governance-Aware AI Sandbox: Design, Implementation, and Lessons Learned
arXiv:2603.03394v2 Announce Type: replace Abstract: Collaborative AI experimentation in industry-academia requires environments that support rapid trials while maintaining controlled access, organisational isolation, and traceable workflows. Although interest in AI sandboxes is increasing, practical guidance on designing and building governance-aware experimentation platforms remains limited. This work designs and operationalizes a governance-aware, multi-tenant AI sandbox that supports structured experimentation and produces reusable evaluation evidence across stakeholders. The sandbox was developed in an industry-academia ecosystem using iteratively validated requirements gathered from industrial partners. The solution adopts a layered reference architecture that separates a multi-tenant presentation layer from a backend control plane and isolates execution and data management concerns into dedicated layers. The sandbox supports governed onboarding, project-based collaboration, controlled access to AI services, and traceable experimentation through approval workflows and audit logging. By structuring experiment context and governance decisions as persistent records, the sandbox enables evaluation evidence to be reused and compared across projects and stakeholders. The development experience yields lessons learned and practical considerations that inform deployment and future evolution of governance-aware sandbox platforms.