Forskningsradar

Science Journals

Peer-reviewade publikationer — 55483 artiklar

Machine Learning Hamiltonians are Accurate Energy-Force Predictors
arXiv:2602.16897v2 Announce Type: replace Abstract: Recently, machine learning Hamiltonian (MLH) models have gained traction as fast approximations of electronic structures such as orbitals and electron densities, while also enabling direct evaluation of energies and forces from their predictions. However, despite their physical grounding, existing Hamiltonian models are evaluated mainly by reconstruction metrics, leaving it unclear how well they perform as energy-force predictors. We address this gap with a benchmark that computes energies and forces directly from predicted Hamiltonians. Within this framework, we propose QHFlow2, a state-of-the-art Hamiltonian model with an SO(2)-equivariant backbone and a two-stage edge update. QHFlow2 achieves $40\%$ lower Hamiltonian error than the previous best model with fewer parameters. Under direct evaluation on MD17/rMD17, it is the first Hamiltonian model to reach NequIP-level force accuracy while achieving up to $20\times$ lower energy MAE. On QH9, QHFlow2 reduces energy error by up to $20\times$ compared to MACE. Finally, we demonstrate that QHFlow2 exhibits consistent scaling behavior with respect to model capacity and data, and that improvements in Hamiltonian accuracy effectively translate into more accurate energy and force computations.
Stability and equilibria of a compressible elastic membrane in Stokes flow
arXiv:2607.03357v1 Announce Type: cross Abstract: We formulate a continuum model for a compressible lipid-bilayer membrane immersed in Stokes flow, replacing exact local area inextensibility by conservation of an areal phospholipid density. The membrane free energy combines Helfrich bending, spontaneous curvature, and a finite area-compression penalty, so that membrane tension becomes a constitutive response to lipid-density variation rather than a Lagrange multiplier enforcing local area conservation. The resulting interfacial stress includes normal elastic forces and tangential Marangoni stresses generated by lipid redistribution; these stresses arise from membrane compressibility and can produce an effective negative tension when the local lipid density exceeds its preferred value. We further derive the linear stability of circular membranes in two dimensions and spherical membranes in three dimensions under full Stokes hydrodynamic coupling. In both cases, bending stabilizes the base shape, while excess lipid density destabilizes it by favoring increased membrane area. The first instability occurs in the lowest nontrivial shape mode, m = 2 in two dimensions and j = 2 in three dimensions. Energy expansions near onset show that the two-dimensional instability is a pitchfork bifurcation, whereas the three-dimensional instability is generically transcritical because prolate and oblate perturbations are geometrically distinct. These results provide a controlled compressible extension of classical vesicle mechanics and directly connect lipid-density variation, membrane tension, hydrodynamic coupling, and shape instability.
Nested-Loop Trajectory-Informed Variational Quantum Solver for Interior-Point OPF
arXiv:2607.03361v1 Announce Type: cross Abstract: Optimal power flow (OPF) solved by an interior-point method (IPM) requires repeatedly solving Newton linear systems. When variational quantum linear solvers (VQLS) are used, each IPM iteration involves an additional nested inner variational optimization loop, which can significantly slow the overall quantum-assisted IPM convergence. To address this challenge, this paper proposes a dual-level trainable quantum IPM framework for OPF that leverages early solver-generated trajectories rather than relying on single-point prediction. The key observation is that early IPM iterates provide informative primal-dual, slack, and barrier-variable evolution about the path to optimality, while early VQLS parameter updates provide useful information about the later variational search. At the quantum-solver level, a trainable parameter model uses a short prefix of the VQLS parameter trajectory to project the remaining variational search toward a lower-cost region. At the OPF-solver level, a second trainable model uses early primal-dual IPM iterates to project a later central path state, which is restored to an admissible point before IPM refinement continues. Simulation studies show that the proposed approach reduces the number of variational updates by up to $95\%$ while maintaining OPF objective values close to the classical IPM reference. A 2-bus demonstration on real quantum hardware is also included to validate the implementation of the proposed workflow.
Towards Understanding Deep Learning Model in Image Recognition via Coverage Test
arXiv:2505.08814v3 Announce Type: replace Abstract: Deep neural networks (DNNs) play a crucial role in the field of artificial intelligence, and their security-related testing has been a prominent research focus. By inputting test cases, the behavior of models is examined for anomalies, and coverage metrics are utilized to determine the extent of neurons covered by these test cases. With the widespread application and advancement of DNNs, different types of neural behaviors have garnered attention, leading to the emergence of various coverage metrics for neural networks. However, there is currently a lack of empirical research on these coverage metrics, specifically in analyzing the relationships and patterns between model depth, configuration information, and neural network coverage. This paper aims to investigate the relationships and patterns of four coverage metrics: primary functionality, boundary, hierarchy, and structural coverage. A series of empirical experiments were conducted, selecting LeNet, VGG, and ResNet as different DNN architectures, along with 10 models of varying depths ranging from 5 to 54 layers, to compare and study the relationships between different depths, configuration information, and various neural network coverage metrics. Additionally, an investigation was carried out on the relationships between modified decision/condition coverage and dataset size. Finally, three potential future directions are proposed to further contribute to the security testing of DNN Models.
Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies
arXiv:2411.07447v5 Announce Type: replace Abstract: LLMs are increasingly used world-wide from daily tasks to agentic systems and data analytics, requiring significant GPU resources. While LLM inference systems are capable of serving millions of requests from multiple users, they often lack theoretical models to determine whether they achieve the performance upper bounds of underlying hardware resources. Beyond online workload serving, merely analyzing existing systems-or developing yet another one-is both GPU-intensive and labor-intensive. This paper provides a comprehensive survey of LLM inference systems, focusing on their cache management policies and availability. We then show that simulations can be an effective tool to save GPU hours in the development and analysis phase of inference systems, revealing useful insights for developing better inference techniques, unlike how existing studies used simulations to find the best parameters inside a given system. Finally, we provide theoretical tools to estimate the optimal performance and formulate new ideas. Based on the theoretical analysis, especially on the cache management in LLM inference, we propose a simple yet effective cache replacement policy that can be easily plugged into existing preemptive schedulers and systems. We show that such a simple policy inspired from database systems can substantially save GPU hours in actual inference systems on online workloads. We share our experience submitting a journal paper to a database venue in November 2025 for anyone considering a similar path.
Surface-exciton enhanced SHG response in few-layer 2H-TMDC
arXiv:2607.05037v1 Announce Type: new Abstract: We explore the nonlinear optical properties of few-layer MoS2 by means of polarization and laser-power-dependent measurements as well as ab initio techniques. While for even layer samples a weak second-harmonic (SH) signal can be attributed to the presence of surface defects or interface effects, our measurements resolve a layer-number dependent signal for odd-layer samples. For the excitation energy of 780 nm, we find that the SH intensity decreases steadily with the layer number. Our simulations demonstrate that this effect cannot be purely attributed to modifications of the band structure, but requires the inclusion of excitonic effects and can be explained by the increasing delocalization of excitons with increasing sample thickness.
Security Analysis of RIS-Assisted Physical-Layer Authentication Over Multipath Channels
arXiv:2607.05042v1 Announce Type: new Abstract: In physical layer authentication, verification of a user's identity is based on the characteristics of the transmission channel through which signals are delivered to the authenticator (Bob). In this paper, we assume that the signals received by Bob pass through a \ac{RIS} (controlled by Bob) and that the legitimate transmitter (Alice) is equipped with one antenna. Conversely, the attacker (Trudy) has multiple antennas and uses precoding to deceive Bob's verification. Assuming that Trudy knows all the channel matrices, we first derive her optimal attack strategy. Then, we analyse the conditions under which the channel estimated by Bob is indistinguishable when either Alice or Trudy is transmitting. When Trudy has a single antenna, we show that the indistinguishability condition cannot be met when the channels to the RIS are the result of propagation over multiple paths. For single-path line-of-sight (LOS) conditions, instead, Trudy can impersonate Alice although transmitting from a different position. We verify these results numerically and assess the security of the considered scenario, even when the indistinguishability conditions cannot be met.
The S-ICDF Dataset: Sionna-Simulated Dynamic Interference Characterization and Direction Finding
arXiv:2607.03411v1 Announce Type: cross Abstract: Jamming and spoofing threaten wireless and satellite navigation by disrupting or manipulating radio frequency (RF) signals, undermining availability, integrity, and trust. Robust interference monitoring (i.e., detection, classification, characterization, and direction finding) is therefore essential to identify and localize anomalous signals. While machine learning (ML) promises improved performance in complex environments, its development and validation depend on large-scale datasets that capture realistic signal and channel variability. Collecting such data in the real world is difficult because intentional jamming is illegal and ground-truth attribution is confounded by propagation, hardware, and environmental effects. To address this gap, we create and publish S-ICDF, a large-scale indoor interference dataset generated with Sionna, a GPU-accelerated simulation library for physical-layer wireless communications. S-ICDF covers 102 interference configurations, including diverse antenna array patterns, bandwidths, and simulation settings such as noise level and reflection depth. We further provide baseline results by benchmarking S-ICDF with classical estimation and direction finding (DF) methods (MUSIC, ESPRIT, and CAPON) and with modern ML approaches. The dataset is publicly available at: https://gitlab.cc-asp.fraunhofer.de/darcy_gnss/sicdf_dataset
Modeling Decision-Making with Will for Cooperation in Social Dilemmas
arXiv:2605.08669v2 Announce Type: replace Abstract: Standard rational actor models often attribute cooperation failures in social dilemmas to insufficient incentives, overlooking the destabilizing effects of continuous utility maximization. To address this, we propose a framework of ``will" defined as a mechanism that persistently pursues goals while ignoring local cost-benefit fluctuations. We formalize the Willed Agents as potential minimizers, distinguishing them from cumulative utility maximization. Dynamical analysis of infinite population demonstrates that willed agents shrink the feasible state space, acting as boundary constraints that accelerate convergence in canonical social dilemmas. Through multi-agent simulations in a spatiotemporal Stag Hunt Game, we show that willed agents function as ``cooperation catalysts", enabling groups to surmount high-risk thresholds where purely utility maximization fails. We find that heterogeneous will strength promotes cooperation, and that agents who autonomously suspend rational re-evaluation can significantly outperform continuous optimizers. These findings suggest that successful cooperation relies on the cognitive capacity to strategically constrain calculation.
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI
arXiv:2605.08678v3 Announce Type: replace Abstract: Modern AI progress has been driven by ML methods that are generalizable across settings and scalable to larger regimes. As large language models demonstrate advanced capabilities in reasoning, coding, and engineering tasks, it is increasingly important to understand whether they can discover such methods rather than only apply existing ones. We introduce MLS-Bench, a benchmark for evaluating whether AI systems can invent generalizable and scalable ML methods. MLS-Bench contains 140 tasks across 12 domains, each requiring an agent to improve one targeted component of an ML system or algorithm and demonstrate that the improvement generalizes across controlled settings and scales. We find that current agents remain far from reliably surpassing human-designed methods, and that engineering-style tuning is easier for them than genuine method invention. We further study the effects of test-time scaling, adaptive compute allocation, and context provision on agents' discovery performance, together with case studies of their behavior. Our analyses suggest that the bottleneck is not only in proposing new methods, but also in the scientific insight needed to plan, validate, and scale claims about them. More search, compute, or context alone does not remove this bottleneck. We build and maintain a community platform for cumulative and comparable iteration, and release the data and code at https://mls-bench.com.
From Retrieval to Synthesis: Repair Literacy and the Domestication of Generative AI
arXiv:2601.20749v2 Announce Type: replace Abstract: How do students develop AI literacy through everyday practice rather than formal instruction? While normative AI literacy frameworks proliferate, empirical understanding of how students actually learn to work with generative AI remains limited. This study analyzes 10,536 ChatGPT messages from 36 undergraduates over one academic year, revealing five use genres -- academic workhorse, emotional companion, metacognitive partner, repair and negotiation, and trust calibration -- that constitute distinct configurations of student-AI learning. Drawing on domestication theory and emerging frameworks for AI literacy, we demonstrate that functional AI competence emerges through ongoing relational negotiation rather than one-time adoption. Students develop sophisticated genre portfolios, strategically matching interaction patterns to learning needs while exercising critical judgment about AI limitations. Notably, repair work during AI breakdowns produces substantial learning about AI capabilities, developing what we term "repair literacy" -- a crucial but underexplored dimension of AI competence. Our findings offer educators empirically grounded insights into how students actually learn to work with generative AI, with implications for AI literacy pedagogy, responsible AI integration, and the design of AI-enabled learning environments that support student agency.
A conservative finite-volume Buckley--Leverett solver with bounded-interval multiwavelet state analysis
arXiv:2603.28981v3 Announce Type: replace Abstract: We develop a conservative finite-volume Buckley--Leverett solver equipped with a bounded-interval multiwavelet state-analysis layer. Because non-capillary Buckley--Leverett transport is a nonlinear hyperbolic conservation law with entropy-admissible shocks, the saturation equation is advanced by a conservative finite-volume method with monotone numerical fluxes. The accepted finite-volume state is then embedded in a bounded-domain multiwavelet hierarchy, reconstructed back to cell averages, and used for multiresolution diagnostics. The formal transport accuracy is therefore governed by the underlying finite-volume discretization, while the multiwavelet layer is used for representation, compression, and front-localization diagnostics. Its purpose is instead to quantify whether the deterministic physical-space saturation state can be represented faithfully, compressed in a controlled manner, and used to identify dynamically active front regions. Validation against reference Buckley--Leverett profiles for a Berea benchmark shows accurate saturation histories, spatial profiles, front-position diagnostics, and mass balance. The multiwavelet reconstruction tracks the internal finite-volume state with essentially exact fidelity. Additional thresholding tests show that a substantial fraction of detail coefficients can be discarded while maintaining small reconstruction errors and negligible global mass defect, and fine-level detail activity localizes the moving displacement front. The resulting formulation provides a conservative and reproducible first stage toward future transport-active adaptive multiwavelet solvers for porous-media flow.
Decentralized Aggregation of LLM Predictions via Wagering Mechanisms
arXiv:2607.04389v1 Announce Type: new Abstract: It is increasingly common to aggregate predictions from multiple LLMs, each with domain expertise or access to private tools and data, to improve collective prediction performance. In decentralized settings, aggregation weights need to be determined without access to models' private information and should remain robust to strategic reporting. We propose a family of advantage-aligned wagering mechanisms for LLM aggregation (WALLA), in which each model reports a prediction and a learned wager, and predictions are aggregated using wagers as weights. WALLA introduces a leave-one-out baseline into the net payout function, yielding three desirable properties: (1) dominant-strategy incentive compatibility of prediction under arbitrary belief structure, (2) advantage--wager alignment, where the optimal wager is proportional to the model's expected score advantage, and (3) prediction-agnostic wager optimization, enabling decentralized learning of wager policies without requiring optimal predictions. We further instantiate two mechanism variants that trade off normality and no-arbitrage while maintaining a bounded worst-case deficit for the mechanism. Experiments on question-answering and forecasting benchmarks across heterogeneous models and private-information settings show that WALLA matches centralized aggregation methods in predictive performance, while simultaneously achieving decentralized learning, advantage-aligned aggregation weights, uncertainty awareness, and incentive-compatible prediction.
Machine-Learning-Enabled Full-State Reconstruction of Fusion Plasmas from Minimal Sensor Measurements
arXiv:2607.04390v1 Announce Type: new Abstract: Plasma in nuclear fusion reactors is only partially observable: diagnostics are constrained by limited access, cost, and the harsh plasma environment, while high-fidelity simulations remain prohibitively costly at reactor-relevant scales to address the observability gap. This paper presents an ML model for reconstructing full-domain plasma states from a small number of accessible measurements. The model combines temporal encoding of sparse sensor histories with spatial decoding into complete plasma-field maps. Demonstrations using high-fidelity kinetic simulation data show that multiple coupled plasma quantities can be reconstructed from only a few density sensors, with robustness to sparse and randomly located probes. The approach provides a route toward diagnostic augmentation, real-time state estimation, and data-driven digital twins for fusion-relevant plasmas.
MMAO-Dyn: A Metabolic Multi-Agent Optimizer for Dynamic Optimization
arXiv:2607.00846v2 Announce Type: replace Abstract: This paper studies whether the Metabolic Multi-Agent Optimizer (MMAO) can be credibly derived into a dynamic-optimization method without replacing its core metabolic control loop by external adaptation modules. The proposed MMAO-Dyn maps private energy, communal budget, role drift, success feedback, and lifecycle turnover to a nonstationary setting in which environmental changes repeatedly invalidate previously useful local structure. We evaluate MMAO-Dyn on an 18-scenario synthetic dynamic continuous benchmark matrix covering shifted sphere, shifted Ackley, and shifted Rastrigin landscapes at $10D$, $20D$, and $30D$, with two change severities and 12 seeds per scenario. The comparison layer includes a generic MMAO variant without dynamic derivation, dynamic random search, dynamic PSO-lite, dynamic DE-lite, and three endogenous ablations. Across the full 216-run matrix, MMAO-Dyn attains mean offline error $28.07$, improving over Generic-MMAO ($29.36$), Dynamic-PSO-lite ($34.65$), Dynamic-DE-lite ($67.09$), and Dynamic-RandomSearch ($111.37$). The gains are clearest in aggregate robustness on sphere and Rastrigin families and in 10-step post-change recovery relative to the generic backbone, whereas the seed-aligned comparison with Dynamic-PSO-lite remains unfavorable in win-loss count and the \texttt{NoMemoryRefresh} ablation stays very close to the full method. We therefore position MMAO-Dyn as a credible family-expansion result for MMAO: the metabolic loop can generate meaningful dynamic behavior, but the strongest current value lies in recovery-oriented resource redistribution rather than in universal dominance or in a fully optimized submechanism design.
Fast Matrix Multiplication meets the Submodular Width
arXiv:2412.06189v5 Announce Type: replace Abstract: One fundamental question in database theory is the following: Given a Boolean Conjunctive Query (BCQ) Q, what is the best complexity for computing the answer to Q in terms of the input database size N? When restricted to the class of combinatorial algorithms, it is known that the best known complexity for any query Q is captured by the submodular width of Q. However, beyond combinatorial algorithms, certain queries are known to admit faster algorithms that often involve a clever combination of fast matrix multiplication and data partitioning. Nevertheless, there is no systematic way to derive and analyze the complexity of such algorithms for arbitrary queries Q. In this work, we introduce a general framework that captures the best complexity for answering any BCQ Q using matrix multiplication. Our framework unifies both combinatorial and non-combinatorial techniques under the umbrella of information theory. It generalizes the notion of submodular width to a new stronger notion called the omega-submodular width that naturally incorporates the power of fast matrix multiplication. We describe a matching algorithm that computes the answer to any query Q in time corresponding to the omega-submodular width of Q. We show that our framework recovers the best known complexities for Boolean queries that have been studied in the literature, to the best of our knowledge, and also discovers new algorithms for some classes of queries that improve upon the best known complexities.
Design of an Electrically Tunable Microtoroid for Frequency Selection of Polarization-Entangled Photons
arXiv:2607.03437v1 Announce Type: cross Abstract: Encoding quantum information into discrete optical frequencies, or "frequency bins," uses different colors of light as additional information channels, allowing each photon to carry more information than polarization alone. We present a computational design for an electrically tunable silica microtoroid that selects desired frequency channels after a polarization-entangled photon pair has been generated without disturbing the photons' polarization entanglement. In the proposed architecture, the 750 nm signal photon passes through the microtoroid, while its entangled 880 nm partner bypasses the resonator and serves as a reference for the selected frequency channel. The principal challenge is resonator birefringence: because horizontally and vertically polarized light resonate at slightly different frequencies, the selected frequency can reveal the photon's polarization state and weaken the quantum correlation between the photon pair. We solve this problem by adding a small lithium-niobate tuning element controlled with a single applied voltage. The voltage shifts the resonator so that it responds almost identically to horizontally and vertically polarized light, reducing the remaining mismatch to only 0.286 optical linewidths across nine frequency channels. The photons remain strongly entangled after passing through the device, with a concurrence of C = 0.969, a Bell-state fidelity of F = 0.981, and a Bell parameter of S_max = 2.785. If the relative timing between the frequency channels is also controlled, the same device can generate a nine-channel polarization-frequency hyperentangled state with an effective dimension of K = 8.97. This computational design provides a compact, electrically tunable bridge between polarization-entangled photon sources and future high-capacity quantum photonic systems.
MechMath Agent Team: LLM Driven Agents for Mathematical Research
arXiv:2607.04394v1 Announce Type: new Abstract: AI reasoning has become a central focus in contemporary artificial intelligence, largely driven by the success of large language models. However, mathematical research, which is characterized by non-linear derivation paths, rigorous logical requirements, and protracted exploration cycles, poses severe challenges for existing reasoning systems. To overcome these limitations, we present the MechMath Agent Team (MMAT), which is a large language model driven agent designed to serve as a co-pilot throughout the full cycle of mathematical research. We design a tripartite Harness Architecture that decouples system responsibilities into Control, Execution, and Augmentation planes, thereby reconciling rigorous logical control with the agility demanded by open-ended research. Building upon this framework, we instantiate three specialized agents: a Knowledge Base Manager, a Natural Language Prover, and a Formal Language Prover, all operating in a closed loop to produce formally certified mathematical proofs. We evaluate MMAT on open problems in Number Theory, Algebraic Complexity Theory, Differential Algebra, Operator Algebra, and Inequalities. Across a two-month deployment, 11 problems have been solved, demonstrating its capacity to act as a co-pilot throughout the entire research cycle. The contributions are threefold: a general decoupled Harness Architecture for multi-agent mathematical reasoning, its concrete instantiation in the MMAT system, and empirical validation on a diverse suite of open problems.
NKI-Agent: Domain-Specific Fine-Tuning and Agentic Tool Use for Neuron Kernel Generation
arXiv:2607.04395v1 Announce Type: new Abstract: Recent agentic approaches to LLM-based kernel generation have achieved impressive results on CUDA. For emerging AI accelerators such as AWS Trainium and Inferentia, automated kernel generation and optimization remain largely unaddressed. Writing kernels for these chips via the Neuron Kernel Interface (NKI) is particularly challenging: developers must navigate a multi-engine architecture, tile-based programming, and explicit data movement across multi-level memory hierarchy. Moreover, no publicly-available training data, benchmarks, or tool-augmented agents exist for this domain. We introduce NKI-Agent, the first system combining domain-specific supervised fine-tuning (SFT) with a compile-verify-fix agent loop for NKI kernel generation. We adapt the existing CUDA-Agent framework to Neuron hardware, curate 6,000 NKI kernel generation tasks for training, and construct NKIBench, a 250-task benchmark across three difficulty levels. Evaluated on real Trn1 hardware, NKI-Agent with Claude Opus 4.8 and a rank-aware system prompt achieves a 77.3% pass rate on the 150-task NKIBench. We show that tool use is critical: Opus 4.8 scores 6% in single-shot mode without agent tools. On a 60-task subset, we show that an SFT-trained Qwen3-Coder-30B-A3B achieves 25.0% pass rate at 1/100th the cost, outperforming Claude Sonnet 4 (15.0%). We also report that Group Relative Policy Optimization (GRPO) with binary compilation reward fails to improve over SFT, providing guidance on reward design for RL-based kernel generation.
Time Series Decomposition using the Fr\'echet Distance
arXiv:2607.04397v1 Announce Type: new Abstract: In this paper, we introduce a new data analysis problem that aims to decompose a set of univariate time series into a small set of $k$ base curves of length at most $l$ such that the sum of Fr\'echet distances of the time series to a ``Fr\'echet combination'' of the base curves is minimized. Here, a Fr\'echet combination allows to combine individually scaled base curves using a $k$-dimensional traversal. We call the problem of finding a set of optimal base curves the Fr\'echet decomposition problem and we consider two variants: (a) the base curves can be arbitrary curves of bounded length and (b) the curves come from a given finite set of candidate curves. We think of the Fr\'echet decomposition problem as a Fr\'echet variant of principal component analysis. For the case of a single base curve we develop a $(1+\varepsilon)$-approximation algorithm for the Fr\'echet decomposition problem. Additionally we give an exact algorithm for the projection distance problem that asks to compute the distance of one given time series to a given set of $k$ base curves. This allows us to design an exact algorithm for the Fr\'echet decomposition problem for general $k$ when curves come from a fixed candidate set.
Robust Counterfactual Explanations under Model Multiplicity Using Multi-Objective Optimization
arXiv:2501.05795v4 Announce Type: replace Abstract: In recent years, explainability in machine learning has gained importance. In this context, counterfactual explanation (CE), which is an explanation method that uses examples, has attracted attention. However, it has been pointed out that CE is not robust when there are multiple machine-learning models with similar accuracy. These problems are important when using machine learning to make safe decisions. In this paper, we propose robust CEs that introduce a new viewpoint -- Pareto improvement -- and a method that uses multi-objective optimization to generate it. To evaluate the proposed method, we conducted experiments using both simulated and real data. The results demonstrate that the proposed method is both robust and practical. This study highlights the potential of ensuring robustness in decision-making by applying the concept of social welfare. We believe that this research can serve as a valuable foundation for various fields, including explainability in machine learning, decision-making, and action planning based on machine learning.
Fast GPU Linear Algebra via Compile Time Expression Fusion
arXiv:2604.22242v2 Announce Type: replace Abstract: We describe the Bandicoot GPU linear algebra toolkit for C++, which prioritises ease of use without compromising efficiency. Bandicoot's API aims for compatibility with the popular Armadillo CPU linear algebra library, enabling easy transition for existing CPU-based codebases. Unlike other GPU-focused toolkits, Bandicoot uses template metaprogramming to generate fused GPU kernels directly at compile-time, yielding efficient kernels that can saturate memory bandwidth. This removes the need for run-time overhead or JIT infrastructure. Empirical results show that Bandicoot outperforms (sometimes by considerable margins) commonly-used linear algebra toolkits including PyTorch, TensorFlow, and JAX.
Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor
arXiv:2507.15903v2 Announce Type: replace Abstract: Empowered by large language models (LLMs), intelligent agents have become a popular paradigm for interacting with open environments to facilitate AI deployment. However, hallucinations generated by LLMs-where outputs are inconsistent with facts-pose a significant challenge, undermining the credibility of intelligent agents. Only if hallucinations can be mitigated, the intelligent agents can be used in real-world without any catastrophic risk. Therefore, effective detection and mitigation of hallucinations are crucial to ensure the dependability of agents. Unfortunately, the related approaches either depend on white-box access to LLMs or fail to accurately identify hallucinations. To address the challenge posed by hallucinations of intelligent agents, we present HalMit, a novel black-box watchdog framework that models the generalization bound of LLM-empowered agents and thus detect hallucinations without requiring internal knowledge of the LLM's architecture. Specifically, a probabilistic fractal sampling technique is proposed to generate a sufficient number of queries to trigger the incredible responses in parallel, efficiently identifying the generalization bound of the target agent. Experimental evaluations demonstrate that HalMit significantly outperforms existing approaches in hallucination monitoring. Its black-box nature and superior performance make HalMit a promising solution for enhancing the dependability of LLM-powered systems.
Experimental Realization of Type-II Quadrupole Topological Insulator
arXiv:2607.05049v1 Announce Type: new Abstract: The discovery of quadrupole topological insulators (QTIs) has spurred extensive research into higher-order topological phases. Recently proposed type-II QTIs exhibit unconventional topological behaviors with 1/2 edge polarization \operatorname{p}_x and zero edge polarization \operatorname{p}_y, due to the inequivalence between Wannier-band and edge-spectrum gap closures, yet their experimental realization remains challenging owing to the long-range and complex off-site hopping terms in their tight-binding model (TBM). Here, we circumvent this difficulty via an optimized Householder tridiagonalization (OHT) mapping that reduces the complex two-dimensional lattices to one-dimensional chains with only negative-real-valued nearest-neighbor hopping terms, greatly facilitating experimental sample fabrication. Using this strategy, we experimentally verify the type-II QTI phase, type-I QTI phase and trivial phase in elastic wave platforms via simple aperiodic plate-beam chain structures, where the plates reflect the on-site potential terms and beams correspond to the off-site hopping terms in the TBM. Our approach provides a versatile route for experimentally exploring more complex and richer topological phenomena based on TBM.
Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny
arXiv:2507.16331v4 Announce Type: replace Abstract: Existing informal language-based (e.g., human language) Large Language Models (LLMs) trained with Reinforcement Learning (RL) face a significant challenge: their verification processes, which provide crucial training signals, are neither reliable nor scalable. In fact, the prevalent large proprietary models could hardly generate verifiable programs. A promising yet largely uncharted alternative is formal language-based reasoning. Grounding LLMs in rigorous formal systems where generative models operate in formal language spaces (e.g., Dafny) enables the automatic and mathematically provable verification of their reasoning processes and outcomes. This capability is pivotal for achieving large-scale, reliable formal software verification. It is a common practice to employ human-annotated chain-of-thought and answers to induce the reasoning and coding capabilities of LLMs. Unfortunately, it becomes unacceptably all-consuming to provide such priors for supervising complex programming tasks. In this work, we systematically explore ways to reduce human annotations with the formal language, Dafny, as the main environment for our pilot study. Our pipeline mainly relies on introducing an automatic and scalable data curation pipeline, and careful RL designs integrated with feedback from the formal language verifier. We introduce DafnyComp, a benchmark of compositional formal programs with auto-formalized specifications for specification reasoning. Our supervised fine-tuning (SFT) stage enables even small models (e.g., 0.5B) to generate syntactically valid and verifiable Dafny code, surpassing proprietary models. RL with regularization further improves performance, achieving stronger generalization to out-of-domain tasks and outperforming all strong baselines on the challenging DafnyComp benchmark.