Forskningsradar

Science Journals

Peer-reviewade publikationer — 55345 artiklar

Generative Adversarial Learning from Deterministic Processes
arXiv:2605.18425v1 Announce Type: new Abstract: Physical AI is being successfully applied to data which does not follow the traditional paradigm of independent and identically distributed (i.i.d.) samples. In fact, physical AI is often trained on data which is not random at all, and is instead derived from chaotic dynamical systems like turbulence. We aim to explain the empirical success of these methods using the example of generative adversarial networks (GANs), whose statistical learning theory under the i.i.d. assumption is generally well understood. We prove that it is possible, using an infinite-dimensional model of generative adversarial learning (GAL), to learn the invariant distribution of a sufficiently chaotic dynamical system from a single deterministically evolving time series of its states or measurements thereof, and give explicit rates for the convergence to the solution in terms of the Jensen-Shannon divergence.
Generalize cross-ratios in n-dimensional Plane-Based Geometric Algebra
arXiv:2605.18398v1 Announce Type: new Abstract: We develop a complete theory of projective cross-ratios in n-dimensional Plane-Based Geometric Algebra (PGA), R(n,0,1), covering geometric objects of every grade: finite and ideal points, hyperplanes, and intermediate flats. For each object type and configuration, we establish an explicit cross-ratio formula, prove that it recovers the appropriate classical invariant, and identify the canonical pairwise measurement operator. A systematic duality analysis further revealed that all eight configurations organize into four dual pairs under the Hodge dual, and that all measurement operators reduce to either the commutator or the commutator dual, depending solely on the geometric configuration rather than on object grade. In each case the formula recovers the appropriate classical invariant: signed distance ratios for parallel configurations and sine cross-ratios for secant ones. These results establish the cross-ratio as a grade-agnostic projective invariant within PGA, and provide a constructive foundation for defining n-dimensional homographies directly from prescribed invariants.
A single multi-configuration Direct Electron Detector for various electron imaging and diffraction-based techniques in SEM
arXiv:2605.18386v1 Announce Type: new Abstract: Addressing the need for efficient and integrated multiscale crystallographic and defect analyses of advanced materials, this paper presents the implementation of a new multi-configuration detection system, integrating a single Timepix3-based direct electron detector (DED) in a scanning electron microscope (SEM). By combining precise translation and rotation movements, this system enables, for the first time, the use of the same detector to realize all principal diffraction geometries. These include conventional Electron BackScatter Diffraction (EBSD), off-axis Reflexion Kikuchi Diffraction (RKD), and Transmission Kikuchi Diffraction (TKD) in on-, off- and near-axis configurations. Furthermore, transitions between all these geometries are accomplished without hardware modification. On the other hand, this work presents efficient reconstruction of electron images using the detector data-driven feature, extending thus its applicability to BackScattered Electron imaging (BSE), Electron Channelling Contrast Imaging (ECCI) and Scanning Transmission Electron Imaging in SEM (STEM-in-SEM) characterizations. High-quality Kikuchi patterns easily indexable were acquired across all geometries as well as micrographs of dislocations in both reflection and transmission modes. This is achieved thanks to the flexibility of the implemented detector, the optimizations made in acquisition parameters, such as energy filtering settings, and the efficiency of the developed custom approach used for electron data post-processing. Through this work, it is demonstrated that with a single DED assisted by an orientable support, it is possible to perform multiple advanced microstructural characterizations of both bulk samples and thin foils in the same SEM.
Beyond Square Roots: Explicit Memory-Efficient Factorization for Multi-Epoch Private Learning
arXiv:2605.18379v1 Announce Type: new Abstract: Correlated-noise mechanisms are among the most promising approaches for improving the utility of differentially private model training, but rigorous guarantees require explicit, analyzable factorizations, and practical deployment requires memory efficiency. Recent works have developed banded inverse factorizations, which address both requirements by exploiting a banded structure in the correlation matrix. The bandwidth controls the size of the noise buffer used to correlate noise across iterations, and thus governs the tradeoff between utility and memory cost. Existing factorizations highlight this tradeoff: DP-$\lambda$CGD achieves high memory efficiency by using only a one-step noise buffer, but this limits its utility gains, while the banded inverse square root (BISR) factorization exploits larger correlation windows and is asymptotically optimal for large bandwidths but performs poorly at low bandwidths. We propose $\gamma$-BIFR, a unified generalization of both factorizations. In the low-memory, low-bandwidth regime, $\gamma$-BIFR significantly improves RMSE, amplified RMSE, and private training performance, while yielding tighter theoretical guarantees for multi-participation error in multi-epoch training.
Smart Contract Security Beyond Detection
arXiv:2605.09124v2 Announce Type: replace Abstract: Smart contract security has progressed from vulnerability detection toward a broader research agenda that includes semantic reasoning, automated repair, adversarial robustness, and real-time exploit detection. This paper develops a capstone-oriented research narrative around four directions: foundation-model-based smart contract semantics and vulnerability reasoning [1], automated smart contract repair with formal guarantees [2], adversarial learning for robust malicious contract and transaction detection [3], and real-time transaction-level exploit detection at blockchain scale [4]. We connect these directions to two recent studies that characterize the current frontier: a diagnostic analysis of where smart contract security analyzers fall short [5] and a scalable real-time system for malicious Ethereum transaction detection [6]. The resulting framework is intended to help students formulate capstone projects that are technically grounded, empirically measurable, and aligned with contemporary smart contract security research.
Deterministic Decomposition of Stochastic Generative Dynamics
arXiv:2605.08794v2 Announce Type: replace Abstract: Modern generative models can be understood as probability transport from a simple base distribution to a target data distribution. Deterministic transport models offer tractable velocity-field parameterizations, whereas stochastic generative models capture richer density evolution through drift and diffusion. Yet when stochastic dynamics are described through deterministic velocity fields, the effects of drift and diffusion are often compressed into a single effective field, obscuring the distinct roles of deterministic evolution and stochastic fluctuation. In this work, we show that the deterministic field \(b_t\) of a stochastic generative process admits a natural transport--osmotic decomposition that separates deterministic transport from stochastic, diffusion-induced effects: \(b_t = u_t + d_t\), where \(u_t\) governs marginal probability transport and \(d_t\) captures an osmotic effect induced by diffusion and determined by the marginal score. Based on this decomposition, we propose Bridge Matching, a flow-based framework for learning decomposed generative dynamics through both marginal and conditional formulations. In generative modeling experiments, we recombine the learned components as \(b_t = u_t + \lambda_d d_t\), showing that the proposed decomposition enables interpretable and controllable sampling by adjusting the osmotic contribution in probability transport.
Energy-Resolved Eigenmode Spectroscopy of 1-D and 2-D Non-Hermitian Skin Effects
arXiv:2605.18272v1 Announce Type: new Abstract: Non-Hermitian lattices can host the non-Hermitian skin effect, a boundary-induced collapse of all bulk eigenstates into exponentially localized edge modes. This effect underlies anomalous bulk-boundary correspondence and remarkable enhancements in non-Hermitian sensing, yet direct energy-resolved access to the eigenmodes of non-Hermitian lattices has remained limited. Here we report band- and energy-resolved eigenmode spectroscopy of skin modes in a frequency synthetic dimension. By introducing strong frequency-domain boundaries in an electro-optically modulated ring resonator, we realize finite non-Hermitian lattices and use laser detuning as a spectroscopic axis for the eigenenergies of the effective Hamiltonian. Site-resolved heterodyne measurements then reconstruct the spatial profile of each mode, revealing boundary-localized skin states throughout the spectrum and their eigenenergy-dependent displacement from the edge. Beyond 1D, the same frequency-boundary architecture, upon incorporating long-range couplings between finite lattices, produces genuine 2D frequency lattices rather than the hitherto-realized folded 1D systems on twisted tubes. In these lattices we observe tunable directional transport and edge localization in two synthetic dimensions. Our results introduce eigenmode spectroscopy as a direct probe of non-Hermitian physics and establish strongly bounded frequency lattices as a flexible platform for Hamiltonian engineering.
LLM-Based Static Verification of Code Against Natural-Language Requirements: An Industrial Experience Report
arXiv:2605.17926v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate requirements specifications, design documents, code, and test cases. In contrast, much less attention has been given to a more difficult assurance problem: statically verifying whether implemented code satisfies requirements written in natural language. Conventional static analysis tools are effective at detecting coding defects and known vulnerability patterns, but they cannot determine whether program behavior matches intended business logic. Detecting such defects requires reasoning over the specification rather than the code alone. Software testing can expose some of these mismatches, but its effectiveness depends heavily on test design, executable artifacts, and runtime environments. This article presents a two-stage LLM-based workflow for addressing this challenge in an intelligent-vehicle cybersecurity case study. In the first stage, an AI-based rule miner extracts verifiable rules from natural-language requirements while explicitly identifying ambiguity, self-contradiction, and other non-verifiable statements. In the second stage, an AI-based code auditor checks implementation evidence against the extracted rules. Instead of asking a single LLM to directly verify code against lengthy natural-language specifications, the workflow introduces a structured intermediate representation to reduce hallucination, output variability, limited explainability, and context loss. The resulting approach is a requirement-aware and semantics-aware form of static analysis that complements software testing. By analyzing requirements and source code without requiring compilation, execution, or runtime environments, the method shifts verification and validation activities left in the development lifecycle. This LLM-based static analysis is also a new approach to addressing the test oracle problem.
AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training
arXiv:2605.17923v1 Announce Type: new Abstract: In video generation models, particularly world models, training large-scale video diffusion Transformers (such as DiT and MMDiT) poses significant computational challenges due to the extreme variance in sequence lengths within mixed-mode datasets. Existing bucket-based data loading strategies typically rely on "equal token length" constraints. This approach fails to account for the quadratic complexity of self-attention mechanisms, leading to severe load imbalance and underutilization of GPU resources. This paper proposes \textit{AdaptiveLoad}, an integrated optimization framework consisting of two core components: (1) A dual-constraint adaptive load balancing system, which eliminates long-sequence bottlenecks by simultaneously limiting memory consumption and computational load ($B \times S^p \le M_{\text{comp}}$); (2) A fused LayerNorm-Modulate CUDA kernel, which utilizes a D-tile coalesced reduction strategy to increase throughput and alleviate memory pressure. Experimental results on the Wan 2.1 world model demonstrate that our method reduces the computational imbalance rate from 39\% to 18.9\%, improves peak VRAM utilization efficiency by 22.7\%, and achieves an overall training throughput increase of 27.2\%.
TabH2O: A Unified Foundation Model for Tabular Prediction
arXiv:2605.18383v1 Announce Type: new Abstract: We present TabH2O, a foundation model for tabular data that performs classification and regression in a single forward pass via in-context learning. TabH2O builds on the TabICL architecture with several key modifications: (1) unified training, a single model handles both classification and regression via a dual-head architecture, eliminating the need for separate models and reducing total pretraining cost; (2) single-stage pretraining, training stability improvements (bounded scalable softmax, inter-stage normalization, learnable residual scaling, logit soft-capping) eliminate the need for multi-stage curriculum learning, enabling training with full-length sequences from the start; and (3) noise-aware pretraining, synthetic datasets include explicit noise dimensions to teach the model robustness to irrelevant features. We evaluate TabH2O v1 (29.2M parameters) on the TALENT benchmark (300 datasets), where it achieves an average rank of 2.55 out of 6 evaluated methods, outperforming tuned CatBoost (4.07), H2O AutoML (4.18), and LightGBM (5.08), competitive with TabPFN v2.6 (2.74), and behind TabICL v2 (2.12), while placing in the top-3 on 81% of the testing datasets across classification and regression tasks.
Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective
arXiv:2604.23267v2 Announce Type: replace Abstract: Large language models (LLMs) operate in two fundamental learning modes - fine-tuning (FT) and in-context learning (ICL) - raising key questions about which mode yields greater language proficiency and whether they differ in their inductive biases. Prior studies comparing FT and ICL have yielded mixed and inconclusive results due to inconsistent experimental setups. To enable a rigorous comparison, we propose a formal language learning task - offering precise language boundaries, controlled string sampling, and no data contamination - and introduce a discriminative test for language proficiency, where an LLM succeeds if it assigns higher generation probability to in-language strings than to out-of-language strings. Empirically, we find that: (a) FT has greater language proficiency than ICL on in-distribution generalization, but both perform equally well on out-of-distribution generalization. (b) Their inductive biases, measured by the correlation in string generation probabilities, are similar when both modes partially learn the language but diverge at higher proficiency levels. (c) Unlike FT, ICL performance differs substantially across models of varying sizes and families and is sensitive to the token vocabulary of the language. Thus, our work demonstrates the promise of formal languages as a controlled testbed for evaluating LLMs, behaviors that are difficult to isolate in natural language datasets. Our source code is available at https://github.com/bishwamittra/formallm.
Verifier-Guided Code Translation via Meta-Step Decoding
arXiv:2605.17626v1 Announce Type: new Abstract: Test-time scaling is an important mechanism for improving large language models, especially on tasks with deterministic verifiers. Code translation is a canonical example: the source program constrains valid outputs, while compilers, type check- ers, and behavioral checks provide exact pass/fail feedback. Existing approaches typically apply these verifiers only after generation, which is inefficient because early errors corrupt the autoregressive context and are rarely corrected later. We introduce Decoding Time Verification (DTV), a framework that treats structural boundaries as meta steps for verifier-guided decoding. DTV interleaves generation with verifier calls under a state-machine controller that enforces valid prefixes, using structural-boundary checks and structure-aware rollback to prevent error propagation while reducing wasted tokens. We evaluate DTV on C-to-Rust and JavaScript-to-TypeScript translation. Using Qwen3-4B as the primary generator under matched token budgets, DTV improves pass rates from 72.3% to 82.0% on C-to-Rust and from 33.3% to 46.0% on JavaScript-to-TypeScript relative to matched self-refinement baselines, while using fewer tokens per case; the same trend largely transfers to Gemma-4-E4B. In the evaluated cost-matched grid, DTV achieves a more favorable pass-rate-cost tradeoff than post-hoc verification or sampling-based scaling. These results show that verifier-guided decoding is an effective use of inference-time compute for code translation.
LitXBench: A Benchmark for Extracting Experiments from Scientific Literature
arXiv:2604.07649v4 Announce Type: replace Abstract: Aggregating experimental data from papers enables materials scientists to build better property prediction models and to facilitate scientific discovery. Recently, interest has grown in extracting not only single material properties but also entire experimental measurements. To support this shift, we introduce LitXBench, a framework for benchmarking methods that extract experiments from literature. We also present LitXAlloy, a dense benchmark comprising 1426 total measurements from 19 alloy papers. By storing the benchmark's entries as Python objects, rather than text-based formats such as CSV or JSON, we improve auditability and enable programmatic data validation. We find that frontier language models, such as Gemini 3.1 Pro Preview, outperform existing multi-turn extraction pipelines by up to 0.37 F1. Our results suggest that this performance gap arises because extraction pipelines associate measurements with compositions rather than the processing steps that define a material.
Diffusional earthquakes and their slip-distance scaling
arXiv:2604.07630v2 Announce Type: replace Abstract: The final size of an earthquake typically cannot be predicted from its ongoing seismic radiation. Expanding observations reveal distinct exceptions, such as slow earthquakes, injection-induced seismicity, and earthquake swarms, in which fault slip has an upper bound. A common thread among these anomalies is the diffusive migration of their active areas. Here, we report a unified scaling relation for these diffusional earthquakes. By tracking prolonged earthquake swarms in Northeast Japan, we constrained the time evolution of their active seismicity areas and cumulative seismic moments. Their moment-duration trajectories coincide with the final states documented for global swarms and induced seismicity across various scales. When plotted as seismic moment versus seismicity area, their trajectories collapse onto those of slow earthquakes, uniformly explained by a diffusional constant-slip model. This constant-slip scaling carves out a unique class of diffusional earthquakes, where the final available seismic energy is predetermined by slip distance.
Coordination Control of Discrete Event Systems under Cyber Attacks
arXiv:2309.11965v4 Announce Type: replace Abstract: In this paper, coordination control of discrete event systems under joint sensor and actuator attacks is investigated. Sensor attacks are described by a set of attack languages using a proposed ALTER model. Several local supervisors are used to control the system. The goal is to design local supervisors to ensure safety of the system even under cyber attacks (CA). The necessary and sufficient conditions for the existence of such supervisors are derived in terms of conditional decomposability, CA-controllability and CA-observability. A method is developed to calculate local state estimates under sensor attacks. Two methods are also developed to design local supervisors, one for discrete event systems satisfying conditional decomposability, CA-controllability and CA-observability, and one for discrete event systems satisfying conditional decomposability only. The approach works for both stealthy and non-stealthy attacks. A practical example is given to illustrate the results.
Learning Fill-in Reduction Ordering via Graph Policy Optimization for Sparse Matrices
arXiv:2605.17362v1 Announce Type: new Abstract: Matrix reordering in large sparse solvers seeks a permutation that minimizes factorization fill-in to reduce memory and computation. Because the minimum fill-in ordering problem is NP-complete and fill-in is implicit in the sparsity pattern, graph-theoretic heuristics are used. Existing reinforcement learning methods either ignore sparsity patterns--missing the global fill-in--or lack local exact fill-in feedback. We propose a graph policy optimization method, modeling fill-ins from global and local views: both the policy and value networks use a multi-hop graph neural backbone to embed global fill-in; the policy further interacts with symbolic factorization over graphs to extract local, step-level fill-ins, and the resulting feedback is aligned with the value network via an adaptive saturation function to improve convergence. On the SuiteSparse Matrix Collection, our method achieves mean reductions of 29.3 in fill-ins and 31.3 in peak memory usage over state-of-the-art baselines.
A Survey on Foundation Models for Personalized Federated Intelligence
arXiv:2505.06907v2 Announce Type: replace Abstract: The rise of large language models (LLMs), such as ChatGPT, Gemini, and Grok, has reshaped the AI landscape. As prominent instances of foundational models (FMs), they exhibit remarkable capabilities in generating human-like content, pushing the boundaries towards artificial general intelligence (AGI). However, their large-scale nature, privacy sensitivity, and substantial computational demands pose significant challenges for personalized customization for end users. To bridge this gap, we present the vision of artificial personalized intelligence (API), which focuses on adapting FMs to individual users while ensuring privacy. As a central enabler of API, we propose personalized federated intelligence (PFI), a new paradigm that not only integrates the privacy benefits of federated learning (FL) with the generalization capabilities of FMs but also places personalization at its core. To this end, we first survey recent advances in FL and FMs that lay the foundation for PFI. We then explore core stages of the PFI pipeline: efficient personalization at the edge, trustworthy adaptation, and adaptive refinement via retrieval-augmented generation. Finally, we highlight future directions for enabling PFI. Overall, this survey aims to lay a foundation for the development of API as a complementary direction to AGI, with PFI as a key enabling paradigm.
Learning-based data-enabled moving horizon estimation with application to membrane-based biological wastewater treatment process
arXiv:2602.13957v3 Announce Type: replace Abstract: In this paper, we propose a data-enabled moving horizon estimation (MHE) approach for a class of nonlinear systems without explicit modeling, by leveraging Koopman operator theory and Willems fundamental lemma. Specifically, the nonlinear system is lifted to a linear parameter-varying Koopman surrogate, in which the lifting functions and scheduling mappings are learned directly from data using neural networks. Willems fundamental lemma is then employed to construct a trajectory-based representation of the Koopman surrogate, which bypasses the explicit identification of the matrices of the Koopman surrogate. Based on this representation, we formulate a convex data-enabled MHE design, which provides real-time estimates of the Koopman surrogate states, from which the states of the original nonlinear system are reconstructed. Sufficient conditions are derived to ensure the stability of the estimation error. The effectiveness of the proposed method is illustrated using a simulated membrane-based biological wastewater treatment process.
Enhanced Ionic Conductivity of confined Ionic-Liquid in Angstrom-scale 2D channels
arXiv:2605.18531v1 Announce Type: new Abstract: Understanding ion-transport under molecular confinement is essential for developing next-generation energy technologies, where ionic motion often occurs within nanoscale or angstrom-scale channels. In this study, we use the model system of 1-ethyl-3-methylimidazolium bis(trifluoromethanesulfonyl)imide ([EMIM]+[TFSI]-) confined within angstrom-scale slit-shaped 2D channels fabricated via van der Waals assembly to exemplify a broader class of confined ionic liquids.This system provides a well-defined platform to unravel generic features of ion transport under extreme confinement. By systematically varying the channel height h, we demonstrate a non-monotonic conductivity dependence on confinement, with a maximum 26.7 S/m at confining height, 1.02 nm, over 30 times of the bulk value for these ionic liquids. The variation of conductivity with confinement arises from structural rearrangements of ionic layers in the slit channel. Enhanced values of conductivity occur under confinements that promote the breakup of ion pairs and larger clusters, thereby increasing the number of free ions. Stronger confinement (h, 0.68 nm) also leads to steric hindrance, lowering conductivity below bulk values. Furthermore, introducing co-solvents with a higher dielectric constant and lower viscosity, such as acetonitrile (ACN), amplifies conductivity to ~145 S/m. Comparative studies using ACN, dimethyl carbonate and diethyl carbonate highlight that both large dielectric constant and low viscosity critically govern ion transport under confinement, as also supported by molecular dynamics simulations. Overall, this work establishes confined [EMIM]+[TFSI]- as a representative system for probing mechanisms of nano- and angstrom-scale ion transport, demonstrating how nanoconfinement and the solvent environment can be systematically tuned to manipulate ionic conductivity at the molecular level.
Capacitated power dominating set problem: a solution approach based on forbidden propagation sets
arXiv:2605.18533v1 Announce Type: cross Abstract: The optimal placement of measurement devices in electrical power systems is commonly modeled through the power dominating set problem. However, in real-world applications, these devices have limited capacities, leading to a capacitated variant of the problem that has received little attention in the literature. In this work, we introduce forbidden propagation sets, novel combinatorial structures that cannot occur simultaneously in any feasible solution. This notion enables a new class of integer linear programming formulations. They combine infection-based variables with exponentially many constraints, while avoiding big-$M$ constraints. We derive structural properties, valid inequalities, and redundancy-breaking constraints, and design an efficient lazy-separation procedure based on cycle detection. Computational experiments on benchmark instances with up to 14,000 vertices show that the proposed method achieves an average execution-time improvement of 1.7x over existing approaches adapted from the literature. Moreover, the results indicate that performance depends not only on network size, but also on capacities.
$\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
arXiv:2605.17862v1 Announce Type: new Abstract: Scaling on-policy distillation (OPD) for large language models (LLMs) confronts a fundamental tension: asynchronous execution is necessary for system efficiency, but structurally deviates from the ideal on-policy objective. To address this challenge, we theoretically decompose the objective discrepancy into rollout drift and supervision drift, capturing staleness in student rollout and teacher context, respectively. Building on this, we introduce a sample-level freshness score that quantifies the reliability of a buffered sample with respect to the on-policy objective. Guided by this signal, we further propose f-OPD, a novel framework that adaptively regulates stale-sample influence and constrains policy drift accumulated under asynchronous training. Across reasoning, tool-use, and coding-agent tasks of increasing interaction horizon, f-OPD consistently achieves task performance comparable to synchronous optimization while largely retaining the throughput advantages of asynchronous execution. Our results establish the first recipe for achieving a performance-efficiency trade-off in OPD, paving the way for long-horizon agentic post-training at scale.
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
arXiv:2604.15851v3 Announce Type: replace Abstract: Differential privacy (DP) has a wide range of applications for protecting data privacy, but designing and verifying DP algorithms requires expert-level reasoning, creating a high barrier for non-expert practitioners. Prior works either rely on specialized verification languages that demand substantial domain expertise or remain semi-automated and require human-in-the-loop guidance. In this work, we investigate whether large language models (LLMs) can automate DP reasoning. We introduce DPrivBench, a benchmark in which each instance asks whether a function or algorithm satisfies a stated DP guarantee under specified assumptions. The benchmark is carefully designed to cover a broad range of DP topics, span diverse difficulty levels, and resist shortcut reasoning through trivial pattern matching. Experiments show that while the strongest models handle textbook mechanisms well, all models struggle with advanced algorithms, revealing substantial gaps in current DP reasoning capabilities. Through further analytic study and failure-mode analysis, we identify several promising directions for improving automated DP reasoning. Our benchmark provides a solid foundation for developing and evaluating such methods, and complements existing benchmarks for mathematical reasoning.
StatQAT: Statistical Quantizer Optimization for Deep Networks
arXiv:2605.17745v1 Announce Type: cross Abstract: Quantization is essential for reducing the computational cost and memory usage of deep neural networks, enabling efficient inference on low-precision hardware. Despite the growing adoption of uniform and floating-point quantization schemes, selecting optimal quantization parameters remains a key challenge, particularly for diverse data distributions encountered during training and inference. This work presents a novel statistical error analysis framework for uniform and floating-point quantization, providing theoretical insight into error behavior across quantization configurations. Building on this analysis, we propose iterative quantizers designed for arbitrary data distributions and analytic quantizers tailored for Gaussian-like weight distributions. These methods enable efficient, low-error quantization suitable for both activations and weights. We incorporate our quantizers into quantization-aware training and evaluate them across integer and floating-point formats. Experiments demonstrate improved accuracy and stability, highlighting the effectiveness of our approach for training low-precision neural networks.
An Omni-Temporal Theory for Hydrodynamic Dispersion and Reaction in Porous Media
arXiv:2505.06063v2 Announce Type: replace Abstract: A frequency-based omni-temporal dispersion theory is developed to capture the transient interplay between diffusion, advection, and reaction during solute transport through porous media. Unlike classical asymptotic dispersion theories, which commonly rely on long-time approximation, the proposed framework simultaneously captures both fast and slow components of dispersion. The theory is formulated by volume averaging the Fourier-transformed pore-scale advection-diffusion equation, yielding four frequency-dependent upscaled transport coefficients for a periodic unit cell: a dispersion tensor, an advection-suppression transfer function, a spectral Sherwood number, and a reactivity-bias vector. These coefficients act as transfer functions that relate microscopic driving forces to corresponding effective fluxes in the frequency domain, enabling prediction of transient transport dynamics in the time domain through inverse Fourier transformation. The utility of the proposed framework is demonstrated by deriving analytical expressions for the transfer functions in Poiseuille flow between parallel plates and through circular tubes, and subsequently using them within a Fast Fourier Transform framework to obtain breakthrough curves. For fast solute pulses between inactive parallel plates, the proposed theory produces breakthrough curves in close agreement with direct numerical simulations, whereas conventional asymptotic theory overpredicts propagation rates by orders of magnitude. Finally, the framework is applied to reactive and non-reactive porous media consisting of periodic arrays of square rods under cross flow, demonstrating the generality and versatility of the proposed omni-temporal theory.
Information-Theoretic Storage Cost in Sentence Comprehension
arXiv:2602.18217v2 Announce Type: replace Abstract: Real-time sentence comprehension imposes a significant load on working memory, as comprehenders must maintain contextual information to anticipate future input. While measures of such load have played an important role in psycholinguistic theories, they have largely been formalized using symbolic grammars, which assign discrete, uniform costs to syntactic predictions. This study proposes a measure of processing storage cost based on an information-theoretic formalization, as the amount of information previous words carry about future context, under uncertainty. Unlike previous discrete, grammar-based metrics, this measure is continuous, probabilistic, theory-neutral, and can be estimated from pre-trained neural language models. The validity of this approach is demonstrated through three analyses in English: our measure (i) recovers well-known processing asymmetries in center embeddings and relative clauses, (ii) correlates with a grammar-based storage cost in a syntactically-annotated corpus, and (iii) predicts reading-time variance in two large-scale naturalistic datasets over and above baseline models with traditional information-based predictors. Our code is available at https://github.com/kohei-kaji/info-storage.