Forskningsradar

Science Journals

Peer-reviewade publikationer — 56237 artiklar

Cross-Cluster Weighted Forests
arXiv:2105.07610v5 Announce Type: replace-cross Abstract: Building trustworthy machine learning algorithms for biological applications requires adapting to data heterogeneity from different sources, batches, distributions, or studies. We propose the 'Cross-Cluster Weighted Forest' (CCWF), an ensembling approach that explicitly leverages heterogeneity in the feature distribution to produce more accurate and more generalizable predictors than the standard Random Forest in cases when data can be naturally clustered. CCWF generalizes the RF architecture to an outer unsupervised layer, supervised subtasks, and ensembling. Specifically it involves unsupervised clustering of the training data, fitting a Random Forest on each cluster, and combining the forests via stacked regression weights that reward cross-cluster generalizability. We provide a theoretical analysis of an analytically tractable forest model showing that cluster-based ensembling is asymptotically more accurate than training a single forest on the full data, with the gain driven by bias reduction. In simulations, we find that CCWF is robust across data-generating regimes and outcome models; furthermore, we explore the influence of data partitioning and ensemble weighting strategies on the benefits of our method. Finally, we apply our approach to cancer molecular profiling and gene expression datasets that are naturally divisible into clusters; in both simulations and real data examples, we illustrate that our approach outperforms classic Random Forest by margins of 30-40%, aligning with our theoretical results. Overall, we show that CCWF provides a statistically grounded prediction algorithm for data spanning multiple domains or sub-populations, a structure common in biological applications.
Strong Spatial Mixing for General 2-Spin Systems: A Unified Approach from Zero-Freeness
arXiv:2401.09317v4 Announce Type: replace-cross Abstract: We study the algorithmic implications of zero-free regions for the partition functions of 2-spin systems. While Barvinok's algorithm yields FPTASes in such regions, the applicability of Weitz's algorithm is limited to parameter regimes where strong spatial mixing (SSM) can be established. It remains open whether Weitz's algorithm can be applied to general zero-free regions, particularly in settings where no standard tree-recurrence-based proof of SSM is known. We establish new SSM results and thereby extend the applicability of Weitz's FPTAS to all currently known zero-free regions of 2-spin systems with pinned vertices. We achieve this through a unified approach to deriving SSM from zero-freeness in the most general settings of 2-spin systems. Our work features two key innovations. 1.Our SSM results cover parts of the celebrated Lee-Yang zero-free region for the ferromagnetic Ising model, where no tree-recurrence-based proof of SSM is currently known or considered feasible. The tree recurrence method typically relies on carefully designed potential functions, the construction and analysis of which can be highly challenging. For ferromagnetic 2-spin systems, it remains an open challenge whether such potential functions can be constructed. We circumvent this difficulty by deriving SSM directly from zero-freeness. 2.The prior approach to deriving SSM from zero-freeness relies on cluster expansions, which are model-specific and known only for a few restricted parameter settings such as the hard-core model near vertex activity $\lambda=1$. We overcome this obstacle by introducing a purely combinatorial approach based on a novel Christoffel-Darboux-type identity that holds universally for 2-spin systems. This provides a broadly applicable framework for handling general 2-spin systems with arbitrary multivariate parameters and zero-free regions of arbitrary shape in a unified manner.
A categorical formulation of Kraus' paradox
arXiv:2403.17961v2 Announce Type: replace-cross Abstract: We give a categorical formulation of Kraus' "magic trick" for recovering information from truncated types. Rather than type theory, we work in Van den Berg-Moerdijk path categories with a univalent universe, and rather than propositional truncation we work with arbitrary cofibrations, which includes truncation as a special case. We show, using Kraus' argument that any cofibration with homogeneous domain is a monomorphism. We give some simple concrete examples in groupoids to illustrate the interaction between homogeneous types, cofibrations and univalent fibrations.
Benchmarking trigonometric continuous-variable gate primitives with trapped ions
arXiv:2607.14085v1 Announce Type: cross Abstract: Hybrid continuous-discrete-variable quantum processors can represent bosonic degrees of freedom directly in oscillator modes, or qumodes, while using qubits for control, readout, and nonlinear operations. Recently proposed trigonometric continuous-variable (CV) gate sets promote periodic functions of oscillator quadratures to elementary operations, making them natural primitives for compact variables, rotor models, lattice gauge theories, and anharmonic dynamics. Here we experimentally demonstrate and benchmark one-qumode cosine gates, and perform a mode-resolved marginal benchmark of two-qumode cosine-gate implementations, on the QSCOUT trapped-ion quantum platform. Our implementation uses collective motional modes of three- and four-ion $^{171}{\rm Yb}^{+}$ chains and realizes finite-step trigonometric-gate circuits through hybrid qubit-qumode operations and conditional phase-space displacements. In contrast to previous theoretical and compilation work, we focus on the gate-level characterization of the trigonometric primitives. We measure Fock-space transition probabilities, study their dependence on gate parameters and Trotter step number, and compare with simulations incorporating thermal initialization and motional dephasing. We also derive ideal gate matrix elements and phase-space diagnostics, connecting the measurements to the non-Gaussian structure generated by these gates. These results establish trigonometric CV gates as reusable building blocks for bosonic Hamiltonian simulations and hybrid quantum algorithms requiring intrinsically non-polynomial operations.
Generalized Neural Distributional Regression
arXiv:2607.14122v1 Announce Type: cross Abstract: We introduce the Generalized Neural Distributional Regression (GNDR) framework, which seamlessly embeds deep neural networks into the parameter space of classical probability distributions. To reconcile the inherent non-identifiability of deep architectures with maximum likelihood theory, we propose a two-step semi-parametric estimation procedure. By isolating the terminal prediction heads and treating the upstream network as a fixed, non-linear basis expansion, GNDR enables the extraction of analytical Fisher Information matrices. This facilitates rigorous uncertainty quantification, generating observation-specific confidence bands and tolerance intervals via the multivariate Delta method. We demonstrate the framework's versatility and superior distributional calibration across diverse data modalities, including overdispersed clinical counts, right-censored transcriptomic survival profiles under a mixture cure framework, and zero-truncated age distributions derived directly from unstructured facial images. The methodology is natively implemented in the open-source Python package \textit{thetaflow}.
Does Multi-Agent Debate Improve AI Feedback on Research Papers?
arXiv:2607.14713v1 Announce Type: cross Abstract: Probably not, at least for meta-analyses in economics. In a pre-registered, identity-masked, within-paper experiment, the authors of 44 meta-analyses ranked three AI reports on their own paper by usefulness for improving it: a single pass by a frontier model against two multi-agent debate tools we built and expected to win. All reports were held to a common length and template. The authors preferred the single pass, by 0.66 rank points over mad-research (95% CI 0.32 to 1.00) and 0.57 over paper-workshop (0.16 to 0.95), though paper-workshop spent roughly thirty times the tokens. Authors who recalled their journal referee report usually placed it first and never last; in a separate exercise, three AI judges almost always placed the real journal referee report last. Among the three AI reports, Gemini (the judge whose model family wrote none of the reports) would have ranked paper-workshop first in the authors' place, reversing the single-pass preference. The reversal warns against substituting an AI judge for the author. We measure perceived usefulness for finished papers; whether AI should referee papers is a separate question.
OvAi Focus: AI-based Multi-class Segmentation of Functional Ovaries and Adnexal Masses in Gynecological Ultrasound
arXiv:2607.14179v1 Announce Type: cross Abstract: Ovarian cancer is the deadliest gynecological malignancy; accurate and objective segmentation of adnexal masses and functional ovaries in ultrasound (US) remains challenging due to operator variability and morphological complexity. We present OvAi Focus (SynDiag s.r.l., Italy), a stand-alone AI software medical device that performs multi-class semantic segmentation of functional ovaries and adnexal masses, distinguishing cystic from solid components. The system was trained and independently validated on a multicenter dataset of 1,081 adult women from 6 centers across Italy and Israel. Segmentation achieved DICE scores of 0.87 (complete lesion), 0.85 (cystic), 0.68 (solid), and 0.62 (functional ovary), in line with or superior to state-of-the-art approaches across heterogeneous acquisition settings.
Ultra-long simulations of collisionless relativistic shocks in front-comoving frame: evidence for a steady state and its properties
arXiv:2607.14235v1 Announce Type: cross Abstract: We present a series of unprecedently long 2D3V PIC simulations of unmagnetized relativistic $e^{-}e^{+}$-pair shocks performed in a front-comoving frame. By implementing a moving-wall boundary condition in the downstream together with continuous injection at the upstream boundary, we maintain a fixed simulation domain size, opening the way to perform substantially longer simulations. Our longest runs extend beyond $100000\,\omega_p^{-1}$, exceeding the duration of the previously published simulations by a factor of several. Across a diverse set of simulations -- varying upstream/downstream lengths, transverse sizes, and particle-per-cell counts -- we find strong evidence that the shock approaches an asymptotic, time-independent state. In the downstream region, the steady state depends only on the upstream temperature at the injection boundary and does not depend on a particular numerical realization. The upstream precursor evolves slower and retains a dependence on the simulation's upstream length, that may be of minor observational consequence, since radiation from astrophysical shocks predominantly originates from the downstream region. We also find that Fermi-type acceleration is limited in energy and a true power-law tail never forms. Another important finding is that the downstream magnetic field has a soliton-like structure, where individual magnetic domains evolve independently, each comprising a compact, highly magnetized core embedded within an extended, weakly magnetized region. The magnetic-field distribution around the centers of these spots has approximately Lorentzian profile.
Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks
arXiv:2512.01208v5 Announce Type: replace Abstract: In standard Transformer architectures, semantic importance is often conflated with activation magnitude, obscuring the geometric structure of latent representations. To disentangle these factors, we introduce PRISM, a complex-valued architecture designed to isolate the computational role of phase. By enforcing a strict unit-norm constraint ($|z| = 1$) and replacing attention with gated harmonic convolutions, the model is encouraged to utilize subtractive interference in the frequency domain to suppress noise, rather than relying on magnitude-based gating. We utilize this constrained regime to study a hybrid architecture -- fusing phase-based routing with standard attention -- which achieves improved parameter efficiency and representation quality compared to baselines in our evaluated settings. Mechanistically, interventional ablations indicate that the model carries substantial task-relevant information in phase: preserving phase largely maintains performance, whereas disrupting phase causes severe degradation. Together, these results suggest that phase-based spectral interference is a usable computational mechanism for neural sequence modeling at the evaluated scale.
Ab initio calculations of $^{229}$Th band-to-band internal conversion rate in $^{229}$ThO$_2$
arXiv:2607.08941v1 Announce Type: cross Abstract: We present an ab initio calculation of the band-to-band internal-conversion rate of the $\hbar\omega_{\rm nuc} \approx 8.35$ eV isomeric transition in $^{229}$ThO$_2$. Because the nuclear transition energy exceeds the electronic band gap of ThO$_2$, the isomer can decay nonradiatively by resonantly promoting a valence electron into the conduction band. We formulate this process as a Brillouin-zone sum over vertical interband transitions weighted by local Th-centered hyperfine matrix elements, which are evaluated directly from all-electron full-potential linearized augmented-plane-wave Bloch spinors. A finite nuclear magnetization model is included to regularize the short-range hyperfine interaction and to account for the Bohr-Weisskopf effect. After applying scissor shifts to span the experimentally reported ThO$_2$ band gaps, we find calculated internal-conversion lifetimes in the range of $1-16~\mu{\rm s}$. The lifetime increases strongly as the band gap approaches $\omega_{\rm nuc}$ because the resonant interband phase space at the nuclear transition energy is reduced. For the larger reported ThO$_2$ gaps, the calculated lifetime is comparable to the measured conversion-electron M\"ossbauer lifetime [Nature 648, 300 (2025)]. Our analysis implies that choosing solid-state hosts with band-gap values slightly lower than $\omega_{\rm nuc}$ can optimize solid-state nuclear clock performance with internal-conversion electron readout.
Renormalisation of Inhomogeneous Random Graphs
arXiv:2607.12459v1 Announce Type: cross Abstract: We consider inhomogeneous random graphs in which vertices are assigned i.i.d.\ random weights, pairs of distinct vertices are connected by an edge independently with a probability that is a bi-variate function of the weights of the vertices, and single vertices are connected to themselves by a self-loop independently with a probability that is a uni-variate function of the weight of the vertex. We apply a renormalisation transformation in which vertices are aggregated into groups of equal size according to a greedy algorithm, namely, distinct groups of aggregated vertices are connected by an aggregated edge if and only if there is at least one edge connecting two constituent vertices across the groups, while a group of aggregated vertices is connected to itself by an aggregated self-loop if and only if there is at least one self-loop at an internal vertex or one edge connecting a pair of distinct internal vertices. We analyse what happens when the renormalisation transformation is iterated. In particular, we show that, starting from appropriately scaled connection functions, the iterated renormalised graphs converge to a two-parameter family of random graphs, acting as an attractor in a universality class. We consider a light-tailed regime, for which the scaling limit is a homogeneous Erd\H{o}s--R\'enyi random graph, and a heavy-tailed regime, for which the scaling limit is an inhomogeneous random graph with stable infinite-mean random weights and an exponential disconnection function. Different scalings are needed for the two regimes. Which of the two regimes prevails depends on the choice of the connection functions and the choice of the law of the random weights.
CoDi -- an exemplar-conditioned diffusion model for low-shot counting
arXiv:2512.20153v2 Announce Type: replace Abstract: Low-shot object counting addresses estimating the number of previously unobserved objects in an image using only few or no annotated test-time exemplars. A considerable challenge for modern low-shot counters are dense regions with small objects. While total counts in such situations are typically well addressed by density-based counters, their usefulness is limited by poor localization capabilities. This is better addressed by point-detection-based counters, which are based on query-based detectors. However, due to limited number of pre-trained queries, they underperform on images with very large numbers of objects, and resort to ad-hoc techniques like upsampling and tiling. We propose CoDi, the first latent diffusion-based low-shot counter that produces high-quality density maps on which object locations can be determined by non-maxima suppression. Our core contribution is the new exemplar-based conditioning module that extracts and adjusts the object prototypes to the intermediate layers of the denoising network, leading to accurate object location estimation. On FSC benchmark, CoDi outperforms state-of-the-art by 15% MAE, 13% MAE and 10% MAE in the few-shot, one-shot, and reference-less scenarios, respectively, and sets a new state-of-the-art on MCAC benchmark by outperforming the top method by 38% MAE. The code is available at https://github.com/gsustar/CoDi
Self-organized defect clustering and concentration-dependent vacancy diffusion in MoS$_2$
arXiv:2607.14951v1 Announce Type: cross Abstract: Sulfur vacancy migration has a crucial impact on electronic transport and the functional behavior of MoS$_2$-based devices such as memristors and memtransistors. According to recent atomistic simulations, vacancy migration proceeds via cooperative, vacancy-assisted sulfur jumps, implying strongly correlated defect dynamics. Here, we investigate the collective behavior of sulfur-vacancy clusters in MoS$_2$ using kinetic Monte-Carlo simulations with transition rates derived from machine learning interatomic potential molecular dynamics simulations. We identify three transport regimes: At low concentrations, vacancies are immobile or confined within small clusters, whereas at high concentrations, classical diffusive transport with a constant diffusion coefficient is observed, and vacancies aggregate into anisotropically extended clusters. A well defined intermediate regime is characterized by clusters merging into a connected, fluctuating network with a concentration-dependent diffusion coefficient. This regime is characterized by a broad distribution of cluster sizes. The strong dependence of the vacancy diffusion coefficient on the average defect concentration provides new insights into the origin of memristive behavior observed in MoS$_2$.
Parallel analog quantum simulation in homogeneous quantum gases
arXiv:2607.14977v1 Announce Type: cross Abstract: Analog quantum simulation offers a powerful way to study strongly correlated quantum systems that are beyond the reach of classical computation. In this context, ultracold atomic gases have been demonstrated to be an exceptionally versatile and well-controlled platform for implementing various quantum Hamiltonians. In this work, we extend this level of control to a multiplexed configuration in which distinct quantum-simulation units are independently controlled and engineered starting from a single atomic cloud. We demonstrate multiplexed operation in two representative settings. First, by shaping box-trap potentials and separately controlling the evaporative cooling trajectories, we prepare subsystems at various temperatures across the superfluid transition of the unitary Fermi gas. Second, we demonstrate parallel quantum simulation of the Josephson Hamiltonian across distinct Josephson-junction quantum simulation units with individually tunable parameters, including local phase control to initialize the dynamics. Our scheme provides a versatile route toward systematic studies of dynamics and transport Hamiltonians in strongly correlated ultracold matter. Moreover, it is readily extendable to a wide range of atomic species, geometries, and dimensionalities.
Atmospheric evolution through outgassing and escape on young molten rocky exoplanets
arXiv:2607.15011v1 Announce Type: cross Abstract: The earliest rocky planet atmospheres are shaped by competition between initial volatile inventories and atmospheric escape. On young magma ocean planets, outgassing competes with atmospheric escape, controlling volatile retention and atmospheric evolution. We investigate how atmospheric escape and replenishment via outgassing during magma ocean crystallization shape rocky planet atmospheres. We extend a coupled interior-atmosphere model to simulate rocky planet evolution during the magma ocean era by incorporating an energy-limited atmospheric escape module. Comparing radiative-convective and prescribed-convective atmospheres, we quantify how atmospheric energy transport affects escape. We explore a wide range of orbital separations, escape efficiencies, oxidation states, and initial volatile inventories to identify regimes where sustained magma-ocean outgassing or escape dominates. We estimate atmospheric loss and compositions for young rocky planets around Sun-like and M-dwarf stars over geologic timescales. Atmospheric escape shortens magma ocean lifetimes by weakening greenhouse insulation. Radiative-convective atmospheres reduce solidification timescales compared to purely convective cases. Volatile dissolution into the magma ocean interacts with escape to chemically fractionate the planetary volatile budget over time by retaining more soluble species. For Earth-mass planets, atmospheres survive if loss rates remain moderate. Mantle redox state remains a key control on retained atmospheric composition: high oxygen fugacity (fO2) yields heavier, H2O- and CO2-rich atmospheres, while low fO2 produces light, H2- or CO-dominated atmospheres, consistent with previous studies. Orbital separation, initial volatile inventory, and stellar type produce diverse evolutionary pathways, from bare rocky planets to magma oceans with thick atmospheres, ranging from H2- to SO2-dominated.
Implicit differentiation of tensor network algorithms
arXiv:2607.15030v1 Announce Type: cross Abstract: The current leading approach to the variational optimization of projected entangled-pair states (PEPS) is based on automatic differentiation, which allows for a convenient evaluation of the energy gradient with respect to the local variational degrees of freedom. However, evaluating the energy gradient not only remains a major computational bottleneck of the optimization procedure, but also suffers from frequent numerical instabilities. In this work, we adopt recent advances in implicit differentiation techniques to address these challenges in PEPS optimization. By reformulating the core step of the gradient computation in terms of a single characteristic equation for the contraction environment, we reduce the cost of the gradient computation and improve its scaling with the problem size. By choosing a suitable parametrization of this characteristic equation based on the intrinsic symmetries of the contraction environment, we can directly remove instabilities from the global gradient computation that would otherwise arise from the derivatives of subroutines of the contraction algorithm. Finally, we demonstrate how this approach drastically simplifies the practical implementation of stable gradient-based PEPS optimization.
ESAR: Event-Based Synthetic Aperture Reconstruction
arXiv:2607.15073v1 Announce Type: cross Abstract: Event cameras report asynchronous polarity events when changes in log--radiance exceed a fixed contrast threshold, producing signed temporal contrast measurements rather than conventional image frames. We formulate monocular event-based imaging as a synthetic-aperture inverse problem for a static ground-domain log--radiance field $\theta \in \mathbb{R}^{N_g}$. Instead of reconstructing a latent pixel-time volume $v \in \mathbb{R}^{N_pN_t}$, we impose the geometric relation $v=P\theta$, where $P$ maps the fixed scene into motion-dependent latent views. Aggregating events over finite time intervals gives the linearized model \[ AP\theta = b+\eta, \] where $A$ is a temporal differencing operator, $b$ contains signed binned event counts, and $\eta$ represents measurement and modeling errors. This decomposition exposes a synthetic-aperture structure: under near-nadir motion, successive projections are approximately shifted views of a common scene, while the composite operator $AP$ remains ill-conditioned because it combines spatial averaging with temporal differencing. We therefore use regularized inversion to recover $\theta$. Numerical experiments on simulated data and real near-nadir Falcon Neuro event data show that the proposed $\theta$-based formulation recovers coherent large-scale spatial structure, relative to dynamic latent-image and learned event-reconstruction baselines, while suppressing fine-scale texture.
Thermodynamics of quantum oscillators
arXiv:2507.04268v2 Announce Type: replace-cross Abstract: In this work, we present a compact analytical approximation for the quantum partition function of systems composed of quantum oscillators. The proposed formula is general and applicable to an arbitrary number of oscillators described by a rather general class of potential energy functions (not necessarily polynomials). Starting from the exact path integral expression of the partition function, we introduce a temperature-dependent Gaussian approximation for the high-temperature propagator and, then, invoke a principle of minimal sensitivity to minimize the error. This leads to a system of coupled nonlinear equations whose solution yields the optimal parameters of the Gaussian approximation. The resulting approximate partition function accurately reproduces thermodynamic quantities such as the free energy, average energy, and specific heat -- even at zero temperature -- with typical relative errors in the range of about 1\%--5\%. The accuracy deteriorates only moderately when the anharmonicity and coupling strengths are increased. We illustrate the performance of our analytical formula with numerical results for systems of up to ten coupled anharmonic oscillators. These results are compared to "exact" numerical results obtained via Hamiltonian diagonalization for small systems and Path Integral Monte Carlo simulations for larger ones.
Converting T1-weighted MRI from 3T to 7T quality using deep learning
arXiv:2507.13782v2 Announce Type: replace-cross Abstract: Ultra-high resolution 7 tesla (7T) magnetic resonance imaging (MRI) provides detailed anatomical views, offering better signal-to-noise ratio, resolution and tissue contrast than 3T MRI, though at the cost of accessibility. We present an advanced deep learning model for synthesizing 7T brain MRI from 3T brain MRI. Paired 7T and 3T T1-weighted images were acquired from 172 participants (124 cognitively unimpaired, 48 impaired) from the Swedish BioFINDER-2 study. To synthesize 7T MRI from 3T images, we trained two models: a specialized U-Net, and a U-Net integrated with a generative adversarial network (GAN U-Net). Our models outperformed two previous state-of-the-art 3T-to-7T models in image-based evaluation metrics. Four blinded MRI professionals judged our synthetic 7T images as comparable in detail to real 7T images, and superior in subjective visual quality to 7T images, due to the reduction of artifacts. Using both SynthSeg and NextBrain, automated segmentations of the synthetic 7T images were more similar to real 7T segmentations than automated segmentations from the 3T images that were used to synthesize the 7T images. Finally, synthetic 7T images showed similar performance to real 3T images in downstream prediction of cognitive status using MRI derivatives (n=3,168). In all, we show that synthetic T1-weighted brain images approaching 7T quality can be generated from 3T images, which may improve image quality and segmentation, without compromising performance in downstream tasks. Future directions, possible clinical use cases, and limitations are discussed.
Mixed-State Phase Transitions in Measurement-Dressed Imaginary-Time Evolution
arXiv:2511.04402v4 Announce Type: replace-cross Abstract: Motivated by the ubiquity of decoherence in quantum hardware and the growing role of imaginary-time evolution (ITE) in quantum algorithms, we investigate how many-body correlations generated by imaginary-time filtering are modified by local decoherence. We introduce measurement-dressed imaginary-time evolution (MDITE), which alternates ITE with projective-measurement channels, producing a competition between low-energy filtering and local dephasing. By developing a new efficient quantum Monte Carlo method, we uncover MDITE mixed-state transitions with spontaneous-symmetry-breaking signatures in the driving of 1D transverse-field Ising and 2D columnar dimerized Heisenberg Hamiltonians in the resulting density matrices. In the continuous limit, the Choi-Jamiolkowski mapping yields a tractable equilibrium description with conformal criticality that qualitatively captures the phase transitions. At finite protocol parameters, however, the four-point correlator violates the conformal cross-ratio form and the critical exponents deviate from their continuous-limit values, signaling the loss of conformal symmetry and richer nonequilibrium criticality. Our results establish MDITE as a controlled setting for exploring mixed-state phases and critical phenomena driven by the interplay between imaginary-time filtering and decoherence.
Micro-macro kinetic flux-vector splitting schemes for the multidimensional Boltzmann-ES-BGK equation
arXiv:2509.21832v2 Announce Type: replace Abstract: The kinetic Boltzmann equation models gas dynamics over a wide range of spatial and temporal scales. Simplified versions of the full Boltzmann collision operator, such as the classical Bhatnagar-Gross-Krook (BGK) and the closely related Ellipsoidal-Statistical-BGK (ES-BGK) operators, can dramatically reduce the computational cost of solving kinetic equations numerically. Classical BGK yields incorrect transport coefficients (relative to the full Boltzmann collision operator) at low Knudsen numbers, whereas ES-BGK captures them correctly. In this work, we develop a finite-volume method based on a micro-macro decomposition of the distribution function, which requires a smaller velocity mesh than direct kinetic methods for low and intermediate Knudsen numbers. The macro portion of the model is a fluid model with a moment closure derived from the heat-flux tensor calculated from the micro portion. The micro portion is obtained by applying to the original kinetic equation a projector into the orthogonal complement of the null space of the collision operator -- this projector depends on the macro portion. In particular, we extend the technique of Bennoune, Lemou, and Mieussens [{\it Uniformly stable schemes for the Boltzmann equation preserving the compressible Navier-Stokes asymptotics, J. Comput. Phys. (2008)}] to two-space dimensions, the ES-BGK collision operator, and problems with reflecting wall boundary conditions. The collision operator in the micro and macro equations is handled via L-stable implicit time discretizations, while the transport terms are computed via kinetic flux vector splitting (for the macro equations) and upwind differencing (for the micro equation). The resulting scheme is applied to various test cases in 1D and 2D. The 2D version of the code is parallelized using MPI, and we present weak- and strong-scaling studies with varying numbers of processors.
Stabilizing Native Low-Rank LLM Pretraining
arXiv:2602.12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose significant computational and memory challenges. Low-rank factorization offers a promising route to reduce training and inference costs, but the community lacks a stable recipe for training models from scratch using exclusively low-rank weights while matching the performance of the dense model. We demonstrate that Large Language Models (LLMs) can be trained from scratch using exclusively low-rank factorized weights for all non-embedding matrices without auxiliary "full-rank" guidance required by prior methods. While native low-rank training often suffers from instability and loss spikes, we identify uncontrolled growth in the spectral norm (largest singular value) of the weight matrix update as the dominant factor. To address this, we introduce Spectron: Spectral renormalization with orthogonalization, which dynamically bounds the resultant weight updates based on the current spectral norms of the factors. Our method enables stable, end-to-end factorized training with negligible overhead. Finally, we establish compute-optimal scaling laws for natively low-rank transformers, demonstrating predictable power-law behavior and improved inference efficiency relative to dense models.
SafeOR-Gym: A Benchmark Suite for Safe Reinforcement Learning Algorithms on Practical Operations Research Problems
arXiv:2506.02255v2 Announce Type: replace Abstract: Most existing safe reinforcement learning (RL) benchmarks focus on robotics and control tasks, offering limited relevance to high-stakes domains that involve structured constraints, mixed-integer decisions, and industrial complexity. This gap hinders the advancement and deployment of safe RL in critical areas such as energy systems, manufacturing, and supply chains. To address this limitation, we present SafeOR-Gym, a benchmark suite of nine operations research (OR) environments tailored for safe RL under complex constraints. Each environment captures a realistic planning, scheduling, or control problems characterized by cost-based constraint violations, planning horizons, and hybrid discrete-continuous action spaces. The suite integrates seamlessly with the Constrained Markov Decision Process (CMDP) interface provided by OmniSafe. We evaluate several state-of-the-art safe RL algorithms across these environments, revealing a wide range of performance: while some tasks are tractable, others expose fundamental limitations in current approaches. SafeORGym provides a challenging and practical testbed that aims to catalyze future research in safe RL for real-world decision-making problems.
TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models
arXiv:2509.24566v2 Announce Type: replace Abstract: Large vision-language models (LVLMs) have achieved impressive performance across a wide range of vision-language tasks, while they remain vulnerable to backdoor attacks. Existing backdoor attacks on LVLMs aim to force the victim model to generate a predefined target pattern, which is either inserted into or replaces the original content. We find that these fixed-pattern attacks are relatively easy to detect, because the attacked LVLM tends to memorize such frequent patterns in the training dataset, thereby exhibiting overconfidence on these targets given poisoned inputs. To address these limitations, we introduce TokenSwap, a more evasive and stealthy backdoor attack that focuses on the compositional understanding capabilities of LVLMs. Instead of enforcing a fixed targeted content, TokenSwap subtly disrupts the understanding of object relationships in text. Specifically, it causes the backdoored model to generate outputs that mention the correct objects in the image but misrepresent their relationships (i.e., bags-of-words behavior). During training, TokenSwap injects a visual trigger into selected samples and simultaneously swaps the grammatical roles of key tokens in the corresponding textual answers. However, the poisoned samples exhibit only subtle differences from the original ones, making it challenging for the model to learn the backdoor behavior. To address this, TokenSwap employs an adaptive token-weighted loss that explicitly emphasizes the learning of swapped tokens, such that the visual triggers and bags-of-words behavior are associated. Extensive experiments demonstrate that TokenSwap achieves high attack success rates while maintaining superior evasiveness and stealthiness across multiple benchmarks and various LVLM architectures.