Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

On Improving Robustness of Deepfake Image Detectors
arXiv:2606.02797v2 Announce Type: replace Abstract: The rapid advancement of Generative AI has introduced remarkable opportunities while simultaneously raising critical concerns regarding content authenticity. While recent work has increasingly focused on improving the generalization of deepfake detectors across unseen generative models, their robustness against adversarial attacks remains limited. In particular, Abdullah et al. (IEEE SP 2024) evaluated eight detectors and demonstrated that most of them exhibit significant performance degradation under adversarial attacks. We also observed the same phenomenon by testing seven most recent state-of-the-art detectors. To address this problem, we propose a unified framework that integrates three complementary design principles without relying on adversarial training data: (i) higher-order statistical modeling in the frequency domain via Discrete Cosine Transform (DCT)-based moment pooling up to fourth order, (ii) content-agnostic feature representations derived from noise residuals, and (iii) cross-scene generalization enforced through patch-level semantic disruption. A key insight underpinning our approach is that adversarial attacks primarily operate on low-order statistics and visual semantics, leaving higher-order residual-frequency characteristics, particularly kurtosis, largely unconstrained. Extensive experiments demonstrate that our method consistently improves robustness across six architecturally diverse detectors. Notably, we achieve up to 88.9% reduction in recall degradation on current adversarial benchmarks, and improve the best-performing recent detector (Yang et al., IEEE CVPR 2025) from 81.9% to 97.15% accuracy under attack. Overall, our method provides a principled, architecture-agnostic approach for improving deepfake detection robustness against current attacks.
A Class of Multipartite Entangled States Based on State Transitions
arXiv:2606.05579v1 Announce Type: cross Abstract: We introduce Transition states (T states), denoted by $\ket{T_k^n}$, as a class of multipartite entangled states characterized by a fixed number of state transitions between adjacent qubits. These states form equal-amplitude superpositions over all states with a specified transition count. Unlike Bell states based on two-qubit correlations, GHZ states characterized by global correlations among all qubits, and W and Dicke states based on fixed numbers of qubit excitations, T states are defined by transition counts along an ordered sequence of qubits. We prove that T states are unitarily equivalent to Dicke states through a chain of CX (controlled-X) operations, thereby establishing a direct correspondence between transition-based and excitation-based representations of multipartite entanglement.
Excited States from Restricted Open Shell Plane-Wave DFT
arXiv:2605.28637v2 Announce Type: replace Abstract: Variational excited-state density functional theory (DFT) enables the calculation of excited states at a cost comparable to ground-state calculations, but single-configuration approaches often suffer from spin contamination. We implement restricted open-shell Kohn-Sham (ROKS) DFT, which recovers spin-pure singlet excitation energies via the variational minimization of a weighted combination of mixed-spin and triplet configurations, within the plane-wave projector augmented-wave framework of VASP. The energy functional is optimized using a preconditioned conjugate-gradient or a direct inversion in the iterative subspace algorithm, and analytical atomic forces are derived. The implementation is validated for eight organic molecules by comparison to the Q-Chem quantum chemistry code, yielding mean deviations of approximately $30\,\mathrm{meV}$. As a solid-state application, we investigate the three lowest lying excitations of MgO with a neutral oxygen vacancy. For a dielectric-dependent hybrid functional, vertical excitation energies from ROKS and time-dependent density functional theory (TDDFT) differ on average by about $0.21\,\mathrm{eV}$. The Franck-Condon shifts deviate on average by $0.14\,\mathrm{eV}$ between the two methods and mass-weighted displacements between the excited states and the ground state by $0.12\,\mathrm{amu}^{1/2}$ Ang. Additional calculations at the PBE level reveal that these properties depend less strongly on the DFT functional for ROKS than for TDDFT. These results demonstrate that ROKS provides excitation energies and excited-state forces with an accuracy similar to TDDFT while retaining the favorable scaling of ground-state DFT, making it a promising approach for affordable excited-state simulations in extended systems.
Mitigating the Curse of Dimensionality in Uniform Convergence of Deep Neural Networks via Smooth Activations
arXiv:2606.05599v1 Announce Type: new Abstract: This paper establishes a theoretical framework for the uniform convergence of smoothly activated deep neural network (DNN) estimators. While standard ReLU networks achieve minimax-optimal rates in the $L^2(P)$ norm for various nonparametric regression tasks, we establish a theoretical lower bound demonstrating that least-squares ReLU estimators can suffer from the curse of dimensionality in their uniform convergence behavior. Motivated by the need for reliable uniform guarantees in downstream tasks requiring worst-case reliability, we address this limitation by analyzing smoothly activated DNNs (smooth DNNs), encompassing both feedforward and residual structures. We establish novel pseudo-dimension bounds, non-asymptotic approximation guarantees, and H\"older-norm bounds for the approximators of these models. Leveraging these results, we derive non-asymptotic uniform convergence rates for smooth DNN estimators across multiple statistical contexts, including Huber, least-squares, quantile, and logistic regression. We prove that smooth DNNs can mitigate the {curse of dimensionality} in uniform convergence by adaptively exploiting the low-dimensional hierarchical composition structure of the target function. Supported by both simulation studies and a real-world application, our results position smooth DNNs as a theoretically grounded and practically viable alternative to ReLU networks for statistical learning tasks requiring uniform guarantees.
Discrete Causal Representations from Heterogeneous Domains: A Bayesian Approach with Social Survey Applications
arXiv:2606.06288v1 Announce Type: cross Abstract: Causal representation learning aims to infer the high-level latent causal concepts that give rise to observed low-level measurements. This is particularly relevant for heterogeneous data from different environments or domains since distribution shifts often arise through sparse, localized changes in some of the underlying causal mechanisms, while other parts of the generative process remain unchanged. Whereas identifiability of causal representations has been studied extensively, practical uncertainty-aware methods and real-world use cases remain less explored. In this work, we propose a Bayesian approach to learning causal representations from multi-environment data, focusing on the case of discrete causal concepts and unknown multi-node soft interventions. To this end, we translate causal assumptions and interpretability desiderata into suitable priors and parametric choices within a hierarchical model. We then devise an inference scheme based on sequential Monte Carlo sampling to approximate the resulting multimodal posterior. We showcase our approach through case studies on social survey data, where latent causal concepts correspond to cultural values or political opinions, measurements to survey responses, and environments to different countries or states. Our model infers meaningful high-level concepts and plausible causal relations among them, demonstrating its utility for learning causal representations of complex real-world data.
Dynamic Coordination Strategy Selection for Enterprise Multi-Agent Systems
arXiv:2606.00804v2 Announce Type: replace Abstract: Enterprise multi-agent systems increasingly expose multiple coordination patterns, but deployments often lack evidence for when to use consensus, debate, synthesis, or a simpler single-agent workflow. This paper evaluates whether coordination strategy should be selected dynamically by problem class rather than fixed globally. We run a frozen matrix of 30 enterprise tasks spanning six industries, five problem classes, four execution conditions, three replications per cell, and four model arms: qwen_local, sonnet, gemma_openrouter, and an auxiliary openai cloud-validation arm. All 1,440 generated outputs are judged by a fixed Sonnet rubric. The main finding is bounded and operationally useful, but it is not the original strict H1. The pre-registered exact-winner/CI criterion is not supported: exact winner identity is unstable across model arms, and several predicted strategies are close to, but not above, the best observed alternative. A weaker near-best routing claim is strongly supported. In every pre-registered model arm and problem class, and again in the auxiliary OpenAI validation arm, the predicted strategy is within 0.10 quality-score points of the best observed condition. Structured compliance verification is the clearest exception to the original mapping: all arms favor single_agent rather than consensus. A pre-registered Kendall's W test finds no reliable difference between Vietnamese-domain and English-domain tasks in how consistently the four coordination conditions are ranked (mean W of 0.20 in both strata; signed-rank p = .85), so H2 is not supported. We conclude that enterprise coordination policy should use dynamic routing as a calibrated default, not as a deterministic winner-selection law.
What's in a Name? Morphological Shortcuts by LLMs in Pharmacology
arXiv:2606.05616v1 Announce Type: new Abstract: The morphological form of a word can often give cues to its meaning, but purely relying on these mappings can lead to overgeneralization in high-stakes domains. In the medical domain, for instance, LLMs can confidently reason about fictitious drugs from their affixes alone (e.g., wugcillin) and generate plausible-looking clinical content. We present a behavioral and mechanistic study of LLM "affix heuristics" in pharmacology. Using fictitious drug names built from real affixes, we show that affix signals alone elicit class-level pharmacological responses. We introduce a framework for identifying whether a model's drug semantics are driven mainly by the affix, the stem, or the drug name as a whole. Applied across 653 drugs, our framework reveals that models often induce drug meaning primarily through affix cues, yet rarely explicitly indicate this reliance, and sometimes incorrectly conflate properties among affix-sharing drugs. Activation patching across models further localizes this behavior to early-mid layers. These findings show that morphological shortcuts pose a subtle but measurable risk to safety.
Generating Graph-Like Logical Rules for Knowledge Graph Reasoning via Diffusion Models
arXiv:2605.30747v2 Announce Type: replace Abstract: Logical rules constitute a cornerstone of knowledge graph (KG) reasoning, valued for their interpretability and ability to model relational patterns. However, existing rule mining methods predominantly focus on simple chain-like rules and therefore neglect the richer relational information encoded in graph-like structures, such as cycles and branches. This limitation is further exacerbated by computational bottlenecks caused by the combinatorial explosion of the search space, which is especially challenging for graph-like rules. Meanwhile, generative approaches such as diffusion models, despite their success in other domains, cannot be directly applied to rule mining because their training objectives are not aligned with the goal of learning high-quality rules, and non-differentiable KG rule quality metrics cannot directly guide model optimization. To address these limitations, we propose GRiD, a framework that reformulates graph-like rule discovery as a discrete generative process conditioned on the target relation. GRiD employs a two-phase training strategy. First, supervised pre-training enables GRiD to capture structural priors from subgraphs sampled from the KG meta-graph. Subsequently, reinforcement learning is applied to fine-tune GRiD through policy gradient optimization guided directly by non-differentiable rule-quality metrics. Experiments on six benchmark datasets show that GRiD achieves competitive performance on KG completion tasks. Ablation studies confirm the efficiency and robustness of GRiD and further show that graph-like rules complement chain-like rules in KG completion. Our code and datasets are available in https://github.com/Haoxiang-Cheng/GRiD.
Minor Ions as a Diagnostic of Solar Wind Heating: Inverted Mass-to-Charge Scaling in Imbalanced Turbulence
arXiv:2606.06340v1 Announce Type: cross Abstract: Alfv\'enic turbulence is vital to powering the solar wind and corona, yet eludes a comprehensive understanding of the kinetic processes by which it dissipates. Minor ions are sensitive tracers of these processes, showing extreme perpendicular temperatures and mass-weighted temperature trends that can either correlate or anticorrelate with mass-to-charge ratio, $A_i/Z_i$. We use a combination of quasilinear theory and 3D hybrid-kinetic simulations to explain these features and their correlations with properties of turbulence in the fast solar wind. When Alfv\'enic turbulence is imbalanced, its cascade to ion-Larmor scales is throttled by the helicity barrier. This barrier ultimately leads to high-frequency proton-cyclotron waves (PCWs), both oblique and parallel, the latter of which produce very flat electric-energy spectra ($\mathcal{E}_{E_{\perp}}\sim k_\parallel^{-\eta}$ with $\eta<2$) over the range of scales that are cyclotron resonant with minor ions. While steeper spectra lead to a positive correlation of heating with $A_i/Z_i$, the shallower spectra cause the dependence to invert, with $Q_i\propto Q_{\mathrm{p}}A_i(A_i/Z_i)^{\eta-2}$. Six simulations of balanced and imbalanced turbulence spanning $\beta_{\rm p0}=\{1,0.3,1/16\}$ corroborate this prediction, showing minor-ion heating rates that follow $(A_i/Z_i)^a$. Minor-ion heating is strongest and most perpendicular in our lowest $\beta_{\rm p0}=1/16$ simulation of imbalanced turbulence, reaching $T_{\perp{\rm O}^{5+}}/T_{\perp{\rm p}}\approx40$ and $T_{\perp{\rm O}^{5+}}/T_{\parallel{\rm O}^{5+}}\approx10$, consistent with low-coronal observations. Future minor-ion measurements should test whether intervals in which minor-ion thermal speeds decrease with increasing mass-to-charge ratio are associated with a history of large cross helicity, enhanced power in parallel PCWs, and a steep transition-range spectrum.
LatentWave: JEPA Pretraining for Wireless Foundation Models
arXiv:2606.06373v1 Announce Type: cross Abstract: Wireless foundation models have emerged as a promising alternative to building separate models for each wireless task. However, existing approaches rely on masked input reconstruction, which can bias representations toward low-level signal details. In this paper, we propose LatentWave, a wireless foundation model pretrained using a Joint-Embedding Predictive Architecture (JEPA) on diverse wireless spectrograms and channel state information (CSI). By predicting masked regions in latent space, LatentWave learns representations that are more transferable out of the box across diverse downstream tasks. The proposed architecture employs per-channel patch embeddings with stochastic channel sampling during pretraining, allowing it to process variable antenna counts and improving usability across heterogeneous wireless configurations. We evaluate LatentWave on four downstream tasks: RF signal classification, 5G NR positioning, beam prediction, and LoS/NLoS classification, comparing against a masked-modeling baseline (WavesFM) pretrained on the same data. Additionally, we show that the masking geometry introduces a task-dependent inductive bias: frequency masking strongly favors channel-related tasks such as positioning and beam prediction, while region masking better preserves discriminability for signal classification.
How abundant are good interpolators?
arXiv:2606.06469v1 Announce Type: cross Abstract: Let $S$ be the set of unit norm linear classifiers $\theta \in \mathbb{R}^d$ which correctly classify every point of a labeled dataset $(X_i,y_i)_{i=1}^n$, $X_i \in \mathbb{R}^d$, $y_i \in \{-1,+1\}$, with a possibly negative margin $\kappa$ fixed in advance. Under two natural data-generating distributions of the $(X,y)$ pairs -- a Gaussian mixture model and a logistic model with Gaussian features -- and in the proportional regime $n/d \to \alpha$ with small enough $\alpha$, we establish a large deviation principle on the event that a point $\theta$ chosen uniformly at random from $S$ achieves a given generalization error, with high probability over the choice of the data. The associated large deviation rate function is deterministic and describes the proportion, at the exponential scale in $d$, of interpolating classifiers having a given desired performance. As a consequence, we establish the following concentration phenomenon: all but an exponentially small fraction of interpolating classifiers have approximately the same generalization performance given by the unique maximizer of this rate function. We numerically compare this maximizer to the performance of empirical risk minimization by gradient descent and to the performance of a natural linear program, both finding a point in $S$, and deduce that in the overparametrized regime of small $\alpha$, these efficient procedures outperform the vast majority of interpolators, pointing to their nontrivial benign overfitting in this setting.
Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification
arXiv:2606.04037v2 Announce Type: replace Abstract: Pre-deployment verification of enterprise artificial intelligence (AI) agents remains a critical gap between large language model (LLM) capability benchmarking and production deployment. Post-deployment monitoring, human-in-the-loop controls, and prompt-level guardrails offer limited assurance once an agent is operating in production. We present an ontology-grounded verification framework -- to our knowledge the first to combine three components: an Agent Operational Envelope formalizing the certification space across permissions, domain constraints, safety properties, governance rules, and autonomy levels; an ontology-to-scenario generation pipeline that derives regulatory, operational, and adversarial test scenarios automatically; and a machine-verifiable Trust Certificate with graduated deployment verdicts. A controlled pilot across four regulated industries (Fintech, Banking, Insurance, Healthcare), instantiated as five industry-by-regulatory-regime cells across the United States and Vietnam (where Vietnam's 2025 AI Law makes such verification legally mandated for financial services), generated 1,800 scenarios evaluated against 125 primary-source regulatory requirements and 25 injected faults. Ontology-grounded generation significantly outperformed the dominant persona-based baseline on regulatory coverage (48.3% versus 33.1%; corrected p_c = .0006) and attained the highest domain specificity (4.77/5.0; p = 2e-6); transparently, its advantage over plain and retrieval-augmented prompting did not survive Bonferroni correction. Cross-validation across three LLM families (Claude Sonnet 4, Qwen 2.5 72B, Gemma 4 26B; 5,400 total scenarios) replicated the persona-versus-ontology pattern. The framework offers a reproducible, regulation-grounded route to pre-deployment assurance for enterprise AI agents, complementing runtime governance with an auditable deployment gate.
Radiation-induced electron spin polarization in ultrarelativistic kinetic turbulence
arXiv:2606.04332v2 Announce Type: replace Abstract: Electron spin polarization in radiative plasmas with ultrarelativistic kinetic turbulence under highly magnetized conditions is investigated using particle-in-cell simulations. We observe that a significant spin polarization can be sustained when the leptons undergo energetic photon emission accompanied by spin flips during the nonequilibrium turbulent evolution. By analyzing the time evolution of spatially dependent spin polarization, we identify an electromagnetic (EM) regime of kinetic turbulence, distinct from the well-known density-dominated regime characterized by vortex currents and magnetic islands. While in the latter regime the spin polarization exists only transiently, in the EM regime significant anisotropic net polarization emerges and persists in non-dissipative scenarios. The correlation between spin signals and turbulence features is leveraged to introduce the characteristic parameter delimiting the EM regime via the ratio of electric and magnetic energy densities and to gain insight into complex plasma turbulence. This study demonstrates the versatility of a spin-resolved study of the plasma turbulence in extreme environments, such as black holes and magnetar magnetospheres.
A Reliable Self-Organized Distributed Complex Network for Communication of Smart Agents
arXiv:2503.07702v3 Announce Type: replace Abstract: Collaboration among distributed agents is fundamental to many complex systems, particularly in communication networks where connectivity must be maintained under energy constraints. In this study, we utilize intelligent agents (nodes) trained through reinforcement learning techniques to establish connections with their neighbors, ultimately leading to the emergence of a large-scale communication cluster. Notably, there is no centralized administrator; instead, agents must adjust their connections based on information obtained from local observations. The connection strategy is formulated using a physical Hamiltonian, thereby categorizing this intelligent system under the paradigm of "Physics-Guided Machine Learning". Agents are trained via a Deep Q-Network using local observations to minimize changes in the Hamiltonian, enabling adaptive decision-making in dynamic environments. Simulation results demonstrate that the proposed collaborative strategy forms robust large-scale communication clusters while reducing transmission energy compared to baseline approaches. The network maintains high connectivity under agent mobility, density variations, node failures, and environmental obstacles, highlighting strong adaptability and resilience. These findings indicate that physics-guided reinforcement learning provides an effective mechanism for distributed topology optimization in emerging IoT and vehicular communication networks.
PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
arXiv:2503.14295v3 Announce Type: replace Abstract: Recent advancements in audio-driven talking face generation have made great progress in lip synchronization. However, current methods often lack sufficient control over facial animation such as speaking style and emotional expression, resulting in uniform outputs. In this paper, we focus on improving two key factors: lip-audio alignment and emotion control, to enhance the diversity and user-friendliness of talking videos. Lip-audio alignment control focuses on elements like speaking style and the scale of lip movements, whereas emotion control is centered on generating realistic emotional expressions, allowing for modifications in multiple attributes such as intensity. To achieve precise control of facial animation, we propose a novel framework, PC-Talk, which enables lip-audio alignment and emotion control through implicit keypoint deformations. First, our lip-audio alignment control module facilitates precise editing of speaking styles at the word level and adjusts lip movement scales to simulate varying vocal loudness levels, maintaining lip synchronization with the audio. Second, our emotion control module generates vivid emotional facial features with pure emotional deformation. This module also enables the fine modification of intensity and the combination of multiple emotions across different facial regions. Our method demonstrates outstanding control capabilities and achieves state-of-the-art performance on both HDTF and MEAD datasets in extensive experiments.
Federating Governance: How Community Rules Scale with Mastodon Instances
arXiv:2606.05069v2 Announce Type: replace Abstract: The rise of decentralized social media platforms like Mastodon and Bluesky highlights the challenge of scaling self-governance and moderation. As communities grow, they face new issues that demand increasingly complex governance structures. However, as moderation is mainly volunteer-driven, there is limited formal guidance on how community rules and moderation practices should evolve with growth. This study investigates how moderation scale with Mastodon instances by analyzing community rules across servers of varying sizes. We categorize these rules to identify key governance priorities and find that these priorities are remarkably consistent across instance sizes: rules addressing problematic content, such as harassment, hate speech, and illegal content, dominate regardless of scale. While smaller communities focus on narrower sets of topics, larger servers maintain a more balanced coverage of a broad range of topics. Our analysis of rule formalization reveals that community size strongly predicts rule development. As instances grow, their rules become more extensive and topically diverse, but also exhibit lower readability and linguistic diversity. In contrast, external federation interactions have a limited role, mainly associated with a broader scope of rules without substantially affecting their diversity or form. These findings highlight the relative influence of internal versus external factors, suggesting that local scaling pressures outweigh network-level dynamics in decentralized social media governance. The scaling pattern observed on Mastodon resemble those previously identified on centralized platforms such as Reddit, suggesting that community size imposes fundamental constraints on self-governance that transcend platform architectures
Pitfalls of Evaluating Language Models with Open Benchmarks
arXiv:2507.00460v3 Announce Type: replace Abstract: Open Large Language Model (LLM) benchmarks, such as HELM and BIG-Bench, provide standardized and transparent evaluation protocols that support comparative analysis, reproducibility, and systematic progress tracking in Language Model (LM) research. Yet, this openness also creates substantial risks of data leakage during LM testing--deliberate or inadvertent, thereby undermining the fairness and reliability of leaderboard rankings and leaving them vulnerable to manipulation by unscrupulous actors. We illustrate the severity of this issue by intentionally constructing cheating models: smaller variants of BART, T5, and GPT-2, fine-tuned directly on publicly available test-sets. As expected, these models excel on the target benchmarks but fail terribly to generalize to comparable unseen testing sets. We then examine task specific simple paraphrase-based safeguarding strategies to mitigate the impact of data leakage and evaluate their effectiveness and limitations. Our findings underscore three key points: (i) high leaderboard performance on limited open, static benchmarks may not reflect real-world utility; (ii) private or dynamically generated benchmarks should complement open benchmarks to maintain evaluation integrity; and (iii) a reexamination of current benchmarking practices is essential for reliable and trustworthy LM assessment.
Defining and classifying models of groups: The social ontology of higher-order networks
arXiv:2507.02758v2 Announce Type: replace Abstract: In complex systems research, the study of higher-order interactions has exploded in recent years. Researchers have formalized various types of group interactions, such as public goods games, biological contagion, and information broadcasting, showing how higher-order networks can capture group effects more directly than pairwise models. However, equating hyperedges-edges involving more than two agents-with groups can be misleading, as it obscures the polysemous nature of ``group interactions''. For instance, many models of higher-order interactions focus on the internal state of the hyperedge, specifying dynamical rules at the group level. These models often neglect how interactions with external groups can influence behaviors and dynamics within the group. Yet, anthropologists and philosophers remind us that external norms, factors, and forces governing intergroup behavior are essential to defining within-group dynamics. In this paper, we synthesize concepts from social ontology relevant to the emerging physics of higher-order networks. We propose a typology for classifying models of group interactions based on two perspectives. The first focuses on individuals within groups engaging in collective action, where shared agency serves as the binding force. The second adopts a group-first approach, emphasizing institutional facts that extend beyond the specific individuals involved. Building on these perspectives, we introduce four dimensions to classify models of group interactions: persistence, coupling, reducibility, and alignment. For the physics of higher-order networks, we provide a hierarchy of nested mathematical models to explore the complex properties of social groups. We highlight social interactions not yet explored in the literature on higher-order networks and propose future research avenues to foster collaboration between social ontology and the physics of complex systems.
Learning to optimize with guarantees: a complete characterization of linearly convergent algorithms
arXiv:2508.00775v2 Announce Type: replace Abstract: The design of many classical optimization algorithms is driven by the certification of linear convergence rates over classes of optimization problems. In this paper, we consider the problem of improving the average-case performance of an algorithm over a specific distribution of problem instances. While this task can be tackled by embedding trainable components into the algorithm updates, a key challenge is to preserve worst-case guarantees across the entire problem class. For classes of composite optimization problems, we show that all linearly convergent algorithms can be parametrized in terms of a baseline linearly convergent algorithm, and a set of trainable, exponentially-decaying modifications to its update rule; crucially, this parametrization excludes all-and only-the algorithms that do not converge linearly. Our results apply to improving the average-case performance of classical algorithms such as gradient descent for nonconvex, gradient-dominated functions; Nesterov's accelerated method for smooth, strongly convex functions; and projected gradient methods for optimization over polyhedral feasible sets. We illustrate how our characterization can be used for learning to optimize with linear convergence and feasibility guarantees. Numerical results showcase benefits over classical optimizers when solving ill-conditioned systems of linear equations and running a model predictive control scheme on a linear dynamical system.
In-Training Defenses against Emergent Misalignment in Language Models
arXiv:2508.06249v3 Announce Type: replace Abstract: Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (EM): Even a small, domain-specific fine-tune can induce harmful behaviors far outside the target domain. Even in the case where model weights are hidden behind a fine-tuning API, this gives attackers inadvertent access to a broadly misaligned model in a way that can be hard to detect from the fine-tuning data alone. We present the first systematic study of in-training safeguards against EM that are practical for providers who expose fine-tuning via an API: We evaluate whether they a) prevent broad misalignment, b) allow narrow misalignment, c) learn well on benign tasks, and d) remain coherent. We investigate five training regularization interventions: (i) KL-divergence regularization toward a safe reference model, (ii) $\ell_2$ distance in feature space, (iii) preventive steering with an evil persona vector, (iv) interleaving training examples from a general instruct-tuning dataset and (v) inoculation prompting. We demonstrate that selecting interleaving data by the perplexity gap between aligned and misaligned models yields the best results overall.
Decomposition Polyhedra of Piecewise Linear Functions
arXiv:2410.04907v2 Announce Type: replace-cross Abstract: In this paper we contribute to the frequently studied question of how to decompose a continuous piecewise linear (CPWL) function into a difference of two convex CPWL functions. Every CPWL function has infinitely many such decompositions, but for applications in optimization and neural network theory, it is crucial to find decompositions with as few linear pieces as possible. This is a highly challenging problem, as we further demonstrate by disproving a recently proposed approach by Tran and Wang [Minimal representations of tropical rational functions. Algebraic Statistics, 15(1):27-59, 2024]. To make the problem more tractable, we propose to fix an underlying polyhedral complex determining the possible locus of nonlinearity. Under this assumption, we prove that the set of decompositions forms a polyhedron that arises as intersection of two translated cones. We prove that irreducible decompositions correspond to the bounded faces of this polyhedron and minimal solutions must be vertices. We then identify cases with a unique minimal decomposition, and illustrate how our insights have consequences in the theory of submodular functions. Finally, we improve upon previous constructions of neural networks for a given convex CPWL function and apply our framework to obtain results in the nonconvex case.
Stochastic Multiscale Reconstruction of Lagrangian Turbulence via Guided Diffusion Models
arXiv:2606.05783v1 Announce Type: new Abstract: Lagrangian turbulence is characterized by intermittent, fat-tailed fluctuations and nontrivial correlations across temporal scales, making a quantitative description of its full multiscale probability distribution a longstanding challenge. A particularly important question is whether unresolved fine-scale fluctuations can be inferred from coarse-grained trajectory information. Here, we address this problem by sampling the conditional distribution of unresolved fluctuations using a diffusion-model prior conditioned on large-scale dynamics obtained through a wavelet-based coarse-graining of Lagrangian trajectories. Using tracer trajectories from direct numerical simulations of homogeneous and isotropic turbulence at $Re_\lambda \simeq 310$, we show that the reconstructed signals recover scale-dependent intermittent statistics, including high-order structure functions, flatness, and local scaling exponents, together with cross-scale temporal correlations between resolved and unresolved fluctuations. The method also reproduces the broad stochastic variability of intermittent acceleration fluctuations conditioned on the same coarse-grained trajectory, whereas Gaussian-process reconstructions in wavelet representation suppress rare events. Our results show that small-scale Lagrangian intermittency can be modeled as a non-Gaussian conditional stochastic process constrained by coarse-scale dynamics and quantitatively reproduced through data-driven generative sampling.
Toward an affordable density-based measure for the quality of a coupled cluster calculation
arXiv:2509.04429v4 Announce Type: replace Abstract: We propose two new diagnostics for the degree to which static correlation impacts the quality of a coupled cluster calculation. The first is the change in the Matito static correlation diagnostic $\overline{I_{ND}}$ between CCSD and CCSD(T), $\Delta I_{ND}[\textrm{(T)}]=\overline{I_{ND}}[\textrm{CCSD(T)}]-\overline{I_{ND}}[\textrm{CCSD}]$. The second is the ratio of the same and of the corresponding change in the total correlation diagnostic $\overline{I_{T}}=\overline{I_{ND}}+\overline{I_{D}}$, i.e., $r_I[(T)]=\Delta I_{ND}[\textrm{(T)}]/\Delta I_{T}[\textrm{(T)}]$. The first diagnostic can be extended to higher-order improvements in the wave function, e.g., $\Delta I_{ND}[\textrm{(Q)}]=\overline{I_{ND}}[\textrm{CCSDT(Q)}]-\overline{I_{ND}}[\textrm{CCSDT}]$. In general, a small $\Delta I_{ND}$[\textrm{level$_1$}] value indicates that at this level$_1$ of theory, the density is converged and any further changes to the energy come from dynamical correlation, while larger $\Delta I_{ND}$[\textrm{level$_2$}] indicates that the density is still not converged at level$_2$ and some static correlation remains. $r_I[(T)]$ is found to be a moderately good predictor for the importance of post-CCSD(T) correlation effects.
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
arXiv:2509.24882v2 Announce Type: replace Abstract: Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models. In this work, we present a systematic analysis of scaling laws for quadratic and diagonal neural networks in the feature learning regime. Leveraging connections with matrix compressed sensing and LASSO, we derive a detailed phase diagram for the scaling exponents of the excess risk as a function of sample complexity and weight decay. This analysis uncovers crossovers between distinct scaling regimes and plateau behaviors, mirroring phenomena widely reported in the empirical neural scaling literature. Furthermore, we establish a precise link between these regimes and the spectral properties of the trained network weights, which we characterize in detail. As a consequence, we provide a theoretical validation of recent empirical observations connecting the emergence of power-law tails in the weight spectrum with network generalization performance, yielding an interpretation from first principles.
Statistical Priors for Implicit Preferences: Decoupling Skill Selection as a Local Harness in Personal Agents
arXiv:2606.05828v1 Announce Type: new Abstract: As Large Language Model (LLM) capabilities advance, locally deployed personal agents relying on API-based remote models and external skills have emerged as a novel paradigm. With the rapid expansion of available skills, enabling personal agents to learn and adapt to implicit user preferences becomes a critical challenge. However, local deployment constraints preclude complex centralized selection algorithms, creating an urgent need for a lightweight local preference harness. This paper explores the implementation of such a harness through a novel architecture that strictly decouples statistical preference learning from semantic intent parsing. Specifically, we leverage localized statistical results to influence and modulate the selection decisions of the remote LLM. Extensive evaluations demonstrate that our decoupled approach achieves the lowest cumulative regret and highest test accuracy, significantly outperforming traditional memory-augmented agents.