Forskningsradar

Science Journals

Peer-reviewade publikationer — 53080 artiklar

Computational Control of Nonlinear Partial Differential Equations Using Machine Learning
arXiv:2604.22414v2 Announce Type: replace-cross Abstract: The numerical reconstruction of controls for nonlinear partial differential equations (PDEs) remains a challenging and relatively underdeveloped problem, despite the extensive literature on controllability theory. In this work, we introduce an operator-decomposed physics-informed neural network framework, called WeightedPINN, for approximating controls in nonlinear PDE settings. The method is designed for both internal and bilinear control problems and incorporates the governing equation, boundary and initial conditions, and terminal control constraints directly into the training objective. The main feature of WeightedPINN is that the different components of the controlled PDE residual are weighted separately. In particular, the time derivative, directional diffusion terms, nonlinear response, and control term are assigned independent adaptive space--time weights, and the same weighted formulation is applied to the boundary, initial, and terminal constraints. This produces a control-aware residual metric that is more sensitive to operator-level imbalance and to the mechanism through which the unknown control enters the equation. We provide a convergence analysis for the proposed method and present numerical experiments for semilinear heat and wave equations with internal and bilinear controls. The high-dimensional experiments demonstrate improved residual-based testing errors compared with the standard PINN baseline, while lower-dimensional manufactured-solution benchmarks show improved direct reconstruction errors for both the state and the control against several adaptive and control-oriented PINN methods. The results suggest that WeightedPINN is particularly effective in regimes where componentwise residual imbalance, anisotropy, variable coefficients, or control-identification sensitivity play a significant role.
Faster Mixing for Triangulations via Transport Flows
arXiv:2605.02067v2 Announce Type: replace-cross Abstract: We prove an $\widetilde O(n^2)$ bound for the relaxation time and the log-Sobolev time (inverse log-Sobolev constant) of the classical triangulation flip chain on a convex $(n+2)$-gon, implying a mixing time of $\widetilde O(n^2)$. The previous state of the art for the mixing time of this chain, due to Eppstein and Frishberg, was $\widetilde O(n^3)$, while the best known lower bound on the mixing time, due to Molloy, Reed, and Steiger, is $\Omega(n^{3/2})$. Our relaxation time bound makes significant progress towards Aldous' conjectured bound of $\Theta(n^{3/2})$ for the relaxation time. We improve upon the analysis of Eppstein and Frishberg by further developing the framework of transport flows introduced in the work of Chen et al. In this light, our results can be seen as a more efficient way of using combinatorial decompositions to obtain functional inequalities for Markov chains. We hope our ideas will find other applications in the future.
GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation
arXiv:2605.25447v1 Announce Type: new Abstract: Generating structured, editable diagrams remains a significant challenge for contemporary large language models, despite their proficiency in general-purpose vector code generation. The primary difficulty lies in the structural fragility of the output; minor errors such as misaligned connector endpoints, text labels overlapping borders, or complex layouts drifting beyond the canvas boundaries render the resulting SVG files functionally unusable for professional applications. To address these issues, we introduce GeoSVG-RL, a specialized reinforcement learning framework designed for layout-constrained text-to-SVG generation. Unlike standard training objectives that rely solely on maximizing token-level likelihood, our approach optimizes the policy against explicit, executable geometric feedback. The model first produces a structured layout plan that serves as a geometric contract for the subsequent generation of the SVG code. This code is then rendered through a browser-backed verifier, enabling the calculation of fine-grained rewards across six critical dimensions: rendering validity, canvas fitting, precise anchor placement, text containment, graph consistency, and code cleanliness. We utilize Group Relative Policy Optimization (GRPO) to refine the model, sampling multiple candidates per prompt to facilitate updates based on relative quality. Starting from a supervised warm-start phase on synthetic data, GeoSVG-RL achieves substantial gains in structural reliability, particularly in arrow-anchor accuracy and text-in-box rates. Quantitative evaluations demonstrate that our method consistently outperforms current state-of-the-art systems in local geometric precision and the preservation of graph connectivity, providing a robust pathway toward automated yet reliable technical illustration.
Localization then Neutralization: Gradient-guided Token Suppression against Visual Prompt Injection Attack
arXiv:2605.25194v1 Announce Type: new Abstract: Adversarial images pose a severe security threat to multimodal large language models through prompt injection. Existing defenses largely lack a principled understanding of the underlying mechanisms and struggle to balance efficiency and defense utility. In this work, we show that successful adversarial attacks do not rely on the entire image uniformly but instead depend on a small subset of critical image tokens. Based on this insight, we propose Gradient Token Masking (GTM), which localizes these tokens via gradient analysis and neutralizes them through masking. We find that attribution based on the first generated token's output probability fails when attacks preserve the predicted token. To overcome this, GTM utilizes the Hidden-State Gradient Norm score for generation-influence attribution under adversarial inputs. We prove that its ranking is consistent with that of the full adversarial loss gradient, providing a theoretical guarantee for accurate localization. Our method requires only a single forward-backward pass to identify and zero out a small number of high-scoring tokens, effectively disrupting the adversarial attack path. Extensive experiments on prompt injection and multimodal jailbreak attacks demonstrate that our approach reduces attack success rates (ASR) to near zero while preserving model utility with negligible computational overhead.
Guess the Unified Model: How Much Can We Recover from Generated Images?
arXiv:2605.25254v1 Announce Type: new Abstract: With unified model-generated images now widespread online, attributing their model of origin offers a path toward transparency and deeper insight into the characteristic behaviors of individual models. Prior work has explored provenance in LLM-generated text, diffusion model images, and datasets, but the separability of unified model-generated images remains an underexplored area. We address this gap by examining separability across corruption, domains, and prompt languages using images generated by seven unified models. We show that model attribution is highly feasible as our model achieves near-perfect accuracy with around 20K images per model. Corruptions and structural perturbations have only a modest effect on attribution performance, and cross-domain generalization reveals that semantic content contributes to separability but is not the dominant signal. Finally, we observe that for most models, prompt language attribution is around chance levels, suggesting minimal language-specific visual signatures. These findings highlight consistent model-specific visual characteristics in unified models outputs and open new directions for tracing and auditing generative image pipelines.
Synchronization of coupled wind turbines
arXiv:2605.25192v1 Announce Type: new Abstract: In the context of renewable energies, wind energy appears as a sustainable alternative to address current environmental and energy challenges. This work studies the synchronization and stability of a network of wind turbines subjected to strong disturbances, by integrating a realistic modeling of wind variability by using the Ornstein-Uhlenbeck stochastic process. The dynamics of each wind turbine are described by a Kuramoto-type equation, while synchronization is analyzed through the time evolution of the phases. Stability is studied by analyzing the basin of attraction to the synchronous solution, namely the set of initial conditions leading to the stable synchronous state. Simulations carried out on various models ranging from an isolated wind turbine with constant power to an isolated wind turbine with variable wind power, reveal that the stability of the system is strongly influenced by inertia, damping, wind speed, wind fluctuation rate, correlation time, and coupling strength. Physically, these parameters control the balance between injected mechanical power, energy dissipation, grid-induced restoring forces, and the temporal structure of wind fluctuations, thereby determining the ability of the wind turbine to absorb perturbations and maintain synchronization under fluctuating wind conditions.
The poset of cancellations induced by gradient dynamics in a filtered Lefschetz complex
arXiv:2311.14364v4 Announce Type: replace-cross Abstract: Motivated by questions about simplification of topology, we take a discrete approach to the dependency of simplifying operations, using methods based on combinatorial gradient dynamics. We interpret the filter in persistent homology as a discrete Morse function. This lets us gradually simplify the dynamics in parallel with space and filter, while preserving homology. As a tool, we use shallow pairs, which are simultaneously birth-death pairs and combinatorial vectors. This allows us to extract topological features by the pairing of cells via persistence and simplify them using combinatorially defined cancellations. The main new concept is the depth poset of birth-death pairs, whose minimal elements are shallow pairs and whose linear extensions are sequences of cancellations that reduce the complex to its essential homology. Cancellations of birth-death pairs in a down set of this poset preserve the other birth-death pairs and the poset dependencies between them. An algorithm that constructs the depth poset in two passes of standard matrix reduction is given and proved correct.
Exchange Monte Carlo for continuous-space Path Integral Monte Carlo simulation
arXiv:2602.05500v2 Announce Type: replace-cross Abstract: We present a novel Exchange Monte Carlo (EMC) method designed for application in continuous-space Path Integral Monte Carlo (PIMC) simulations at finite temperature. Traditional PIMC methods for bosonic systems suffer from long autocorrelation times, particularly when measuring observables affected by particle permutations, such as the winding number. To address this issue, we introduce an exchange update scheme that facilitates replica transitions between different interaction regimes, significantly accelerating Monte Carlo dynamics-especially for global observables sensitive to permutation effects. Furthermore, we incorporate Stochastic Potential Switching (SPS) to efficiently decompose interactions, substantially enhancing computational efficiency for long-range interatomic pair potentials such as the Lennard-Jones and Aziz potentials.
Injecting Image Guidance into Text-Conditioned Diffusion Models at Inference
arXiv:2605.25191v1 Announce Type: new Abstract: Text-to-image diffusion models like Stable Diffusion generate high-quality images from text, but lack a way to inject visual guidance (e.g. sketches, styles) at inference without retraining. Existing methods either require computationally expensive fine-tuning or rely on style transfer techniques that risk semantic misalignment with textual prompts. We introduce Visual Concept Fusion (VCF), the first method offering dual conditioning on both an image and text prompt at inference time without any concept-specific training. VCF enables visual concept injection into Stable Diffusion by aligning CLIP image features with the text embedding space. VCF consists of three components: (1) a lightweight aligner that maps image tokens to the text embedding manifold using InfoNCE and cross-attention reconstruction losses, (2) a fusion strategy that preserves both textual and visual semantics, and (3) an optional Prompt-Noise Optimization (PNO) module for test-time refinement. Our experiments demonstrate that VCF successfully transfers visual attributes including style, composition, and color palette from reference images while maintaining prompt adherence. Quantitative results show a trade-off between text alignment (CLIP score) and visual correspondence (LPIPS), with VCF outperforming baselines in reference fidelity.
Locality Matters for Training-Free Audio Token Compression in Audio-Language Models
arXiv:2605.25179v1 Announce Type: new Abstract: Audio-language models (ALMs) are increasingly used for audio captioning, question answering, and open-ended audio understanding, but their inference cost remains high when audio inputs are represented as long prefix-token sequences. These audio prefixes consume context budget, increase memory usage, and make deployment harder in resource-constrained or latency-sensitive settings. Existing training-free audio-token reduction methods mainly rely on fixed pooling or score-based pruning. Fixed pooling is content-agnostic, while score-based pruning can preserve isolated salient tokens but discard nearby acoustic context. We propose Local Temporal Bipartite Merging (LTBM), a training-free encoder-space compression method that merges similar nearby audio tokens under an explicit temporal window constraint. Beyond introducing LTBM, we use a controlled Global Merge variant to isolate whether temporal locality itself is a useful inductive bias for audio-token compression. Experiments on AudioCaps, Clotho, and MMAU with Qwen2-Audio show evidence of a task-dependent locality effect: locality-aware merging is more favorable for captioning at several compression settings, especially under stronger compression, while global matching is more competitive for multiple-choice audio understanding. A cross-backbone validation on Audio Flamingo 3 further supports the captioning-side advantage of locality-aware merging under moderate and aggressive compression.
Spurious Stationarity and Hardness Results for Bregman Proximal-Type Algorithms
arXiv:2404.08073v3 Announce Type: replace-cross Abstract: Bregman proximal-type algorithms (BPs), such as mirror descent, have become popular tools in machine learning and data science for exploiting problem structures through non-Euclidean geometries. In this paper, we show that BPs can get trapped near a class of non-stationary points, which we term \emph{spurious stationary points}. Such stagnation can persist for any finite number of iterations if the gradient of the Bregman kernel is not Lipschitz continuous, even in convex problems. The root cause lies in a fundamental contrast in descent behavior between Euclidean and Bregman geometries: While Euclidean gradient descent ensures sufficient decrease near any non-stationary point, BPs may exhibit arbitrarily slow decrease around spurious stationary points. As a result, commonly used Bregman-based stationarity measure, such as relative change in terms of Bregman divergence, can vanish near spurious stationary points. This may misleadingly suggest convergence, even when the iterates remain far from any true stationary point. Our analysis further reveals that spurious stationary points are not pathological, but rather occur generically in a broad class of nonconvex problems with polyhedral constraints. Taken together, our findings reveal a serious blind spot in Bregman-based optimization methods and calls for new theoretical tools and algorithmic safeguards to ensure reliable convergence.
Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
arXiv:2512.15605v4 Announce Type: replace Abstract: Autoregressive models (ARMs) currently constitute the dominant paradigm for large language models (LLMs). Energy-based models (EBMs) represent another class of models, which have historically been less prevalent in LLM development, yet naturally characterize the optimal policy in post-training alignment. In this paper, we provide a unified view of these two model classes. Taking the chain rule of probability as a starting point, we establish an explicit bijection between ARMs and EBMs in function space, which we show to correspond to a special case of the soft Bellman equation in maximum entropy reinforcement learning. Building upon this bijection, we derive the equivalence between supervised learning of ARMs and EBMs. Furthermore, we analyze the distillation of EBMs into ARMs by providing theoretical error bounds. Our results provide insights into the ability of ARMs to plan ahead, despite being based on the next-token prediction paradigm.
One-shot Conditional Sampling: MMD meets Nearest Neighbors
arXiv:2509.25507v2 Announce Type: replace-cross Abstract: How can we generate samples from a conditional distribution that we never fully observe? This question arises across a broad range of applications in both modern machine learning and classical statistics, including image post-processing in computer vision, approximate posterior sampling in simulation-based inference, and conditional distribution modeling in complex data settings. In such settings, compared with unconditional sampling, additional feature information can be leveraged to enable more adaptive and efficient sampling. Building on this, we introduce Conditional Generator using MMD (CGMMD), a novel framework for conditional sampling. Unlike many contemporary approaches, our method frames the training objective as a simple, adversary-free direct minimization problem. A key feature of CGMMD is its ability to produce conditional samples in a single forward pass of the generator, enabling practical one-shot sampling with low test-time complexity. We establish rigorous theoretical bounds on the loss incurred when sampling from the CGMMD sampler, and prove convergence of the estimated distribution to the true conditional distribution. In the process, we also develop a uniform concentration result for nearest-neighbor based functionals, which may be of independent interest. Finally, we show that CGMMD performs competitively on synthetic tasks involving complex conditional densities, as well as on practical applications such as image denoising and image super-resolution.
Bayesian Estimation of Spectroscopic Parameters: Application to the Atomic Nitrogen Bound-Bound System
arXiv:2605.25368v1 Announce Type: new Abstract: Atomic nitrogen bound-bound radiation is a major component of the radiative heat flux on hypersonic vehicles entering nitrogen-dominated atmospheres, yet its prediction is limited by substantial parametric uncertainty in the published Einstein coefficients and Stark broadening coefficients. In the present study, these spectroscopic parameters are inferred and their uncertainty is quantified through Bayesian inversion of equilibrium spectral radiance measured in the NASA Ames Electric-Arc Shock Tube for two shots of the Test 62 campaign at shock speeds of 10.32 and 10.72 km/s. The inference is restricted to the post-shock equilibrium region, where the Boltzmann assumption closes the species population degree of freedom. The residual uncertainty in the post-shock temperature and species number densities is incorporated as a coupled nuisance parameter distribution. A hybrid principal component analysis and polynomial chaos expansion surrogate model and a likelihood formulated jointly over the two shots enable tractable Markov chain Monte Carlo sampling across multiple wavelength regions. Eighteen parameters in total, ten Einstein coefficients and eight Stark broadening coefficients, are inferred across eight wavelength regions, with posterior uncertainties significantly reduced relative to the prior literature bands. Forward propagation of the joint posterior through the stagnation-line flow field around a 3 m radius sphere at entry velocities of 10, 12, and 14 km/s demonstrates a reduction in the standard deviation of the predicted radiative heat flux by approximately a factor of five compared with the prior, in particular at 14 km/s, it drops from 10.4 to 1.94 W/cm$^{2}$.
Sampling Distributions as Regularization in Learned Inverse Problems
arXiv:2605.25177v1 Announce Type: new Abstract: Neural networks have emerged as effective tools for solving ill-posed inverse problems. In many scientific applications, however, observational training data are insufficient, and learned inverse operators must instead be trained on synthetic data generated from the forward model. This requires specifying unknown parameters in the forward model and solving the model to generate synthetic observations. Typically, the unknown parameters are sampled from a prescribed probability distribution. Here, we show that this sampling strategy is not a neutral preprocessing step, but instead defines an implicit regularization operator. This result follows from the fact that the learned inverse operator minimizes empirical risk together with the classical result that conditional expectation minimizes mean-square error. We present theoretical results for the implicit regularization operator in both infinite- and finite-data settings, including Physics Informed Neural Networks (PINNs). These results are demonstrated numerically on three inverse problems of increasing complexity: a 1D linear Fredholm integral equation, a 1D nonlinear subsurface interface inversion, and a 2D nonlinear cross-well seismic traveltime tomography problem. Across all three problems, three distinct sources of regularization are identified in the learned operator: prior sampling, architectural, and physics-informed regularization. A mismatched sampling distribution is shown to degrade reconstruction quality in ways that neither more expressive architectures nor augmented physics residuals can fully correct. The results demonstrate that the sampling distribution should be chosen with the same care as a classical regularization functional and provide a practical framework for implementing more sophisticated regularization operators using neural networks.
Arnoldi-Enhanced Multivariate Hermite Interpolation of Manifold-Valued Data
arXiv:2605.25176v1 Announce Type: new Abstract: This paper presents a robust enhancement of the Tangent space Hermite Interpolation (THI) method for manifold-valued data by integrating the multivariate Arnoldi process. To circumvent the inherent numerical instability of multivariate confluent Vandermonde matrices, we use a $G$-Arnoldi-based recurrence to construct a discrete orthogonal polynomial basis directly on the tangent space. The method generates better numerical conditioning for high-order approximations. We analyze the convergence rates for both $C^0$ and $C^1$ errors in the multivariate setting. When only function values are used, the $C^0$ approximation error decays as $\mathcal{O}\left(\sqrt{M} n^{-m}\right)$. For the $C^1$ error without derivative data, the rate becomes $\mathcal{O}\left(\sqrt{M} h^{-1} n^{-m}\right)$, where $h$ is the fill distance of the sampling set. When derivative data are additionally available, the $C^1$ error is $\mathcal{O}\left(\sqrt{M} n^{-(m-1)}\right)$. In all cases, $n$ is the polynomial degree, $m$ denotes the regularity of the target function, and $M$ is the number of sampling points. Importantly, as $n$ increases, the required number of points $M$ must also increase. This reveals the interplay among approximation order, sampling density ($M$), fill distance ($h$), dimension ($d$), and the regularity ($m$) of the target function. Extensive numerical experiments conducted on the special orthogonal group $SO(3)$ and the unit sphere $S^2$ show that the Arnoldi-enhanced THI method outperforms the Kriging-based approaches in terms of both computational efficiency and accuracy.
Discrepancy Minimization Improves Cross-Hospital Robustness in Digital Pathology
arXiv:2605.25175v1 Announce Type: new Abstract: Pathology foundation models (PFMs) have advanced rapidly in recent years and support training classifiers for a range of histopathology tasks. However, their robustness across hospitals remains limited: performance often degrades when training a classifier on data from one hospital and evaluating it on another target hospital. We address this challenge by fine-tuning PFMs with a local maximum mean discrepancy (LMMD) objective that applies to two settings: domain adaptation, where unlabeled target-hospital data is available, and domain generalization, where target-hospital data is unavailable at all. Experiments at both the patch- and slide-level show consistent improvements across multiple PFMs and tasks.
Bell Correlations from Prepared Coherence in Entangled Dirac Wavepackets
arXiv:2511.12258v2 Announce Type: replace-cross Abstract: Bell correlations are usually formulated for an ideal spin singlet, for which the Bell--CHSH combination reaches the maximal quantum value \(B=-2\sqrt{2}\), independent of detector separation. Here we derive Bell correlations from a more general physical state: an antisymmetrized pair of entangled Dirac wavepackets with source-prepared amplitude and phase coherence. The propagated branches are sampled locally by spatially separated endpoint detectors, yielding a separation-dependent CHSH value \(B(Z)\). For a fixed CHSH analyzer geometry, the zero-separation, full-overlap limit gives \[ B(0)=-2\sqrt{2}, \] independent of the preparation parameters. At large detector separation, once the direct branch-overlap contribution is suppressed, the surviving Bell--CHSH value approaches the prepared-coherence kernel \[ B(\infty)=\mathcal{K}_{\rm coh} = -\sqrt{2}\left[1+\sin(2\theta)\cos\chi\right]. \] Thus the asymptotic Bell value is controlled by the coherence fixed at the source through the amplitude balance \(\theta\) and relative phase \(\chi\). Bell violation is therefore a phase-sensitive local readout of prepared nonseparable Dirac-wave coherence: it rules out separable classical probability, but does not by itself require superluminal causation. In this wave-realist account, Bell correlations retain their full quantum content while remaining compatible with relativistic causal locality.
Learning Treatment Effects during Resource Allocation via Priority-Queue Randomization
arXiv:2605.25169v1 Announce Type: new Abstract: Public service programs often allocate limited resources under uncertainty about their benefits, creating a need for randomization to support credible evaluation. In practice, however, applicants commonly enter waitlists where resources are prioritized toward individuals judged to have higher need through tiered priority queues, making direct randomization difficult. Motivated by this, we develop an experimental design framework for learning treatment effects while treating those most in need where incoming applicants are randomized into priority queues based on their assessed risk scores. Treatments are then provided across queues in priority order and first-in-first-out within queue as budget becomes available. Our contributions are two-fold. First, we characterize what causal effects are identified under this priority-queue allocation. When arrivals are exogenous, treatments are conditionally randomized, and hence standard estimands are identified; when arrivals are endogenous, queue randomization instead provides an instrument for treatment, identifying local treatment effects induced by the queuing process. Second, we develop optimized queue-assignment designs that trade off statistical efficiency against prioritizing higher-need applicants. We show in the process that, despite dependence in treatment assignments induced by the design, usual iid efficiency bounds remain well-justified design objectives. We illustrate the proposed designs using data from a housing allocation program in a large U.S. county.
Multi-Alignment Contrastive Learning for Enzyme--Reaction Retrieval
arXiv:2512.08508v2 Announce Type: replace-cross Abstract: Identifying enzymes that catalyze target biochemical reactions is a key step in computational enzyme discovery and biocatalyst design. Recent representation-learning methods formulate this problem as enzyme--reaction matching, where paired enzymes and reactions are embedded into a shared space. However, most existing approaches primarily rely on pairwise enzyme--reaction supervision and make limited use of the relationships within reaction sets or enzyme families. This work introduces a multi-alignment contrastive learning framework for biochemical retrieval. The framework jointly models cross-domain compatibility between enzymes and reactions and within-domain relationships induced by functional annotations. In addition, a Gromov--Wasserstein-inspired regularization objective encourages geometric consistency between the learned enzyme and reaction representation spaces. By combining pairwise catalytic supervision with higher-order relational alignment, the model captures both direct enzyme--reaction associations and broader functional organization. We evaluate the approach on enzyme virtual screening and bidirectional enzyme--reaction retrieval tasks. Experiments on EnzymeMap show improved early-recognition performance under BEDROC and enrichment-factor metrics compared with strong contrastive baselines. On ReactZyme, the method achieves consistent gains across time-based, enzyme-similarity, and reaction-similarity splits, demonstrating robustness to unseen enzymes and unseen reactions. Ablation studies further indicate that within-domain alignment, functional supervision, and the geometric regularization term each contribute to the observed improvements. These results suggest that modeling multiple forms of alignment can improve contrastive retrieval models for enzyme discovery, reaction annotation, and related computational biology applications.
Shortest paths in planar domains with hyperbolic type metrics
arXiv:2512.17211v2 Announce Type: replace-cross Abstract: We study planar domains $G$ equipped with a hyperbolic type metric and approximate geodesics that join two points $x,y \in G$ and their lengths. We present an algorithm that enables one to approximate the shortest distance in polygonal domains taken with respect to the quasihyperbolic metric. The method is based on Dijkstra's algorithm, and we give several examples demonstrating how the algorithm works and analyze its accuracy. We experimentally demonstrate several previously theoretically observed features of geodesics, such as the relationship between hyperbolic and quasihyperbolic distance in the unit disk. We also investigate bifurcation of geodesics and the connection of this phenomenon to the medial axis of the domain.
Generalizable Video Quality Assessment via Weak-to-Strong Learning
arXiv:2505.03631v5 Announce Type: replace Abstract: Video quality assessment (VQA) seeks to predict the perceptual quality of a video in alignment with human visual perception, serving as a fundamental tool for quantifying quality degradation across video processing workflows. The dominant VQA paradigm relies on supervised training with human-labeled datasets, which, despite substantial progress, still suffers from poor generalization to unseen video content. In this work, we explore weak-to-strong (W2S) learning as a new paradigm for advancing VQA without reliance on human-labeled datasets. We first provide empirical evidence that a straightforward W2S strategy allows a strong student model to not only match its weak teacher on in-domain benchmarks but also surpass it on out-of-distribution (OOD) benchmarks, revealing a distinct weak-to-strong effect in VQA. Building on this insight, we propose a novel framework that enhances W2S learning from two aspects: (1) integrating homogeneous and heterogeneous supervision signals from diverse VQA teachers -- including off-the-shelf VQA models and synthetic distortion simulators -- via a learn-to-rank formulation, and (2) iterative W2S training, where each strong student is recycled as the teacher in subsequent cycles, progressively focusing on challenging cases. Extensive experiments show that our method achieves state-of-the-art results across both in-domain and OOD benchmarks, with especially strong gains in OOD scenarios. Our findings highlight W2S learning as a principled route to break annotation barriers and achieve scalable generalization in video quality assessment. Our data and code will be available at https://github.com/clh124/W2S-VQA.
Integrated photon-pair sources on periodically poled thin-film lithium tantalate
arXiv:2605.24988v1 Announce Type: new Abstract: Chip-integrated photon-pair sources based on spontaneous parametric down-conversion (SPDC) have emerged as a promising solution for scalable quantum light generation. Thin-film lithium tantalate (TFLT) is a compelling $\chi^{(2)}$ platform, combining strong nonlinearity with a high optical-damage threshold, weak photorefractive response, and ferroelectricity that enables quasi-phase matching. However, SPDC-based photon-pair generation on TFLT has not yet been demonstrated. Here, we combine high-quality periodic poling with low-loss nanophotonic waveguides to realize photon-pair sources on TFLT in both traveling-wave and resonant configurations. In periodically poled straight waveguides, we achieve broadband photon-pair generation with high efficiency ($2.1~\mathrm{GHz}~\mathrm{mW}^{-1}$) and coincidence-to-accidental ratio (up to $3.8\times10^{5}$). We further confirm high-purity single-photon operation via heralded second-order correlation ($g^{(2)}_\mathrm{H}(0) = 0.0018 \pm 0.0002$) and high-fidelity time-energy entanglement through Franson interference (visibility of $98.9 \pm 0.5\%$). In periodically poled racetrack resonators, we map out a broad quantum frequency comb spanning the telecom C- and L-bands. By isolating individual frequency-correlated pairs, we measure a high spectral brightness of $11~\mathrm{GHz}~\mathrm{mW}^{-1}~\mathrm{GHz}^{-1}$. These results are competitive with the state of the art across $\chi^{(2)}$ integrated platforms, positioning TFLT as a strong contender for integrated quantum light sources, with applications in wavelength-multiplexed quantum communications and photonic quantum information processing.
DNA end tethering through break-induced DNA--protein condensation
arXiv:2605.24987v1 Announce Type: new Abstract: Cells deploy robust mechanisms to repair DNA damage, safeguarding genomic stability and cellular health, but the physical principles underlying these processes remain incompletely understood. Experiments show \emph{in vitro} that upon a DNA double-strand break, a DNA--protein condensate can tether the broken DNA ends before they disperse away, a critical step for subsequent repair biochemistry. However, it remains puzzling how such condensation reliably achieves spatiotemporal localization at the break site and captures both broken ends despite intrinsic stochasticity. Here, we propose that broken DNA ends can trigger a conversion of proteins from a soluble state to a condensate-competent state. Combining this idea with Brownian dynamics simulations and theory, we propose a physical mechanism for reliable DNA-end tethering. Simulations show that such break-induced conversion can drive local DNA--protein condensation with two possible outcomes: successful or failed tethering. To rationalize this, we construct an effective free energy landscape, identify the corresponding stationary states, and demonstrate that tethering is governed by a kinetic competition between polymer relaxation and condensation dynamics. Together, our study shows that DNA end-dependent conversion, coupled with DNA--protein condensation, can reliably tether broken DNA ends.
A Guided Tour of Modern Domain Decomposition: From Schwarz Iterations to Robust Preconditioners and HPC Implementations
arXiv:2605.24982v1 Announce Type: new Abstract: Domain decomposition methods (DDMs) provide a unifying framework for the scalable numerical solution of partial differential equations. Originating from Schwarz's alternating method, they have evolved into a rich family of algorithms that combine local robustness with global convergence acceleration and natural parallelism. Over the past decades, domain decomposition has played a central role in enabling large-scale simulations in numerous applications. This chapter presents an overview of modern DDMs, with a particular emphasis on scalable preconditioning techniques for challenging problems, including indefinite and high-frequency regimes. We revisit the fundamental concepts - overlapping decompositions, partition of unity, additive and restricted Schwarz formulations - and explain their algebraic interpretations. We then clarify their role as preconditioners in Krylov subspace solvers and discuss the necessity of coarse space corrections for scalability. Beyond a the survey aspect, the chapter distills key theoretical insights and practical design principles that have emerged over the past twenty years. Special attention is given to robust coarse spaces (GenEO, DtN-based approaches) and high-performance implementations. The goal is to provide both a coherent overview of the field and a concise, practice-oriented guide for readers seeking to understand and apply domain decomposition methods without navigating the entire literature.