Forskningsradar

Science Journals

Peer-reviewade publikationer — 53899 artiklar

MolPIF: A Parameter Interpolation Flow Model for Molecule Generation
arXiv:2507.13762v4 Announce Type: replace Abstract: Motivation: Structure-based drug design (SBDD) has advanced with deep generative models, but bridging the gap between continuous atomic coordinates and discrete atom types remains a challenge. Current approaches, such as diffusion and flow matching models, often fail to unify these heterogeneous modalities, relying on separate strategies or ill-fitting Euclidean metrics for discrete variables. This lack of a consistent framework limits generative models' ability to capture the geometric and chemical structure of protein-ligand complexes. Results: We present MolPIF, a parameter interpolation flow mechanism designed to unify the generation of continuous and discrete molecular variables. Unlike traditional flow models that operate in sample space, MolPIF interpolates between distributions in the parameter space, theoretically recovering Wasserstein-2 optimal transport for continuous coordinates and establishing Fisher-Rao geodesics for discrete atom types. We further incorporate a geometry-enhanced learning strategy to improve the capture of atomic contexts. Extensive evaluations on the CrossDocked2020 dataset demonstrate that MolPIF outperforms baselines in binding affinity, chemical validity, geometric fidelity and chemical space coverage. Additionally, MolPIF exhibits versatility in lead optimization and offers flexible prior distribution selection (such as Laplace), establishing a robust paradigm for SBDD. Availability: Source code is freely available at https://github.com/BLEACH366/MolPIF. Supplementary information: Supplementary data are available at Bioinformatics.
Silicon-on-sapphire metasurfaces generate arrays of dark and bright traps for neutral atoms
arXiv:2601.01038v2 Announce Type: replace Abstract: We demonstrated crystalline silicon-on-sapphire (c-SOS) metasurfaces that convert a Gaussian beam into arrays of complex optical traps, including arrays of optical bottle beams that trap atoms in dark regions interleaved with bright tweezer arrays. The high refractive index and indirect band gap of crystalline silicon makes it possible to design high-resolution near-infrared ($\lambda>700$ nm) metasurfaces that can be manufactured at scale using CMOS-compatible processes. Compared with active components like spatial light modulators (SLMs) that have become widely used to generate trap arrays, metasurfaces provide an indefinitely scalable number of pixels, enabling large arrays of complex traps in a very small form factor, as well as reduced dynamic noise. To design metasurfaces that can generate three-dimensional bottle beams to serve as dark traps, we modified the Gerchberg-Saxton algorithm to enforce complex-amplitude profiles at the focal plane of the metasurface and to optimize the uniformity of the traps across the array. We fabricated and measured c-SOS metasurfaces that convert a Gaussian laser beam into arrays of bright traps, dark traps, and interleaved bright/dark traps.
CSV-ViT: A Vision Transformer with the Variable-sized Cortical Supervertices for Detection of Alzheimer's Disease Pathologies
arXiv:2605.26514v1 Announce Type: new Abstract: Confirming Alzheimer's disease (AD) typically relies on positron emission tomography (PET), which remains costly and invasive, motivating the use of structural MRI-based prescreening. Deep learning on non-Euclidean manifolds, particularly brain cortical surfaces, faces significant challenges due to the data's spherical topology. Recent surface models have enabled learning from cortical surface data; however, imposing face-based uniform patches often causes duplicate vertices at patch boundaries. In general, many surface-based models are limited in their awareness of the region of interest (ROI), which can result in non-cortical regions, such as the medial wall, being included. We propose a cortical surface tokenization that performs ROI-preserving, vertex-based, variable-sized patch partitioning. We refer to these cortical surface patches as cortical supervertices (CSVs). Building on this representation, we design the CSV Vision Transformer (CSV-ViT), a variable-size patch-tolerant Vision Transformer that uses padding and a mask-aware patch embedding. We used T1-weighted MRI and evaluated our framework by classifying AD-related status into three categories: AD diagnosis, amyloid positivity, and tau positivity. Across the experiments, CSV-ViT achieved higher classification performance than recent surface-based models. The results suggest that the proposed CSV-ViT may support MRI-based prediction of AD-related status prior to PET or CSF confirmation.
Non-equilibrium exciton dynamics in tailored molecular potentials of Rydberg ion crystals
arXiv:2605.21250v2 Announce Type: replace Abstract: Trapped ions excited to high-lying electronic states combine strongly coupled collective vibrational and electronic degrees of freedom with long-ranged interparticle interactions. These ingredients enable the quantum simulation of biochemical processes, associated with the dynamics of excitons in non-perturbative parameter regimes. The key feature of such a quantum simulator are electronic-state-dependent molecular potential surfaces which can be strongly coupled. This allows to shed light on a variety of mechanisms underlying exciton transport. We illustrate this in a system of three trapped ions, which is amenable to an ab initio treatment. Given that ion traps can be routinely prepared with hundreds of ions, these quantum simulators can immediately realise scenarios which are inaccessible by current numerical methods.
SLA-Aware Traffic Steering in Hybrid TN-NTN 5G Backhaul: A Potential Game Approach
arXiv:2605.26673v1 Announce Type: new Abstract: The integration of Non-Terrestrial Networks (NTN) with Terrestrial Networks (TN) is a key enabler for resilient 5G-Advanced and future 6G backhaul infrastructures. However, managing traffic across these highly asymmetric links remains a significant routing challenge, as systems must support heterogeneous network slices with conflicting service-level agreements (SLAs) while selectively utilizing costly NTN resources. This paper presents a computationally lightweight SLA-aware traffic-steering framework for a hybrid TN-NTN backhaul that models the load-balancing problem as an exact potential game. This mathematical foundation inherently enables decentralized coordination between uplink and downlink load-balancing agents without control-message overhead. By formulating traffic steering as a coupled optimization problem, per-slice (or per-user group) traffic fractions are dynamically distributed across terrestrial and satellite paths based on utility functions that capture throughput, latency, packet loss, and SLA penalties. The resulting game admits a pure Nash equilibrium, ensuring stable and predictable traffic adaptation under non-stationary load conditions. The framework is evaluated on a geographically distributed 5G testbed, using bidirectional traffic generated for five representative slices. Experimental results show that the proposed controller significantly outperforms heuristic and conventional baselines, reducing SLA violations to 1.7% for V2X and 0.7% for the emergency slice while completely eliminating them for video, IoT, and best-effort traffic.
A gauge identity for interscale transfer in inhomogeneous turbulence
arXiv:2512.20653v4 Announce Type: replace Abstract: Local interscale energy transfer in Large Eddy Simulation (LES) is typically diagnosed using the subgrid-scale (SGS) production, $\Pi^{SGS}$. In this work, an exact algebraic gauge identity is derived, demonstrating that $\Pi^{SGS}$ is composed of a kernel-integrated increment-based transfer, $\Pi^{inc}$, and the divergence of a spatial transport current, $\nabla \cdot J$. This identity was verified to machine precision ($10^{-16}$) using the analytical multi-harmonic Womersley solution. Further evaluation was conducted via Direct Numerical Simulation (DNS) of turbulent channel flow at $Re_\tau \approx 1000$. It was observed that $\nabla \cdot J$ dominates $\Pi^{SGS}$ in the near-wall region. The results suggest that $\Pi^{SGS}$ is not a unique proxy for the local cascade in inhomogeneous flows. A new framework is thus provided for the interpretation of interscale transfer diagnostics in wall-bounded transport phenomena.
PersLitEval: Fine-grained Benchmark and Evaluation of LLMs on Persian Literature Questions
arXiv:2605.27015v1 Announce Type: new Abstract: Despite impressive multilingual capabilities, large language models (LLMs) remain poorly evaluated on literary knowledge in non-English languages. We introduce PersLitEval, a benchmark of 4,514 Persian literature multiple-choice questions across eight fine-grained categories spanning spelling, literary devices, grammar, vocabulary, word formation, and conceptual understanding, sourced from materials for the Konkur university entrance examination. We evaluate six LLMs across ten prompting strategies, revealing striking category-level disparities across three tiers of task difficulty: models reach higher accuracy on conceptual similarity tasks but struggle with formal linguistic analysis, with spelling and word formation proving the hardest across all models. Prompting strategy has a significant impact on performance, with explained few-shot examples yielding the best results, particularly on formal linguistic categories. An error analysis identifies three failure modes: semantic comprehension gaps, formal linguistic knowledge gaps, and counting/enumeration errors, suggesting that different categories require different improvement strategies.
A Technical Policy Blueprint for Trustworthy Decentralized AI
arXiv:2512.11878v3 Announce Type: replace Abstract: Decentralized AI systems, such as federated learning, can play a critical role in further unlocking AI asset marketplaces (e.g., healthcare data marketplaces) thanks to increased asset privacy protection. Unlocking this big potential necessitates governance mechanisms that are transparent, scalable, and verifiable. However current governance approaches rely on bespoke, infrastructure-specific policies that hinder asset interoperability and trust among systems. We are proposing a Technical Policy Blueprint that encodes governance requirements as policy-as-code objects and separates asset policy verification from asset policy enforcement. In this architecture the Policy Engine verifies evidence (e.g., identities, signatures, payments, trusted-hardware attestations) and issues capability packages. Asset Guardians (e.g. data guardians, model guardians, computation guardians, etc.) enforce access or execution solely based on these capability packages. This core concept of decoupling policy processing from capabilities enables governance to evolve without reconfiguring AI infrastructure, thus creating an approach that is transparent, auditable, and resilient to change.
Fast Quadratic Manifold Learning For Nonlinear Dimensionality Reduction in Large-scale Systems using Riemannian Optimization
arXiv:2605.26039v2 Announce Type: replace Abstract: The effectiveness of dimensionality reduction with quadratic manifolds hinges on the choice of a reduced basis and the associated quadratic correction terms. Existing approaches typically rely on subspaces spanned by the leading principal components of the training data. Although optimal for linear approximation, such bases are inherently suboptimal for quadratic manifold learning. Greedy basis-selection methods can significantly improve the representational capacity of quadratic manifolds by searching over a larger pool of candidate principal components, but the combinatorial cost limits the basis sizes that can be used in practice. This work proposes FastQM, an approach that treats the identification of an optimal quadratic approximation as a continuous optimization problem on the Stiefel manifold. By rotating the reduced basis within a candidate span of singular vectors, FastQM learns an ideal coordinate alignment tailored to quadratic manifold approximation. A feature-space formulation ensures that the optimization cost scales independently of the full state-space dimension. The efficacy of the proposed method is demonstrated on a turbulent airfoil-wake large-eddy simulation.
Vibroacoustic Underwater Noise from Fixed and Floating Offshore Wind Turbines
arXiv:2605.26699v1 Announce Type: new Abstract: Anthropogenic underwater noise from offshore wind turbines is a growing environmental concern, particularly with the large-scale deployment of bottom-fixed and floating devices. This study presents a physics-based vibroacoustic framework to predict operational underwater noise emissions from offshore wind turbines and compares monopile-supported and floating configurations for a 10 MW turbine. The methodology combines time-domain aero-hydro-servo-elastic simulations with a frequency-domain acoustic formulation based on equivalent dipole sources and Green's function solutions, accounting for underwater confinement between the free surface and seabed through the method of images. Results show that floating configurations exhibit enhanced low-frequency acoustic emissions, producing up to 15% higher OASPL than the monopile structures under equivalent water depths for frequencies below 10 Hz due to additional rigid-body motions, while monopile structures radiate more efficiently at higher frequencies associated with drivetrain excitations. Significant differences in the spatial distribution and directivity of the radiated sound field are also observed, with floating platforms displaying more complex three-dimensional radiation patterns and stronger direction-dependent variations, reaching approximately 20-25 dB in the 100-1000 Hz band, compared with the smoother and nearly axisymmetric response of monopile configurations. Water depth strongly influences propagation regimes and overall sound levels, with shallow-water floating configurations showing variations of up to 7% in OASPL relative to deep-water cases. The proposed framework enables quantification of vibro-acoustic noise and provides a predictive tool for assessing underwater acoustic impacts during the design phase, supporting environmentally informed offshore wind turbine design and future regulatory and monitoring strategies.
A Logical View of GNN-Style Computation and the Role of Activation Functions
arXiv:2512.19332v2 Announce Type: replace Abstract: We study the numerical and Boolean expressiveness of MPLang, a declarative language that captures the computation of graph neural networks (GNNs) through linear message passing and activation functions. We begin with A-MPLang, the fragment without activation functions, and give a characterization of its expressive power in terms of walk-summed features. For bounded activation functions, we show that (under mild conditions) all eventually constant activations yield the same expressive power - numerical and Boolean - and that it subsumes previously established logics for GNNs with eventually constant activation functions but without linear layers. Finally, we prove the first expressive separation between unbounded and bounded activations in the presence of linear layers: MPLang with ReLU is strictly more powerful for numerical queries than MPLang with eventually constant activation functions, e.g., truncated ReLU. This hinges on subtle interactions between linear aggregation and eventually constant non-linearities, and it establishes that GNNs using ReLU are more expressive than those restricted to eventually constant activations and linear layers.
Integrating Network and Attack Graphs for Service-Centric Impact Analysis
arXiv:2507.00637v3 Announce Type: replace Abstract: Cyberattacks on enterprise networks exploit complex dependencies among infrastructure, services, and applications, which challenge traditional analysis methods that focus on attack paths or network topology in isolation. In this study, we introduce a novel probabilistic multilayer modelling framework, based on influence propagation in networks, that integrates attack graphs with the communication network topology, enabling a service-centric impact analysis of cyberattacks. Our method captures both the vulnerability exploitability and network connectivity, allowing us to assess the likelihood of attack propagation and cumulative impacts across interconnected services. By integrating standard vulnerability metrics (such as CVSS) with the network-level connectivity probabilities, the framework provides a cohesive view of the dynamics of cyberattacks. We validate this approach using a realistic case study of an enterprise network, demonstrating its ability to determine critical nodes, vulnerabilities, and service dependencies that significantly influence attack outcomes. Our findings show that integrating network and attack graph perspectives offers more actionable insights into risk assessment and mitigation planning, advancing the analysis of cyberattacks in complex networked environments.
Shedding Light on Dark Matter at the LHC with Machine Learning
arXiv:2509.15121v2 Announce Type: replace-cross Abstract: We investigate a WIMP dark matter (DM) candidate in the form of a singlino-dominated lightest supersymmetric particle (LSP) within the $Z_3$-symmetric Next-to-Minimal Supersymmetric Standard Model (NMSSM). This framework gives rise to regions of parameter space where DM is obtained via co-annihilation with nearby higgsino-like electroweakinos and DM direct detection~signals are suppressed, the so-called ``blind spots''. On the other hand, collider signatures remain promising due to enhanced radiative decay modes of higgsinos into the singlino-dominated LSP and photons, rather than into leptons or hadrons. Compared to MSSM scenarios with light bino- and wino-like electroweakinos, the NMSSM allows for final states with multiple photons arising from cascade radiative decays, providing a distinctive collider signature. This motivates searches for radiatively decaying neutralinos, however, these signals face substantial background challenges, as the decay products are typically soft due to the small mass-splits ($\Delta m$) between the LSP and the higgsino-like coannihilation partners. We apply a data-driven Machine Learning (ML) analysis that improves sensitivity to these subtle signals, offering a powerful complement to traditional search strategies to discover a new physics scenario. Using an LHC integrated luminosity of $100~\mathrm{fb}^{-1}$ at $14~\mathrm{TeV}$, the method achieves a $5\sigma$ discovery reach for higgsino masses up to $225~\mathrm{GeV}$ with $\Delta m\!\lesssim\!12~\mathrm{GeV}$, and a $2\sigma$ exclusion up to $285~\mathrm{GeV}$ with $\Delta m\!\lesssim\!20~\mathrm{GeV}$. These results highlight~the power of collider searches to probe DM candidates that remain hidden from current~direct detection experiments, and provide a motivation for a search by the LHC collaborations using ML methods.
From Scores to Gibbs Correctors: Accelerating Uniform-Rate Discrete Diffusion Models
arXiv:2605.27352v1 Announce Type: new Abstract: Discrete diffusion models have achieved strong empirical performance in text and other symbolic domains, but, especially for uniform-rate models, they often require many steps to generate a single sample. Existing acceleration methods either rely on training additional quantities or suffer from slow mixing. In this work, we propose a novel Gibbs-based corrector for discrete diffusion models, termed Gibbs-Accelerated Discrete Diffusion (GADD). GADD leverages the structure of the concrete score function to construct Gibbs posterior likelihoods directly, without requiring any additional training beyond standard score estimation. We show that GADD achieves an overall sampling complexity of $\mathcal{O}(\mathrm{polylog} (\varepsilon^{-1}))$, yielding the first such rate for diffusion-based samplers for uniform-rate discrete diffusion models. We also conduct numerical experiments demonstrating the practical advantages of GADD across synthetic data, zero-shot text sampling, and zero-shot conditional music generation. These results corroborate the theory and show that GADD consistently improves sample quality and wall-clock efficiency over standard baselines, including vanilla Euler methods and CTMC correctors. Beyond this, our theoretical analysis introduces a novel framework for analyzing predictor-corrector methods in discrete diffusion models, which may be of independent interest. Unlike existing approaches that rely on the Girsanov change-of-measure technique, our method is based on an induction argument that tracks error propagation across predictor iterations while accounting for inaccuracies in the corrector updates.
The Sensation Modulating Network:Haltability as the architectural ground for object-directed phenomenology
arXiv:2605.26856v1 Announce Type: cross Abstract: Cognitive science remains split between cognitivism - which accounts for recursion and language but cannot ground formal symbols in meaning - and 4E approaches - which ground cognition in the body but rarely specify the body's architecture in enough detail to support generativity. We argue the impasse stems from an incomplete account of the embodied agent's architecture, and propose one: the Sensation Modulating Network (SMN), the cognitive agent conceived as the whole body, organized at every anatomical scale by opponent dynamics, built from Sensation Modulators that sense and act through one substrate, paired into Coordinated Action Zones routed by a body-wide broadcast network. Three commitments give the SMN its purchase. Haltability - the recruitment of antagonistic affordance into co-activated equilibrium - provides the architectural locus that object-directed phenomenology, in Husserl's sense, requires: opponency enables co-activation, co-activation enables halt, halt enables attention, attention enables intentional directedness, with no module added on top. The dual-signal property of self-modulatable action patterns (SMAPs) makes the self/world distinction a structural feature of the wiring rather than a category the agent applies. And a four-level action-pattern hierarchy - Basal, Haltable, Negotiable, Transactional - gives a single trajectory from autonomic regularity to public conventionalization, locating the conditions for grammar-grounded generativity as architectural transitions. The SMN reconciles the cognitivism-4E debate: recursion lives in the modifiable dynamics of Negotiable Action Patterns, embodiment in the opponent substrate that supports them. A tentative formalism and eight predicted registers (seven testable, one hypothetical), with reference simulations, are given in an appendix.
Scalable Algorithm for Dynamic Quasi-clique Detection
arXiv:2605.26235v1 Announce Type: new Abstract: Identifying dense subgraphs known as quasi-cliques is pivotal in numerous graph mining tasks across domains such as social networks, biology, and e-commerce. While prior work has developed efficient algorithms for quasi-clique detection in static graphs, real-world networks are inherently dynamic, where edges appear and disappear continuously. This renders static methods inefficient and ill-suited for real-time analysis. In this paper, we initiate the study of the Dynamic Maximum Quasi-Clique Problem (DMQCP), which aims to maintain and update the largest quasi-clique in a graph under streaming graph updates. We propose DMI, a novel MinHash-based dynamic framework that supports fast, high-quality approximate maintenance of quasi-cliques. DMI leverages two update-efficient hashing schemes, i.e., $l$-buffered $k$-MinHash and Bottom-$k$ MinHash, to maintain candidate quasi-cliques incrementally. To ensure robustness and reduce bias, we further design a batch reconstruction strategy to periodically rebuild the candidate set, guaranteeing both stability and adaptability under frequent updates. Extensive experiments on real-world and synthetic datasets show that DMI achieves up to four orders of magnitude speedup over static baselines, while preserving solution quality. As a side product, we also propose a framework NSF that primarily uses the neighbor-search technique to maintain quasi-clique candidates while edge updating. This work establishes the first efficient algorithmic framework for dynamic quasi-clique extraction, enabling scalable and real-time dense subgraph mining in evolving networks.
Cost of Structural Learning Under Censored Feedback: A Threshold-Bandit Approach
arXiv:2605.27076v1 Announce Type: new Abstract: In many multi-agent applications, tasks yield rewards only when executed by a coalition meeting an unknown size threshold; otherwise, feedback is fully censored. This censorship creates an identifiability problem: agents cannot distinguish stochastic failure from insufficient coordination. We formalize this setting as the Threshold-Activated Cooperative Multi-Armed Bandit (TAC-MAB) and analyze it under both centralized and decentralized coordination. We show that a centralized algorithm (C-TAC) achieves cumulative regret O(log T), decomposed into a structural-search term that captures the cost of resolving feasibility under censored feedback and a statistical-monitoring term for value estimation. We then introduce D-TAC, a decentralized event-triggered protocol in which agents synchronize only when their structural beliefs change. Empirically, D-TAC achieves a 23x reduction in communication relative to the centralized baseline while preserving feasibility alignment under conservative belief fusion. These results characterize the coordination cost of learning under censored feedback and show that near-centralized communication efficiency is achievable without continuous synchronization.
A Symmetric Unified Transport and Charge Model for Metal-Oxide-Semiconductor Field-Effect Transistor from Diffusive to Ballistic Regimes
arXiv:2604.23541v3 Announce Type: replace-cross Abstract: This paper presents a symmetric unified transport (UT) compact model for metal-oxide-semiconductor field-effect transistors (MOSFETs) that bridges drift-diffusion (DD) and ballistic transport (BT) regimes. The proposed model self consistently accounts for both current and charge across the DD-BT transition. Quantum capacitance and carrier transport are incorporated into the charge density formulation. Drain side velocity saturation and the source side thermal velocity limit are unified within a single framework using a physically motivated high field scattering length, enabling accurate modeling from DD square law behavior to the ballistic limit. In addition, a physical channel charge and capacitance model is developed to capture capacitance reduction in the quasi-ballistic regime, which is not considered in standard compact models. The model is verified using theoretical analysis and experimental data from MOSFETs with multiple channel lengths, achieving accurate fitting using only physically motivated model parameters. The formulation is continuous and symmetric, and it passes both DC and AC symmetry tests.
Local imperfect feedback control in non-equilibrium biophysical systems enabled by thermodynamic constraints
arXiv:2507.07295v2 Announce Type: replace-cross Abstract: How biological networks achieve robust control despite relying on imperfect, local information remains an important open question. Here, we identify thermodynamic constraints that can curtail non-equilibrium steady-state responses so severely that even crude, local feedback rules can achieve globally stable control without requiring precise network design or global information. Specifically, using Markov jump processes as a general framework for biophysical dynamics, we derive general non-equilibrium response constraints showing that for many classes of rate perturbations, steady-state responses have fixed signs across all driving strengths, so that near-equilibrium responses predict far-from-equilibrium behavior regardless of system complexity. These constraints clarify several biological phenomena: monotonicity is thermodynamically guaranteed whenever a perturbation acts on a single transition rate, and non-monotonic responses, as observed for example in transcription factor regulation, arise only when an input simultaneously modulates multiple rates. Even in this case, we identify a graph-theoretic concept termed ``coherence'' that allows for a restoration of monotonicity. We show how coherence naturally and generally emerges in classic biophysical models of adaptation, including E. coli chemotaxis, and transcription factor regulation when biological constraints on network parameterization are included. We next show that, within a control-theoretic framework, these constraints guarantee that simple linear feedback on small subsets of kinetic rates achieves globally stable tracking and adaptation without coordinated manipulation of many variables. For systems with one regulator, local stability implies global stability for arbitrary network topologies without fine tuning, revealing that non-equilibrium thermodynamics fundamentally constrains biochemical network responses.
MATCHA: Matching Text via Contrastive Semantic Alignment
arXiv:2605.27345v1 Announce Type: new Abstract: Reliable evaluation is essential for understanding large language model (LLM) performance, yet today's go-to metrics, namely token-overlap scores (e.g., ROUGE) and embedding-based measures (e.g., BERTScore), often misjudge semantic similarity of documents. Our study shows that both token-overlap metrics and embedding-based metrics routinely assign nearly identical scores to texts that directly contradict each other, thereby potentially masking fundamental errors. We introduce MATCHA, an automatic metric that jointly rewards semantic agreement with a reference and penalizes contradictions. MATCHA employs a dual-view perspective that measures (i) proximity to the gold text and (ii) distance from an adversarially generated counterfactual contradiction. In eight public benchmarks, MATCHA outperforms popular metrics, compared with human annotations on question-answering, image caption generation, natural language inference, summarization, and semantic textual similarity tasks. On the TruthfulQA dataset (i.e., a dataset without a training set, where no embedding-based metrics could locally train on), this improvement in terms of matching texts with a reference reaches 18.38% over ROUGE-L and 20.82% over BERTScore. Both quantitative comparison and qualitative human assessments confirm the efficacy and validity of MATCHA and uncover fundamental weaknesses in pre-existing metrics. Compared with 23 embedding models, including top state-of-the-art ones, used as a metric similar to BERTScore, MATCHA remains the most accurate in distinguishing correct from incorrect statements solely based on a reference. Our code and metric are publicly available (https://github.com/Siran-Li/MATCHA).
VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
arXiv:2510.17759v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) extend large language models with visual reasoning, but their multimodal design also introduces new, underexplored vulnerabilities. Existing multimodal red-teaming methods largely rely on brittle templates, focus on single-attack settings, and expose only a narrow subset of vulnerabilities. To address these limitations, we introduce VERA-V, a variational inference framework that recasts multimodal jailbreak discovery as learning a joint posterior distribution over paired text-image prompts. This probabilistic view enables the generation of stealthy, coupled adversarial inputs that bypass model guardrails. We train a lightweight attacker to approximate the posterior, allowing efficient sampling of diverse jailbreaks and providing distributional insights into vulnerabilities. VERA-V further integrates three complementary strategies: (i) typography-based text prompts that embed harmful cues, (ii) diffusion-based image synthesis that introduces adversarial signals, and (iii) structured distractors to fragment VLM attention. Experiments on HarmBench and HADES benchmarks show that VERA-V consistently outperforms state-of-the-art baselines on both open-source and frontier VLMs, achieving up to 53.75% higher attack success rate (ASR) over the best baseline on GPT-4o. We include the code on the project page available here: https://github.com/kxwhiowo/VERA-V
Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion
arXiv:2605.26383v1 Announce Type: new Abstract: Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and large intra-class appearance variations. Objects may leave and re-enter the field of view, and the large diversity of instances with limited annotations makes supervised ReID difficult to scale, motivating zero-shot approaches. We study zero-shot object ReID on the EPIC-Kitchens benchmark, where the goal is to match active food and kitchen-tool instances across frames using only pre-trained visual features. We first evaluate five state-of-the-art feature extractors, including Vision-Language Models (VLMs) - CLIP, DINOv2, DreamSim, I-JEPA, and SAM3 - and show that zero-shot methods fail, with the best baseline achieving only 45.3% mAP. We then propose an Enhanced SAM3 ReID Pipeline, a zero-shot multi-stage method built around SAM3 segmentation as the core component. Stage 1 uses SAM3 to suppress background clutter. Stage 2 fuses embeddings from SAM3, DINOv2, and CLIP into a single L2-normalized descriptor. Stage 3 augments cosine similarity with mask-shape IoU for geometric consistency, and Stage 4 applies k-reciprocal re-ranking. The full pipeline improves performance by 7.5% mAP to 52.8%.
LLMs versus the Halting Problem: Characterizing Program Termination Reasoning
arXiv:2601.18987v5 Announce Type: replace Abstract: Determining whether a program terminates is a central problem in computer science. Turing's Halting Problem established termination as undecidable, showing that no algorithm can universally determine termination for all programs and inputs. Hence, verification tools approximate termination, sometimes failing to prove or disprove; these tools rely on problem specific architectures, and are usually tied to particular programming languages. Recent advances in LLMs raise a natural question: To what extent can they reason about program termination? We evaluate frontier LLMs on a diverse set of C programs from the International Competition on Software Verification (SV Comp) 2025. Our results show that GPT-5 and Claude Sonnet 4.5 achieve scores comparable to top ranked verification tools (with test time scaling). However, while models often correctly infer whether programs terminate, they frequently fail to construct a witness as formal proof, revealing a gap between semantic recognition and symbolic proof generation. Performance further degrades as code length increases. To analyze this gap, we introduce a divergence precondition formulation that characterizes non termination conditions as logical constraints. We hope these findings motivate future research on real-world termination benchmarks, neuro-symbolic approaches that combine LLMs with symbolic verification methods, and, more broadly LLM reasoning on other undecidable problems.
Nonlinear spectral clustering with C++ GraphBLAS
arXiv:2605.26975v1 Announce Type: new Abstract: Nonlinear reformulations of the spectral clustering method have gained a lot of recent attention due to their increased numerical benefits and their solid mathematical background. However, the estimation of the multiple nonlinear eigenvectors is associated with an increased computational cost. We present an implementation of a direct multiway spectral clustering algorithm in the $p$-norm, for $p\in(1,2]$, using a novel C++ GraphBLAS API. The key operations are expressed in linear algebraic terms and are executed over the resulting sparse matrices and dense vectors, parameterized in the algebra pertinent to the computation. We demonstrate the effectiveness and accuracy of our shared-memory algorithm on several artificial test cases. Our numerical examples and comparative results against competitive methods indicate that the proposed implementation attains high quality clusters in terms of the balanced graph cut metric. The strong scaling capabilities of our algorithm are showcased on a range of datasets with up to $8$ million nodes and $48$ million edges.
EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning
arXiv:2601.03471v3 Announce Type: replace Abstract: Reliable epidemiological reasoning requires synthesizing study evidence to infer disease burden, transmission dynamics, and intervention effects at the population level. Existing medical question answering benchmarks primarily emphasize clinical knowledge or patient-level reasoning, yet few systematically evaluate evidence-grounded epidemiological inference. We present EpiQAL, the first diagnostic benchmark for epidemiological question answering across diverse diseases, comprising three subsets built from open-access literature. The three subsets progressively test factual recall, multi-step inference, and conclusion reconstruction under incomplete information, and are constructed through a quality-controlled pipeline combining taxonomy guidance, multi-model verification, and difficulty screening. Experiments on fifteen models spanning open-source and proprietary systems reveal that current LLMs show limited performance on epidemiological reasoning, with multi-step inference posing the greatest challenge. Model rankings shift across subsets, and scale alone does not predict success. Chain-of-Thought prompting benefits multi-step inference but yields mixed results elsewhere. EpiQAL provides fine-grained diagnostic signals for evidence-grounding, inferential reasoning, and conclusion reconstruction.