arXiv:2605.16031v1 Announce Type: new
Abstract: Ventilated acoustic silencers combing sound attenuation with high ventilation are pivotal for advanced noise control. However, balancing attenuation, bandwidth, openness, and thickness remains a high-dimensional challenge. Here, we report a physics-aware machine-learning-driven inverse design framework for ultra-open acoustic silencers (UAS). By leveraging Green's function-based parameterization, we physically decouple the design space into spectral and radial parameters, ensuring physical interpretability while reducing complexity. We introduce a two-stage forward prediction architecture that captures broadband envelopes and sharp resonant features via a coarse-to-fine strategy. Coupled with a population-based, hybrid-objective parallel (PHP) inverse strategy, our framework enables rapid exploration of non-convex landscapes, identifying hundreds of optimized candidates within seconds. Crucially, this framework uncovers hidden linear design rules that govern high-performance monolithic designs, acting as geometric proxies for optimal impedance-matching. We experimentally validate a family of prototypes: UAS-2 demonstrates the monolithic limit with high ventilation ratio, while UAS-3 demonstrates versatility in multi-mode interactions. To circumvent the trade-off ceiling of single-unit resonators, a parallel-composite architecture (UAS-4) is introduced to enhance performance through spatial interference distribution. Results confirm a broadband bandwidth exceeding 830 Hz achieved with an ultra-thin profile (0.1-0.2{\lambda}) and 80% ventilation. This work establishes a data-driven paradigm for discovering design principles in functional metamaterials.
Science Journals
arXiv:2605.16026v1 Announce Type: new
Abstract: Compositional speech-to-speech translation (S2ST) systems built upon speech large language models (SpeechLLMs) have recently shown promising performance. However, existing S2ST systems often either neglect source-language information or encode it through a language-as-label paradigm, representing each source language as an independent flat embedding. Such a design overlooks systematic linguistic structure shared across languages, which may limit data-efficient multilingual adaptation when supervised S2ST data are scarce. To address this issue, we propose S2ST-Omni 2, a many-to-one compositional S2ST framework that systematically reformulates multilingual language conditioning from flat language labels to structured typological priors. Specifically, S2ST-Omni 2 revisits language conditioning at three levels: typology-informed hierarchical language encoding for structured source-language representation, dynamically-gated language-aware Dual-CTC for content-adaptive acoustic modulation, and typology-aware LLM prompting for decoder-side linguistic guidance. Experiments on CVSS-C show that S2ST-Omni 2 achieves superior average performance among representative S2ST approaches across BLEU, COMET, ASR-BLEU, and BLASER 2.0 under the adopted evaluation protocol. Ablation studies indicate that the proposed representation-level, acoustic-level, and decoding-level strategies provide complementary benefits. Moreover, controlled data-budget analyses and a Japanese-to-English evaluation using only approximately 3 hours of supervised training data suggest that explicit typological priors provide useful inductive biases for data-efficient multilingual S2ST.
arXiv:2605.16011v1 Announce Type: new
Abstract: Adaptive learning refers to educational technologies that track learners' learning progress and adapt the instructional process based on individual learners' learning performance. It is increasingly recognized as critical for developing an effective learning support tool. Vision language models (VLMs) have seen adoption in mathematics education, and students have been using them as learning aids for personalized instruction. However, it is unknown whether VLMs have the ability to adapt to different learner profiles when providing mathematical instructions. Current VLMs lack a systematic evaluation framework for this adaptivity to different learner profiles in mathematics tutoring tasks. To address this gap, we draw on the learner model from the adaptive learning framework (Shute and Towle, 2018) and propose a learner model-based rubric. Our rubric formalizes adaptivity assessment into three aspects: cognitive aspects, motivational aspects, and complexity. We also evaluate two additional dimensions of VLM responses: correctness (of answers and solutions) and quality (of the response itself). Our experimental results show measurable differences in adaptivity across models and also reveal that current VLMs struggle to consistently produce learner model-based instructional responses, especially when receiving limited learner information.
arXiv:2605.16009v1 Announce Type: new
Abstract: Local navigation is one of the fundamental problems in robot navigation, and numerous approaches have been proposed over the years, including methods such as the Dynamic Window Approach, Model Predictive Control, and more recently, Control Barrier Functions and machine learning based techniques. While these methods perform well in simple environments, many of them rely on optimization or learning based procedures that can struggle in more complex scenarios. In contrast, this article proposes a more geometric algorithmic approach that enables a local navigation method with faster computation times and longer planning horizons. The proposed method is based on the computation of a sequence of circular regions from a local LiDAR scan that expand in the direction of the goal and capture free local navigable space. The proposed method was implemented in the ROS2 framework and evaluated in a simulated environment.
arXiv:2603.25101v2 Announce Type: replace-cross
Abstract: We present a formulation of the problem of finding the smallest T -Count circuit that implements a given unitary as a binary search over a sequence of continuous minimization problems, and demonstrate that these problems are numerically solvable in practice. We reproduce best-known results for synthesis of circuits with a small number of qubits, and push the bounds of the largest circuits that can be solved for in this way. Additionally, we show that circuit partitioning can be used to adapt this technique to be used to optimize the T -Count of circuits with large numbers of qubits by breaking the circuit into a series of smaller sub-circuits that can be optimized independently.
arXiv:2605.16008v1 Announce Type: new
Abstract: Plaque assays remain the gold standard readout of virus infectivity; however, plaque counting from plate images is labor-intensive and prone to inter-operator variability. We present an end-to-end, computer-aided workflow for cytopathic effect-based virus titration directly from laboratory plaque assay images. The proposed approach combines two models derived from the Segment Anything Model (SAM): a SAM2-based well-segmentation module that localizes assay wells across heterogeneous imaging conditions, and a SAM-based plaque-segmentation model that detects and enumerates plaques within each well. The method was evaluated on a mixed dataset comprising private plaque assay images of Mayaro virus and Coxsackievirus B3, together with public Vaccinia virus images from the VACVPlaque dataset. The pipeline outputs per-well plaque counts, automatically computes plaque-forming units per milliliter (PFU/mL), and is integrated into a web-based platform that allows users to review results and organize experiments. On held-out plates (17 from MAYV/CVB3 and 22 from VACV), the workflow generalized across two plate formats (6-well and 12-well) and showed strong agreement with manual annotations (Pearson correlation coefficients of 0.92 for MAYV/CVB3 and 0.88 for VACV). Automated plaque counts were further compared with annotations from four independent experts, demonstrating high concordance. The proposed system will be open sourced and publicly released upon acceptance of this manuscript to enable reproducible, scalable, and audit-ready plaque assay analysis while substantially reducing manual annotation effort.
arXiv:2605.16012v1 Announce Type: cross
Abstract: BEDT-TTF-based organic conductors host a number of ground states, tuned by electron repulsion from Mott and charge ordered insulators to superconductors. Knowing charge distribution on the molecular sites in the insulating state of these materials is a key to understanding the origin of these ground states. We survey and discuss the C=C stretching modes in BEDT-TTF based molecular conductors. These molecular vibrations are extremely crucial in characterization of charge-ordered insulators, and are recently linked to superconductivity in some compounds. Focusing on the known examples of BEDT-TTF$^{+0.5}$ salts, we analyse the reliability of the C=C stretching modes for the determination of charge ordering and absolute site charge. Considering the charge-ordered states, a prominent shift in frequency of 141 cm$^{-1}$ per elementary charge $e$ for $\nu_{27}(b_{1u})$ and 98 cm$^{-1}$$e$ for $\nu_2$($a_g$) can be clearly realised, however, the distribution resulting from different compounds span over 20 cm$^{-1}$. For nominal BEDT-TTF$^{+0.5}$ compounds, the distribution of the resonance also extends around 20 cm$^{-1}$, yielding an unexpected large uncertainty of $\Delta\rho~\approx~(~\pm~0.045)e$, which is presumably due to the influence of small differences in the structure. This highlights the limitations of charge-frequency relations to detect small deviations in absolute charge values on molecular lattice sites, and emphasises on the use of the relations to estimate charge-ordering, rather than absolute site charge.
arXiv:2605.16001v1 Announce Type: new
Abstract: A broadcast on a connected graph is a function f that assigns each vertex v an integer f(v) with 0 <= f(v) <= ecc(v) where ecc(v) denotes the eccentricity of v. A vertex u hears a broadcasting vertex v (with f(v)>0) if u is at distance at most f(v) from v. Beyond the classical broadcast domination problem, where every vertex is required to hear at least one vertex, two variants raise intriguing combinatorial and algorithmic questions. In an independent broadcast, no broadcasting vertex hears another broadcasting vertex, while a broadcast packing requires that every vertex hears at most one broadcasting vertex. The corresponding problems Broadcast Independence and Broadcast Packing ask for broadcasts of values at least k under these constraints, where the value is the sum of the broadcast values. We initiate a systematic study of the parameterized complexity of such problems. We prove that Broadcast Independence and Broadcast Packing are FPT parameterized by the treewidth plus the diameter of G, with a family of dynamic-programming algorithms over nice tree decompositions. We obtain as a corollary that both problems are FPT parameterized by k and the treewidth of G and XP for treewidth only. The latter result shows that the known algorithm for trees (Bessy and Rautenbach, DAM 2022) can indeed be extended to bounded treewidth graphs. On the negative side, we show that Broadcast Independence is W[1]-hard parameterized by the pathwidth of G. Note that this result completes the picture for parameter k and treewidth for Broadcast Independence since it is known to be W[1]-hard for k only. We complement these results by showing that a weighted version of both problems, where the input comes with a weight function on the edges, is W[1]-hard parameterized by the vertex cover of G. Finally, we provide a constant-factor approximation algorithm parameterized by treewidth for Broadcast Independence.
arXiv:2605.15999v1 Announce Type: new
Abstract: This paper introduces a motion planning framework to plan morphology and trajectory for morphing quadrotors under extremely constrained environments. We develop a novel obstacle avoidance cost function for nonlinear model predictive control (MPC) that enables navigation through extremely narrow gaps under limited perception from a 2D LiDAR. Classical artificial potential field-based costs typically have a high cost in narrow passages, artificially blocking the navigable path. In contrast, we propose a smooth exponential obstacle cost that preserves low traversal cost within narrow gaps while maintaining strong collision avoidance behavior. The formulation avoids hard activation thresholds and introduces a cost reduction factor to reduce the cost within narrow passages. Direct use of 2D LiDAR measurements in MPC allows navigation around arbitrarily shaped obstacles. The method is embedded within an acados-based nonlinear MPC framework. Simulation and experimental results demonstrate successful traversal of narrow corridors where typical repulsive cost functions would fail. The approach provides a computationally efficient and practical solution for navigating through tight spaces while maintaining safety from the obstacles. While we are implementing the framework on the morphing quadrotors, the cost function formulation is general-purpose for any mobile robot application, and is not limited to the morphing quadrotors. The implementation code is available at \href{https://github.com/harshjmodi1996/morphocopter_mpc}{Github Repo} and a short video is available at \href{https://zh.engr.tamu.edu/wp-content/uploads/sites/310/2026/03/MPC_MorphoCopter_video.mp4}{Video Link}.
arXiv:2605.15998v1 Announce Type: new
Abstract: Coherent perfect absorbers (CPAs) have recently attracted considerable attention due to their ability to enhance light--matter interaction. By exploiting interference, CPAs enable even weakly absorbing materials to achieve complete absorption under appropriate excitation conditions. Generalizing this concept to the simultaneous absorption of arbitrary multimode input states remains challenging, however, since conventional implementations typically operate only for a single or a very small number of input channels. Here, we propose a compact realization of a multimode coherent perfect absorber based on a gradient-index (GRIN) fiber. Using the self-imaging property of the fiber, the bulky free-space architecture of previous approaches is replaced by a monolithic waveguiding platform that supports near-degenerate rephasing of many spatial modes. We show that standard GRIN profiles optimized for minimal intermodal dispersion enable highly efficient absorption of complex multimode fields, with field-of-view reflectivities well below \(1\%\) for realistic parameters. This approach provides a practical and scalable route toward efficient multimode absorption in fiber-based and integrated photonic systems, with potential applications in light harvesting, optical control, and imaging.
arXiv:2605.15964v1 Announce Type: new
Abstract: Aerial vision-language navigation (VLN) requires agents to follow natural-language instructions through closed-loop perception and action in 3D environments. We argue that aerial VLN can be formulated as a prediction-driven world-action problem: the agent should anticipate latent world evolution and act according to the predicted consequences. To this end, we propose WorldVLN, the first autoregressive world action model for aerial VLN. Unlike full-sequence video-generation world models that generate an entire visual clip, WorldVLN adapts a latent autoregressive video backbone to predict short-horizon world-state transitions and directly decodes them into executable waypoint actions. After each action segment is executed, newly received observations are encoded back into the autoregressive context, enabling closed-loop world-action prediction. We further introduce a two-stage training framework that first grounds the video prior in instruction-conditioned navigation dynamics and then develops Action-aware GRPO, the first reinforcement learning method tailored to autoregressive WAMs, to optimize waypoint decisions through their downstream rollout consequences. On public outdoor and indoor benchmarks, WorldVLN consistently outperforms existing Vision-Language-Action baselines with 12\%+ success-rate gains and larger advantages on challenging cases. It further transfers zero-shot to real drone deployment, suggesting that the proposed WorldVLN offers a promising route for spatial action tasks. Demos and code are available at https://embodiedcity.github.io/WorldVLN/.
arXiv:2605.15962v1 Announce Type: new
Abstract: Website Fingerprinting (WFP) has traditionally focused on inferring which website a user visits from encrypted traffic metadata such as packet sizes and timing. In this paper, we identify and quantify a new privacy risk in modern web settings: an adversary can infer a user's persona using only packet-length and inter-arrival-time sequences. To study this risk at scale, we build an LLM-driven multi-agent browsing framework that enforces controllable persona constraints while a computer-use agent interacts with real websites and collects corresponding encrypted traffic traces. We formalize persona fingerprinting under both closed-set and open-world settings and further evaluate whether persona information is already embedded in representations learned by existing WFP models and can be amplified at low cost. Across 10 modern websites and 15 personas (plus an open-world class), persona inference achieves about 84% accuracy on mixed-site traffic; moreover, a lightweight multi-task objective can boost persona accuracy to around 80% while retaining strong site classification performance (about 93% baseline). Our results show that, on modern websites, encrypted traffic metadata can leak not only which site a user visits, but also how they browse and who is browsing.
arXiv:2605.16041v1 Announce Type: cross
Abstract: Machine learning systems increasingly make life-changing decisions about individuals, such as loan approvals, hiring, and cheating detection, raising a pressing question: how can individuals respond to negative decisions made by these opaque systems? While explainable artificial intelligence (XAI) has largely focused on algorithmic recourse -- helping individuals change their features to obtain a desired outcome -- the parallel problem of algorithmic contestability -- helping individuals review and correct erroneous algorithmic decisions -- has received far less attention, despite its central ethical and legal importance. We trace this neglect to the absence of clear formal definitions and a systematic operationalization of contestability as an algorithmic problem. To address it, we propose an operational definition of contestability as a natural complement to recourse: contestability starts from the presumption that a decision may be incorrect and focuses on identifying evidence to challenge and potentially overturn it, whereas recourse assumes the decision is valid and instead provides pathways for changing it. We show that standard XAI explanations, such as counterfactuals, LIME, or Anchors, even when combined with human intuitions about decision continuity or monotonicity, reveal only errors in the neighborhood of the individual, but provide insufficient grounds for overturning the decision at hand. Going thus beyond traditional XAI, we identify three types of evidence warranting reversal according to the decision maker's own ethical standards: predictive multiplicity, incorrect feature values, and neglected overruling evidence. We argue that these render decisions normatively indefensible and thus successfully contestable. Finally, we analyze how existing EU legislation connects to our framework and argue that individuals already hold some legal rights to these forms of evidence.
arXiv:2605.15959v1 Announce Type: new
Abstract: Physics-informed neural networks (PINNs) are powerful surrogates for differential equations but are notoriously difficult to train due to spectral bias, stiffness, and poor accuracy on high-frequency or multiscale solutions. Adversarial training based on generative adversarial networks (GANs) has recently gained surprisingly strong empirical results in improving training, but the underlying mechanisms remain elusive. To this end, we propose a new analysis framework for adversarially trained PINNs, based on the key observation of how the discriminator in GANs can influence the training dynamics of PINNs. The framework first provides a much needed theoretical grounding to why and when adversarial training is effective in PINNs, then presents a unified analysis of GANs variants in such training, and finally leads to a new, practical, efficient training algorithm for PINNs. Empirical results demonstrate that our method can significantly reduce the pathology of PINNs training, thereby providing better models with superior performances, often several magnitudes more accurate than alternative methods.
arXiv:2605.14324v2 Announce Type: replace-cross
Abstract: We analyze convergence rates of norm-minimization-based outer approximation algorithms for convex vector optimization when the scalarization uses an $\ell_p$ norm with $p \in (1,\infty)$. While the Euclidean case ($p=2$) achieves the optimal rate $O(k^{2/(1-q)})$, the behavior under general $\ell_p$ norms has remained open. A direct approach via the modulus of smoothness yields only the weaker exponent $\min(p,2)$, which degrades for $1 < p < 2$. We prove that the Hausdorff approximation error satisfies $\delta_H(P_k, A) = O(k^{2/(1-q)})$ for \emph{every} $p \in (1,\infty)$, where $q$ is the number of objectives and $k$ is the iteration count. The proof introduces a Euclidean intermediary technique that exploits the ambient inner product structure of $\R^q$ to obtain a quadratic bound on the hyperplane distance, bypassing the $\ell_p$ smoothness limitation; norm equivalence then converts this to any $\ell_p$ metric at the cost of only a dimension-dependent constant, not a loss of exponent. Numerical experiments confirm the $p$-independent rate predicted by the theory.
arXiv:2605.15958v1 Announce Type: new
Abstract: Energy system models are increasingly dependent on representative climate input. Yet, a fundamental mismatch persists between the hundreds of simulated years often used in climate science and the handful of years that computationally demanding power system models can process. Current practice, including ENTSO-E's European Resource Adequacy Assessment, relies on climate year selections that have not been validated against explicit representativeness criteria. This risks biased investment decisions and blind spots for plausible weather conditions. This study proposes simulated annealing as an optimisation method for selecting representative subsets of complete climate years from large climate ensembles. Representativeness is quantified using the seasonal sliced Wasserstein distance, a metric from optimal transport theory that captures representativeness on marginal distributions, inter-variable correlations, and seasonal structure simultaneously. We evaluate simulated annealing against the alternative methods random search, filtered random search, and K-Medoids clustering across three test cases spanning the Netherlands and Europe, using 180 climate years from the Pan-European Climate Database as a reference. Simulated annealing consistently produces the most representative subsets and outperforms all compared methods. Simulated annealing achieves an effective sample size four to five times the actual subset size. The resulting subsets are roughly 2.5--3.5 times more representative than current ENTSO-E practice. The method is application-agnostic and its output can serve as a validated climate data input to any subsequent (energy) impact study.
arXiv:2504.00663v2 Announce Type: replace
Abstract: Existing hardware-aware NAS (HW-NAS) methods typically assume access to precise information circa the target device, either via analytical approximations of the post-compilation latency model, or through learned latency predictors. Such approximate approaches risk introducing estimation errors that may prove detrimental in risk-sensitive applications. In this work, we propose a two-stage HW-NAS framework, in which we first learn an architecture controller on a distribution of synthetic devices, and then directly deploy the controller on a target device. At test-time, our network controller deploys directly to the target device without relying on any pre-collected information, and only exploits direct interactions. In particular, the pre-training phase on synthetic devices enables the controller to design an architecture for the target device by interacting with it through a small number of high-fidelity latency measurements. To guarantee accessibility of our method, we only train our controller with training-free accuracy proxies, allowing us to scale the meta-training phase without incurring the overhead of full network training. We benchmark on HW-NATS-Bench, demonstrating that our method generalizes to unseen devices and searches for latency-efficient architectures by in-context adaptation using only a few real-world latency evaluations at test-time.
arXiv:2605.15957v1 Announce Type: new
Abstract: Vector search (VS) is now available in most database engines. However, while vector search is a common feature in AI/ML/LLMs where the dominant computing platforms are GPUs, existing database engines operate on CPUs even when implementing vector search. This raises the question of whether integrating vector processing on GPUs as part of the engine would be a better design. In this paper, we explore this question in detail. First, we extend the TPC-H benchmark with vector data (from text and images) and propose a number of representative SQL+VS queries. Second, we develop a modular execution engine that can run SQL+VS queries across CPU and GPU. Third, we perform extensive experiments on a number of deployments: running the SQL+VS queries across CPU and/or GPU, with data residing in CPU or GPU memory, with existing indices and novel, optimized versions, as well as across different GPUs and interconnects (PCIe, NVLink). The results provide actionable and counter-intuitive insights on how to run such queries over CPUs and GPUs. For instance, the relational components benefit much more from running on the GPU than the vector search part. In addition, when the vector search involves moving data and indexes, using the GPU is not the best option, even with fast interconnects. Thus, we develop an alternative organization of vector index and embeddings that reduces the size of the index, making GPU-based vector search more competitive. With these improvements, the final result is that both the relational and vector search components are faster on the GPU, particularly on fast interconnects, in contrast with the architecture used in existing engines.
arXiv:2605.16070v1 Announce Type: cross
Abstract: Antibody-based therapeutics-including antibody-drug conjugates (ADCs), bispecific antibodies, and novel formats-are reshaping oncology, yet key determinants of efficacy, safety, and manufacturability frequently emerge after conjugation and formulation. We argue that computational biophysics provides an underexploited framework to address this gap by connecting molecular interactions to biological outcomes. We highlight how molecular dynamics, coarse-grained simulations, and free energy calculations reveal how conjugation site, linker chemistry, and drug-antibody ratio reshape conformational landscapes. We emphasize structural coupling between antibody, linker, and payload, with implications for antigen binding, internalization, and developability. We propose that integrating physics-based modeling into development pipelines-alongside experimental validation-can reduce empirical iteration and de-risk translation. As force fields, and hybrid physics-machine-learning methods improve, this field is poised to become a central driver of next-generation ADC design.
arXiv:2605.09403v2 Announce Type: replace
Abstract: Architectural choices inside the Transformer feedforward network (FFN) block do not merely affect the block itself; they reshape the computations learned by the rest of the model. We study this effect in one-layer Transformers trained on digit addition with carry, modular arithmetic, and histogram counting. Comparing dense FFNs, gated linear units (GLUs), mixture-of-experts (MoE), and MoE-GLUs, we find that sparse MoE routing can shift computation from FFN to attention, with the strongest ablation-visible effect on carry-based addition. We decompose this redistribution into reduced per-token FFN capacity and sparse partitioning across experts. Critically, frozen random routing nearly matches learned routing, suggesting that redistribution is driven largely by architectural sparsity rather than router-learned specialization. As a secondary finding, GLU-style multiplicative gating rotates task-relevant Fourier structure out of the per-neuron basis and into distributed subspaces, making neuron-level interpretability less informative while preserving structured computation. We validate these conclusions with random-routing, narrow-FFN, and top-2 MoE controls, plus parameter-matching, activation-function, and width-scaling analyses. Together, these results show that local FFN design choices can have nonlocal consequences for Transformer computation.
arXiv:2605.16101v1 Announce Type: cross
Abstract: Measurement is an important scientific activity. In most of science, including classical physics, is may be understood as a way of finding out about the physical world and representing the results numerically. No-go theorems show that measurement of quantum observables is not like that: the recorded outcome is typically created rather than revealed in a quantum measurement, in which case there is no objective fact about the observable's prior value. Other no-go theorems show that unitary quantum theory can generally neither explain nor even represent a unique recorded outcome, thereby threatening that outcome's objectivity. Methodological norms inherent in quantum physical practice nevertheless institute the objectivity, not only of unique recorded outcomes of quantum measurements, but also of non-quantum features of the world that physicists and other scientists take their models to represent.
arXiv:2605.15498v1 Announce Type: new
Abstract: From a new perspective, this paper rederives Lagrange's equations. By applying the chain rule of differentiation, the intrinsic relationship between the momentum theorem and the kinetic energy theorem is first established. Subsequently, expressing the differential form of energy conservation in an arbitrary coordinate system and performing suitable differential operations yields Lagrange's equations. Generalized forces and generalized displacements are shown to be component representations of forces and displacements in a chosen coordinate system. Consequently, the essence of Lagrange's equations is identified as the transformation of the kinetic energy theorem into the momentum theorem via the chain rule for composite functions, thereby revealing how energy conservation constructs momentum conservation.
arXiv:2605.15901v1 Announce Type: new
Abstract: Diffusion geometry is a manifold learning framework that uses random walks defined by Markov transition matrices to characterize the geometry of a dataset at multiple scales. We use diffusion geometry for neural representations, incorporating tools from multi-view learning into this field for the first time. Our key technical observation is that a broad class of similarity measures based on representational similarity matrices (RSMs) admits a closed-form equivalent formulation in terms of row-stochastic Markov matrices, opening the door to manipulations from diffusion geometry. As a first application, we develop multi-scale variants of Centered Kernel Alignment and Distance Correlation, which utilise the $t^{th}$ power of the underlying transition matrix to probe the data geometry at adjustable diffusion scales. Going further, we introduce variants of these measures which fuse the Markov matrices of several layers via alternating diffusion into a single operator that captures the network's joint sample geometry, allowing similarity to be computed across multiple layers and shifting the comparison from layer-to-layer to network-to-network. We perform extensive numerical experiments, evaluating our measures on the Representational Similarity (ReSi) benchmark comprising 14 architectures trained on 7 datasets across three different domains. Our methods achieve SoTA results in accuracy and output correlation for both language and vision tasks across different models. We furthermore show SoTA performance on an additional benchmark evaluating on out-of-distribution data.
arXiv:2602.22918v3 Announce Type: replace
Abstract: Vision-language models (VLMs) can read text from images, but where does this optical character recognition (OCR) information enter the language processing stream? We investigate the OCR routing mechanism across three architecture families (Qwen3-VL, Phi-4, InternVL3.5) using causal interventions. By computing activation differences between original images and text-inpainted versions, we identify architecture-specific OCR bottlenecks whose dominant location depends on the vision-language integration strategy: DeepStack models (Qwen) show peak sensitivity at mid-depth (about 50%) for scene text, while single-stage projection models (Phi-4, InternVL) peak at early layers (6-25%), though the exact layer of maximum effect varies across datasets. The OCR signal is remarkably low-dimensional: PC1 captures up to 72.9% of variance. Crucially, principal component analysis (PCA) directions learned on one dataset transfer to others, demonstrating shared text-processing pathways. Surprisingly, in models with modular OCR circuits (notably Qwen3-VL-4B), OCR removal can improve counting performance (up to +6.9 percentage points), suggesting OCR interferes with other visual processing in sufficiently modular architectures.
arXiv:2605.15900v1 Announce Type: new
Abstract: The performance of tip-enhanced optical microscopy is often limited by inefficient coupling of the excitation field to the plasmonic tip apex, as well as by thermal drift and optical aberrations. Here, we demonstrate that adaptive wavefront shaping based on Zernike mode provides a practical approach to achieving robust near-field optimisation at the tip apex. Using a sequential feedback algorithm, initially using the near-field signal, we narrow the illumination point-spread function and suppress sidelobes. This demonstrates that Zernike-mode control can be used for both aberration correction and field engineering. In tip-enhanced Raman measurements of a Janus MoSSe monolayer, conventional near-field optimisation increases the signal intensity by around 1.4 fold. A second optimisation step based directly on the Raman-band intensity yields a further 5 to 15 fold enhancement, depending on the specific tips used. These results establish a systematic, optics-based strategy for optimising tip fields, providing a transferable framework for improving tip-enhanced and related near-field spectroscopies.