arXiv:2606.11885v2 Announce Type: replace-cross
Abstract: We characterize the quasi-stationary distribution (QSD) of the bond directed-percolation line of the Domany--Kinzel automaton using a matrix-product-state representation of the probability distribution, obtained by projecting out the absorbing state and iterating the transfer matrix. Unlike moment- or sampling-based methods, this yields the full conditional distribution and direct access to information-theoretic diagnostics. The spatial structure of the QSD changes sharply across the transition: the active phase is bulk-like with finite density, whereas in the inactive phase the surviving activity collapses into a single flock -- the smallest interval containing all active sites -- occupying a vanishing fraction of the chain. Throughout the inactive phase the bipartite mutual information of the QSD equals the entropy of a single binary choice -- whether the flock lies to the left or right of the cut -- so the surviving clusters together encode just one bit of positional information, corresponding to a single effective cluster.
Science Journals
arXiv:2508.10051v2 Announce Type: replace-cross
Abstract: Infrastructure shapes societies and scientific discovery. Traditional scientific infrastructure, often static and fragmented, leads to issues like data silos, lack of interoperability and reproducibility, and unsustainable short-lived solutions. Our current technical inability and social reticence to connect and coordinate scientific research and engineering lead to inefficiencies and impede progress. With AI technologies changing how we interact with the world around us, there is an opportunity to transform scientific processes. Neuroscience's exponential growth of multimodal and multiscale data, together with its urgent clinical relevance, demands an adaptive infrastructure that can expose computable states, coordinate across systems, and improve through use. Using neuroscience as a stress test, this perspective argues for a paradigm shift: infrastructure must evolve into a dynamic, AI-aligned ecosystem to accelerate science. Building on several existing principles for data, collective benefit, and digital repositories, I recommend operational guidelines for implementing these principles to create this dynamic ecosystem, aiming to foster a decentralized, self-learning, and self-correcting system where humans and AI can collaborate seamlessly. Addressing the chronic underfunding of scientific infrastructure, acknowledging diverse contributions beyond publications, and coordinating global efforts are critical for this transformation. A coordinating role, even more than analysis, is where AI becomes transformative rather than merely assistive. By prioritizing an intelligent infrastructure as a central scientific instrument for knowledge generation, we can overcome current limitations, accelerate discovery, ensure reproducibility and ethical practices, and ultimately translate neuroscientific understanding into tangible societal benefits, setting a blueprint for other scientific domains.
arXiv:2509.00078v2 Announce Type: replace-cross
Abstract: The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question. While end-to-end approaches promise theoretical advantages, cascaded systems (CSs) continue to outperform them in language understanding tasks, despite being constrained by sequential processing latency. In this work, we introduce ChipChat, a novel low-latency CS that overcomes traditional bottlenecks through architectural innovations and streaming optimizations. Our system integrates streaming (a) conversational speech recognition with mixture-of-experts, (b) state-action augmented LLM, (c) text-to-speech synthesis, (d) neural vocoder, and (e) speaker modeling. Implemented using MLX, ChipChat achieves sub-second response latency on a Mac Studio without dedicated GPUs, while preserving user privacy through complete on-device processing. Our work shows that strategically redesigned CSs can overcome their historical latency limitations, offering a promising path forward for practical voice-based AI agents.
arXiv:2606.21632v2 Announce Type: replace-cross
Abstract: Molecular dynamics simulation of plasma-surface interactions requires an interatomic potential that is simultaneously accurate, computationally efficient, and able to describe many elements and bonding types in reactive systems. In principle, a foundation model for machine-learned interatomic potential (MLIP) can meet these demands. We explore the use of the Universal Models for Atoms (UMA) model, developed by Meta FAIR, for the interactions of oxygen plasma species on a multilayer of WS$_2$, a promising 2D material. Starting from the pretrained uma-s-1p1 model under the Open Catalyst 2020 (OC20) task, we apply an iterative fine-tuning loop with maximally diverse configuration sampling using Smooth Overlap of Atomic Positions (SOAP) and Farthest Point Sampling (FPS); DFT labeling at the PBE+D3+$U$+spin level; and fine-tuning on energy, force, and stress labels. Even in the absence of fine-tuning, the pretrained model reproduces the production-scale observables of interest, namely, chemisorbed S and O coverage under 15eV O$^+$ and O$_2^+$ bombardment. These results were obtained without spin polarization and Hubbard $U$ correction. Nonetheless, fine-tuning reduces the energy and force mean absolute error (MAE) to $4.5\times10^{-3}$eV/atom and $0.076$eV/angstrom, respectively.
arXiv:2606.23722v2 Announce Type: replace-cross
Abstract: In recent times there has been growing interest in Raman optical activity (ROA) for its label free detection of absolute configuration, conformation, and stereochemical structure in chiral biosamples and drug molecules. Since ROA signals are generally small, techniques such as stimulation by a probe beam can be used to enhance the signal strength. However, with a classical probe, the measurement precision is still fundamentally limited by its shot noise. To solve this problem we propose the use of two-mode squeezed vacuum and show that it can achieve sub-shot noise limited measurement sensitivity. Using quantum estimation theory, we derived the quantum Fisher information and the quantum Cram\'er-Rao bound (QCRB) for stimulated ROA measurement to quantify the precision enhancement. This improvement comes from photon-number correlations which suppress the intensity fluctuation common to both modes. We further show that balanced detection of the output intensity difference is a practical measurement scheme that approaches the QCRB and becomes optimal in the small-chirality limit. This opens a promising path toward more sensitive Raman chiroptical spectroscopy of weak and photosensitive samples.
arXiv:2512.16636v2 Announce Type: replace
Abstract: Latent diffusion models (LDMs) achieve state-of-the-art image synthesis, yet their reconstruction-style denoising objective provides only indirect semantic supervision: high-level semantics emerge slowly, requiring longer training and limiting sample quality. Recent works inject semantics from Vision Foundation Models (VFMs) either externally via representation alignment or internally by jointly modeling only a narrow slice of VFM features inside the diffusion process, under-utilizing the rich, nonlinear, multi-layer spatial semantics available. We introduce REGLUE (Representation Entanglement with Global-Local Unified Encoding), a unified latent diffusion framework that jointly models (i) VAE image latents, (ii) compact local (patch-level) VFM semantics, and (iii) a global (image-level) [CLS] token within a single SiT backbone. A lightweight convolutional semantic compressor nonlinearly aggregates multi-layer VFM features into a low-dimensional, spatially structured representation, which is entangled with the VAE latents in the diffusion process. An external alignment loss further regularizes internal representations toward frozen VFM targets. On ImageNet 256x256, REGLUE consistently improves FID and accelerates convergence over SiT-B/2 and SiT-XL/2 baselines, as well as over REPA, ReDi, and REG. Extensive experiments show that (a) spatial VFM semantics are crucial, (b) non-linear compression is key to unlocking their full benefit, and (c) global tokens and external alignment act as complementary, lightweight enhancements within our global-local-latent joint modeling framework. The code is available at https://github.com/giorgospets/reglue .
arXiv:2603.16750v2 Announce Type: replace
Abstract: We present thermopneumatic pixels (TPPs) -- low-profile pixels and arrays that generate dynamic tactile feedback. These devices are thin, fast, reconfigurable, and output localized transient displacements at each pixel. Their parsimonious design -- a layered architecture without internal moving parts -- and low-voltage ($\lesssim$10 V) operation may facilitate practical integration in a wide variety of interfaces. Each TPP converts brief electrical pulses into transient air pressure increases in an internal cavity, yielding out-of-plane forces and displacements for tactile feedback. We demonstrate TPPs that output displacements of 1 mm and forces exceeding 1 N, with millisecond response times, in packages that are less than 3 mm thick. Force and displacement increase with pixel surface area, facilitating tailorability. The pixels can also generate oscillating feedback at pulse rates up to 300 Hz range. We report designs for compact arrays of pixels at 4 mm spacing, and simple pulse driving architectures using miniature transistors driven by microcontrollers. We characterize the mechanical, dynamic, and thermal response of TPPs, and their robustness and consistency over tens of thousands of cycles. We report perceptual experiments on spatial localization and intensity as a function of driving power. Together, these results establish thermopneumatic pixels as a compact, adaptable tactile technology that blends performance and practicality.
arXiv:2604.22328v2 Announce Type: replace
Abstract: Driven by the transition towards a climate-neutral energy system, accurate energy time series forecasting is critical for planning and operations. Yet, it remains a dataset-specific task, requiring comprehensive training data, limiting scalability, and resulting in high model development and maintenance effort. Recently, foundation models aiming to learn generalizable patterns via extensive pretraining have shown strong performance in multiple prediction tasks. Despite their success and strong potential in energy forecasting, a systematic, use-case-differentiated evaluation is still missing. We address this gap by presenting the Foundation Models in Energy Time Series Forecasting (FETS) benchmark. We (1) provide a structured overview of energy forecasting use cases along three main dimensions, i.e., stakeholders, attributes, and data categories, (2) curate 54 datasets across 9 data categories, guided by typical stakeholder interests, and (3) benchmark foundation models against task-specific machine learning across different forecasting settings. In our benchmark study, covariate-informed zero-shot foundation models perform best in aggregate, with Chronos-2 attaining the lowest overall median NRMSE (0.472), closely followed by TiRex-2 (0.474). Both perform better than XGBoost (0.611) and random forest (0.696), although they were trained task-specifically on the full historic target data. Further analysis reveals a strong correlation between predictive performance and spectral entropy. Performance saturates beyond a certain context length and improves with aggregation level, e.g., for national load, district heating, and power grid data. Overall, with the lowest median error, limited data requirements, and low inference and hardware demands, foundation models reduce development and maintenance effort, emerging as scalable and generalizable energy forecasting solutions.
arXiv:2511.08451v3 Announce Type: replace-cross
Abstract: Proximal methods such as the Alternating Direction Method of Multipliers (ADMM) are effective at solving constrained quadratic programs (QPs). To tackle infeasible QPs, slack variables are often introduced to ensure feasibility, which changes the structure of the problem, increases its size, and slows down numerical resolution. In this letter, we propose a simple ADMM scheme to tackle QPs with slack variables without increasing the size of the original problem. The only modification is a slightly different projection in the z-update, while the rest of the algorithm remains standard. We prove that the method is equivalent to applying ADMM to the QP with additional slack variables, even though slack variables are not added. Numerical experiments show speedups of the approach.
arXiv:2607.12420v3 Announce Type: replace-cross
Abstract: The quantum Hall effect establishes that topology can fix a material response to integer multiples of fundamental constants when an energy gap isolates the relevant symmetry-protected electronic states. Whether such universal quantization can also emerge in gapless matter, where topological bands coexist with a continuum of metallic excitations, has remained a fundamental question in the field of quantum materials. Chiral topological semimetals provide a unique setting in which to explore this principle; when optical transitions are confined to a single chiral node, the resulting circular photogalvanic effect is predicted to be quantized by the topological charge of the node. In real materials, however, this nonlinear optical phenomenon has remained experimentally elusive, obscured by trivial band transitions, insufficient energy separation between node pairs, and their relative positions with respect to the Fermi level. Here we observe a quantized circular photogalvanic effect in the chiral topological semimetal Rh0.95Ni0.05Si. Band engineering via Ni substitution opens a photon-energy window dominated by interband optical transitions at the {\Gamma}-point multifold node. This allows circularly polarized near- to mid-infrared pulses to drive a helicity-odd terahertz response that manifests three hallmarks of quantization: a sharp onset, a photon-energy-independent plateau governed by the magnitude of the monopole charge, and an abrupt long-wavelength cutoff imposed by Pauli blocking. Our work thus establishes an all-optical analogue of the quantum Hall effect and a new paradigm to realize topological quantization in gapless matter.
arXiv:2601.13359v3 Announce Type: replace
Abstract: Prefill attacks are an effective and low-cost jailbreaking method, as they directly insert an acceptance sequence (e.g., "Sure, here is...") at the start of an LLM's output and lead the model to continue the response. We make two contributions to this prior work. First, we show that an unsophisticated adversary can improve the well-known prefill attacks by ensembling a small number of prefill variants. Running three easy-to-generate prefills yields a combined attack success rate (ASR) of 22%, 90%, and 99% on Gemma-7B, Llama-3.1-8B, and Qwen3-8B respectively, an up to 38 percentage point improvement over the standard "Sure, here's..." prefill and up to 82 percentage points over our reproduction of GCG (Zou et al., 2023). Second, we introduce "sockpuppetting", a hybrid attack that optimizes an adversarial suffix placed inside the "assistant" message block of the chat template, rather than within the user prompt. The rolling variant of this attack, RollingSockpuppetGCG, increases prompt-agnostic ASR by up to 64 percentage points over our universal GCG baseline on Llama-3.1-8B. An ablation indicates that part of this gain stems from the choice of acceptance sequence rather than suffix placement alone (Appendix F). Both findings highlight the need for defences against output-prefix injection in open-weight models. Code: https://gitlab.com/asendotsinski/sockpuppetting
arXiv:2603.14355v2 Announce Type: replace
Abstract: Safety tuning through supervised fine-tuning and reinforcement learning from human feedback has substantially improved the robustness of large language models. However, it typically suppresses rather eliminates unsafe behaviors, leaving rare but critical failures hidden in the long tail of the output distribution. While most red-teaming work emphasizes adversarial prompt search, we show that these hidden risks can be systematically exposed through diverse response generation. Specifically, we show that, for a fixed safety-critical prompt, increasing the number and diversity of sampled responses monotonically raises the jailbreak success rate. To efficiently uncover these failures, we propose Progressive Diverse Population Sampling (PDPS). This approach replaces naive, large-scale IID sampling with a multi-stage expansion-and-selection strategy that generates a compact, semantically diverse set of responses at a substantially lower computational cost. Across multiple jailbreak benchmarks and open-source LLMs, PDPS achieves attack success rates comparable to large-scale IID sampling while using only 8%-29% of the computational cost, and outperforms IID sampling and Diverse Beam Search by 26%-40% under limited-response budgets, while uncovering a broader and more semantically diverse range of failure modes. Critically, this diversity translates directly into more effective safety hardening: when integrated into an RLHF-based safety-tuning pipeline, PDPS-generated unsafe responses yield 33% and 41% greater reductions in ASR than those generated by IID sampling and Diverse Beam Search, respectively. Finally, we show that while input-space prompt optimization methods fall short of output-space exploration when used in isolation, combining input-space perturbation with diversity-driven output-space exploration covers a wider range of failure modes more efficiently than either paradigm alone.
arXiv:2510.03949v4 Announce Type: replace-cross
Abstract: Simulating the kinetic Langevin dynamics is a popular approach for sampling from distributions, where only their unnormalized densities are available. Various discretizations of the kinetic Langevin dynamics have been considered, where the resulting algorithm is collectively referred to as the kinetic Langevin Monte Carlo (KLMC) or underdamped Langevin Monte Carlo. Specifically, the stochastic exponential Euler discretization, or exponential integrator for short, has previously been studied under strongly log-concave and log-Lipschitz smooth potentials via the synchronous Wasserstein coupling strategy. Existing analyses, however, impose restrictions on the parameters that do not explain the behavior of KLMC under various choices of parameters. In particular, all known results fail to hold in the overdamped regime, suggesting that the exponential integrator degenerates in the overdamped limit. In this work, we revisit the synchronous Wasserstein coupling analysis of KLMC with the exponential integrator. Our refined analysis results in Wasserstein contractions and bounds on the asymptotic bias that hold under weaker restrictions on the parameters, which assert that the exponential integrator is capable of stably simulating the kinetic Langevin dynamics in the overdamped regime, as long as proper time acceleration is applied.
arXiv:2511.03826v4 Announce Type: replace-cross
Abstract: Accurate and efficient registration of whole slide images (WSIs) is essential for high-resolution, nuclei-level analysis in multi-stained tissue slides. We propose a novel coarse-to-fine framework CORE for accurate nuclei-level registration across diverse multimodal whole-slide image (WSI) datasets. The coarse registration stage leverages prompt-based tissue mask extraction to effectively filter out artefacts and non-tissue regions, followed by global alignment using tissue morphology and accelerated dense feature matching with a pre-trained feature extractor. From the coarsely aligned slides, nuclei centroids are detected and subjected to fine-grained rigid registration using a custom, shape-aware point-set registration model. Finally, non-rigid alignment at the cellular level is achieved by estimating a non-linear displacement field using Coherent Point Drift (CPD). Our approach benefits from automatically generated nuclei that enhance the accuracy of deformable registration and ensure precise nuclei-level correspondence across modalities. The proposed model is evaluated on three publicly available WSI registration datasets, and two private datasets. We show that CORE outperforms current state-of-the-art methods in terms of generalisability, precision, and robustness in bright-field and immunofluorescence microscopy WSIs
arXiv:2607.10645v2 Announce Type: replace
Abstract: An LLM agent's public behaviour reveals little about its social reasoning: an agent that votes correctly may be guessing, and an agent that lies well leaves no trace of what it actually believes. We present MafiaScope, an open testbed that turns the social deduction game Mafia into a measurement instrument for machine Theory of Mind. It distinguishes whether an agent lost because it misread the game or because it failed to act on a correct assessment, a distinction that is invisible from outcomes and dialogue transcripts alone. After every public utterance, each agent privately answers structured probe questions whose responses never re-enter the game and are scored against the ground truth known to the engine. An interactive visualizer replays games from the perspective of an individual agent's beliefs, displays timeline-aligned accuracy and calibration, and supports counterfactual replay from any recorded step. In a case study across two model families comprising tens of thousands of parsed probe responses, we find that stated confidence is poorly calibrated, agents overestimate how often they are suspected by a factor of 1.5, and single-vote counterfactual replays rarely change game outcomes: outcome flips occur primarily when the agent had already formed a correct belief state, whereas decisions made under an incorrect model of the world remain largely unchanged under resampling. The engine, visualizer, recorded games, and counterfactual replay corpus are released under an open-source licence. Code: https://github.com/karpovilia/mafiascope. Live demo: https://karpovilia.github.io/mafiascope/. Screencast: https://vimeo.com/1208920221.
arXiv:2607.15097v2 Announce Type: replace
Abstract: All-in-one image restoration aims to recover clean images degraded by multiple corruption types using a single unified model. Existing methods typically rely on image-level prompts or shared guidance to handle diverse degradations. However, such a paradigm becomes inadequate when degradations are spatially heterogeneous or even coexist in mixed forms within a single image. Yet spatially adaptive guidance alone is not sufficient, since accurate restoration also requires each spatial query to reliably aggregate complementary information from local neighborhoods and global contexts. To this end, we propose QuReC, a unified framework for all-in-one image restoration. QuReC consists of a Degradation-Guided Query Reconstruction Module (DQRM) and a Local-Global Response Calibration Module (LGRCM). Specifically, DQRM matches each spatial query against a degradation prototype space to reconstruct a query-specific degradation-aware representation, thereby providing fine-grained spatially adaptive restoration guidance. To further stabilize this query-wise matching process, we introduce a weakly supervised prototype matching learning strategy to improve optimization stability and degradation semantic consistency. Meanwhile, LGRCM performs local-global dual-branch aggregation and calibrates the aggregated responses with learnable priors, improving the reliability of feature aggregation and the coordination between local detail modeling and global context modeling. Extensive experiments demonstrate that QuReC achieves superior performance on multiple all-in-one image restoration benchmarks. The code is released at https://github.com/zhoushen1/QuReC.
arXiv:2607.16051v2 Announce Type: replace
Abstract: We present the Loopie series, consisting of two Mixture-of-Experts (MoE) models: a 20B-parameter model with 2B active parameters and a 6B-parameter model with 0.6B active parameters. Looped Transformers have long faced a challenge: given an N times increase in pre-training compute, increasing the parameter count by a factor of N usually outperforms looping a model N times. Loopie addresses this challenge. Extensive ablation studies, including comparisons with a vanilla 30B-A3B model, show that Loopie substantially outperforms vanilla Transformer baselines trained with the same compute budget. With a novel post-training method, Loopie develops strong reasoning abilities and achieves frontier-level reasoning performance.
arXiv:2607.15400v2 Announce Type: replace
Abstract: Falls among older adults are a major safety challenge, but continuous monitoring is difficult to sustain. Video captures fall-related posture and motion, yet deployment is limited by privacy, computation, and bandwidth. Supervised pose estimation is anatomically interpretable but vulnerable to occlusion and partial body visibility. We propose a privacy-preserving framework that replaces RGB transmission with compact motion representations based on unsupervised keypoints and predictive temporal modeling. Local processing performs segmentation and keypoint extraction; variational recurrent prediction and sequence classification then detect falls from observed and forecasted motion. We evaluate the framework on the UR Fall Detection and Human Fall datasets using random, subject-disjoint, and occlusion-based splits. Under random splits, neither representation consistently dominates, suggesting that standard protocols may hide meaningful differences. Under subject-disjoint evaluation, supervised keypoints show a statistically significant advantage, but performance varies by subject: they perform better when anatomical landmarks are visible, whereas unsupervised keypoints are more robust to occlusion and partial visibility, though they produce more false positives for complex activities. Under occlusion-based evaluation, supervised keypoints miss nearly half of all falls, while unsupervised keypoints retain strong sensitivity and substantially outperform them. Their anatomical independence allows spatial anchors to adapt to visible body structure rather than fail on absent landmarks. The gap widens under bandwidth constraints, where supervised localization errors compound through the temporal model. These findings show that representation choice should reflect expected visual conditions and that unsupervised keypoints offer an advantage when body visibility is compromised.
From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems
arXiv:2607.15459v2 Announce Type: replace
Abstract: A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that reproduces its behaviour and that a person can read, a logic engine can run, and an optimizer can edit. We present a three-stage post-hoc transformation that extracts a frozen proximal policy optimization teacher, induces an ordered rule list from its decisions in the manner of classical relational learning, and emits the result as a Prolog program whose every decision is executed by an off-the-shelf logic engine; a subsequent expansion stage edits the rule base and accepts an edit only when policy evaluation certifies a return increase. We prove four guarantees. A return-loss bound makes the distilled program a machine-checkable certificate in a finite Markov decision process, and the expansion loop improves monotonically and terminates. For the continuous-observation setting we answer whether the conversion is possible at all: the propositional threshold instantiation converts the network to arbitrary fidelity as the resolution B grows, with disagreement O(1/B) and a return gap that closes at the same rate, and a matching lower bound shows the cost is exponential in the observation dimension for an oblique decision boundary. Empirically, on a two-room key-and-door task with 16,944 reachable states the expanded Prolog program attains exact optimal return in every seed and, in a budget-capped regime, exceeds the stochastic teacher on exact return in ten of ten seeds. On three continuous-control tasks the emitted program substitutes the network, matching the neural teacher within noise on Acrobot with eleven clauses and recovering about 97% of its return on CartPole, while on the finer-control LunarLander it recovers only partially, exactly the ceiling the exponential lower bound predicts.
arXiv:2511.11771v2 Announce Type: replace-cross
Abstract: We study the localization problem in quantum stochastic mechanics. We start from the Edwards model for a particle in a bath of scattering centers and prove static localization of the ground state wavefunction of the particle in a one dimensional square well coupled to Dirac delta like scattering centers in arbitrary but fixed positions. We see how the localization increases for increasing coupling $g$ and increasing number of scattering centers at constant density. Then we choose the scattering centers positions as pseudo random numbers with a uniform probability distribution and observe an increase in the localization of the average of the ground state over the many positions realizations. We discuss how this averaging procedure is consistent with a picture of a particle in a Bose-Einstein condensate of of non interacting boson scattering centers interacting with the particle with Dirac delta functions pair potential. We then study the dynamics of the ground state wave function. We conclude with a discussion of the affine quantization version of the Lax model which reduces to a system of contiguous square wells with walls in arbitrary positions independently of the coupling constant $g$.
arXiv:2607.15523v2 Announce Type: replace
Abstract: Scene-centric visualization systems expose semantic components, such as marks, encodings, layouts, and axes, as first-class objects that can be directly manipulated. Existing interaction abstractions, however, are largely based on event streams, signals, and data selections rather than semantic scene components. This mismatch makes interactions involving scene components less natural to specify and limits the expressive power of scene-centric visualization systems. We present Interactive Mascot, a scene-centric interaction grammar for data visualizations. Interactive Mascot extends scene-centric representations for static visualizations by modeling interactive behavior as information flow among four interaction components (trigger, responder, evaluator, and updater) and two forms of context (event context and state context). To realize these semantics, we introduce a dependency-graph execution model that systematically transforms interaction specifications into executable dependency graphs using reusable graph patterns associated with semantic visualization components. We implement Interactive Mascot in the JavaScript library Mascot$.$js and evaluate its expressiveness, performance, and usability. Interactive Mascot naturally covers Vega-Lite's interaction design space while additionally supporting stateful interactions, direct manipulation of scene components, and freeform selection. It achieves runtime performance comparable to Vega-Lite, and a qualitative user study shows that the grammar is learnable and usable for interaction authoring.
arXiv:2512.06429v2 Announce Type: replace-cross
Abstract: We propose a qubit-oscillator platform based on the motional states of two interacting atoms in an optical tweezer. By stroboscopically modulating an engineered trap with tunable anharmonicity, we implement a complete set of bosonic operations and their qubit-controlled counterparts with high fidelity. This motional control enables accurate detection of magnetic dipolar interactions with $\sim10$ Hz sensitivity in one second, reaching sub-Hz resolution within a few minutes in a $20\times20$ tweezer array under realistic experimental imperfections. Our approach establishes a versatile platform for motional quantum control of two atoms, with applications to spin-boson physics and precision sensing of interaction potentials and trapping environments.
arXiv:2602.24056v2 Announce Type: replace-cross
Abstract: Spectral density functions quantify how environmental modes couple to quantum systems and govern their open dynamics. Inferring such frequency-dependent functions from time-domain measurements is an ill-conditioned inverse problem. Here, we use exactly solvable spin-boson models with pure-dephasing and amplitude-damping channels to reconstruct spectral density functions from noisy simulated data. First, we introduce a parameter estimation approach based on machine learning regressors to infer Lorentzian and Ohmic-like spectral density parameters, quantifying robustness to noise. Second, we show that a cosine transform inversion yields a physics-consistent spectral prior estimation, which is refined by a constrained neural network enforcing positivity and correct asymptotic behaviour. Our neural network framework robustly reconstructs structured spectral densities by filtering simulated noisy signals and learning general functional dependencies.
arXiv:2607.10754v2 Announce Type: replace
Abstract: Category-level object pose estimation is a crucial yet challenging task in both academia and industry, and has achieved remarkable success by leveraging keypoint-based correspondence paradigms. However, most existing methods increasingly rely on stronger feature learning while overlooking whether the established correspondences are geometrically stable across diverse perturbations. This often results in fragile pose recovery under intra-class shape variations and occlusions. To tackle this challenge, we develop a novel Triangle-Invariant Geometric Consistency Learning for Category-Level Object Pose Estimation (TriCons-Pose) to anchor stable keypoints and aggregate pose-invariant cues, yielding reliable canonical mapping and accurate pose estimation. Specifically, a Structure-Consistent Keypoint Detector (SCKD) is designed to identify robust keypoints by enforcing cross-view structural consistency via normalized pairwise distance matching. Moreover, we propose a Pose-Invariant Geometric Aggregator (PIGA) to augment keypoint representations by injecting triangle-based pose-invariant descriptors into a local-to-global attention mechanism. The proposed framework is optimized using standard objective functions while incorporating an additional geometry consistency loss. Extensive experiments on REAL275, CAMERA25, and HouseCat6D datasets demonstrate the effectiveness of the proposed approach.
arXiv:2602.18265v3 Announce Type: replace-cross
Abstract: First-passage times are often the most relevant aspect of a complex Markovian network because they signify when information processing has resulted in a definite decision. Previous studies have shown that for kinetic proofreading networks in the limit of large network size the first-passage time distribution converges either to a delta or to an exponential distribution. Remarkably, these two forms correspond to the two extreme distributions of minimal and maximal entropy for a fixed mean, respectively. Here we build on the connection between first-passage times and graph theory to show that these two limits are not model-specific, but arise generically in Markovian networks from the distribution of the eigenvalues of the generator matrix. A deterministic peak emerges when infinitely many eigenvalues contribute, while the exponential limit arises from a single dominant eigenvalue. We also show that the exponential limit emerges robustly for reversible networks when the mean first-passage time from the initial state to the target state becomes much larger than the mean first-passage time in the reverse direction. In contrast, the deterministic limit is not obtained from a simple reversal of this condition, but follows from a non-vanishing conductance or a mean-residual lifetime of the process which becomes small compared to the mean first-passage time in the long-time limit. This reveals a fundamental asymmetry between the two regimes. Our theoretical analysis is illustrated and validated by computer simulations of one-step master equations and random networks.