Forskningsradar

Science Journals

Peer-reviewade publikationer — 57198 artiklar

Degree-Four Vector-Coordinate SoS Cannot Detect the MUB Upper Bound
arXiv:2606.13903v1 Announce Type: cross Abstract: We prove a degree-four Sum-of-Squares lower bound for the standard vector-coordinate formulations of mutually unbiased bases. For every dimension $d$ and every proposed number $m$ of bases, we construct a degree-four pseudoexpectation satisfying the orthonormality constraints and the cross-unbiasedness constraints in the quartic equality formulation. The construction is expectation over $m$ independent Haar-random orthonormal bases. We also prove that the same pseudoexpectation satisfies the degree-four localizing constraints for the natural $2\times 2$ Hermitian semidefinite formulation of the cross-coherence inequalities. Consequently, degree-four vector-coordinate SoS cannot refute the existence of $m$ mutually unbiased bases, even when $m>d+1$. In particular, under the two vector-coordinate encodings explicitly described in Randomstrasse101 Open Problem 23, degree-four SoS cannot prove that seven mutually unbiased bases do not exist in $\mathbb C^6$. We contrast this with a centered projector-coordinate Gram formulation, where degree-four SoS already recovers the elementary upper bound $m\le d+1$, giving a simple separation between vector-coordinate and projector-coordinate degree-four relaxations.
Binary Black Hole Parameter Estimation with Hybrid CNN-Transformer Neural Networks
arXiv:2606.13941v1 Announce Type: cross Abstract: The detection of gravitational waves has revolutionized our ability to explore fundamental aspects of the Universe. Traditionally, modeled gravitational-wave signals have been identified using template-based matched filtering, followed by coincidence analysis across multiple detectors in the signal-to-noise ratio time series. Recent advances in Machine Learning and Deep Learning have sparked growing interest in their application to both signal detection and parameter estimation. In this study, a hybrid Deep Learning strategy is proposed that leverages the effectiveness of Transformer encoders alongside well-established Convolutional Neural Network architectures in an attempt to estimate the intrinsic and extrinsic parameters of non-precessing binary black hole systems. The primary focus of this work is point estimation, producing single best-fit values for each parameter rather than full posterior distributions. This method is evaluated on both simulated signals embedded in Gaussian noise and real gravitational-wave events, and it demonstrates strong predictive performance and robustness across key astrophysical parameters.
Implementation of two-qubit Rydberg operations on neutral Rb-87 atoms in systems with different intermediate states
arXiv:2606.13975v1 Announce Type: cross Abstract: This work presents an experimental setup for implementing two-qubit operations on neutral atoms ($^{87}$Rb) with the possibility of using two different Rydberg excitation schemes. One of them uses 5P$_{1/2}$ as the intermediate level and applies the second-stage beam locally to the addressed atoms. The second scheme uses the 6P$_{3/2}$ level; in this scheme, the particles to be entangled are moved to a separate zone through which both Rydberg beams pass. The advantages and limitations of both schemes are analyzed. Based on numerical modeling performed with a Julia package developed by the authors, it is demonstrated that the spatial configuration has a greater effect on quantum-operation fidelity than the choice of intermediate level. An experimental implementation of the scheme using the 6P$_{3/2}$ level is demonstrated, making it possible to achieve a two-qubit operation fidelity of 94%.
Adaptive Nucleus Truncation for Long-Form Reasoning
arXiv:2606.13982v1 Announce Type: cross Abstract: Sampling plays an important role in long-form language-model reasoning. Over thousands of decoding steps, small changes in the candidate token set can compound into different reasoning trajectories, stability profiles, and final answers. Existing truncation methods such as top-$p$, min-$p$, and fixed top-$n\sigma$ sampling improve over unrestricted sampling, but they rely on fixed thresholds that cannot adapt to changes in entropy, task difficulty, training stage, or generation budget. We introduce Adaptive Nucleus Truncation Sampling (ANTS), which extends top-\(n\sigma\) sampling from a fixed decoding rule into an adaptive rollout-control mechanism for long-form generation. ANTS selects standardized neighborhoods around the maximum logit before temperature scaling, adapts the truncation width using an entropy-conditioned controller, and retains a no-truncation fallback arm to stabilize training when truncation becomes unsafe. On a 33B-total / 4B-active sparse Mixture-of-Experts reasoning model, ANTS improves average performance over percentage-based benchmarks by +1.9, +3.8, and +5.2 points at 8K, 16K, and 32K generation budgets, respectively. The strongest gains appear on instruction following and mathematical reasoning, with IFBench improving by more than 10 points at 32K and AIME 2025 improving by 7 points. Code generation reveals an important budget interaction. On Codeforces, ANTS trails the baseline at 8K, but reverses this gap and substantially improves ELO at 16K and 32K. These results suggest that sampler design should be treated not just as a decoding hyperparameter, but as part of how we stabilize and scale long-budget reasoning.
Geometric Domain Adaptation via Optimal Transport for Linear Regression in R^2
arXiv:2606.14023v1 Announce Type: cross Abstract: Optimal Transport has become recently a powerful method for domain adaptation by aligning source and target distributions. We study a supervised domain adaptation problem where source and target domains are related by a rotation or a translation or a homothety in $\mathbb{R}^2$. We prove that the optimal transport map recovers the underlying map when using a $p-$norm cost with $p \geq 2$. Based on this insight, we develop a method combining $K-$means and optimal transport to estimate the underlying map, enabling adaptation of linear regression models when target data is scarce. Simulations demonstrate improved performance over baseline methods. Rather than relying on highly expressive deep learning architectures, we focus on classical machine learning models to emphasize interpretability and theoretical insight. This perspective allows us to explicitly characterize the role of optimal transport in recovering geometric transformations such as rotations, translations, and homotheties. Our contributions include a theoretical result linking optimal transport and rotations, translations and homothecies in $\mathbb{R}^2$, and a practical method for adaptation in linear regression offering both conceptual clarity and applied value in domain adaptation tasks in this space.
Anytime-Valid Confirmation of Label-Shift Corrections
arXiv:2606.14028v1 Announce Type: cross Abstract: In small-batch scientific deployments, labeled target outcomes may be too scarce for reliable shift estimation even when unlabeled target inputs are available. We address the complementary setting where the practitioner has a pre-specified label-shift correction from domain knowledge and asks whether incoming labeled outcomes support it. We show that the per-observation likelihood ratio between a label-shift-corrected predictive and the source predictive is a conditional e-value, so its running product is a nonnegative martingale and Ville's inequality yields an anytime-valid confirmation rule. The log martingale equals the cumulative negative log-predictive density (NLPD) gap between the source and the corrected predictive, converting routine model monitoring into a formal sequential test. Rejection means the incoming data support the posited correction relative to the source predictive, but it is not a precise estimate of the degree of shift. Closed forms are available for GP sources with Gaussian label-shift ratios. GP regression simulations validate Type I control, finite-sample power, miscalibration sensitivity, and the small-batch advantage of a reliable prior over label-based re-estimation.
Field-selective criticality in 2D melting revealed by multi-field Lee-Yang zeros
arXiv:2606.14075v1 Announce Type: cross Abstract: How a two-dimensional solid melts remains unsettled after 60 years of study, as theory, model systems, simulations, and atomic-resolution experiments continue to suggest conflicting scenarios. The same transition can appear continuous or abrupt depending on how it is observed, where this ambiguity is especially acute in confined water. Here we study bilayer water under nanoconfinement and ask not only where its phase boundaries lie, but how the system responds to the two fields that drive them: temperature and lateral pressure. Using Lee-Yang zeros together with enhanced sampling, we find that some phase boundaries are field-selective: the two responses can differ either in continuity itself, or in how strongly they are rounded in finite systems. This distinction changes the two-step melting picture. The solid--hexatic transition is field-selective first-order, with the density channel remaining unusually rounded, whereas the hexatic--liquid transition becomes a conventional first-order transition once larger cells reveal a hidden bimodal enthalpy distribution. This framework organizes the apparent disagreement among confined-water simulations, hard-disk models and AgI experiments by identifying which thermodynamic channel each probe sees.
Diffusion-driven autocatalytic dynamics on a sphere
arXiv:2606.14133v1 Announce Type: cross Abstract: We study the collective dynamics of independent particles that diffuse outside a spherical surface, on which they are replicated with a prescribed catalytic rate. In spatial dimensions three and higher, the transient nature of diffusion creates the competition between autocatalytic and escape events, thus leading to a rich phase diagram between subcritical (extinction), critical (steady-state), and supercritical (growth) regimes at long times. The rotational symmetry of the domain and an explicit form of the single-particle diffusion propagator allow us to obtain the statistics of the population size (i.e., the number of particles). In this way, we analyze the mean population size, its variance and higher-order moments, as well as the full distribution. In particular, we obtain a fully explicit form of the distribution at long times and describe a slow, power-law approach to this steady-state limit.
Uniform-in-time error estimates for McKean-Vlasov SDEs with common noise and stochastic algorithms
arXiv:2606.14170v1 Announce Type: cross Abstract: In this work, by construct an asymptotic coupling by reflection, we first explore the uniform-in-time estimate on probability distance for two measure-valued processes induced by a McKean-Vlasov SDE with common noise and an interacting particle system, where the drift terms are dissipative merely in the long distance. As direct applications of this estimate, we establish the uniform-in-time error estimates for the numerical solutions derived via backward/tamed/adaptive Euler-Maruyama methods. Moreover, as another direct application, the uniform-in-time conditional propagation of chaos is quantified.
Spectrum Aware Illumination Estimation Using Multispectral Image
arXiv:2606.14248v1 Announce Type: cross Abstract: Multispectral (MS) imaging extends beyond conventional RGB imaging by capturing more spectral bands, thereby improving illuminant spectrum estimation (ISE). However, existing methods often fail to fully exploit spectral information, resulting in suboptimal performance under diverse lighting conditions and across different sensor domains. Hence, we propose a deep learning framework with a spatio-spectral feature extraction block, which incorporates spectral attention mechanisms to enhance spectral correlation and preserve illuminant-relevant spatial features. Through the inclusion of an illuminant prior (IP), our approach prioritizes specific channels that provide more meaningful information in an MS image. We also propose a spectral-domain transform across different MS sensor spaces. The results demonstrate that illuminant spectra learned in high-dimensional sensor spaces can be effectively transformed to various lower-dimensional camera sensor spaces without any additional training. To facilitate evaluation, we introduce a real-world MS dataset containing high-dimensional ground-truth illumination spectra captured under diverse lighting conditions. Through extensive experiments, we demonstrate that our method achieves superior accuracy compared to existing models, thus providing a practical solution for real-world ISE. The code and dataset are available at https://github.com/hyejin5/Spectrum-Aware-Illumination-Estimation-Using-Multispectral-Image.
Operator Calculus for Population-Based Optimization: A Mean-Field Convergence Theory
arXiv:2606.14289v1 Announce Type: cross Abstract: Population-based and distributional optimization methods, from evolution strategies and consensus-based optimization to covariance-matrix adaptation and stochastic gradient methods viewed as distributional dynamics, are widely used for nonconvex or black-box problems, yet their convergence analyses remain fragmented across algorithm-specific techniques. We introduce an operator calculus in which a broad class of such methods, after choosing an appropriate state space and, where necessary, augmenting the state by memory or strategy variables, is described as a composition of three elementary operators (mutation, selection, and recombination) acting on probability measures. Under explicit stability and regularity conditions, the composite operator admits a pre-generator whose continuous-time limit is a transport-reaction-jump (TRJ) PDE that preserves the operator splitting. On this foundation we establish a modular Lyapunov principle. If a state-space Lyapunov function both dissipates under the full generator and controls the relevant search-space gauges, then the state-space Lyapunov functional and the induced search errors decay exponentially. The additive generator structure allows dissipation estimates to be assembled operator by operator, providing a toolkit for certifying convergence of composite mean-field algorithms.
Counting contiguous superregular $4 \times 4$ matrices
arXiv:2606.14296v1 Announce Type: cross Abstract: This short paper has two goals. First, explaining a simple procedure (which is essentially folklore) that, sometimes, makes it possible to obtain a formula for the number of solutions to a system of multivariate polynomial inequalities over a finite field. Second, applying that procedure to prove a formula for the number of contiguous superregular $4 \times 4$ matrices over a finite field. The formula was previously conjectured by Appuswamy, Bazzani, Connelly, Ekaireb, Congero, and Zeger [Probability of super-regular matrices and MDS codes over finite fields, arXiv:2603.20983]. In addition, the same procedure is used to provide formulas for the number of contiguous superregular $3 \times 4$, $3 \times 5$, and $3 \times 6$ matrices over a finite field.
Quantum-Classical Hierarchical Equations of Motion
arXiv:2606.14363v1 Announce Type: cross Abstract: We develop a quantum-classical hierarchical equations of motion (QC-HEOM) approach for simulating non-Markovian open quantum systems. The method combines the ensemble-averaged classical path reference of the quantum-classical path integral formalism with a hierarchy of auxiliary quantum influence functionals. By incorporating thermal fluctuations through an ensemble average over reference trajectories, the hierarchy is required to represent only the residual quantum memory associated with the imaginary part of the bath response function. Consequently, unlike conventional hierarchical equations of motion, QC-HEOM does not require Matsubara or Pad\'e expansions of the thermal kernel and exhibits only weak temperature dependence of the hierarchy size. Furthermore, because thermal fluctuations are supplied through reference classical trajectories, the framework naturally extends beyond harmonic baths and enables the incorporation of anharmonic and molecular environments through externally generated trajectories. We derive the formalism and demonstrate its exactness for a harmonic bath. Applications to an asymmetric spin-boson model and the seven-site Fenna--Matthews--Olson complex illustrate the accuracy of QC-HEOM. It reproduces benchmark quasi-adiabatic path integral and hierarchical equations of motion results while requiring substantially fewer auxiliary objects, particularly at low temperatures. These results establish QC-HEOM as an efficient framework for treating residual quantum memory in quantum-classical descriptions of open-system dynamics. The separation of thermal fluctuations from residual quantum memory through the use of Wigner trajectories provides an approximate route toward hierarchical treatments of complex anharmonic environments that are inaccessible to conventional HEOM approaches.
Certification of the genuine resolution of photon number resolving detectors
arXiv:2606.14365v1 Announce Type: cross Abstract: Photon-number-resolving (PNR) detectors are essential components of photonic quantum technologies, yet thus far, no practical metric exists to certify how many photons they can genuinely resolve in a single measurement. Here we introduce an operational framework for quantifying the capability of a PNR detector to distinguish between different numbers of photons, i.e. its genuine resolution. In turn, we develop a practical and scalable protocol for certifying the genuine resolution of a detector, which is based on coherent state probes. We apply the method to a 28-pixel photon-number-resolving superconducting nanowire single-photon detector (PNR-SNSPD) and certify genuine four-outcome resolution. Our work highlights the critical requirements in terms of detector efficiency towards achieving high genuine resolution. This approach provides an operational benchmark for PNR detectors and fills a crucial gap in the characterization of photonic quantum devices.
Machine-learned particle flow as a foundation model for collider physics
arXiv:2606.14373v1 Announce Type: cross Abstract: The workflow from particle collision to physics analysis passes through a series of reconstruction steps that are traditionally modular and disconnected, with no shared representation linking low-level detector data to high-level analysis tasks. We show that casting event reconstruction as a machine learning problem naturally produces such a shared representation. We repurpose a machine learning model trained for particle-flow reconstruction (MLPF) to perform three distinct analysis tasks: jet flavor identification, jet energy regression, and missing momentum regression. By appending the per-particle latent representations learned during reconstruction as additional input features, we substantially improve over baselines that use kinematic features alone. We further demonstrate that a single linear layer trained using only the latent representations achieves competitive performance against state-of-the-art baseline architectures, and outperforms the baseline for missing momentum regression with approximately 35 times fewer parameters. These results demonstrate that the latent representations learned during reconstruction encode essential physics information needed for downstream analysis, establishing MLPF as a foundation model and offering a concrete step toward an end-to-end pipeline from detector data to physics analysis.
The Future of Computing for Materials Science Challenges
arXiv:2606.14387v1 Announce Type: cross Abstract: Materials discovery increasingly relies on the coordinated use of theory, computation, experiment, data-driven methods, and emerging quantum technologies, yet the full potential of these tools is realised only when they operate within workflows that reflect the complexity of real systems. This perspective summarises current capabilities, limitations, and opportunities across these domains, drawing on contributions from academia, industry, and national laboratories to identify the scientific and structural requirements for more reliable and efficient discovery. Classical simulations provide broad coverage across design spaces, while experimental measurements reveal degradation, heterogeneity, and kinetic processes that determine performance under realistic conditions. Machine learning accelerates exploration when supported by well-curated datasets with clear provenance and uncertainty quantification, and quantum computing offers promising routes into correlated electronic behaviour when aligned with properties that influence engineering decisions. Collectively, these insights highlight the need for reproducible workflows, shared data standards, realistic benchmarks, and a research culture that prepares scientists to work across paradigms. By integrating these methodological and organisational elements, the community can move toward discovery processes that deliver robust predictions, support confident decision making, and shorten the path from conceptual design to deployable materials.
Repeater-Assisted Massive MIMO Downlink Performance with Calibration Errors
arXiv:2606.14412v1 Announce Type: cross Abstract: Reciprocity-based downlink beamforming is imperative for a scalable time-division duplex massive multiple-input multiple-output~(MIMO) deployment. Specifically, for a dual-antenna repeater-assisted massive MIMO system, a mismatch between forward and reverse path gains at the repeater can exacerbate the overall calibration error between the user equipments (UEs) and the base station (BS), which potentially also contains calibration errors of their individual radio-frequency chains. This paper models the effects of such calibration errors, underpins the relations between the uplink and downlink channels for repeater-assisted systems with calibration errors clubbed with the over-the-air channel estimation errors, and derives analytical expressions of the downlink spectral efficiency. The presented results can then be simplified to several special cases, underscoring situations wherein such errors can become pronounced.
Implications of the Reciprocity Theorem for Reconfigurable Intelligent Surfaces
arXiv:2606.14486v1 Announce Type: cross Abstract: Reciprocity between a transmitter and receiver is a foundational requirement in wireless communications. A few recent works have suggested that reciprocity is broken under reflection by reconfigurable intelligent surfaces (RIS) when the reflection phase becomes incident angle dependent. In this work, we rigorously show that these claims are based on the use of idealized reflection coefficients that ignore mutual coupling between heterogeneous unit cells, surface-truncation effects, and structural scattering contributions from the RIS. Full-wave electromagnetic simulations of transmit/receive antennas and a finite-size RIS implemented via a particular unit cell design are performed to quantitatively demonstrate that reciprocity holds even in the presence of incident-angle dependent reflection phases. To show this, we calculate two-port antenna scattering parameters and evaluate the electromagnetic reciprocity integral to support our claims.
Bacterial adhesion to curved surfaces in fluid flow
arXiv:2606.14543v1 Announce Type: cross Abstract: Minimising bacterial surface adhesion and subsequent biofilm formation in industrial and medical settings requires understanding how bacteria are transported and adhere to complex surface geometries in the presence of non-uniform flow. In this paper, we consider the transport of a dilute suspension of motile bacteria through a corrugated two-dimensional channel with perfectly adhesive walls. We asymptotically analyse the diffusive boundary layer that forms in high velocity flows using a curvilinear coordinate system based on the fluid streamfunction, presenting a similarity solution to the diffusivity-varying diffusion-type equation that arises. From this solution, we derive an analytical expression for the bacterial adhesion rate as a function of surface arclength and the spatially varying wall shear rate. Our model predicts that bacterial adhesion becomes localised on curved surfaces, with bacteria showing preferential adhesion to wall `peaks' at lower shear rates and preferential adhesion to wall `valleys' at higher shear rates. More broadly, our results highlight how spatially varying flows generated by complex geometries can lead to localised bacterial adhesion, with potential implications for both enhancing and minimising biofilm formation.
Scaling native entanglement generation in layered semiconductors with quasi-phase matching
arXiv:2606.14553v1 Announce Type: cross Abstract: Efficient generation of entangled photons typically relies on spontaneous parametric down-conversion (SPDC) in phase-matched macroscopic nonlinear media. However, generating entanglement under phase-matching constraints requires additional bulk optics or interferometers. In contrast, ultrathin van der Waals semiconductors - such as transition metal dichalcogenides (TMDs) - exhibit strong enough optical nonlinearities for SPDC to be observed from subwavelength-thick media, thereby bypassing conventional phase-matching constraints. In this microscopic domain, the intrinsic crystal symmetry governs the nonlinear optical response, enabling the native generation of polarization-entangled photon pairs. However, generating these states efficiently has been fundamentally restricted by the material's coherence length ($L_c$), which limits the attainable conversion efficiency. Here, we investigate periodically-poled TMDs (PPTMDs) designed to scale up this interaction via quasi-phase matching. We demonstrate that mechanically flipping the sign of the nonlinearity at precise intervals of $L_c$ introduces quasi-phase matching, that scales the pair-production rate while preserving the pristine, symmetry-generated polarization entanglement, with fidelities exceeding 99%. Backed by a rigorous theoretical model, our work clarifies the interplay between crystal symmetry and propagation effects in thin nonlinear media, providing a new avenue for engineering quantum light in nanophotonic systems.
Application of Artificial Intelligence and Machine Learning in Libraries: A Systematic Review
arXiv:2112.04573v2 Announce Type: replace Abstract: As the concept and implementation of cutting-edge technologies like artificial intelligence and machine learning has become relevant, academics, researchers and information professionals involve research in this area. The objective of this systematic literature review is to provide a synthesis of empirical studies exploring application of artificial intelligence and machine learning in libraries. To achieve the objectives of the study, a systematic literature review was conducted based on the original guidelines proposed by Kitchenham et al. (2009). Data was collected from Web of Science, Scopus, LISA and LISTA databases. Following the rigorous/ established selection process, a total of thirty-two articles were finally selected, reviewed and analyzed to summarize on the application of AI and ML domain and techniques which are most often used in libraries. Findings show that the current state of the AI and ML research that is relevant with the LIS domain mainly focuses on theoretical works. However, some researchers also emphasized on implementation projects or case studies. This study will provide a panoramic view of AI and ML in libraries for researchers, practitioners and educators for furthering the more technology-oriented approaches, and anticipating future innovation pathways.
Learning optimal policies from event logs through reinforcement learning: a comparison of deep and MDP-based approaches
arXiv:2303.09209v2 Announce Type: replace Abstract: Prescriptive Process Monitoring is an emerging area within Process Mining that focuses on recommending actions to optimize business outcomes. Most existing works prescribe pre-defined interventions, i.e., sets of actions applied to ongoing process executions to achieve a specific objective or Key Performance Indicator (KPI). In contrast, only a few approaches have explored learning and evaluating optimal behavioral policies, i.e., general strategies that determine the best sequence of actions to maximize a desired KPI. In this paper, we address the problem of learning optimal behavioral policies by proposing an AI-based approach that learns an optimal policy directly from historical process executions using Reinforcement Learning (RL) to recommend the best actions for optimizing a KPI. To this end, we employ two RL techniques. The first is a classical model-based approach that extends previous work by the authors through the construction of a Markov Decision Process (MDP) capturing process behavior. The second is a model-free technique based on offline Deep RL. Unlike state-of-the-art work, we aim to minimize the use of domain knowledge and learn optimal policies directly from historical event data. This allows us to learn when to apply interventions and discover effective ones directly from data. Moreover, we target complex scenarios involving external actors, where the process owner controls only part of the activities. We adopt a data-driven Business Process Simulation (BPS) environment to evaluate the learned policies. Results show that both methods improve the targeted KPI with similar effectiveness, while the model-based approach outperforms offline Deep RL in computational efficiency.
Automating Boundary Filling in Cubical Type Theories
arXiv:2402.12169v5 Announce Type: replace Abstract: When working in a proof assistant, automation is key to discharging routine proof goals such as equations between algebraic expressions. Homotopy type theory allows the user to reason about higher structures, such as topological spaces, using higher inductive types (HITs) and univalence. Cubical type theory provides computational support for HITs and univalence. A difficulty when working in cubical type theory is dealing with the complex combinatorics of higher structures, an infinite-dimensional generalisation of equational reasoning. To solve these higher-dimensional equations consists in constructing cubes with specified boundaries. We develop a simplified cubical language in which we isolate and study two automation problems: contortion solving, where we attempt to "contort" a cube to fit a given boundary, and the more general Kan solving, where we search for solutions that involve pasting multiple cubes together. Both problems are difficult in the general case-Kan solving is even undecidable-so we focus on heuristics that perform well on practical examples. Our language encompasses different variations of cubical type theory which differ in their "contortion theory", i.e., the class of contortions they support. We provide a solver for the contortion problem for the most complex contortion theories currently being researched, the Dedekind and De Morgan contortions, by utilizing a reformulation of contortions in terms of poset maps. We solve Kan problems using constraint satisfaction programming, which is applicable independently of the underlying contortion theory. We have implemented our algorithms in an experimental Haskell solver that can be used to automatically solve many goals a user of cubical type theory might face. We illustrate this with a case study establishing the Eckmann-Hilton theorem using our solver, as well as various benchmarks.
The Noisy Work of Uncertainty Visualisation Research: A Review
arXiv:2411.10482v4 Announce Type: replace Abstract: Better representation of the uncertainty in a data visualisation is a focus of recent research activity. A problem with the current literature is that there is a lack of clarity about the definition of uncertainty and what it means to represent it in a plot. This confusion results in a significant amount of conflicting results in the literature, especially in experiments that assess the effectiveness of different uncertainty representations. In this review, we summarise the current literature, provide workable definitions, and illustrate these definitions with examples. In doing so, we ask what it really takes to achieve transparency in statistical graphics. It is hoped that it will be useful for guiding new graphics methodology and experimental research.
MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models
arXiv:2502.10886v3 Announce Type: replace Abstract: Entity state tracking is a necessary component of world modeling that requires maintaining coherent representations of entities over time. Previous work has benchmarked entity tracking performance in purely text-based tasks. We introduce MET-Bench, a multimodal entity tracking benchmark designed to evaluate the ability of vision-language models to track entity states across modalities. Using three domains, we assess how effectively current models integrate textual and image-based state updates. Our findings reveal a significant performance gap between text-based and image-based entity tracking. We empirically show this discrepancy primarily stems from deficits in visual reasoning rather than perception. We further show that explicit text-based reasoning strategies improve performance, yet limitations remain, especially in long-horizon multimodal tasks. We apply reinforcement learning to improve entity tracking in open-source VLMs. This yields substantial in-modality gains, but does not transfer robustly across input modalities. Our results highlight the need for improved multimodal representations and reasoning techniques to bridge the gap between textual and visual entity tracking.