Forskningsradar

Science Journals

Peer-reviewade publikationer — 56237 artiklar

Smaller and Faster 3DGS via Post-Training Dictionary Learning
arXiv:2605.30396v1 Announce Type: new Abstract: 3D Gaussian Splatting (3DGS) is a promising neural scene representation for real-time rendering, but trained models often suffer from large memory footprints, limiting deployment on less powerful devices. Existing compression techniques often lead to architectures with several additional trainable parameters. While achieving outstanding compression ratios, they introduce noticeable drops in image quality. In this work, we introduce the first dictionary-learning-based compression framework for 3DGS. The proposed post-training compression pipeline can be deployed in virtually any 3DGS model without the need for re-training or modifications to existing 3DGS models. Our compression framework is straightforward to implement, yet provides significant compression capabilities, preserves image quality, and improves real-time rendering performance. Across 13 benchmark scenes, our approach achieves an average compression ratio of 3.95x, 3.10x, and 4.55x when applied to 3DGS, 3DGS-MCMC, and PixelGS, respectively. This yields consistent rendering speedups of 23.3%, 24.3%, and 25.3%, while maintaining image quality.
SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
arXiv:2605.31148v1 Announce Type: new Abstract: Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into actions in everyday 3D environments. Although recent vision-language models (VLMs) have shown promising performance on observation-conditioned spatial perception and reasoning tasks, it remains unclear whether they can build coherent spatial understanding, act upon it, and refine their actions through multi-turn feedback. To study this problem, we introduce \textbf{SpatialAct}, a simulator-grounded benchmark for probing \textit{action-conditioned spatial reasoning} in 3D scenes. Starting from the most challenging setting, Multi-turn Interactive Refinement, we further design its decomposed counterpart, Single-step Error Detection and Fix, together with five fundamental spatial ability tasks to diagnose the underlying causes of model failures. Experiments reveal a clear reasoning-to-action gap: current VLMs can perform well on isolated spatial reasoning tasks, but struggle to maintain coherent spatial beliefs and produce reliable actions during multi-turn feedback, substantially underperforming humans. These results suggest that current VLM agents still lack robust spatial state tracking under action-induced environment changes, even when low-level control is abstracted away.
Non-Hermitian fluctuations enable model-free particle manipulation
arXiv:2605.31151v1 Announce Type: new Abstract: Contactless manipulation of microscopic matter is central to applications ranging from the isolation of circulating tumor cells in liquid biopsies to the removal of microplastics from environmental water. Electromagnetic approaches are particularly attractive because fields can be structured within compact microfluidic systems using either light or simple electrode architectures. However, precise manipulation requires calibrated models of the field distribution and accurate knowledge of the properties of both the object and the surrounding medium, which limits applicability to well-characterized, static systems. Here we show that energy dissipation itself provides sufficient information for deterministic particle control. Instead of relying on explicit field calibration, our approach exploits an original relationship between particle position, energy dissipation, and electromagnetic body forces, which can be accessed experimentally through variations of conductance matrices. By extracting force-shaping voltage patterns from these measurements, we demonstrate fully automated closed-loop manipulation of silica microbeads in one and two dimensions, including in the presence of other freely moving particles in a disordered background. These results establish a pathway toward deterministic force control by deliberately measuring and exploiting the non-Hermitian response of the system to engineer electromagnetic momentum transfer. This framework expands micromanipulation into realistic, dynamically evolving environments, where wave-matter interactions cannot be fully pre-characterized or eliminated through design.
BIAS-ID: A Framework for Analyzing Transformation Biases in AI-Generated Image Detectors
arXiv:2605.31153v1 Announce Type: new Abstract: Given the surge of harmful AI-generated imagery online, reliably distinguishing authentic images from generated ones has become an urgent research topic. While many proposed detection methods perform well under controlled settings, they often collapse when tested on real-world data. A potential root cause are subtle biases in the detectors' training data. As a result, detectors may rely on spurious correlations instead of learning true forensic artifacts. While a recent line of work has identified the problem, there is not yet an established protocol to evaluate how biased a detector actually is. In this work, we therefore take a step back: First, we discuss what it means for a detector to be biased, and how this differs from a lack of robustness. Second, we propose BIAS-ID, a transparent framework for analyzing and quantifying the presence of transformation biases in AI-generated image detectors. We validate our framework by performing an evaluation of six detectors across two datasets, revealing that several state-of-the-art detection methods are strongly affected by biases. Our results highlight the importance of bias-aware evaluation for developing reliable AI-generated image detectors.
Orbital Angular Momentum Locking via Bound States in the Continuum
arXiv:2605.31154v1 Announce Type: new Abstract: Optical vortices are electromagnetic fields twisting around a phase singularity, resulting in quantized orbital angular momentum (OAM). When such vortices are formed by evanescent hybrid light-matter quasiparticles known as polaritons, they are referred to as polaritonic vortices (PVs). The nanometer-scale topologically robust features of such PVs promise to enable applications for lasing and thermal emission at deeply subwavelength scales. However, many conventional techniques are prone to producing multimode PVs due to poor mode selectivity, resulting in OAM mixing that degrades vortex purity and limits their performance for high-fidelity optical information encoding and multi-dimensional imaging. To overcome this limitation, we introduce a platform that generates deeply subwavelength PVs through quasi-bound states in the continuum (qBICs) in dielectric metasurfaces. In contrast to existing approaches, the qBIC intrinsically locks the PV to a single OAM and makes it robust against the polarization state of the excitation, including linear, elliptical and circular polarization. We experimentally realize qBIC-driven PVs through the interference of hyperbolic phonon polaritons (HPhPs) in hexagonal boron nitride by exploiting the highly uniform out-of-plane electric fields generated by the photonic qBIC, characterized via scattering scanning near-field optical microscopy. This results in HPhPs with a wavelength of around 30-40 smaller than the incident light, thereby enabling ultra-dense packing of multiple robust PVs with distinct OAM. Our platform brings PVs to the photonic chip scale, enabling applications in structured optical information transfer and communications.
Free energy Estimation on Any State Space
arXiv:2605.31063v1 Announce Type: cross Abstract: Free energy estimation is a fundamental yet challenging problem, from physics to statistics. Classical approaches rely on thermodynamic transformations, ranging from direct estimation, quasistatic integration, to finite-time averaging. Recent work [He and Du et al., 2025] learns neural transports to significantly accelerate the efficiency in the finite-time regime. In this paper, we generalize this framework to arbitrary state spaces. Building on this view, we develop a generalized neural transport learning approach for efficient estimation. Experiments validate the effectiveness and efficiency of the proposed method beyond continuous settings, extending to discrete and multimodal spaces as well as autoregressive settings. Beyond free energy estimation, we establish algebraic identities and reveal a group-theoretic structure linking infinitesimal time reversal and generalized Doob's $h$-transforms, showing that their compositions form a generalized dihedral group.
$q$-Exponential Random Graphs: higher-order networks from simple constraints
arXiv:2605.31209v1 Announce Type: new Abstract: Exponential Random Graphs (ERGs) are among the most widely used network models, derived as principled least-bias graph ensembles that maximize Shannon entropy under constraints on the expected values of given structural properties. However, it has been recently (re)discovered that, in the absence of additional information privileging Shannon entropy, the most agnostic inferential construction should maximize the broader class of Uffink entropies. The resulting entropy-maximizing distribution changes from the exponential (Boltzmann-Gibbs) to the so-called q-exponential one. Since maximizing Shannon entropy may produce an unjustified independence between degrees of freedom, here we investigate how the most popular ERGs with independent edges (namely, the Erdos-Renyi and configuration models) generalize to higher-order q-Exponential Random Graphs with dependent edges in the non-Shannon case, while keeping their defining constraints (number of links and degree sequence, respectively) unchanged. We find features, such as a phase transition between sparse and dense regimes, that are absent in the original ERGs but typical of higher-order networks, plus novel phenomena such as richer assortativity and clustering profiles, which allow for the coexistence of link sparsity and triadic closure. These results show that higher-order networks do not necessarily require higher-order constraints, as they naturally arise from simpler ones in a framework that is even more agnostic than Shannon's.
A Hierarchical Spatiotemporal Action Tokenizer for In-Context Imitation Learning in Robotics
arXiv:2604.15215v3 Announce Type: replace Abstract: We present a novel hierarchical spatiotemporal action tokenizer for in-context imitation learning. We first propose a hierarchical approach, which consists of two successive levels of vector quantization. In particular, the lower level assigns input actions to fine-grained subclusters, while the higher level further maps fine-grained subclusters to clusters. Our hierarchical approach outperforms the non-hierarchical counterpart, while mainly exploiting spatial information by reconstructing input actions. Furthermore, we extend our approach by utilizing both spatial and temporal cues, forming a hierarchical spatiotemporal action tokenizer, namely HiST-AT. Specifically, our hierarchical spatiotemporal approach conducts multi-level clustering, while simultaneously recovering input actions and their associated timestamps. Finally, extensive evaluations on multiple simulation and real robotic manipulation benchmarks show that our approach establishes a new state-of-the-art performance in in-context imitation learning.
HiERO-StepG @ Ego4D Step Grounding Challenge: hierarchical activity understanding enables zero-shot step grounding
arXiv:2605.31227v1 Announce Type: new Abstract: Procedural activities follow well-defined structures: whether we consider a cooking recipe or a mechanic repairing a car, these activities naturally decompose in a hierarchy of steps and sub-steps. Traditional approaches for step grounding require extensive annotations and scale poorly. Instead, we argue that such hierarchical structure can emerge naturally from uncurated videos of human activities through recurring patterns of co-occurring actions and activities. Our approach builds on HiERO, a weakly-supervised representation learning approach that maps close in the feature space actions that are functionally related to each other, leveraging only fine-grained action-level narrations. In this feature space, procedure steps can be detected by a simple clustering, with no additional task-specific fine-tuning. For the Ego4D Step Grounding challenge, we augment this approach by ensuring fine and coarse level agreement in step assignments, enforcing strict temporal monotonicity of the grounded steps and post-processing the detected steps to reduce the impact of noisy predictions. We call this approach HiERO-StepG and it achieves 56.27 % on the R@1 (IoU = 0.3) metric on the global leaderboard at submission time, ranking second while being completely zero-shot and not requiring procedure-specific annotations. Project page: https://github.com/andreazenotto/HiERO-StepG.
Diffusion Models Preferentially Memorize Prototypical Examples or: Why Does My Diffusion Model Love Slop?
arXiv:2605.30642v1 Announce Type: new Abstract: Generative models have a persistent limitation: their tendency to memorize training data can create legal liabilities and erode creative diversity. Understanding which samples are memorized in whole or in part, and under what conditions, therefore remains an important open problem. Here we answer the question "Are atypical or rare samples memorized first?" in the negative. We train diffusion models on strings generated according to the production rules of the Random Hierarchy Model (RHM), and find that samples composed of common substrings are preferentially memorized. This holds true even if the training data consists of entirely unique samples, indicating that deduplication at the data point level does not provide a meaningful privacy guarantee. Correspondingly we predict, then observe, delayed memorization for fat-tailed datasets (i.e., those with more atypical samples). This effect is amplified when fat-tails are introduced into high-level production rules. These together suggest that dataset diversity, particularly at higher levels of abstraction, plays an important role in staving off memorization. Finally, we identify an intermediate regime of partial memorization in which common substrings are learned first and subsequently overproduced during generation. If training is stopped in this regime, models will exhibit the reversion-to-the-mean blandness often derided as "slop".
Convergence of Steepest Descent and Adam under Non-Uniform Smoothness
arXiv:2605.30648v1 Announce Type: new Abstract: Recent work has analyzed the convergence of first-order methods under non-uniform smoothness assumptions that better model the loss landscape in machine learning tasks. We generalize this assumption to objectives whose curvature is an affine function of the objective value. This property is satisfied by a broad class of problems, including logistic regression, generalized linear models with a logistic link function, softmax policy gradient in reinforcement learning, and a class of neural networks. Under this assumption and gradient domination conditions, we establish a general convergence rate for the steepest descent method, and deterministic, diagonal variants of RMSProp and Adam. Our results imply that for logistic regression on separable data and the softmax policy gradient objective, sign GD converges linearly and is provably faster than GD. Furthermore, we show that for a class of two-layer neural networks on separable data, RMSProp and Adam can converge at a linear rate with a constant step-size and momentum parameter. Finally, we present a lower bound demonstrating that, under our assumption, RMSProp and Adam are provably faster than AdaGrad, AMSGrad, gradient descent, and heavy-ball momentum.
Private Noise and Public Error in Collective Information Acquisition
arXiv:2605.30522v1 Announce Type: new Abstract: Collective information acquisition requires groups to combine personal evidence with social information while remaining coupled to the external state. Communication noise can affect this process, but the role of noise remains unclear. In an online experiment, 600 participants worked in four-person human groups estimating a room temperature across 25 rounds while receiving either faithful social information, comprehension noise in which each receiver saw independently perturbed social information, or production noise in which perturbations were stored before display and could be seen by multiple receivers. The thermometer cue was objectively veridical, but its reliability was subjectively uncertain and the unitless 50--250 room-temperature range created a task-induced conflict between displayed evidence and everyday temperature expectations. Production-noise groups spent more rounds tightly clustered around a wrong value than comprehension-noise groups (\(p=0.016\), group-level permutation). Production noise more often created a wrong common signal (\(p=0.025\), Fisher's exact test) and made that signal persist across more rounds (\(p=0.004\), permutation). Dynamic update models showed that production noise was not more harmful because people followed peers more strongly, but because the same peer influence acted on more correlated production-noise perturbations. Exploratory human analyses linked the mechanism to psychological patterns while a GPT-agent experiment clarified a boundary condition: GPT agents registered uncertainty through reduced confidence without reproducing human-scale production-noise vulnerability. Overall, noise did not simply degrade collective information acquisition. Comprehension noise could sometimes improve correction relative to the faithful control, whereas production noise could turn perturbations into common evidence and stabilize consensus on error.
TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories
arXiv:2605.31308v1 Announce Type: new Abstract: Agent benchmarks increasingly record rich interaction trajectories, yet evaluation often reduces each rollout to a pass rate or reward score. We introduce TraceGraph, a graph-based framework that turns released multi-model agent trajectories into shared decision landscapes. For each task, TraceGraph builds a graph over observable action-observation states from pooled rollouts before model identity is introduced. It then overlays outcome-informed productive cores and trap regions, and summarizes each rollout with three events: Access, Trap exposure, and Repair. Across trajectories spanning five benchmark splits, TraceGraph profiles reveal navigation differences hidden by aggregate scores and show that splits differ in whether they reward avoiding traps or recovering from them. The same TraceGraph landscape also motivates a trap-aware recovery pipeline for SWE-bench: aruntime detector fires on states matching historical trap regions, then lightweight continuation policies are evaluated from the same prefix. On fired states, the best pooled single-factor policy raises official resolved rate from 40.4% to 43.5% on the per-provider fired subset and from 41.0% to 44.8% on common-fired instances, with provider-specific active components. Overall, TraceGraph provides a process vocabulary for asking what agent benchmarks test, where models diverge on a shared landscape, and how failure regions can guide downstream improvement.
AR Forcing: Towards Long-Horizon Robot Navigation World Model
arXiv:2605.31314v1 Announce Type: new Abstract: The diffusion based robot navigation world models are typically trained using parallel supervision, while autoregressive inference is employed during path planning. This results in a distribution shift between training and inference, which destabilizes the performance over long-horizon prediction. We propose AR Forcing, an autoregressive training strategy, which integrates the standard diffusion loss into the autoregressive training loop. At each step, the model uses its own predictions to update the context and optimize the single step noise prediction objective, thereby explicitly exposing the model to the inference state distribution during training. Our method does not require additional discriminators or distribution-matching losses, retains the original diffusion framework and sampler, and is easy to integrate. Experiments on multi-domain navigation datasets (RECON, SCAND, HuRoN, TartanDrive) show that compared with strong baselines, AR Forcing improved the consistency of generated images during long-horizon navigation and the accuracy of predicted trajectories, enhancing robustness of the model in complex known and unknown environments. We will release the code soon.
Generalized Intention Modeling in Multi-Agent Reinforcement Learning
arXiv:2605.31318v1 Announce Type: new Abstract: Modeling an opponent's intent is critical for effective decision-making in non-cooperative, competitive, and general-sum multi-agent reinforcement learning. Existing opponent modeling methods encode intent using an embedding derived from episode information chosen a priori, such as the opponent's next action or a future environment state, and use this to guide the ego-agent's behavior. These approaches assume that the chosen information is universally representative of intent; however, we show empirically that this is not the case as intentions are often task- and environment-dependent. To address this, we introduce a task-adaptive opponent modeling framework that learns a performance-driven mixture of multiple intent representations. We further introduce a new intention representation that maximizes mutual information with the ego-agent's future returns, thereby capturing opponent information that is most directly relevant to performance. Our approach consistently matches or exceeds the performance of state-of-the-art baselines across diverse tasks and yields insights into when and why different opponent modeling strategies succeed.
Constrained Multi-Objective Reinforcement Learning with Max-Min Criterion
arXiv:2605.31388v1 Announce Type: new Abstract: Multi-Objective Reinforcement Learning (MORL) extends standard RL by optimizing policies with respect to multiple, often conflicting, objectives. While max-min MORL has emerged as an effective approach for promoting fairness, its applicability remains limited, particularly when constraints must be incorporated. In this paper, we propose a MORL framework that integrates the max-min criterion with explicit constraint satisfaction. We establish a theoretical foundation for the proposed framework and validate the resulting algorithm through convergence analysis and experiments in tabular settings. We further demonstrate the practical relevance of our approach in simulated building thermal control, multi-objective locomotion control, and greenhouse-gas-emission-aware traffic management. Across these domains, our method effectively balances fairness and constraint satisfaction in multi-objective decision-making.
Learning to Perceive the World Through Control: Empowerment-Based Representation Learning
arXiv:2605.30656v1 Announce Type: new Abstract: In many practical reinforcement learning environments, observations are far higher-dimensional than the variables that matter for control. In this work, we ask: can we learn representations that capture only control-relevant features of the environment? We study this question through the empowerment objective, which maximizes an agent's influence over the environment and is widely used for unsupervised skill learning. We show that empowerment agents induce two distinct representations -- forward and backward -- that capture complementary aspects of the state, and both of which are invariant to control-irrelevant features. Thus, empowerment maximization leads agents to learn an implicit, control-centric model of the world. Our analysis highlights the importance of learning representations through interaction rather than from passive datasets: interaction aimed at maximizing control is essential for learning useful invariance properties, a perspective that aligns closely with the causal learning literature.
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models
arXiv:2605.30394v1 Announce Type: new Abstract: This paper introduces Code Bench, a benchmark capable of evaluating Large Language Models (LLMs) concise code generation abilities in 60 programming languages. Based on code golf, a recreational programming competition focused on minimal character or byte solutions, the benchmark provides a distinctive measure of LLMs ability to produce efficient, concise code. Unlike existing benchmarks limited by fixed problem sets and language coverage, CodeGolf Bench leverages the code.golf platform to provide new problems and live human performance baselines. Evaluation of nine LLMs on Python and C++ tasks demonstrates that reasoning models significantly outperform non-reasoning models, achieving best average percentile of 70.97%. This performance gap is particularly pronounced in C++, highlighting reasoning's importance for languages with strict syntax requirements. Non-reasoning models struggle more with efficiency optimization across both languages, with best percentiles significantly lower than reasoning counterparts. CodeGolf Bench offers a dynamic framework for evaluating LLM code generation capabilities against evolving human performance on code golf.
Spectral density estimation for normal matrices
arXiv:2605.31430v1 Announce Type: new Abstract: The spectral density estimation problem asks for an algorithm that, given an $n\times n$ matrix $A$, outputs a probability measure that is a good approximation to the uniform distribution on the eigenvalues of $A$, called the spectral density of $A$. This paper considers the setting where $A$ is a large normal matrix that is accessible only through matrix-vector product queries. We provide an algorithm that makes just $m$ matrix-vector queries to $A$ and returns, with high probability, a measure within earth mover's distance $O(1/m+\log m/{\sqrt n})$ of the true spectral density of $A$. We provide a complementary lower bound that any algorithm producing an $\varepsilon$-approximation to the true spectral density for large matrices must make $\Omega(1/\varepsilon)$ matrix-vector queries. The lower bound holds even for the more restricted case of real symmetric input matrices. In combination with our upper bound, it shows that spectral density estimation is essentially no harder for complex normal matrices than for real symmetric matrices.
NumLeak: Public Numeric Benchmarks as Latent Labels in Foundation Models
arXiv:2605.30393v1 Announce Type: new Abstract: Public numeric benchmarks appear in pretraining, so an evaluation that conditions on a date may be measuring memorized recall rather than out-of-sample skill. We introduce NumLeak, a measurement framework that combines API-boundary probes on production models with a white-box controlled validation on an open causal LM. Top-tier frontier LLMs recall the Fama-French market excess return at 3-seed pooled Pearson r=0.97-0.99 while staying within 0.15 within-25bps on the five sibling factors; comparable fidelity appears on U.S. unemployment, CPI inflation, and NOAA temperature. On a recent-release holdout, parse rate collapses to 21-57% but r stays at approximately 0.99 on months answered, the refuse-or-recall asymmetry a memorized channel predicts. The white-box experiment reproduces the dose-response, and logprob ranking detects memorization that open-ended generation misses, implying closed-API black-box probes understate the channel. A Sonnet "date to market-sentiment" regression that correlates with true Mkt-RF at r=0.74 collapses to r=0.02 once the model's own recall is residualized out. A one-line system-prompt defense blocks 99.8% of a non-adaptive single-turn suffix attack set at near-zero utility cost on conceptual and historical-narrative queries
Human-Alignment, Calibration, and Activation Patterns in Large Language Model Uncertainty
arXiv:2605.30675v1 Announce Type: new Abstract: Uncertainty Quantification is a large and growing subfield of large language model behavioral analysis. Primarily to recognize and combat hallucination, the field has largely focused on measuring and improving calibration, the accuracy of uncertainty judgments to task efficacy. In this work, we investigate the relatively underexplored question of how similar large language model uncertainty is to human uncertainty. We investigate the presence and strength of human-similar uncertainty signals, deemed uncertainty alignment, in large language model overt behavior and internal activation patterns. We identify whether the models show evidence of simultaneous alignment and calibration on a variety of datasets covering both multiple choice and open ended factual recall. And we characterize the effect of instruct fine-tuning on each of these facets.
Universal Decision Learners
arXiv:2605.30694v1 Announce Type: new Abstract: Many theories of decision making -- planning, reinforcement learning, causal intervention, online learning, and game-theoretic equilibrium -- turn local information into globally coherent behavior. This paper proposes a common categorical formulation: a Universal Decision Learner (UDL) extends a partially specified decision functor from observed contexts to new contexts by a pair of universal constructions. Left Kan extensions express rollout, aggregation, and candidate generation; right Kan extensions express consistency, constraint satisfaction, and fixed-point semantics. The central claim is not that every decision problem has the same algorithm, but that many decision formalisms instantiate the same universal problem: extend local behavioral data canonically, then characterize the globally coherent extensions. We give the abstract UDL construction, prove its universal comparison property, define Kan-invariant behavioral equivalence and minimal abstractions, and show how Bellman equations, planning recursions, causal interventions, online regret, and equilibria arise as special cases. The supplementary material develops the reinforcement-learning specialization in more detail.
Geometry-Aware Control Barrier Functions for Collision Avoidance via Bernstein Polynomial Approximations
arXiv:2605.30696v1 Announce Type: new Abstract: Safe navigation often relies on well-defined conditions based on the shape of robots and obstacles, and can be challenging when they have irregular geometries. While Control Barrier Functions (CBFs) offer an efficient mechanism to enforce safe set forward invariance, common shape surrogates (e.g., spheres or super-ellipsoids) either are overly conservative in unstructured scenes or require many local primitives, which inflates constraint counts and degrades real-time performance. In this paper, we introduce a novel geometry-aware Control Barrier Function (CBF) based on Bernstein-Polynomial Signed Distance Fields (BP-SDFs). It provides a unified way to represent the obstacles and robots, so as to represent the barrier function with a unified minimum distance. Benefiting from the differentiability of the Bernstein polynomials, one can easily enforce the control constraints in a closed loop. We validate the method's efficiency and performance to guarantee safety in single-robot navigation and heterogeneous multi-robot collision avoidance via simulations under different environments.
High-speed mid-infrared single-photon upconversion spectrometer
arXiv:2605.30701v1 Announce Type: new Abstract: Sensitive and fast mid-infrared (MIR) spectroscopy is highly attractive in a variety of applications including astronomical observation, pharmaceutical synthesis, and environmental monitoring. However, the performance of conventional MIR spectrometers has long been hindered by the limited sensitivity of narrow-bandgap detectors and/or the deficient brightness of broadband light sources. Here, we devise and implement an ultra-sensitive and broadband MIR upconversion spectrometer, which integrates a supercontinuum source covering 1.5-4.2 $\mu$m based on a silicon nitride nanophotonic waveguide. High-efficiency and low-noise nonlinear frequency upconversion is realized based on coincidence pulsed pumping with spectro-temporal optimization, which enables to leverage silicon detectors for facilitating MIR single-photon spectroscopy at 0.2 photons/nm/pulse. Furthermore, the upconversion-based array spectrometer is manifested with high-speed spectral acquisition rates beyond 200 kHz, which is about ten-fold faster than the state-of-the-art scan rates for FTIR-based spectrometers at a comparable spectral resolution. The achieved features of broadband spectral coverage, single-photon sensitivity, and sub-MHz refreshing rate might open up new possibilities in infrared transient spectral measurements in combustion analysis, high-throughput sorting and reaction tracking, among others.
Mid-infrared single-photon 3D imaging
arXiv:2605.30702v1 Announce Type: new Abstract: Active mid-infrared (MIR) imagers capable of retrieving three-dimensional (3D) structure and reflectivity information are highly attractive in a wide range of biomedical and industrial applications. However, the infrared 3D imaging at low-light levels is still challenging due to the deficiency of sensitive and fast MIR sensors. Here we propose and implement a MIR time-of-flight imaging system that operates at single-photon sensitivity and femtosecond timing resolution. Specifically, back-scattered infrared photons from a scene are optically gated by delay-controlled ultrashort pump pulses through nonlinear frequency upconversion. The upconverted images with time stamps are then recorded by a silicon camera to facilitate the 3D reconstruction with high lateral and depth resolutions. Moreover, an effective numerical denoiser based on spatiotemporal correlation allows us to reveal the object profile and reflectivity under photon-starving conditions with a detected flux below 0.05 photons/pixel/second. The presented MIR 3D imager features with high detection sensitivity, precise timing resolution, and wide-field operation, which may open new possibilities in life and material sciences.