arXiv:2506.01883v3 Announce Type: replace Abstract: Training deep learning models on single-cell datasets with hundreds of millions of cells requires loading data from disk, as these datasets exceed available memory. While random sampling provides the data diversity needed for effective training, it is prohibitively slow due to the random access pattern overhead, whereas sequential streaming achieves high throughput but introduces biases that degrade model performance. We present scDataset, a PyTorch data loader that enables efficient training from on-disk data with seamless integration across diverse storage formats. Our approach combines block sampling and batched fetching to achieve quasi-random sampling that balances I/O efficiency with minibatch diversity. On Tahoe-100M, a dataset of 100 million cells, scDataset achieves more than two orders of magnitude speedup compared to true random sampling while working directly with AnnData files. We provide theoretical bounds on minibatch diversity and empirically show that scDataset matches the performance of true random sampling across multiple classification tasks and model architectures.
Science Journals
arXiv:2605.29142v2 Announce Type: replace Abstract: Observations from the RAPID array at 26.5$^\circ$N indicate a linear decline in the Atlantic Meridional Overturning Circulation (AMOC) over the past two decades, linked to contrasting boundary changes: a weakening western-boundary contribution that is partly compensated by strengthening at the eastern boundary. Yet it remains unclear whether this partial compensation reflects a basin-wide adjustment or a regional feature, and what processes drive it. Here we use a high-resolution ocean model to investigate the spatial structure and underlying mechanisms of the AMOC change across the mid-latitude North Atlantic. The model reproduces a meridionally coherent decline in western-boundary deep overturning transport together with a partially compensating strengthening at the eastern boundary, consistent with observations at 26.5$^\circ$N. These opposing trends arise from a vertically coherent ocean bottom pressure trend shaped by two competing drivers: rising coastal sea level and decreasing interior density. Through geostrophic balance, this mechanism produces partial boundary compensation across latitudes, yielding a basin-wide AMOC decline throughout the mid-latitude North Atlantic.
Properties of Adjoint Solutions of the Full-potential Equations for Two-Dimensional Subcritical Flow
arXiv:2506.16886v4 Announce Type: replace Abstract: The adjoint full-potential equations are studied for two-dimensional (2D) steady subcritical flows. In contrast with the incompressible case, explicit closed-form solutions are generally not available in the compressible setting, so the emphasis is placed here on the underlying structure. Using the Green's-function approach and the relation between the adjoint full-potential and compressible adjoint Euler equations, we identify the adjoint potential and stream function with linear combinations of the Euler adjoint variables associated with point mass and vorticity sources. For lift-based cost functions, the corresponding adjoint solutions contain two unknown functions that encode the effect of perturbations to the Kutta condition. We show that these functions obey the linearized full-potential equations, are linked by generalized Cauchy-Riemann equations, and reduce in the incompressible limit to the Poisson kernel of the Laplacian on the exterior of the circle and its harmonic conjugate. Their properties are examined analytically and through numerical adjoint solutions. Finally, a continuous formulation of the Kutta condition for the adjoint full-potential equations is discussed and interpreted in terms of singular boundary forcing, Green-function kernels, and an equivalent Lagrange-multiplier formulation.
Federated Client Selection under Partial Visibility: A POMDP Approach with Spatio-Temporal Attention
arXiv:2605.11752v2 Announce Type: replace Abstract: Federated learning relies on effective client selection to alleviate the performance degradation caused by data heterogeneity. Most existing methods assume full visibility of all clients at each communication round. However, in large-scale or edge-based deployments, the server can only access a subset of clients due to communication, mobility, or availability constraints, resulting in partial visibility where only a subset of clients is observable for aggregation in each communication round. In this paper, we formulate federated client selection under partial visibility as a Partially Observable Markov Decision Process (POMDP) and propose a Spatial-Temporal attention-based reinforcement learning framework. By integrating historical global models and client identity embeddings, the proposed method captures both the temporal contexts of training and the persistent characteristics of clients. Experimental results across multiple datasets demonstrate that our approach achieves superior performance compared to existing baselines in heterogeneous and partially visible settings, validating its effectiveness in addressing the challenges of incomplete observations in practical federated learning systems.
arXiv:2605.12608v2 Announce Type: replace Abstract: Object detection in adverse weather is critical for the safety of autonomous vehicles; however, the scarcity of labelled, real-world foggy data remains a significant bottleneck. In this paper, we propose Clear2Fog (C2F), an end-to-end, physics-based pipeline that simulates fog on clear-weather datasets while ensuring cross-modal consistency across camera and LiDAR. C2F combines monocular depth estimation with a novel atmospheric light estimation method to improve the physical consistency of synthetic fog generation while reducing structural artifacts and chromatic biases observed in existing frameworks. Utilising a training set of 270,000 images from the Waymo Open Dataset, we conduct an extensive data efficiency study to investigate whether environmental diversity can reduce dataset scale requirements and improve model generalisation under varying fog conditions. Our findings reveal that models trained on mixed-density fog datasets at 75% scale achieve comparable detection performance to those trained on fixed-density datasets at 100% scale, reducing synthetic training data requirements by 25%. We observe that this efficiency trend is consistent across two representative detector architectures. Furthermore, we investigate the sim-to-real transfer by using C2F-generated data as a pre-training foundation before fine-tuning on real-world fog data. We demonstrate that, within the evaluated settings, a relative 10x increase in the default fine-tuning learning rate reduces the negative transfer caused by standard fine-tuning, achieving up to a 1.17 mAP point improvement beyond the real-only baseline. Overall, this work demonstrates the value of diverse synthetic fog as a pre-training tool for real-world adaptation.
arXiv:2607.00461v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are often constrained by a language-space bottleneck, forcing complex visual reasoning into discrete tokens which can lose perceptual nuance. A promising alternative is continuous latent reasoning, where the goal is to discover implicit reasoning pathways that bridge the multimodal query and the final answer. However, this introduces a severe train-inference mismatch: a training-time posterior, conditioned on the ground-truth answer, can exploit answer-dependent shortcuts. Standard variational training then forces the inference-time prior to mimic a posterior that has access to information unavailable at test time, leading to poor performance. To address this, we propose Asymmetric Mutual Variational Learning (AMVL), a framework that resolves this mismatch via a bidirectional calibration objective. A forward KL divergence trains the target-agnostic prior to match the posterior, while a novel reverse KL divergence simultaneously regularizes the posterior, preventing it from collapsing into inference-incompatible regions and mitigating this ``answer leakage''. We provide theoretical analysis formalizing this leakage as prior contamination and prove that our dual-KL objective reduces it. We instantiate AMVL in a latent-integrated MLLM and show that it consistently outperforms strong discrete and latent-reasoning baselines, improving the average score on the complex BLINK benchmark by +10.83 and achieving gains of up to +32.00 on individual reasoning tasks, with analyses confirming improved latent-space stability.
arXiv:2607.00959v1 Announce Type: new Abstract: Audio-driven talking head synthesis has achieved impressive progress in lip synchronization and visual quality, yet generating expressive emotional avatars with controllable intensity remains challenging, especially under real-time constraints. In this paper, we present GaussianEmoTalker, an audio-driven framework for real-time emotional talking head synthesis based on 3D Gaussian Splatting. Instead of directly predicting the final emotional avatar from speech, we formulate emotional animation as a neutral-to-emotional residual deformation problem. GaussianEmoTalker first constructs an identity-specific neutral talking space with GaussianBlendshapes, which provides high-fidelity Gaussian attributes and phoneme-synchronized neutral motion. It then predicts an emotion-conditioned residual deformation by combining mesh displacement cues, audio features, emotion categories, and intensity encodings. To fuse these heterogeneous signals, we introduce a spatial-audio-emotion attention module that estimates the offsets of Gaussian attributes for expressive and temporally stable rendering. Extensive experiments demonstrate that GaussianEmoTalker achieves competitive video quality, accurate lip synchronization, controllable emotional expression, and real-time rendering compared with recent emotional talking head methods. Our project page is available at https://njust-yang.github.io/GaussianEmoTalker.github.io/
arXiv:2510.06452v3 Announce Type: replace Abstract: Recent advances in Large Language Models (LLMs) have introduced a new paradigm for software development, where source code is generated from natural language prompts. While this paradigm significantly boosts development productivity, building complex, real-world software systems remains challenging because natural language offers limited control over the code generation process. Inspired by the historical evolution of programming languages toward higher levels of abstraction, we advocate for a high-level abstraction language that gives developers greater control over LLM-assisted code writing. To this end, we propose Code Semantic Zooming (CodeZoom), a novel approach based on pseudocode that allows developers to iteratively explore, understand, and refine code across multiple layers of semantic abstraction. In a within-subjects user study (n=26), our method matches a state-of-the-art coding agent, Claude Code, on usability while producing a large effect on code comprehension: over 90% of participants reported feeling more in control of design decisions when using CodeZoom compared to using Claude Code.
arXiv:2510.08735v2 Announce Type: replace Abstract: Kinetic instabilities develop when a system's distribution function deviates from thermal equilibrium in such a way that allows free energy from that distribution to drive resonant modes. These instabilities occur in many systems, such as fusion and astrophysical plasmas, neutral fluids, and self-gravitating systems. Motivated by the case of Alfv\'enic instabilities driven by minority populations of energetic particles in tokamak plasmas, we consider a kinetic instability in the presence of sources and sinks and with perturbative drive far from its instability threshold, where a theoretical description of the time evolution has not yet been established. In cases with steady-state saturation levels, we find that the mode first evolves linearly, driven by the positive distribution gradient, then undergoes a fast, strongly nonlinear transition where the distribution function slope is completely flattened around the resonance. The system then evolves in a weakly nonlinear regime driven by the balance of the wave drive and dissipation until reaching saturation. The strongly nonlinear transition is sufficiently fast that it can be treated as occurring instantaneously, and by considering a time-local approximation for the distribution after the flattening has occurred, we find a closed-form analytical solution for the mode amplitude in the weakly nonlinear phase. A compact piecewise-continuous solution for the entire time evolution of the mode amplitude is therefore constructed. This result is shown to agree closely with nonlinear kinetic simulations and is derived within a framework common to other physical systems such as galactic discs and viscous fluids.
arXiv:2607.00471v1 Announce Type: new Abstract: Low Earth orbit (LEO) satellite constellations are emerging as a backbone for global 6G connectivity, where independent tenant slices share orbital infrastructure, each requiring an ordered chain of security virtual network functions (VNFs). Because onboard computation and networking are scarce, slices cannot be given dedicated VNFs. They must share instances on the same satellites, enlarging the attack surface and exposing tenants to cross-slice side-channel risk. This exposure shifts continually as visibility, orbital motion, and the inter-satellite topology change in time (epochs), making VNF migration a structural necessity that couples resource efficiency, service continuity, and security isolation into a single problem. We formulate this security- and migration-aware security function chain (SFC) placement as a multi-slice mixed-integer linear programming (MILP) whose core is a co-location risk model, grounded in ISO/NIST principles and supported by analytic bounds, in which we separate avoidable migrations from those forced by orbital motion. Because the joint program scales quadratically with the cross-slice co-location terms, we develop an alternating direction method of multipliers (ADMM)-inspired penalized per-slice best response decomposition that recasts the coupling as a linear per-slice penalty, yielding independent subproblems through sequential (S-ADMM) and parallel, collision-repaired (P-ADMM) schedules. Simulations over a Walker-Delta satellite constellation show that the proposed framework eliminates co-location risk, reduces SFC migrations, and sustains full delay compliance, while remaining feasible within the per-epoch budget for slice counts where the monolithic security-aware MILP is intractable.
arXiv:2607.00996v1 Announce Type: new Abstract: We determine with high accuracy the energy of the inner-shell transition $1s^2 2s^2~{}^1\mathrm{S}_0 \rightarrow 1s~2s^2~2p_{3/2}~{}^1\mathrm{P}_1$ ${}^{16}\mathrm{O}_{K\alpha}^{4+}$ at $554.372(3)~\mathrm{eV}$ ($\lambda$ = $22.36480(12)~\unicode{x212B}$) as well as its small shift of $2.2 \pm 1.3~\mathrm{meV}$ ($\Delta \lambda$ = $0.089(52)~\mathrm{m}\unicode{x212B}$) for the ${}^{18}\mathrm{O}$ isotope. This transition blends with a $K_\alpha$ line of $\mathrm{O}^{5+}$ used in astrophysical diagnostics, potentially affecting its reliability. In contrast to our experimental uncertainty of $\pm 3~\mathrm{meV}$, advanced electronic structure predictions for this four-electron system, including quantum electrodynamic (QED) corrections on the order of $100~\mathrm{meV}$, still scatter by more than $\pm 250~\mathrm{meV}$. Ions generated and stored in an electron beam ion trap were excited at the ELETTRA synchrotron facility with monochromatic soft x rays, with photon energies corrected by an additional spectrometer. Upon resonant excitation of $\mathrm{O}^{4+}$ and subsequent autoionization, we separate the photoions of each isotope by a time-of-flight measurement. This way, we resolve soft x-ray isotopic shifts of a few meV, obtain very accurate data on an essential astrophysical ion, and test calculations down to the level of QED contributions.
arXiv:2512.23365v4 Announce Type: replace Abstract: Recent progress in Multimodal Large Language Models (MLLMs) has enabled 3D scene understanding and spatial reasoning directly from multi-view images, without requiring explicit 3D reconstructions. Nevertheless, key challenges that frequently arise in real-world environments, such as partial visibility, occlusion, and low-overlap conditions that require reasoning from fragmented visual cues, remain under-explored. To address these limitations, we propose a scalable multi-view data generation and annotation pipeline that constructs realistic spatial reasoning QAs, resulting in SpatialMosaic, a comprehensive instruction-tuning dataset with 2M QA pairs. We further introduce SpatialMosaic-Bench, a challenging benchmark for evaluating multi-view spatial reasoning under complex and diverse scenarios, consisting of 1M QA pairs across 11 tasks with both multiple-choice and numerical-answer formats. Our dataset spans both indoor and outdoor scenes, enabling comprehensive evaluation across diverse real-world scenarios. In addition, we provide a practical baseline for multi-view settings by integrating geometry encoders into VLMs for improved cross-view consistency and spatial grounding. Extensive experiments demonstrate that our dataset effectively enhances spatial reasoning under challenging multi-view conditions, validating the effectiveness of our data generation pipeline in constructing realistic and challenging QAs.
arXiv:2606.23974v2 Announce Type: replace Abstract: The Lawson Machine 26 (LM26) at General Fusion has demonstrated compressional heating of a spherical tokamak deuterium plasma as it was compressed by an imploding solid lithium liner. Results from the first 11 compression shots on LM26 are presented, the highest-performing of which show more than a 3x increase in $T_e$, a 10x increase in $n_e$, and a 10x increase in $B_{pol}$ within the plasma driven by 3x radial compression. The experimental device and instrumentation are reviewed in detail, followed by observations about the liner trajectory and evolution of plasma properties, including increases in emission of neutrons, X-rays, and visible radiation. Observations from fast-camera images during compression provide context for interpreting the spatial structure of plasma-wall interaction. Overviews of relevant models and analysis are presented. Diagnostic data are used to reconstruct the experimental equilibrium state in computational framework as a function of time. The results build confidence in the stability and transport analyses that support the primary conclusions. Trends across the full set of 11 compression shots are presented, and detailed examinations of the high-performance shots are given individually. The central conclusions of the integrated physics model specifically indicate that compressional heating was achieved in this set of experiments, as evidenced by the balance of heating power from compression, Ohmic heating from plasma current, and losses to the boundary needed to match the experimental data. A majority of the temperature rise is attributable to compressional heating. An increase in neutron flux is also observed during compression. The results provide a basis for planned improvements to the LM26 facility that will enable the compression of magnetized plasma to increasingly higher densities and temperatures.
arXiv:2510.27285v4 Announce Type: replace Abstract: Concept erasure methods aim to remove specific unsafe target concepts in diffusion models while preserving image generation utility. To address the vulnerability that erased concepts can be easily recovered under adversarial attacks, adversarial concept erasure methods integrate adversarial optimization into the concept erasure process. However, existing adversarial concept erasure methods face a trade-off between robustness and computational cost. We attribute this to adversarial optimization techniques that use random samples to approximate the adversarial objective function. Adversarial optimization that uses a small number of samples fails to produce adversarial embeddings that accurately capture the target concept space. To mitigate this limitation, we propose Semantic-Guided Adversarial Optimization, which uses a single sample to produce adversarial embeddings that better capture the target concept space. We also propose Semantic-Guided Concept Erasure, which automatically maps the target concept to a semantically similar surrogate. Extensive experiments on not-safe-for-work content, artistic styles, and object-related concepts demonstrate that our method, S-GRACE (Semantic-Guided Robust Adversarial Concept Erasure) achieves state-of-the-art erasure robustness and superior image generation utility, with significantly lower computational cost than existing methods. Our code is available at https://github.com/Qhong-522/S-GRACE.
arXiv:2603.00198v2 Announce Type: replace Abstract: Token reduction accelerates long-video vision--language models (VLMs), but existing methods target Transformers, where reduction is treated as token pruning. We study token reduction in hybrid Mamba--Transformer VLMs and find that it is \emph{stateful}: Mamba layers maintain a recurrent state that accumulates information from earlier tokens, allowing discarded tokens to persist, so reduction behaves more like compression than dropping.We support this view with a representation-based probing method measuring how much information from discarded tokens is retained, and analyze layer-wise sparsity and cross-layer importance stability. Our findings show importance is sparse within layers but unstable across layers, making aggressive early pruning unreliable while hybrids remain robust to later reduction.Motivated by this, we propose a hybrid-aware token reduction framework with a low-to-high progressive schedule and a unified query-conditioned importance score for attention and Mamba layers. For Mamba, excluding the position-dependent decay from the recurrence produces a stronger selection signal. Across long-video benchmarks, our method achieves $3.8{\times}$--$4.2{\times}$ prefilling speedups at a 25% token budget while maintaining near-baseline accuracy and improving with light finetuning. Hybrid models benefit from aggressive reduction, improving both efficiency and accuracy, whereas Transformers exhibit the standard trade-off. Our method also outperforms prior baselines on the same hybrid backbone and combines effectively with visual redundancy reduction methods.
arXiv:2511.02644v2 Announce Type: replace Abstract: We study computable probably approximately correct (CPAC) learning, where learners are required to be computable functions. It had been previously observed that the Fundamental Theorem of Statistical Learning, which characterizes PAC learnability by finiteness of the Vapnik-Chervonenkis (VC-)dimension, no longer holds in this framework. Recent works recovered analogs of the Fundamental Theorem in the computable setting, for instance by introducing an effective VC-dimension. Guided by this, we investigate the connection between CPAC learning and recursively enumerable representable (RER) classes, whose members can be algorithmically listed. Our results show that the effective VC-dimensions can take arbitrary values above the traditional one, even for RER classes, which creates a whole family of (non-)examples for various notions of CPAC learning. Yet the two dimensions coincide for classes satisfying sufficiently strong notions of CPAC learning. We then observe that CPAC learnability can also be characterized via containment of RER classes that realize the same samples. Furthermore, it is shown that CPAC learnable classes satisfying a unique identification property are necessarily RER. Finally, we establish that agnostic learnability can be guaranteed for RER classes, by considering the relaxed notion of nonuniform CPAC learning.
arXiv:2511.03591v2 Announce Type: replace Abstract: Safe multi-agent motion planning (MAMP) under task-induced constraints is a critical challenge in robotics. Many real-world scenarios require robots to navigate dynamic environments while adhering to manifold constraints imposed by tasks. For example, service robots must carry cups upright while avoiding collisions with humans or other robots. Despite recent advances in decentralized MAMP for high-dimensional systems, incorporating manifold constraints remains difficult. To address this, we propose a manifold-constrained Hamilton-Jacobi reachability (HJR) learning framework for decentralized MAMP. Our method solves HJR problems under manifold constraints to capture task-aware safety conditions, which are then integrated into a decentralized trajectory optimization planner. This enables robots to generate motion plans that are both safe and task-feasible without requiring assumptions about other agents' policies. Our approach generalizes across diverse manifold-constrained tasks and scales effectively to high-dimensional multi-agent manipulation problems. Experiments show that our method outperforms existing constrained motion planners and operates at speeds suitable for real-world applications. Video demonstrations are available at https://youtu.be/RYcEHMnPTH8 .
arXiv:2507.14706v2 Announce Type: replace Abstract: Detecting fraudulent credit card transactions remains a significant challenge, due to the extreme class imbalance in real-world data and the often subtle patterns that separate fraud from legitimate activity. Existing research commonly attempts to address this by generating synthetic samples for the minority class using approaches such as GANs, VAEs (Variational Autoencoders), or hybrid generative models. However, these techniques, particularly when applied only to minority-class data, tend to result in overconfident classifiers and poor latent cluster separation, ultimately limiting real-world detection performance. In this study, we propose the Causal Prototype Attention Classifier (CPAC), an interpretable architecture that promotes class-aware clustering and improved latent space structure through prototype-based attention mechanisms and we couple it with the encoder of a Variational Autoencoder-Generative Adversarial Network (VAE-GAN) in order to achieve improved latent cluster separation moving beyond post-hoc sample augmentation. We compared CPAC-augmented models to traditional oversamplers, such as SMOTE, as well as to state-of-the-art generative models, both with and without CPAC-based latent classifiers. Our results show that classifier-guided latent shaping with CPAC delivers superior performance, achieving an F1-score of 93.74% and recall of 92.85%, along with improved latent cluster separation. Further ablation studies and visualizations provide deeper insight into the benefits and limitations of classifier-driven representation learning for fraud detection. The codebase for this work can be found at the following link: https://github.com/claudiunderthehood/VAEGAN-CPAC.git.
arXiv:2511.04500v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as decision-making agents in high-stakes domains and as imitators of human behavior in the social and behavioral sciences. Yet how closely LLMs mirror human decision-making remains poorly understood. This gap is critical: misalignment could produce harmful outcomes in practice, while failure to replicate human behavior renders LLMs ineffective as social simulators. Here, we address this gap by replicating large-scale game-theoretic experiments and by introducing a systematic prompting and probing framework for machine-behavioral evaluation. We test three open models typically used to power agents (Llama, Mistral, and Qwen). Across 121 dyadic games spanning four classical game types, Llama reproduces human cooperation patterns with high fidelity, while Qwen aligns closely with Nash equilibrium predictions. Characterizing models through behavioral phenotyping, we find that humans and Llama share an envious decision profile, while Qwen and Mistral exhibit different profiles. An attention-based analysis of payoff salience reveals Llama processes payoff information in a structured, layer-dependent manner absent in Qwen and Mistral, suggesting a mechanistic basis for its closer alignment with human behavior. Population-level behavioral replication is achieved without persona-based prompting, simplifying the simulation process. Extending the experimental parameter space beyond the original human-tested games, we generate and preregister testable hypotheses for novel game configurations. Our findings demonstrate appropriately configured LLMs can replicate aggregate human behavioral patterns, exhibit human-like decision phenotypes, and enable systematic exploration of unexplored experimental spaces, offering a complementary approach to traditional behavioral research that generates new empirical predictions about human social decision-making.
arXiv:2607.00484v1 Announce Type: new Abstract: Optimizing vaccine prioritization is often treated as the default policy response when vaccine supply is limited. Yet optimized prioritization carries administrative, ethical and communication costs, motivating an upstream question: whether differences among vaccine allocations can alter epidemic outcomes enough to make optimization epidemiologically necessary. We show that optimization is not always worth pursuing: in some regimes, vaccination markedly reduces epidemic burden, but many feasible allocation rules perform almost equally well, making the necessity of optimization low. We quantify this necessity as the range of epidemic outcomes generated by different allocations under fixed supply and show that it is governed by competition between vaccinating high-contact groups to slow transmission and vaccinating groups that benefit most directly: necessity is low when these protection routes are balanced and high when one dominates. Increasing transmission intensity changes this balance and drives a transition in the optimal allocation from transmission-focused prioritization toward direct protection. Different prevention objectives exhibit distinct transition thresholds, creating regimes in which optimizing one objective substantially compromises another, thereby revealing when the choice of prevention target matters most. This framework reframes vaccine prioritization as a prior decision problem, identifying when optimization is warranted, when simpler rules suffice, and when prevention goals conflict.
arXiv:2510.02598v2 Announce Type: replace Abstract: Si-PIN detectors can be microstructured to achieve angular-selective particle detection capabilities, which we call active Transverse Energy Filter (aTEF). The microstructuring consists of a honeycomb structure of deep hexagonally-shaped holes with active silicon side walls, while the bottom of the holes is made insensitive to ionizing radiation. The motivation for this kind of detector arises from the need to distinguish background electrons from signal electrons in a spectrometer of MAC-E filter type. We have demonstrated the angular-dependent detection efficiency of self-fabricated aTEF prototypes in a test setup using an angular-selective photoelectron source to illuminate the detector from various incidence angles.
arXiv:2606.29963v2 Announce Type: replace Abstract: The structural vulnerabilities of point cloud-based 3D object detectors remain poorly understood. Prior work has studied adversarial robustness primarily on isolated 3D object models, while recent LiDAR spoofing attacks target richer and more realistic driving scenes but focus mainly on physical realizability rather than understanding detector behavior or attack efficiency. In this work, we investigate how LiDAR-based detectors rely on spatial evidence in complex scenes and whether these reliance patterns can be exploited to induce failures more efficiently. To this end, we propose an explainability-guided adversarial analysis methodology. We introduce the Saliency-LiDAR (SALL) method, which aggregates Integrated Gradient attributions across scenes to produce universal saliency maps for LiDAR-based 3D object detectors. Guided by these maps, we design the Explainability-aware Frustum Attack (EFA), which selectively perturbs only the most influential frustums rather than uniformly attacking entire object regions. Experiments on KITTI and nuScenes, across detectors such as PointPillars and SECOND, show that EFA reduces detection recall by more than 15 percentage points while requiring 25-50% fewer perturbed frustums than the state-of-the-art non-saliency-aware baseline. These findings reveal that modern 3D detectors concentrate discriminative evidence in a small subset of spatial regions, exposing a structural robustness vulnerability in current LiDAR perception systems. Our code is released at https://github.com/SecMindLab/Saliency_LiDAR.
arXiv:2607.00491v1 Announce Type: new Abstract: Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input. Existing what-if tasks typically vary the observer while keeping the scene fixed. Can VLMs instead predict the consequences of hypothetically moving or rotating an object? We introduce MindEdit-Bench, a benchmark of six spatial reasoning tasks built from three-photo smartphone triplets of newly captured indoor scenes via an automatic in-the-wild 3D scene-graph extraction pipeline. Four tasks probe perception and perspective transformation over observed structure; two new tasks, L4 (spatial editing) and L5 (cross-view visibility editing), probe object-level counterfactual reasoning, where correct answers are absent from all input images. Each question provides 8-24 structured answer choices, enabling answer-letter-level diagnosis of spatial and fallback errors. The benchmark covers 120 private indoor scenes not drawn from public datasets, reducing public-data pretraining-overlap risk. Across 15 VLMs on 1,003 human-verified questions, task-wise mean VLM accuracy is only 8%-31%, versus 81%-97% human majority-vote accuracy. The pooled human--best-VLM gap is 53 pp, with at least 39 pp on every task. The structured answer space further reveals non-uniform failures, including weaker camera-depth-axis inference and fallback behavior on difficult visibility-editing cases.
arXiv:2511.17731v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential in multimodal large language models (MLLMs) remains largely untapped, hindered by the absence of large-scale datasets that capture the rich, spatially grounded reasoning intrinsic to visual understanding. Existing visual-CoT resources are typically small, domain-specific, or lack the human-like stepwise structure necessary for compositional visual reasoning. In this paper, we introduce VisReason, a large-scale dataset designed to advance visual Chain-of-Thought reasoning. VisReason comprises 489K annotated examples spanning four diverse domains, each featuring multi-round, human-like rationales that guide MLLMs through interpretable visual reasoning steps. Building upon this, we curate VisReason-Pro, a 165K subset produced with a stronger expert-level GPT annotator, enriched with detailed reasoning traces and 3D spatial grounding via depth-informed annotations. Fine-tuning the state-of-the-art Qwen2.5-VL model on VisReason and VisReason-Pro yields substantial improvements in step-by-step visual reasoning accuracy, interpretability, and cross-benchmark generalization. These results demonstrate that VisReason equips MLLMs with more systematic and generalizable reasoning capabilities. We envision VisReason as a cornerstone for cultivating human-like visual reasoning, paving the way toward the next generation of multimodal intelligence.
arXiv:2603.06852v2 Announce Type: replace Abstract: Sparse-view computed tomography (CT) is critical for reducing radiation exposure to patients. Recent advances in radiative 3D Gaussian Splatting (3DGS) have enabled fast and accurate sparse-view CT reconstruction. Despite these algorithmic advancements, practical reconstruction fidelity remains fundamentally bounded by the quality of the captured data, raising the crucial yet underexplored problem of X-ray active view selection. Existing active view selection methods are primarily designed for natural-light scenes and fail to capture the unique geometric ambiguities and physical attenuation properties inherent in X-ray imaging. In this paper, we present Perturbed Gaussian Ensemble, an active view selection framework that integrates uncertainty modeling with sequential decision-making, tailored for X-ray Gaussian Splatting. Specifically, we identify low-density Gaussian primitives that are likely to be uncertain and apply stochastic density scaling to construct an ensemble of plausible Gaussian density fields. For each candidate projection, we measure the structural variance of the ensemble predictions and select the one with the highest variance as the next best view. Extensive experimental results on arbitrary-trajectory CT benchmarks demonstrate that our density-guided perturbation strategy effectively eliminates geometric artifacts and consistently outperforms existing baselines in progressive tomographic reconstruction under unified view selection protocols.