Forskningsradar

Science Journals

Peer-reviewade publikationer — 64018 artiklar

Reexamination of collisional ionization cross sections including double photoionization processes
arXiv:2607.00875v1 Announce Type: new Abstract: Collisional ionization (CI) cross sections in dense plasmas remain difficult to constrain due to uncertainties in plasma conditions and the overlapping spectral signatures of competing atomic processes. The use of x-ray free electron lasers (XFELs) to both heat and probe solid-density targets has significantly advanced the field by eliminating assumptions about ion density. However, questions remain regarding collisional cross sections, suprathermal electron evolution and competing atomic processes. In this work, we revisit experimental data from XFEL-heated aluminum, previously analyzed using collisional radiative models that did not treat the degenerate electron distribution and atomic processes self consistently. We present a new analysis using BibBarT which dynamically evolves non-thermal electron populations and explicitly includes degeneracy effects. Furthermore, we incorporate an important atomic process recently observed in plasma state that mimic signatures of CI, shake-off. Our results show that including shake-off processes improves agreement with observed emission features, and lowering recombination rates further improves the agreement with data -- indicating a possible overestimate of three-body recombination in these conditions.
Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning
arXiv:2606.24548v3 Announce Type: replace Abstract: Text-to-image (T2I) generation models have achieved remarkable progress in producing visually realistic images from natural language prompts. Yet it remains unclear whether their success reflects genuine causal understanding or sophisticated pattern matching over visual-textual correlations. Inspired by Russell's inductivist turkey, we introduce Counterfactual-World (CF-World), a counterfactual benchmark designed to investigate whether text-to-image models can generate images under rules that systematically contradict real-world priors. CF-World organizes each scenario into three progressive levels: factual generation under ordinary world knowledge, explicit counterfactual generation with direct visual instructions, and implicit counterfactual generation requiring causal deduction from altered rules. We evaluate both open-source and closed-source T2I models using a Vision Language Model (VLM)-based evaluator (CF-Eval). Furthermore, we introduce two metrics: Prior Resistance Rate (PRR), which measures a models' ability to overcome entrenched real-world priors, and Reasoning Retention Rate (RRR), which assesses whether models can maintain reasoning-dependent counterfactual generation without explicit visual cues. Experiments show that all models exhibit sharp degradation from factual to counterfactual settings. Further analyses suggest that these failures arise because current T2I models encode world knowledge and visual appearances as tightly coupled patterns. Consequently, their heavy reliance on frequent visual co-occurrences within the training data forces them to default to familiar commonsense priors when tasked with rendering counterfactual worlds.
OmniFall: From Staged Through Synthetic to Wild, A Unified Multi-Domain Dataset for Robust Fall Detection
arXiv:2505.19889v3 Announce Type: replace Abstract: Visual fall detection models are usually trained on small, staged datasets. Their real-world utility remains unclear; such data lacks diversity and evaluation protocols differ from paper to paper. We propose OmniFall, a unified benchmark of 15k videos (80 hours) with frame-level annotations in a single 16-class taxonomy. It spans three domains: OF-Staged unifies eight staged datasets with cross-subject and cross-view splits; OF-Synthetic adds 12k videos (17 h) with controlled demographic and environmental diversity; and OF-In-the-Wild provides a test-only set of genuine accident videos. We evaluate fine-tuned models as well as much larger zero-shot multimodal LLMs. On in-the-wild fall events, both do comparably well. The clinically critical fallen state is where they part: zero-shot models keep confusing fallen with lying, whereas models fine-tuned on synthetic data with explicit fallen-state scenes do substantially better. We release the unified annotations, the synthetic data, and the in-the-wild test set to foster the development of fall and fallen-state detectors for uncontrolled environments. Dataset: https://hf.co/datasets/simplexsigil2/omnifall
Toward a Unified Security and Privacy Framework for AI-Native 6G Networks
arXiv:2607.01019v1 Announce Type: new Abstract: Sixth Generation (6G) communication networks are expected to evolve into AI-native, highly autonomous ecosystems that integrate communication, computing, sensing, and artificial intelligence. While these capabilities enable unprecedented connectivity and intelligent services, they also create a highly heterogeneous security and privacy landscape that cannot be addressed through isolated, technology-specific solutions. This paper presents a comprehensive survey of security and privacy in AI-native 6G networks from a cross-layer perspective. We first examine the fragmentation of existing security and privacy approaches across emerging technologies, network architectures, AI systems, and standardization efforts, motivating the need for a unified security and privacy framework. Building upon this framework, we develop a cross-layer threat taxonomy encompassing infrastructure, network and architectural, AI, privacy, and security management domains, and analyze representative threats across key AI-native 6G technologies. Furthermore, we map these threats to corresponding cross-layer countermeasures, including standards harmonization as a security function, and identify critical research gaps and future priorities for secure, interoperable, and trustworthy AI-native 6G ecosystems. Finally, we discuss future research directions toward realizing secure, privacy-preserving, resilient, and globally interoperable 6G networks. This survey provides researchers, practitioners, and standardization communities with a holistic foundation for the design, evaluation, and deployment of trustworthy AI-native 6G systems.
Wasserstein Contraction of Coordinate Ascent Variational Inference
arXiv:2605.30253v3 Announce Type: replace-cross Abstract: We study the non-asymptotic contraction in Wasserstein distance of the sequential, parallel, and random-scan coordinate ascent variational inference algorithms. This is shown to hold under a functional smoothness condition of the optimality maps and a transportation-information inequality at their fixed points. Our results are sharp and general, and as opposed to those based on global strong log-concavity assumptions, they allow for local convergence on smooth, non-smooth, and discrete manifolds, including within the context of data augmentation. We consider many applications in statistical physics and Bayesian statistics. These include pairwise Markov Random field models such as Ising and Curie-Weiss, unbalanced Bayesian Gaussian Mixture Models, high-dimensional Bayesian Probit Regression, and high-dimensional Logistic Regression with P\'olya--Gamma random variables (i.e. Jaakkola-Jordan's algorithm). In many of these models, these represent the first available convergence results of their kind.
Exponential Low-Regularity Parareal Algorithms for Nonlinear Schr\"odinger Equations
arXiv:2607.00384v1 Announce Type: new Abstract: The parareal algorithm is one of the most widely studied parallel-in-time methods for the numerical approximation of time-dependent problems. For non-diffusive equations, however, standard parareal methods may converge slowly or even become unstable due to the absence of damping, while nonlinear interactions can transfer and amplify phase errors across Fourier modes. In this work, we consider the nonlinear Schr\"odinger equation (NLS) as a representative non-diffusive model and analyze parareal algorithms with an exact fine propagator, with particular emphasis on the design of suitable coarse propagators. We establish a general convergence framework, valid for solutions with limited regularity, under stability and local truncation error assumptions on the coarse propagator. These assumptions are verified for selected exponential low-regularity integrators designed for one-dimensional quadratic and cubic NLS equations, which achieve optimal approximation orders without derivative loss. To the best of our knowledge, this is the first construction of parareal algorithms for NLS equations that are provably linearly convergent, with a contraction factor proportional to the coarse time-step size even for solutions of limited regularity. Numerical experiments on quadratic, cubic, and quintic NLS equations demonstrate rapid convergence and improved performance over parareal variants using classical coarse propagators, including Lie and Strang splitting methods and first- and third-order exponential Runge--Kutta integrators.
Self-selected phase-matched second harmonic generation in nonlinear optical materials: from phenomenon to applications
arXiv:2607.00657v1 Announce Type: new Abstract: Self-selected phase-matched second harmonic generation is introduced as an all-optical probe of refractive-index dispersion in birefringent nonlinear optical materials. Rather than requiring wavelength or angular tuning, the exposure with a spectrally broad, intense ultrashort pulse allows the material to self-select the fundamental spectral component that satisfies the type-I noncritical phase-matching condition. This produces a narrow peak in the second harmonic spectrum whose position is governed by the refractive indices and is therefore highly sensitive to material parameters that affect the optical dispersion. We demonstrate the application of this phenomenon for the optical inspection of stoichiometry and temperature gradients in technologically relevant lithium niobate, as well as composition inhomogeneities in newly grown lithium niobate-tantalate solid solutions. These results establish self-selected phase-matched second harmonic generation as a rapid, non-contact method for inspecting nonlinear optical materials, with potential relevance for bulk crystals, wafers, and thin-film platforms.
Physics Informed Neural Networks for Nonlinear Delay Differential Equations
arXiv:2607.00380v1 Announce Type: new Abstract: In this paper we propose a novel physics-informed neural network framework for solving general first-order delay differential equations. Our approach combines a differentiable history switch, a trial-solution formulation that explicitly enforces history constraints, and a segmented collocation strategy to stabilize gradient propagation across large temporal domains. The method enables a scalable and physics-consistent approximation of delay differential equation solutions while maintaining continuity across subintervals. Numerical experiments demonstrate the effectiveness of the proposed method.
Sparse LDL, LU, and inverse butterfly factorization via tree decomposition
arXiv:2504.20305v5 Announce Type: replace Abstract: While linear systems over general fields can be solved in matrix-multiplication time, the complexity of symmetric triangular factorization has received relatively little formal study. We give dense and sparse LDL algorithms for symmetric matrices over an arbitrary field. Both algorithms leverage pivoted (rank-revealing) LU on off-diagonal blocks of a saddle-point form of a general symmetric matrix. For an $n\times n$ matrix, this yields an $O(n^\omega)$ dense LDL algorithm, where $n\times n$ matrix multiplication is assumed to cost $O(n^\omega)$ with $\omega>2$. For sparse matrices whose graph has treewidth $\tau$, we provide an implicit LDL in $O(n\tau^{\omega-1})$ time, and an explicit LDL whenever the rank deficiency is $O(\tau)$. We give analogous results for sparse LU via a standard off-diagonal embedding. We also obtain bounds on work, storage, and parallel-depth in terms of the dense $\tau\times\tau$ kernels executed at each bag in a tree decomposition. Finally, in the full-rank bounded-treewidth setting, we prove that $A^{-1}$ has complementary low-rank structure and admits an exact butterfly factorization with rank $O(\tau)$.
Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data Annotations
arXiv:2504.20490v2 Announce Type: replace Abstract: The Single-Program Multiple-Data (SPMD) paradigm provides a unified abstraction to annotate various parallel dimensions in distributed deep learning (DL) training. With SPMD, users can write training programs from the viewpoint of a single device, and the system will automatically deduce the tensor sharding and communication patterns. However, with the recent development in large-scale DL models, distributed training exhibits spatial and temporal workload heterogeneity, arising from both device disparities (e.g., mixed hardware, failures) and data variations (e.g., uneven sequence lengths). Such heterogeneity violates SPMD's assumption of symmetric workload partitioning, which restricts its ability to express and optimize heterogeneous parallel strategies effectively. To address this, we propose HSPMD within the Hetu v2 system to achieve general and scalable DL training. HSPMD extends SPMD's declarative annotations to support asymmetric sharding and composes standard communication primitives for hierarchical communication, all while retaining the simplicity of a single-device programming model. HSPMD handles spatial heterogeneity through progressive graph specialization, enabling device-specific execution logic, and addresses temporal heterogeneity via dynamic graph switching. Evaluations on (a) heterogeneous devices, (b) unstable devices, and (c) mixed-length data scenarios show that HSPMD matches or outperforms specialized systems, providing a flexible and efficient solution for modern distributed DL training.
HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning
arXiv:2606.29126v2 Announce Type: replace Abstract: Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat messages as flat dense vectors detached from the structure of the observations they summarize. This design overlooks an important source of inductive bias in many cooperative environments, where observations naturally follow a hierarchy such as groups and entities. We propose \textsc{HiComm}, a plug-in communication module that grounds messages in the sender's hierarchical observation. \textsc{HiComm} is receiver-driven: the receiver issues a query, and the hierarchy is resolved through a three-stage decoding process that first selects a group, then a sender, and then an entity within that group, returning the corresponding feature slice as the message. This converts communication from unstructured vector transmission into structured information retrieval over the sender's observation hierarchy. We instantiate this mechanism with Straight-Through Gumbel-Softmax for differentiable discrete selection and a lightweight shared projection design that attaches to standard MARL pipelines. Experiments across cooperative MARL tasks with different observation structures and coordination demands show that \textsc{HiComm} matches or outperforms representative learned communication baselines while reducing communication volume by up to $23\times$ per receiver per episode.
Mobile Base Station Positioning in Smart Ports Based on Kriged Sparse Measurements and Obstacle Inference
arXiv:2607.00709v1 Announce Type: new Abstract: Smart-port wireless networks suffer from dynamic radio blockage caused by container stacks and industrial structures, challenging efficient mobile integrated access and backhaul (MIAB) deployment. Existing approaches rely on obstacle maps, geometry information, or computationally intensive propagation models that limit adaptability. This paper presents DOCKING, a radio environment map (REM)-driven framework that converts sparse radio measurements into optimization-ready obstacle representations for MIAB deployment. The framework infers propagation-relevant obstacle abstractions from reconstructed REMs, eliminating the need for obstacle-geometry databases while relying only on known network parameters and sparse measurements. Reference signal received power (RSRP) and signal-to-interference-plus-noise ratio (SINR) observations are reconstructed using Ordinary Kriging (OKG), and dominant attenuation regions are approximated by compact cuboidal blockage models. The inferred geometry feeds a backhaul-aware optimization that determines MIAB placement, user equipment (UE) association, and backhaul selection. Under realistic smart-port conditions, REM reconstruction achieves prediction errors below 3 dB at the 90th percentile using only 15% spatial sampling, while obstacle characterization exceeds 85% true-positive coverage. Capacity gains reach 150% in sparse deployments, and a fast Genetic Algorithm converges within 5-15 s per network snapshot. A field campaign using real measurements validates the workflow, showing throughput trends consistent with optimization predictions. Results demonstrate that sparse radio measurements provide sufficient environmental awareness for practical obstacle-aware MIAB deployment in obstruction-prone industrial environments.
Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark
arXiv:2607.00710v1 Announce Type: new Abstract: Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones. This is especially limiting for small and medium-sized labs and startups that cannot afford to misallocate scarce resources. We argue that impactful dataset creation begins with a diagnosis: whether a research question is blocked by a data problem or an evaluation problem, and proceeds by selecting the minimal data operator(s) that closes the resulting gap, recording new data only when no cheaper operator(s) suffices. We analyze the evolution of major autonomous driving (AD) datasets through this lens and distill a strategic framework spanning gap identification, operator choice, sensor suite design, and annotation strategy. We ground the framework in a running case study of our KITScenes dataset family. The datasets are available at: https://kitscenes.com/
Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption
arXiv:2607.00712v1 Announce Type: new Abstract: Autoregressive (AR) streaming models have emerged as a powerful paradigm for long video generation. However, the linearly growing Key-Value (KV) cache poses a significant bottleneck, leading to memory overload and degraded inference throughput. A common compression method is to drop redundant KV tokens, which often breaks long-range dependencies, resulting in temporal flickering and identity loss. In this paper, we propose Instance-Specific Parametric Absorption (ISPA), a novel framework that shifts the KV cache compression from discarding to distilling. The core idea is to transit a subset of layers from Full-Attention (F-Layers) to memory-efficient Local-Attention (L-Layers) by "absorbing" historical context into the model's weights. Specifically, during a brief warmup phase, ISPA monitors the output discrepancy between global and local attention. At the transition point, we solve a closed-form least-squares problem to compute an instance-specific weight modulation that compensates for the missing history. Experiments across architectures (1.3B to 14B) demonstrate that ISPA can remove up to 50\% of the KV cache with near-lossless visual quality. We hope this perspective encourages future work to explore parametric memory consolidation beyond external token-level cache management for streaming generative models.
Self-conditioned Flow Map Language Models via Fixed-point Flows
arXiv:2607.00714v1 Announce Type: new Abstract: Self-conditioning is a core technique that enhances continuous flow-based language models, where the model learns to denoise generated text by conditioning on its own denoising estimate. While empirically successful, its performance improvements are poorly understood. Moreover, there is growing interest in the use of few-step generators based on flow maps, for which how to leverage self-conditioning is unclear. Here, we show that flow language models with self-conditioning solve a fixed-point iteration that bootstraps the performance of the learned denoiser. We use this viewpoint to formulate fixed-point flows, a two-dimensional class of self-conditioned flows, where the first dimension represents the flow process and the second represents the fixed-point iteration. We show that fixed-point flows define valid flow maps, and show that they can be distilled from self-conditioned flow models by compressing both fixed-point iterations and the flow process, the former with fixed-point distillation and the latter with flow map distillation. Our resulting flow map language model, FMLM$^\star$, outperforms state-of-the-art self-conditioned models and few-step models in one- and few-step generation on OpenWebText. Code is available at https://github.com/Ugness/self-conditioned-fmlm.
Fractal-Fractional HIV Dynamics with Mittag-Leffler Kernel: Analysis, Stability, and Numerical Simulations
arXiv:2607.00061v1 Announce Type: new Abstract: In this paper, a fractal--fractional HIV model with the Mittag--Leffler kernel is proposed using the Atangana--Baleanu--Caputo operator to capture the memory and hereditary properties of the disease dynamics. The existence and uniqueness of the solutions are investigated using suitable analytical techniques, and the Hyers--Ulam stability analysis is carried out to verify the stability behavior of the proposed system. For the numerical simulations, the Newton polynomial approximation method together with the Atangana--Toufik numerical scheme is employed to obtain approximate solutions for different parameter settings. Furthermore, several visualization techniques, including sensitivity heatmap representation and tornado diagram analysis, are utilized to study the influence of model parameters on the HIV dynamics. The obtained numerical results demonstrate that the proposed fractal--fractional framework provides an effective and reliable approach for analyzing the transient and long-term behavior of HIV transmission dynamics.
Structure-preserving dynamical low-rank approximation for parametric elastic guided waves
arXiv:2606.30469v2 Announce Type: replace Abstract: Elastic guided waves are widely used in Structural Health Monitoring (SHM). In many-query settings, the computational cost of high-fidelity simulations motivates the use of projection-based reduced order modeling (ROM). However, the transport-dominated and dispersive nature of guided waves challenges static linear subspaces. In addition, preserving the Hamiltonian structure of the equations for energy conservation necessitates dedicated projection techniques. While the Dynamical Low Rank Approximation (DLRA) has proven effective for other wave equations, its application to elastic guided waves in SHM has remained unexplored. In this work, we introduce a structure-preserving parametric ROM framework that leverages the DLRA in an off-line/on-line strategy. During the off-line stage, a time-dependent symplectic reduced basis is constructed from training simulations. For a simplified class of parameter dependencies, we derive a closed-form solution of the nonlinear basis evolution equation. This analytical result yields a closed-form, energy-preserving reduced propagator during wave propagation, eliminating on-line time integration after the loading phase. We validate our approach on a 2D elasticity problem featuring dispersive guided waves interacting with a damage. The results demonstrate high compression ratios (rank $\sim 10-30$), low full field reconstruction errors ($\sim 10^{-3}-10^{-2}$), speedups of two to three orders of magnitude, and excellent long-time energy conservation.
KGS-GCN: Kinematics-Driven Gaussian Splatting and Probabilistic Topology for Skeleton-Based Action Recognition
arXiv:2603.16943v2 Announce Type: replace Abstract: Skeleton-based action recognition is widely applied in sensor-based systems, including human-computer interaction and intelligent surveillance. However, typical sensors produce sparse and discrete joint coordinates, often leading to the loss of fine-grained spatiotemporal information during dynamic movements. Furthermore, predefined physical topologies restrict modeling potential long-range dependencies. To address these challenges, we propose KGS-GCN, which integrates kinematics-driven Gaussian splatting and probabilistic topology within a graph convolutional network. A Gaussian splatting module constructs anisotropic covariance matrices by extracting instantaneous joint velocity vectors, rendering sparse skeleton sequences into multi-view continuous heatmaps rich in spatiotemporal semantics. Additionally, a probabilistic topology construction strategy transcends physical connectivity limitations by utilizing the Bhattacharyya distance to quantify statistical correlations between joint Gaussian distributions, generating an adaptive prior adjacency matrix. Finally, the lightweight multi-view rendering branch and topological GCN backbone are unified through a visual context gating mechanism, enabling seamless fusion of continuous dynamic cues with structural priors while maintaining high computational efficiency, requiring only 1.4M parameters and 1.3 GFLOPs. Extensive experiments on multiple benchmark datasets demonstrate that KGS-GCN significantly enhances the modeling of complex spatiotemporal dynamics and achieves competitive performance at low computational cost, establishing an efficient paradigm for improving the perceptual robustness of low-fidelity sensor data.
Robust Operational Space Control with Conformal Disturbance Bounds for Safe Redundant Manipulation
arXiv:2607.00424v1 Announce Type: new Abstract: Redundant robotic manipulators operating in constrained and human-interactive environments require accurate task-space tracking together with rigorous safety guarantees under dynamic uncertainties. Classical operational space computed torque controller (OSCTC) relies on accurate dynamic models and degrades in the presence of disturbances. In contrast, the data-driven paradigm of residual learning approximates disturbances as functions learned from full-state measurements, which are often noisy in practice, lack rigorous theoretical guarantees, and introduce additional design complexity. This paper proposes a robust OSCTC framework that integrates an extended state observer (ESO) with conformal prediction to combine model-based robustness and data-driven adaptability. The ESO estimates lumped disturbances directly in operational space without requiring full-state measurements as in residual learning, and a robust control barrier function (CBF) is constructed to enforce safety under uncertainty. However, robust CBFs require a known disturbance-variation bound to guarantee absolute safety, which often leads to conservatism in practice. To address this limitation, we further employ a sliding-window conformal prediction mechanism to estimate the bound online in a distribution-free manner, thereby achieving practical probabilistic safety guarantees. Experiments on a 7-DoF Franka Research 3 manipulator demonstrate millimeter-level tracking accuracy and real-time safe control at 1~kHz under various disturbances.
HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding
arXiv:2607.00428v1 Announce Type: new Abstract: CLIP (Contrastive Language-Image Pre-training) has become a de facto paradigm for image-text alignment, but it struggles with long-context descriptions (>77 tokens) due to absolute positional encoding and pretraining on short captions. In long contexts, sentences are often reordered, summarized, or partially omitted. Although prior works extend CLIP with longer positional encodings, they often suffer from degraded image-text alignment under such text perturbations. We attribute this limitation to the Euclidean contrastive objective, which enforces strict one-to-one matching and lacks explicit mechanisms for modeling hierarchical relationships between global context and its constituent elements. To address this issue, we propose HyFL-CLIP, a hyperbolic fine-tuning framework that distills the well-established text-image alignment learned in Euclidean CLIP into hyperbolic space via cross-manifold similarity distillation, leveraging its geometry to capture hierarchical and entailment relations. Our method models hierarchical semantics by linking summarized token-wise features, long-context descriptions, constituent short textual components, and images, capturing part-whole relationships via hyperbolic entailment with Einstein midpoint aggregation. Experiments on diverse benchmarks, including long-context cross-modal retrieval, cross-modal retrieval with caption perturbations, intra-modality retrieval, and short-text cross-modal retrieval, show that HyFL-CLIP achieves more robust long-context understanding. In particular, it yields up to 19.5% improvement in long-text cross-modal retrieval under textual perturbations over the best prior method. We also show HyFL-CLIP can be seamlessly integrated into other model frameworks by applying it to Stable Diffusion XL (SDXL).
Caption Bottleneck Models
arXiv:2607.00578v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) provide interpretability by routing predictions through a layer of human-understandable concepts. However, defining an optimal concept set for a specific dataset remains an open challenge. Existing approaches rely on expensive expert annotations or LLM-generated lists based solely on class names. Even "open-vocabulary" variants typically depend on static concept sets, which restrict discovery and introduce label bias. Furthermore, traditional CBMs often suffer from information leakage, where unmodeled visual features bypass the bottleneck and compromise the integrity of the explanations. To overcome these limitations, we propose Caption Bottleneck Models (CaBM), a framework that circumvents the need for predefined concept sets by replacing rigid concept layers with free-form natural language. By representing images via LMM-generated captions and training a classifier strictly on this text, CaBM ensures a leakage-free architecture by construction. Additionally, by analyzing the text classifier post-training, CaBM autonomously discovers high-quality, dataset-specific concepts. Our results across fine- and coarse-grained benchmarks demonstrate that CaBM achieves competitive accuracy while preserving interpretability without the constraints of external dictionaries or manual labeling.
CellPrior-Net: Prior-Guided Nuclei Detection and Classification for H&E Whole-Slide Images
arXiv:2607.00802v1 Announce Type: new Abstract: Accurate nuclei detection and classification in hematoxylin and eosin (H and E) whole-slide images (WSIs) is a key task in computational pathology, particularly for quantitative analysis of the tumor microenvironment. However, this task remains highly challenging due to variations in nuclei morphology, staining procedures, scanners, organs, magnifications, and WSI artifacts. In addition, many existing pipelines rely on computationally demanding architectures and post-processing procedures, making gigapixel WSI analysis time consuming. In this work, CellPriorNet (CP Net) is proposed, an efficient nuclei detection and classification pipeline that utilizes a lightweight convolutional neural network architecture and hematoxylin (H) channel as prior information to enhance nuclei-aware feature learning. Extensive benchmarking was conducted against state of the art pipelines on 8 public and private datasets (total:10.4M nuclei) obtained from different organs, scanners, magnifications, and clinical centers. Experimental results demonstrate that CP Net achieves comparable performance while significantly reducing inference time. Furthermore, CellQuant Net was introduced, an end to end nuclei quantification pipeline, that integrates a quality assessment (QA) model to exclude regions with artifacts, followed by CP-Net cell detection and classification. The pipeline is publicly available on GitHub, and provides a potentially efficient and scalable framework for downstream computational pathology applications.
Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold
arXiv:2607.00647v1 Announce Type: new Abstract: Training-free guidance (TFG) steers a pretrained diffusion model toward a desired attribute at inference. To be effective, this guidance must be applied from the earliest, high-noise steps of sampling. Because its objective (a classifier or energy) is defined on clean images, $\epsilon$- and $v$-prediction models must first estimate the clean image $\hat{x}$ from the noisy state at each step, and the accuracy of that estimate determines how easily guidance drifts off the data manifold. $x$-prediction, a recent alternative, outputs the clean image directly, removing this source of error even at high noise. This is our motivation. We provide a theoretical analysis of how each prediction target shapes this accuracy, and introduce guided-class FID (Child FID), a metric that exposes the manifold damage standard evaluation misses. Experiments on a new fine-grained bird benchmark and on style transfer confirm that $x$-prediction keeps guided samples on the manifold most reliably, making it the strongest foundation for training-free guidance. Code is available at https://github.com/ManLuML/on-manifold-tfg
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
arXiv:2605.31597v3 Announce Type: replace Abstract: Measuring structured object understanding in vision foundation models remains challenging due to inconsistent evaluation protocols and limited part-level supervision. Semantic correspondence (SC) evaluates this capability by testing whether object parts can be matched across instances and categories under large variations in appearance, viewpoint, and geometry. To enable a systematic SC evaluation, we introduce SOCO, a new benchmark for Semantic Object Correspondence that introduces a taxonomy of correspondence types and provides consistent, functionally meaningful keypoint annotations across 100 categories and over 1M correspondence pairs. In addition, SOCO includes keypoint language descriptions, enabling the evaluation of large vision-language models (LVLMs) and their fine-grained part-level understanding. Comprehensive experiments reveal that (i) vision foundation backbones encode strong semantic structure but transfer correspondences poorly across related categories and only partially capture object-part position, (ii) LVLMs are stronger at text-prompted part localization than at visual-reference cross-image matching, exposing a gap between language-grounded localization and fine-grained visual correspondence, and (iii) correspondence performance predicts performance on dense downstream tasks, including segmentation, tracking, 3D pose estimation, and 3D detection, more strongly than ImageNet classification. Together, these findings position SOCO as a benchmark for structured, part-level representation quality in vision and multimodal foundation models.
Learning-based control of a single-DOF Aero system
arXiv:2607.00640v1 Announce Type: new Abstract: This paper presents a learning-based control framework that integrates feedback linearization with reinforcement learning for the adaptive control of nonlinear mechatronic systems. The control law is derived using Lyapunov stability analysis, ensuring closed-loop stability in the presence of modeling uncertainties and external disturbances. Feedback linearization serves as the main control framework, while a reinforcement learning component estimates and compensates for unmodeled dynamics and disturbances online. The learning module is based on the REINFORCE-with-baseline algorithm, which improves learning efficiency by reducing the variance of policy-gradient estimates and enabling stable policy updates during adaptation. The proposed controller is evaluated on a single-degree-of-freedom rotor-based AERO system. Results from simulations demonstrate accurate trajectory tracking, fast adaptation, and strong robustness against parameter variations and external disturbances. Overall, the proposed approach combines the analytical guarantees of Lyapunov-based control with the adaptability of reinforcement learning, providing an effective solution for controlling nonlinear mechatronic systems.