arXiv:2607.00875v1 Announce Type: new Abstract: Collisional ionization (CI) cross sections in dense plasmas remain difficult to constrain due to uncertainties in plasma conditions and the overlapping spectral signatures of competing atomic processes. The use of x-ray free electron lasers (XFELs) to both heat and probe solid-density targets has significantly advanced the field by eliminating assumptions about ion density. However, questions remain regarding collisional cross sections, suprathermal electron evolution and competing atomic processes. In this work, we revisit experimental data from XFEL-heated aluminum, previously analyzed using collisional radiative models that did not treat the degenerate electron distribution and atomic processes self consistently. We present a new analysis using BibBarT which dynamically evolves non-thermal electron populations and explicitly includes degeneracy effects. Furthermore, we incorporate an important atomic process recently observed in plasma state that mimic signatures of CI, shake-off. Our results show that including shake-off processes improves agreement with observed emission features, and lowering recombination rates further improves the agreement with data -- indicating a possible overestimate of three-body recombination in these conditions.
Science Journals
arXiv:2606.24548v3 Announce Type: replace Abstract: Text-to-image (T2I) generation models have achieved remarkable progress in producing visually realistic images from natural language prompts. Yet it remains unclear whether their success reflects genuine causal understanding or sophisticated pattern matching over visual-textual correlations. Inspired by Russell's inductivist turkey, we introduce Counterfactual-World (CF-World), a counterfactual benchmark designed to investigate whether text-to-image models can generate images under rules that systematically contradict real-world priors. CF-World organizes each scenario into three progressive levels: factual generation under ordinary world knowledge, explicit counterfactual generation with direct visual instructions, and implicit counterfactual generation requiring causal deduction from altered rules. We evaluate both open-source and closed-source T2I models using a Vision Language Model (VLM)-based evaluator (CF-Eval). Furthermore, we introduce two metrics: Prior Resistance Rate (PRR), which measures a models' ability to overcome entrenched real-world priors, and Reasoning Retention Rate (RRR), which assesses whether models can maintain reasoning-dependent counterfactual generation without explicit visual cues. Experiments show that all models exhibit sharp degradation from factual to counterfactual settings. Further analyses suggest that these failures arise because current T2I models encode world knowledge and visual appearances as tightly coupled patterns. Consequently, their heavy reliance on frequent visual co-occurrences within the training data forces them to default to familiar commonsense priors when tasked with rendering counterfactual worlds.
arXiv:2505.19889v3 Announce Type: replace Abstract: Visual fall detection models are usually trained on small, staged datasets. Their real-world utility remains unclear; such data lacks diversity and evaluation protocols differ from paper to paper. We propose OmniFall, a unified benchmark of 15k videos (80 hours) with frame-level annotations in a single 16-class taxonomy. It spans three domains: OF-Staged unifies eight staged datasets with cross-subject and cross-view splits; OF-Synthetic adds 12k videos (17 h) with controlled demographic and environmental diversity; and OF-In-the-Wild provides a test-only set of genuine accident videos. We evaluate fine-tuned models as well as much larger zero-shot multimodal LLMs. On in-the-wild fall events, both do comparably well. The clinically critical fallen state is where they part: zero-shot models keep confusing fallen with lying, whereas models fine-tuned on synthetic data with explicit fallen-state scenes do substantially better. We release the unified annotations, the synthetic data, and the in-the-wild test set to foster the development of fall and fallen-state detectors for uncontrolled environments. Dataset: https://hf.co/datasets/simplexsigil2/omnifall
arXiv:2607.00560v1 Announce Type: new Abstract: We develop a multilevel stochastic-gradient neural solver for boundary integral equations of the second kind. The unknown density is represented by a multilayer perceptron, trained by minimizing the Nystr\"om-discretized residual on a ladder of refining quadrature grids, each level warm-started from the parameters of the previous one. Each step requires only dense matrix-vector products on mini-batches of collocation rows and network passes, operations that map directly onto GPU hardware. The residual contraction is governed by the empirical neural tangent kernel (NTK), the discrete sample of a single continuum kernel. On a fixed grid, training stalls once the residual concentrates in modes the network contracts slowly, the plateau described by the frequency principle; a spectral analysis explains, and experiments confirm, how refining the quadrature resolves more of the continuum kernel's spectrum and returns these modes to the optimizer's reach. Spectral bias, elsewhere an obstruction to neural network solvers, thus serves as the smoother of a multigrid-type iteration, with quadrature refinement in place of coarse-grid correction. Under a uniform regularity bound on the network, the total work is a constant multiple of the work on the finest grid, and the uniform conditioning of the discrete second-kind operator leaves the NTK as the sole rate-determining spectrum while converting the training residual into an a posteriori error bound. Experiments on interior Dirichlet Laplace/Poisson problems and exterior Neumann Helmholtz problems, using both parametric and signed-distance surface representations, demonstrate the effectiveness and efficiency of the proposed method compared with GMRES at comparable tolerances.
arXiv:2607.01019v1 Announce Type: new Abstract: Sixth Generation (6G) communication networks are expected to evolve into AI-native, highly autonomous ecosystems that integrate communication, computing, sensing, and artificial intelligence. While these capabilities enable unprecedented connectivity and intelligent services, they also create a highly heterogeneous security and privacy landscape that cannot be addressed through isolated, technology-specific solutions. This paper presents a comprehensive survey of security and privacy in AI-native 6G networks from a cross-layer perspective. We first examine the fragmentation of existing security and privacy approaches across emerging technologies, network architectures, AI systems, and standardization efforts, motivating the need for a unified security and privacy framework. Building upon this framework, we develop a cross-layer threat taxonomy encompassing infrastructure, network and architectural, AI, privacy, and security management domains, and analyze representative threats across key AI-native 6G technologies. Furthermore, we map these threats to corresponding cross-layer countermeasures, including standards harmonization as a security function, and identify critical research gaps and future priorities for secure, interoperable, and trustworthy AI-native 6G ecosystems. Finally, we discuss future research directions toward realizing secure, privacy-preserving, resilient, and globally interoperable 6G networks. This survey provides researchers, practitioners, and standardization communities with a holistic foundation for the design, evaluation, and deployment of trustworthy AI-native 6G systems.
arXiv:2605.30253v3 Announce Type: replace-cross Abstract: We study the non-asymptotic contraction in Wasserstein distance of the sequential, parallel, and random-scan coordinate ascent variational inference algorithms. This is shown to hold under a functional smoothness condition of the optimality maps and a transportation-information inequality at their fixed points. Our results are sharp and general, and as opposed to those based on global strong log-concavity assumptions, they allow for local convergence on smooth, non-smooth, and discrete manifolds, including within the context of data augmentation. We consider many applications in statistical physics and Bayesian statistics. These include pairwise Markov Random field models such as Ising and Curie-Weiss, unbalanced Bayesian Gaussian Mixture Models, high-dimensional Bayesian Probit Regression, and high-dimensional Logistic Regression with P\'olya--Gamma random variables (i.e. Jaakkola-Jordan's algorithm). In many of these models, these represent the first available convergence results of their kind.
arXiv:2607.00384v1 Announce Type: new Abstract: The parareal algorithm is one of the most widely studied parallel-in-time methods for the numerical approximation of time-dependent problems. For non-diffusive equations, however, standard parareal methods may converge slowly or even become unstable due to the absence of damping, while nonlinear interactions can transfer and amplify phase errors across Fourier modes. In this work, we consider the nonlinear Schr\"odinger equation (NLS) as a representative non-diffusive model and analyze parareal algorithms with an exact fine propagator, with particular emphasis on the design of suitable coarse propagators. We establish a general convergence framework, valid for solutions with limited regularity, under stability and local truncation error assumptions on the coarse propagator. These assumptions are verified for selected exponential low-regularity integrators designed for one-dimensional quadratic and cubic NLS equations, which achieve optimal approximation orders without derivative loss. To the best of our knowledge, this is the first construction of parareal algorithms for NLS equations that are provably linearly convergent, with a contraction factor proportional to the coarse time-step size even for solutions of limited regularity. Numerical experiments on quadratic, cubic, and quintic NLS equations demonstrate rapid convergence and improved performance over parareal variants using classical coarse propagators, including Lie and Strang splitting methods and first- and third-order exponential Runge--Kutta integrators.
arXiv:2607.00657v1 Announce Type: new Abstract: Self-selected phase-matched second harmonic generation is introduced as an all-optical probe of refractive-index dispersion in birefringent nonlinear optical materials. Rather than requiring wavelength or angular tuning, the exposure with a spectrally broad, intense ultrashort pulse allows the material to self-select the fundamental spectral component that satisfies the type-I noncritical phase-matching condition. This produces a narrow peak in the second harmonic spectrum whose position is governed by the refractive indices and is therefore highly sensitive to material parameters that affect the optical dispersion. We demonstrate the application of this phenomenon for the optical inspection of stoichiometry and temperature gradients in technologically relevant lithium niobate, as well as composition inhomogeneities in newly grown lithium niobate-tantalate solid solutions. These results establish self-selected phase-matched second harmonic generation as a rapid, non-contact method for inspecting nonlinear optical materials, with potential relevance for bulk crystals, wafers, and thin-film platforms.
arXiv:2607.00380v1 Announce Type: new Abstract: In this paper we propose a novel physics-informed neural network framework for solving general first-order delay differential equations. Our approach combines a differentiable history switch, a trial-solution formulation that explicitly enforces history constraints, and a segmented collocation strategy to stabilize gradient propagation across large temporal domains. The method enables a scalable and physics-consistent approximation of delay differential equation solutions while maintaining continuity across subintervals. Numerical experiments demonstrate the effectiveness of the proposed method.
arXiv:2504.20305v5 Announce Type: replace Abstract: While linear systems over general fields can be solved in matrix-multiplication time, the complexity of symmetric triangular factorization has received relatively little formal study. We give dense and sparse LDL algorithms for symmetric matrices over an arbitrary field. Both algorithms leverage pivoted (rank-revealing) LU on off-diagonal blocks of a saddle-point form of a general symmetric matrix. For an $n\times n$ matrix, this yields an $O(n^\omega)$ dense LDL algorithm, where $n\times n$ matrix multiplication is assumed to cost $O(n^\omega)$ with $\omega>2$. For sparse matrices whose graph has treewidth $\tau$, we provide an implicit LDL in $O(n\tau^{\omega-1})$ time, and an explicit LDL whenever the rank deficiency is $O(\tau)$. We give analogous results for sparse LU via a standard off-diagonal embedding. We also obtain bounds on work, storage, and parallel-depth in terms of the dense $\tau\times\tau$ kernels executed at each bag in a tree decomposition. Finally, in the full-rank bounded-treewidth setting, we prove that $A^{-1}$ has complementary low-rank structure and admits an exact butterfly factorization with rank $O(\tau)$.
arXiv:2504.20490v2 Announce Type: replace Abstract: The Single-Program Multiple-Data (SPMD) paradigm provides a unified abstraction to annotate various parallel dimensions in distributed deep learning (DL) training. With SPMD, users can write training programs from the viewpoint of a single device, and the system will automatically deduce the tensor sharding and communication patterns. However, with the recent development in large-scale DL models, distributed training exhibits spatial and temporal workload heterogeneity, arising from both device disparities (e.g., mixed hardware, failures) and data variations (e.g., uneven sequence lengths). Such heterogeneity violates SPMD's assumption of symmetric workload partitioning, which restricts its ability to express and optimize heterogeneous parallel strategies effectively. To address this, we propose HSPMD within the Hetu v2 system to achieve general and scalable DL training. HSPMD extends SPMD's declarative annotations to support asymmetric sharding and composes standard communication primitives for hierarchical communication, all while retaining the simplicity of a single-device programming model. HSPMD handles spatial heterogeneity through progressive graph specialization, enabling device-specific execution logic, and addresses temporal heterogeneity via dynamic graph switching. Evaluations on (a) heterogeneous devices, (b) unstable devices, and (c) mixed-length data scenarios show that HSPMD matches or outperforms specialized systems, providing a flexible and efficient solution for modern distributed DL training.
arXiv:2606.29126v2 Announce Type: replace Abstract: Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat messages as flat dense vectors detached from the structure of the observations they summarize. This design overlooks an important source of inductive bias in many cooperative environments, where observations naturally follow a hierarchy such as groups and entities. We propose \textsc{HiComm}, a plug-in communication module that grounds messages in the sender's hierarchical observation. \textsc{HiComm} is receiver-driven: the receiver issues a query, and the hierarchy is resolved through a three-stage decoding process that first selects a group, then a sender, and then an entity within that group, returning the corresponding feature slice as the message. This converts communication from unstructured vector transmission into structured information retrieval over the sender's observation hierarchy. We instantiate this mechanism with Straight-Through Gumbel-Softmax for differentiable discrete selection and a lightweight shared projection design that attaches to standard MARL pipelines. Experiments across cooperative MARL tasks with different observation structures and coordination demands show that \textsc{HiComm} matches or outperforms representative learned communication baselines while reducing communication volume by up to $23\times$ per receiver per episode.
arXiv:2607.00709v1 Announce Type: new Abstract: Smart-port wireless networks suffer from dynamic radio blockage caused by container stacks and industrial structures, challenging efficient mobile integrated access and backhaul (MIAB) deployment. Existing approaches rely on obstacle maps, geometry information, or computationally intensive propagation models that limit adaptability. This paper presents DOCKING, a radio environment map (REM)-driven framework that converts sparse radio measurements into optimization-ready obstacle representations for MIAB deployment. The framework infers propagation-relevant obstacle abstractions from reconstructed REMs, eliminating the need for obstacle-geometry databases while relying only on known network parameters and sparse measurements. Reference signal received power (RSRP) and signal-to-interference-plus-noise ratio (SINR) observations are reconstructed using Ordinary Kriging (OKG), and dominant attenuation regions are approximated by compact cuboidal blockage models. The inferred geometry feeds a backhaul-aware optimization that determines MIAB placement, user equipment (UE) association, and backhaul selection. Under realistic smart-port conditions, REM reconstruction achieves prediction errors below 3 dB at the 90th percentile using only 15% spatial sampling, while obstacle characterization exceeds 85% true-positive coverage. Capacity gains reach 150% in sparse deployments, and a fast Genetic Algorithm converges within 5-15 s per network snapshot. A field campaign using real measurements validates the workflow, showing throughput trends consistent with optimization predictions. Results demonstrate that sparse radio measurements provide sufficient environmental awareness for practical obstacle-aware MIAB deployment in obstruction-prone industrial environments.
arXiv:2607.00710v1 Announce Type: new Abstract: Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones. This is especially limiting for small and medium-sized labs and startups that cannot afford to misallocate scarce resources. We argue that impactful dataset creation begins with a diagnosis: whether a research question is blocked by a data problem or an evaluation problem, and proceeds by selecting the minimal data operator(s) that closes the resulting gap, recording new data only when no cheaper operator(s) suffices. We analyze the evolution of major autonomous driving (AD) datasets through this lens and distill a strategic framework spanning gap identification, operator choice, sensor suite design, and annotation strategy. We ground the framework in a running case study of our KITScenes dataset family. The datasets are available at: https://kitscenes.com/
Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption
arXiv:2607.00712v1 Announce Type: new Abstract: Autoregressive (AR) streaming models have emerged as a powerful paradigm for long video generation. However, the linearly growing Key-Value (KV) cache poses a significant bottleneck, leading to memory overload and degraded inference throughput. A common compression method is to drop redundant KV tokens, which often breaks long-range dependencies, resulting in temporal flickering and identity loss. In this paper, we propose Instance-Specific Parametric Absorption (ISPA), a novel framework that shifts the KV cache compression from discarding to distilling. The core idea is to transit a subset of layers from Full-Attention (F-Layers) to memory-efficient Local-Attention (L-Layers) by "absorbing" historical context into the model's weights. Specifically, during a brief warmup phase, ISPA monitors the output discrepancy between global and local attention. At the transition point, we solve a closed-form least-squares problem to compute an instance-specific weight modulation that compensates for the missing history. Experiments across architectures (1.3B to 14B) demonstrate that ISPA can remove up to 50\% of the KV cache with near-lossless visual quality. We hope this perspective encourages future work to explore parametric memory consolidation beyond external token-level cache management for streaming generative models.
arXiv:2607.00714v1 Announce Type: new Abstract: Self-conditioning is a core technique that enhances continuous flow-based language models, where the model learns to denoise generated text by conditioning on its own denoising estimate. While empirically successful, its performance improvements are poorly understood. Moreover, there is growing interest in the use of few-step generators based on flow maps, for which how to leverage self-conditioning is unclear. Here, we show that flow language models with self-conditioning solve a fixed-point iteration that bootstraps the performance of the learned denoiser. We use this viewpoint to formulate fixed-point flows, a two-dimensional class of self-conditioned flows, where the first dimension represents the flow process and the second represents the fixed-point iteration. We show that fixed-point flows define valid flow maps, and show that they can be distilled from self-conditioned flow models by compressing both fixed-point iterations and the flow process, the former with fixed-point distillation and the latter with flow map distillation. Our resulting flow map language model, FMLM$^\star$, outperforms state-of-the-art self-conditioned models and few-step models in one- and few-step generation on OpenWebText. Code is available at https://github.com/Ugness/self-conditioned-fmlm.
arXiv:2607.01103v1 Announce Type: new Abstract: Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-asa-Judge approaches. Whether such evaluators replicate clinical calibration and caution, however, remains untested. We introduce MedQADE, the first standardised open-response clinical benchmark for German, a major clinical language lacking native evaluation infrastructure, comprising 3,800 items annotated by ten practising physicians and nine Large Language Model (LLM) evaluators. The top-performing evaluator model, Gemini 3 Flash, reached alignment consistent with the physician ceiling (\k{appa} = 0.694 vs. \k{appa} = 0.709), though wide confidence intervals limit interpretation. Despite this statistical alignment, automated evaluators exhibited near-absent clinical metacognition: physicians scaled abstention with item difficulty, while frontier models assigned definitive scores in every case. We additionally quantified systematic lineage-dependent biases, where models preferentially scored architectural siblings, an effect independent of language. These results show that statistical alignment does not ensure clinical caution, and that evaluator independence requires explicit verification.
arXiv:2607.00263v1 Announce Type: cross Abstract: The Chromatic Sum problem asks, given a graph $G$ and an integer $k$, whether $G$ admits a colouring $c$ with sum $\sum_{v\in V}c(v) \leq k$. We study the complexity of Chromatic Sum on graph classes defined by some set of forbidden graphs. First, we show that three known frameworks fully classify the complexity of Chromatic Sum on $HH$-minor-free graphs and $HH$-topological-minor-free graphs for any set of graphs $HH$, and on $HH$-subgraph-free graphs for any finite set of graphs $HH$. To show this, we prove a new NP-completeness result for Chromatic Sum on certain subdivisions of planar subcubic graphs. Next, we consider other containment relations. We formalise a novel framework of problems that are NP-complete for planar graphs as well as for graphs of bounded independence number. For every problem in this framework, we obtain an almost complete complexity classification on $H$-induced-minor-free graphs, $H$-induced-topological-minor-free graphs, and $H$-free graphs for every graph $H$. We show that Chromatic Sum belongs to this framework, as do several other problems. We also define a more fine-grained framework for the induced subgraph relation. We apply this to obtain a complete complexity classification for Chromatic Sum on $H$-free graphs, as well as for several other problems. We justify the choice of this framework by proving that Chromatic Sum is NP-complete for graphs of clique-width at most $3$. This result complements a known polynomial-time result for graphs of clique-width at most $2$.
arXiv:2607.00061v1 Announce Type: new Abstract: In this paper, a fractal--fractional HIV model with the Mittag--Leffler kernel is proposed using the Atangana--Baleanu--Caputo operator to capture the memory and hereditary properties of the disease dynamics. The existence and uniqueness of the solutions are investigated using suitable analytical techniques, and the Hyers--Ulam stability analysis is carried out to verify the stability behavior of the proposed system. For the numerical simulations, the Newton polynomial approximation method together with the Atangana--Toufik numerical scheme is employed to obtain approximate solutions for different parameter settings. Furthermore, several visualization techniques, including sensitivity heatmap representation and tornado diagram analysis, are utilized to study the influence of model parameters on the HIV dynamics. The obtained numerical results demonstrate that the proposed fractal--fractional framework provides an effective and reliable approach for analyzing the transient and long-term behavior of HIV transmission dynamics.
arXiv:2606.30469v2 Announce Type: replace Abstract: Elastic guided waves are widely used in Structural Health Monitoring (SHM). In many-query settings, the computational cost of high-fidelity simulations motivates the use of projection-based reduced order modeling (ROM). However, the transport-dominated and dispersive nature of guided waves challenges static linear subspaces. In addition, preserving the Hamiltonian structure of the equations for energy conservation necessitates dedicated projection techniques. While the Dynamical Low Rank Approximation (DLRA) has proven effective for other wave equations, its application to elastic guided waves in SHM has remained unexplored. In this work, we introduce a structure-preserving parametric ROM framework that leverages the DLRA in an off-line/on-line strategy. During the off-line stage, a time-dependent symplectic reduced basis is constructed from training simulations. For a simplified class of parameter dependencies, we derive a closed-form solution of the nonlinear basis evolution equation. This analytical result yields a closed-form, energy-preserving reduced propagator during wave propagation, eliminating on-line time integration after the loading phase. We validate our approach on a 2D elasticity problem featuring dispersive guided waves interacting with a damage. The results demonstrate high compression ratios (rank $\sim 10-30$), low full field reconstruction errors ($\sim 10^{-3}-10^{-2}$), speedups of two to three orders of magnitude, and excellent long-time energy conservation.
arXiv:2606.30514v2 Announce Type: replace Abstract: Human image animation, which aims to generate a video of a reference subject following a provided action sequence, has received increasing research interest. With the development of diffusion-based/flow-based video foundation models, existing animation works have began to upgrade the guidance information from 2D skeleton/pose to 3D modeling conditions. Despite achieving reasonable results, these approaches face challenges in synthesizing trajectory-controllable human motion within natural scene under changed camera views. In this work, we present a scene-adaptive human image animation framework that controls both human motion and camera trajectories within a reconstructed 3D environment for video generation. To achieve this, we first develop a ground-adaptive 3D motion retargeting approach to enable user-friendly motion trajectory control adapting to the changes of elevations of ground and orientations automatically. Then we design a viewpoint-adaptive latent fusion mechanism to inject point-cloud geometric priors through scene-visibility masking into the generative process, providing precise guidance of viewpoint changes under camera control. Experiments on two standard human image animation benchmark datasets demonstrate remarkable improvements of our method over the state of the arts in related video generation metics. Project page: https://robinhood256100.github.io/web-disp
arXiv:2603.16943v2 Announce Type: replace Abstract: Skeleton-based action recognition is widely applied in sensor-based systems, including human-computer interaction and intelligent surveillance. However, typical sensors produce sparse and discrete joint coordinates, often leading to the loss of fine-grained spatiotemporal information during dynamic movements. Furthermore, predefined physical topologies restrict modeling potential long-range dependencies. To address these challenges, we propose KGS-GCN, which integrates kinematics-driven Gaussian splatting and probabilistic topology within a graph convolutional network. A Gaussian splatting module constructs anisotropic covariance matrices by extracting instantaneous joint velocity vectors, rendering sparse skeleton sequences into multi-view continuous heatmaps rich in spatiotemporal semantics. Additionally, a probabilistic topology construction strategy transcends physical connectivity limitations by utilizing the Bhattacharyya distance to quantify statistical correlations between joint Gaussian distributions, generating an adaptive prior adjacency matrix. Finally, the lightweight multi-view rendering branch and topological GCN backbone are unified through a visual context gating mechanism, enabling seamless fusion of continuous dynamic cues with structural priors while maintaining high computational efficiency, requiring only 1.4M parameters and 1.3 GFLOPs. Extensive experiments on multiple benchmark datasets demonstrate that KGS-GCN significantly enhances the modeling of complex spatiotemporal dynamics and achieves competitive performance at low computational cost, establishing an efficient paradigm for improving the perceptual robustness of low-fidelity sensor data.
arXiv:2606.30879v2 Announce Type: replace Abstract: Accurately assessing a programmer's skill level is critical for hiring, team composition, and performance evaluation in the software industry. Conventional methods, such as coding tests or interviews, often fail to capture the full spectrum of cognitive abilities underlying programming expertise. This study explores using electroencephalography (EEG) and machine learning to investigate neural correlates of programming skill. We analyzed an existing EEG dataset recorded during code comprehension from 37 programmers with 1 to 30 years of experience (8.1 +/- 6.3 years) to examine relationships between neural activity and expertise. Additionally, we conducted classification experiments using Random Forest classifiers with diverse features for binary (experts vs. novices) and multi-class (experts, intermediates, novices) setups. We identified EEG features and brain regions associated with programming expertise. Specifically, EEG entropy showed the strongest correlation with skill level. Furthermore, experts' brains were characterized by highly localized centro-frontal activation, whereas frontal activation in other groups was part of a more distributed network. Regarding classification, our setup achieved an average accuracy of 91.83% (binary) and 78.15% (multi-class) in stratified 10-fold cross-validation, while leave-one-subject-out validation achieved 85.00% and 58.80%, respectively. Individual frequency bands outperformed full-spectrum analyses, and both program comprehension and resting-state data yielded strong results. These findings demonstrate that EEG features effectively capture neural correlates across different skill levels and highlight the potential of neural data to complement traditional methods of skill assessment.
arXiv:2607.00424v1 Announce Type: new Abstract: Redundant robotic manipulators operating in constrained and human-interactive environments require accurate task-space tracking together with rigorous safety guarantees under dynamic uncertainties. Classical operational space computed torque controller (OSCTC) relies on accurate dynamic models and degrades in the presence of disturbances. In contrast, the data-driven paradigm of residual learning approximates disturbances as functions learned from full-state measurements, which are often noisy in practice, lack rigorous theoretical guarantees, and introduce additional design complexity. This paper proposes a robust OSCTC framework that integrates an extended state observer (ESO) with conformal prediction to combine model-based robustness and data-driven adaptability. The ESO estimates lumped disturbances directly in operational space without requiring full-state measurements as in residual learning, and a robust control barrier function (CBF) is constructed to enforce safety under uncertainty. However, robust CBFs require a known disturbance-variation bound to guarantee absolute safety, which often leads to conservatism in practice. To address this limitation, we further employ a sliding-window conformal prediction mechanism to estimate the bound online in a distribution-free manner, thereby achieving practical probabilistic safety guarantees. Experiments on a 7-DoF Franka Research 3 manipulator demonstrate millimeter-level tracking accuracy and real-time safe control at 1~kHz under various disturbances.
arXiv:2607.00428v1 Announce Type: new Abstract: CLIP (Contrastive Language-Image Pre-training) has become a de facto paradigm for image-text alignment, but it struggles with long-context descriptions (>77 tokens) due to absolute positional encoding and pretraining on short captions. In long contexts, sentences are often reordered, summarized, or partially omitted. Although prior works extend CLIP with longer positional encodings, they often suffer from degraded image-text alignment under such text perturbations. We attribute this limitation to the Euclidean contrastive objective, which enforces strict one-to-one matching and lacks explicit mechanisms for modeling hierarchical relationships between global context and its constituent elements. To address this issue, we propose HyFL-CLIP, a hyperbolic fine-tuning framework that distills the well-established text-image alignment learned in Euclidean CLIP into hyperbolic space via cross-manifold similarity distillation, leveraging its geometry to capture hierarchical and entailment relations. Our method models hierarchical semantics by linking summarized token-wise features, long-context descriptions, constituent short textual components, and images, capturing part-whole relationships via hyperbolic entailment with Einstein midpoint aggregation. Experiments on diverse benchmarks, including long-context cross-modal retrieval, cross-modal retrieval with caption perturbations, intra-modality retrieval, and short-text cross-modal retrieval, show that HyFL-CLIP achieves more robust long-context understanding. In particular, it yields up to 19.5% improvement in long-text cross-modal retrieval under textual perturbations over the best prior method. We also show HyFL-CLIP can be seamlessly integrated into other model frameworks by applying it to Stable Diffusion XL (SDXL).