Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation
arXiv:2606.08737v1 Announce Type: new Abstract: World action models inherit the predictive capability of world models, enabling action generation to be guided by anticipated future observations. However, they rely primarily on vision and often fail in contact-rich manipulation, where critical cues arise from physical interaction. In this paper, we propose Dream-Tac, a unified Tactile-World Action Model that jointly models actions, future visual observations, and tactile dynamics. Specifically, Dream-Tac introduces (i) contact-gated visuotactile fusion to selectively integrate tactile signals and (ii) a contact-aware attention bias to better regulate cross-modal interactions during manipulation. To support real-time deployment, we further design a dual-level acceleration strategy, reformulating the contact-aware bias to preserve the fused attention path during training and introducing cache-based diffusion acceleration at inference, achieving up to 2.9$\times$ faster training and 1.8$\times$ faster inference. Across six contact-rich manipulation tasks, Dream-Tac improves action accuracy by 31.7\% on average, demonstrating the effectiveness of unified visuotactile world modeling.Code is available at https://github.com/LYFCLOUDFAN/Dream-Tac.
Studies of PHP with CASCO code and its experimental validation
arXiv:2606.09333v1 Announce Type: new Abstract: We discuss here two major issues related to the steady functioning of the pulsating (oscillating) heat pipe (PHP): the effect of the surface properties and stopovers. They are studied with the CASCO simulation software (Code Avanc{\'e} de Simulation du Caloduc Oscillant: Advanced PHP Simulation Code in French) version 4. Its experimental validation against two different prototypes is presented. The first is used also to study the effect of the nucleation barrier (the wall superheating necessary for the bubble nucleation) that reflects the wall wettability and roughness. An optimal value of the nucleation barrier is found where the thermal resistance achieves a minimum for a given evaporator power. The functioning regime is continuous showing pressure waves propagating along all the PHP channel. The stopover regime is observed both for small and large barriers. The second experimental setup (PHP Smart Loop) is used to study the stopover regime. It is found that it is characterized by a chaotically repeating sequence of fast pressure growth (corresponding to oscillations) followed by a slower pressure decay during a stopover. The decrease of the thermal resistance with heating load is explained by a decrease of the stopover time caused by a faster liquid film shrinking.
ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning
arXiv:2606.02802v2 Announce Type: replace Abstract: Large language models (LLMs) exhibit strong natural-language reasoning abilities for clinical decision support, but struggle to effectively model structured longitudinal electronic health records (EHRs). In contrast, EHR foundation models can learn predictive patient representations, yet lack interpretable language-based reasoning. To bridge this gap, we propose ChatHealthAI, a multimodal reasoning framework that aligns structured EHR representations from a pretrained EHR foundation model with the semantic space of a frozen LLM through a task-aware resampler. By integrating longitudinal patient representations with refined clinical event descriptions, ChatHealthAI enables clinically grounded natural-language reasoning while maintaining accurate patient prediction. We evaluated ChatHealthAI on three clinical predictive tasks from the EHRSHOT benchmark. Results show that ChatHealthAI improves reasoning quality and interpretability while preserving competitive predictive performance. These findings highlight the potential of integrating EHR foundation models with pretrained LLMs for interpretable clinical prediction.
Breakdown of Adiabatic Scaling and Noise-Induced Functional Synchronization in Deeply Quiescent Excitable Systems
arXiv:2605.06692v3 Announce Type: replace-cross Abstract: Coherence resonance (CR) characterizes noise-induced regularity in excitable systems, yet its evaluation in quiescent biological media is often obscured by flattened energy landscapes and complex nonlinear dynamics. In this study, we investigate the stochastic dynamics of a 3D Sherman-Rinzel-Keizer (SRK) model driven by multiplicative Feller noise. We show that traditional extremal evaluations of CR encounter a "bathtub effect", a broad resonance valley that can lead to statistical inaccuracies. To address this, we propose a logarithmic centroid extraction method, which filters out stochastic jitter and recovers the underlying adiabatic Kramers scaling with high linearity. Furthermore, we identify the physical boundary where this adiabatic approximation breaks down under the strong-noise limit. Extending our analysis to gap-junction coupled systems, we observe a noise-induced transition from sub-threshold physiological shivering (characterized by statistical correlation but negligible functional output) to macroscopic functional synchronization. Our results provide a mathematical framework for extracting optimal noise intensities in broad energy valleys and offer insights into how quiescent biological systems utilize stochastic fluctuations for functional recovery.
A Topological Soliton Model for Ball Lightning: Theory and Numerical Verification with the 3D Gross-Pitaevskii Equation
arXiv:2605.09851v3 Announce Type: replace-cross Abstract: Ball lightning remains one of the most enigmatic atmospheric phenomena, characterized by its long lifetime, ability to penetrate materials, and stable spherical structure. Here we propose a novel theoretical framework interpreting ball lightning as a three-dimensional projection of a high-dimensional topological soliton. The system is described by a nonlinear Schr\"odinger equation with attractive interactions, stabilized by a non-zero topological charge. Through comprehensive numerical simulations of the three-dimensional Gross-Pitaevskii equation, we verify the model's core predictions: (1) long-lived stability protected by topological invariants, (2) low transmission probability due to wavefunction orthogonality, and (3) energy and size scales consistent with observational data. The soliton lifetime $\tau \sim \hbar/\Gamma$ naturally explains the observed second-scale durations. Our work provides a self-consistent physical explanation for ball lightning while offering concrete pathways for experimental realization of three-dimensional topological solitons in Bose-Einstein condensates and nonlinear optical systems. This theoretical framework gains additional support from recent experimental breakthroughs in laboratory generation of ball-lightning-like structures.
FXplorer: A Map-Based Interface for Exploratory Audio Effect Design
arXiv:2606.08286v1 Announce Type: new Abstract: Audio effects (FX) shape sound in contemporary music practice. However, most interfaces present them as discrete modules and parameters that favor targeted adjustment over exploratory listening. This separation can make it difficult to build intuition about the broader space of possible transformations or to move fluidly between searching and refinement. We present FXplorer, an interface that organizes audio effects within a perceptually informed 2D space, allowing sound transformations to be browsed as a continuous landscape rather than as isolated presets. By combining established spatial interaction approaches and interpretable DAW-style controls with recent embedding-based machine learning methods for similarity and semantic search, the system brings exploration and parameter refinement into a single workspace. FXplorer supports composition, production, or performance by allowing users to edit and interpolate between effect presets interactively.
Trustworthy Smart Fabs via Professional Proxies: Scaling Safe and Sustainable by Design (SSbD) through Industrial Data Spaces
arXiv:2606.09227v1 Announce Type: new Abstract: The convergence of the 2026 European Union Safe and Sustainable by Design (SSbD) framework, Corporate Sustainability Due Diligence Directive (CSDDD), and Carbon Border Adjustment Mechanism (CBAM) introduce a severe governance bottleneck for advanced semiconductor manufacturing facilities ("Smart Fabs"). Regulatory compliance demands have surpassed the capacity of manual corporate reporting, creating a direct conflict between multi-stakeholder transparency and corporate data privacy. This paper addresses this challenge by introducing a zero-trust socio-technical orchestration framework that operationalizes a six-layer SSbD reference architecture within trustworthy industrial data spaces. We propose a shift from reactive automation to autonomous governance through "Professional Proxies"-role-based agentic workflows executing within hardware-isolated trust zones. Structured as an interoperable network protocol stack, the framework coordinates an automated, five-step "relay race" between Facility, Process Engineering, and Finance proxy teams to align factory-floor yield models with macro-level sustainability mandates. By executing Virtual Metrology (VM) predictions and Federated Machine Learning (FML) inside hardware-rooted Trusted Execution Environments (TEEs), this architecture resolves the Data Sovereignty Paradox, demonstrating how fabs can export cryptographically signed compliance tokens via International Data Spaces (IDS) connectors without exposing proprietary process recipes. Ultimately, this framework provides technology managers with a verifiable, evidence-based pathway toward resilient, net-zero Industry 5.0 ecosystems.
OASIS: From Simulation Data Collection to Real-World Humanoid Loco-Manipulation
arXiv:2606.08548v1 Announce Type: new Abstract: Recent progress in robot manipulation has been largely driven by learning from large-scale demonstrations. For humanoid robot loco-manipulation tasks, however, existing data sources force an unsatisfying tradeoff between trajectory quality and scalability. Real-world teleoperation provides the highest-quality trajectories but requires dedicated physical space and time-consuming scene resets. Simulation offers an alternative way out of this dilemma: it can produce clean, embodiment-aligned data at scale without any physical hardware. In this paper, we propose OASIS, a simulation-data-driven framework for humanoid loco-manipulation. OASIS automatically reconstructs realistic object assets from real-world images using a 3D generative model. Based on these assets, trajectories are first collected through teleoperation in simulation, and then augmented under diverse domain randomizations in a post-processing stage. With the resulting simulation data, we further design a hierarchical visuomotor policy for humanoid loco-manipulation. Extensive experiments on the real humanoid robot show that, under zero-shot deployment, the policy trained on our simulation data achieves higher success rates on most tasks than that trained on real-robot teleoperation data, owing largely to the broad lighting and environmental variations covered by our simulation rendering, which real-robot data fails to capture. The project page is available at https://oasis-humanoid.github.io/.
Investigation of Thick-GEM detectors fabricated in India for muography application
arXiv:2606.08664v1 Announce Type: new Abstract: Muography, commonly known as muon tomography, is a passive, non-destructive imaging technique that utilizes naturally occurring cosmic-ray muons to visualize the internal density structures of large, static, or inaccessible objects. In course of developing a small prototype muography system for material identification that relies upon the multiple Coulomb scattering of muons in matter, we explored the possible use of Thick-GEM detector as muon tracking device. It is a 5-20 fold scaled-up version of traditional GEM technology, that has become increasingly popular in recent years, owing to its mechanical robustness, cost-effective production, and excellent position sensing capabilities. A few prototypes of this detector of dimension $40\,\mathrm{mm} \times 48\,\mathrm{mm}$ with variation in other design parameters, were manufactured from a local industry. Subsequent to conditioning, detailed characterization of the detectors was performed to validate their suitability in muography applications. To identify the optimal operating region, gain variation was studied under various voltage configurations for both single and double-stage configurations. Experimental measurement of muon detection efficiency across the entire operating range yielded a maximum efficiency of 99.5\% in both cases. Using a collimated Fe$^{55}$-source, the best spatial resolution was determined to be 30 $\mu$m for both single and double-stage operation.
Language as a Sensor: Calibrated Spatial Belief Estimation in 3D Scenes from Natural Language
arXiv:2606.08666v1 Announce Type: new Abstract: Robots deployed in human-centric environments routinely receive natural-language descriptions of spatial information ("I left my backpack on the table") that reference parts of the world beyond their perceptual field of view. Traditional metric-semantic mapping ignores this signal, while off-the-shelf multimodal models remain limited in 3D spatial reasoning and are not directly amenable to fusion with other sensor modalities. To convert language observations into a calibrated spatial distribution, we train a Language Sensor Model (LSM) that maps each utterance and its scene-graph context to a multimodal distribution, with mixture weights encoding referential ambiguity (e.g., "which table") and component covariances encoding spatial uncertainty (e.g., where "on the table" the target lies). We then introduce VL-Map (Vision-Language Metric-Semantic Mapping), a probabilistic framework that treats these language predictions as stochastic observations and fuses them with onboard perception within a unified belief map. On the VLA-3D benchmark as well as on a real-world mobile robot, LSM is the only language predictor whose covariance estimates remain within the calibrated regime; fused into VL-Map, it leads to more accurate predictions of the target object location (~70% more probability mass on the true target compared to the strongest foundation-model baseline).
Learning to Solve Generative ODEs Beyond the Linear Span
arXiv:2606.08672v1 Announce Type: new Abstract: Diffusion and flow generative models sample by integrating a learned ODE, but high quality still requires many sequential model evaluations. Solver learning reduces this cost by adapting scalar coefficients, timesteps, or both, while keeping the backbone model fixed. In this work, we identify a structural bottleneck in this update family: each step remains span-limited. Since the scalar-coefficient update lies in the span of buffered velocity evaluations, it can fit only the in-span component while leaving any out-of-span residual unreachable by scalar recombination alone. We propose SpanLift, a lightweight neural solver that augments scalar-coefficient updates with a spatial residual operator. SpanLift keeps a fixed base solver as an in-span prior and learns a spatial residual operator over the state and velocity buffer. The operator is trained by endpoint teacher matching, preserves the pretrained backbone, and adds no model NFEs. Empirically, the learned correction transfers across base solvers and is predominantly out-of-span. Across pixel-space diffusion, latent flow matching, and precipitation nowcasting, SpanLift achieves state-of-the-art few-step sampling. With only 3 NFE, it improves CIFAR-10 FID from 8.16 to 5.69 and ImageNet FID from 17.37 to 11.83.
ClinicalAligner26AM: A Cross-Lingual Aligner for Dataset Translation; Evidences from the MultiClinCorpus Shared Task
arXiv:2606.08673v1 Announce Type: new Abstract: Word-level cross-lingual alignment is central to annotation projection, translation auditing, and cross-lingual faithfulness estimation, yet existing neural aligners are rarely adapted to specialized domains. In this paper, we introduce ClinicalAligner26AM, a large-context multilingual aligner model for biomedical and clinical text initialized from ClinicalEncoder26AM. Our training recipe is inspired by AWESoME Align. We build our soft alignment target by sharpening with Sinkhorn-Knop optimal transport a cost matrix established for parallel clinical texts and conversations through the fusion of sentence-level, phrase-level, and token-level signals. We distill this sharpened alignment matrix directly into our student aligner, by encouraging its naive cosine-based token similarity scores to match this target. At inference time, we project source-span scores through the learned token alignment matrix and decode the longest valid high-scoring span in the target text, optionally supported by MultiClinNER predictions summarized in Appendix B. We evaluate CA26AM on the MultiClinCorpus shared task, which projects Spanish clinical entity annotations into six target languages. Our two submitted systems ranked respectively first and second across all languages and entity types, with character-weighted F1 scores above 0.95 in nearly all settings.
FAME: Forecastability-Aware Mixture of Experts for Heterogeneous Time Series Forecasting
arXiv:2606.08896v1 Announce Type: new Abstract: Large-scale retail and industrial forecasting systems contain many heterogeneous time series whose lifecycle, sparsity, volatility, seasonality, spectral patterns, and contextual sensitivity differ substantially. A single forecasting model rarely performs well across all regimes, while dense ensembles increase inference cost and provide limited insight into expert suitability. This paper studies forecastability-aware expert routing: learning how data characteristics determine the suitability of forecasting experts. We propose \method{}, a sparse mixture-of-experts framework that represents each series with a multidimensional forecastability fingerprint, mines expert-suitability targets from validation performance, and trains a cost-aware sparse router to activate a small budgeted set of experts for each series. Using a production-scale vending-machine sales dataset from Shandong New Beiyang (SNBC), where the forecasting component has been integrated into the replenishment-planning pipeline, together with public retail benchmarks, we show that expert suitability varies systematically across data regimes. On the industrial dataset with 5,000+ machines and 60M+ transactions, \method{} Top-2 reduces MSE by 12.4\% over the strongest single expert, LightGBM, while executing 1.92 experts per series on average. The deployed component produces demand forecasts, while inventory-oriented gains are estimated by an offline replay simulator under a fixed replenishment policy rather than by online intervention. The framework turns heterogeneous sales forecasting from heuristic model selection into data mining of forecastability patterns and expert specialization. Code is available at https://github.com/hit636/FAME
Rise regimes of freely rising droplets with a moderate viscosity ratio
arXiv:2606.08575v1 Announce Type: new Abstract: The dynamics of buoyant droplets rising freely in a large body of an immiscible liquid is investigated numerically for a moderate drop-to-fluid viscosity ratio $\mu^\ast$. We focus on toluene droplets rising in clean water, for which $\mu^\ast=0.62$, and vary the radius over $0.5\,\text{mm}\leq R\leq3.0\,\text{mm}$. Direct numerical simulations are performed in imposed axisymmetric and fully three-dimensional configurations. As $R$ increases, the system displays a rich sequence of rise regimes. Starting from steady vertical rise with an axisymmetric disturbance flow, it first undergoes an internal flow instability associated with an azimuthal mode $m=2$, leading to a biplanar-symmetric wake and reduced terminal speed. This state is followed by a steady oblique regime, in which the $m=1$ mode also becomes unstable and coexists with the $m=2$ mode. At larger radii, the path becomes nearly vertical again before the flow enters an $m=2$ rotating-wave regime, where the wake drifts azimuthally at an approximately constant angular velocity. For still larger droplets, persistent shape oscillations and vortex shedding lead to fully three-dimensional chaotic paths. Simulations initialised from finite-amplitude asymmetric states further reveal several multistable size ranges, in which distinct terminal states coexist depending on the initial condition. Taken together, these findings show that the path instability of moderate-viscosity-ratio droplets differs fundamentally from that of bubbles and solid particles: in most regimes encountered here, axisymmetry breaking is initiated within the droplet, highlighting the central role of the internal flow instability in shaping the subsequent wake structure, rise speed and droplet dynamics.
Slice-Profile-Enabled Phase Distribution Graphs for MRI Simulation
arXiv:2606.09233v1 Announce Type: new Abstract: MRI simulation often separates two descriptions that are both essential for realistic sequence analysis: Bloch dynamics for waveform-resolved radiofrequency (RF) excitation, and phase-graph methods for coherence-pathway evolution. Extended Phase Graph (EPG) models provide pathway tracking, and Phase--Distribution Graphs (PDG) extend this idea to spatially resolved $k$-space simulation, but existing PDG formulations rely on hard-pulse RF mixing that is \emph{order-local}: the RF pulse mixes $F_n^+$, $F_n^-$, and $Z_n$ at a fixed coherence order $n$, without coupling different $k_z$ orders. This work introduces a unified Bloch-resolved PDG framework for slice-profile-aware MRI simulation. A scanner-rasterized sequence is partitioned into RF-sensitive Bloch spans and non-RF phase-graph spans. For each unique RF span, Bloch dynamics are solved on a slice grid to obtain a spatially varying propagator $R(z)$. Its Fourier coefficients $\mathcal{R}_{\Delta}$, indexed by slice-order offset $\Delta$, are compiled into the PDG state graph as sparse cross-order coupling in $k_z$. Graph growth is controlled by retaining the dominant Fourier coefficients and pruning low-contribution PDG states. This retains PDG pathway history and voxel-wise image formation while incorporating shaped slice-selective and off-resonant RF behavior. Experiments show close agreement with direct one-dimensional Bloch slice-profile evolution through repeated excitations, while retaining only a few hundred active PDG states. Image simulations further illustrate slice-position dependence, fat-suppression behavior, measured three-dimensional $B_0$ field maps, and comparison with scanner data. The proposed framework enables sequence-consistent simulation and signal formation understanding in regimes where RF physics, spatial encoding, object heterogeneity, and echo-pathway formation interact.
Closing the Sim-to-Real Gap: An Evaluation Framework for Autonomous Cyber Defense Configuration of Commercial EDR
arXiv:2606.08168v1 Announce Type: new Abstract: Leading commercial endpoint detection and response (EDR) products have shifted from operator-configured rule sets to multi-component systems where autonomous AI components operate alongside, and increasingly in place of, operator-deployed policies. Autonomous defense agents using commercial EDR as their hardening tool are no longer tuning a passive tool, but a black-box autonomous system capable of making vendor-specific decisions. We present the first evaluation framework for autonomous defense agents hardening commercial EDR. We instantiate it in a Game of Active Directory (GOAD) lab with Horizon3.ai's NodeZero as the autonomous pentester and Microsoft Defender XDR as the EDR. We run a sample benchmark of defense agents with two large language model (LLM) backbones (Claude Sonnet 4.6 and Cisco Foundation-Sec-8B). We report three lessons learned that neither simulation nor open-source-EDR evaluation can surface: (i) commercial EDR telemetry is engineered for Security Operations Center (SOC) analyst workflows rather than scientific benchmarking; (ii) the importance of per-policy attribution to separate defense agent actions from autonomous EDR actions; and (iii) the EDR's autonomous behavior varies during the evaluation window. Together, these findings highlight a sim-to-real gap for enterprise defense and motivate evaluation methodology for benchmarking autonomous defense agents in environments with black-box, autonomous tools.
The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In
arXiv:2606.08172v1 Announce Type: new Abstract: Large language models (LLMs) increasingly mediate high-stakes interactions in finance, medicine, and mental-health support, yet users have limited control over how these systems communicate. We frame interaction style as a governance object: provider-side alignment not only blocks harmful content, but also stabilizes communicative defaults that shape users' epistemic distance, relational expectations, and capacity to opt out of emotionalized or anthropomorphic interaction. We introduce a deterministic multi-agent evaluation pipeline for measuring prompt steerability and style drift in long-horizon dialogue. The study replays 100 frozen user-only scripts across four domains and three runnable persona conditions: default, sarcastic, and cold, using three generator models, yielding 90,000 assistant replies scored by a human-calibrated LLM judge on harmfulness, negative emotion, inappropriateness, empathic language, anthropomorphism, and refusal behavior. A fourth harmful persona is evaluated separately as a safety-gating test. The paper contributes a reproducible method for quantifying whether prompt-specified styles remain stable over time and a governance framework distinguishing safety gating, civility steering, and affective default lock-in. Overall, we show that prompt steerability and regression-to-default are observable indicators of provider control over communicative form, with implications for pluralism, autonomy, and democratic agency in human-LLM interaction.
Quaternion Maximum-Volume Submatrix Selection with Applications to Multichannel Imaging and Visual Data
arXiv:2606.08175v1 Announce Type: new Abstract: Low-rank approximation based on selected rows and columns is a useful alternative to singular value decompositions when the goal is an interpretable and compact matrix representation. A standard way to choose these rows and columns is the maximum-volume principle: it selects submatrices with large volume, which usually leads to stable interpolation coefficients and accurate CUR-type approximations. In this paper, we study this idea for quaternion matrices. This setting is natural for color images, three-dimensional motion data, and multi-channel signals, but requires care because quaternion multiplication is noncommutative. We define quaternion maximum-volume submatrix selection using quaternion singular values and the Study determinant. We then derive quaternion rank-one update formulas and use them to build two selection procedures: a greedy square-core method for row and column replacement, and a rectangular method that enlarges a selected row set until the interpolation coefficients are controlled. We prove that successful row and column swaps increase the quaternion volume of the selected square core when the exact quaternion inverse is used. We also connect the stopping criterion with quasi-dominance, prove an exact quaternion CUR identity in the full-rank case, and derive an interpolation stability bound. For the rectangular case, we derive an append-row pseudoinverse update and show how it gives a natural right preconditioner for overdetermined quaternion least-squares problems. Finally, we illustrate the methods on three applications: quaternion CUR approximation of RGB images, RectMaxVol-based preconditioning for ill-conditioned quaternion least-squares systems, and row selection in quaternion motion-capture data. The experiments show that the proposed quaternion MaxVol and RectMaxVol methods provide stable and efficient selection routines.
TextEconomizer: Enhancing Lossy Text Compression with Denoising Transformers and Entropy Coding
arXiv:2606.08184v1 Announce Type: new Abstract: Lossy text compression reduces data size while preserving core meaning, making it well-suited for summarization, automated analysis, and digital archives. Despite the dominance of transformer-based models in language modeling, integrating context vectors and entropy coding into Sequence-to-Sequence (Seq2Seq) generation remains underexplored. A key challenge lies in identifying the most informative context vectors from encoder output and incorporating entropy coding to enhance storage efficiency while maintaining high-quality outputs, even under noisy text. We introduce TextEconomizer, an encoder-decoder framework paired with a transformer neural network that reduces variable-sized inputs by 50% to 80% without prior knowledge of dataset dimensions. Our model achieves competitive compression ratios via entropy coding while delivering near-perfect text quality, assessed by BLEU, ROUGE, METEOR, and semantic similarity scores. TextEconomizer operates with approximately 153x fewer parameters than comparable models, achieving a 5.39x compression ratio without sacrificing semantic quality. We also evaluate an LSTM-based autoencoder achieving a state-of-the-art 67x compression ratio with 196x fewer parameters, and LLaMAFormer, a modified transformer with 263x fewer parameters than ICAE while maintaining competitive text quality. TextEconomizer significantly surpasses existing transformer-based models in balancing memory efficiency and high-fidelity outputs, marking a breakthrough in lossy compression with optimal space utilization.
CHROMA: Detecting AI-Generated Images through Inter-Channel Color-Space Correlations
arXiv:2606.08864v1 Announce Type: new Abstract: The rapid adoption of diffusion and large-scale generative models has made it increasingly challenging to distinguish synthetic imagery from real photographs. While automated detectors have been proposed, their generalization to unseen generators remains brittle. To address this limitation, we investigate inter-channel color correlations, a lightweight and underexploited forensic cue. We first demonstrate that LPIPS, a widely used perceptual metric, exhibits inconsistent responses to perturbations that selectively alter channel dependence across different color-space parameterizations, indicating that cross-channel statistics are not uniformly constrained by common perceptual training objectives. Motivated by this, we analyze the distributions of pairwise inter-channel correlation features across multiple color spaces. Our analysis reveals systematic, generator-specific differences in these distributions, with RGB and Lab color spaces providing the most apparent separation between real and generated images. Building on this, we introduce Chroma, a detector of AI-generated images which augments standard RGB inputs with inter-channel correlation maps and employs a fixed CNN backbone trained with a modest computational budget. We assess its robustness under both single-generator training and a limited multi-generator supervision regime, where only a few samples from additional generators are available. Across a standard benchmark protocol, correlation-augmented inputs improve real-vs-generated discrimination and robustness, yielding performance competitive with recent detectors while maintaining a simple architecture and training procedure. Code is available at https://github.com/JPSoteloSilva/CHROMA
Frequency-Domain Latent Attention Gating for Cross-Domain Token Aggregation
arXiv:2606.08191v1 Announce Type: new Abstract: Token aggregation is a common bottleneck in models that map token representations to sample-level predictions, yet most pooling methods operate only in the original token domain. We propose FLaG, a plug-in aggregation module that transforms token representations with the real FFT, summarizes spectral components with learnable latent queries, applies a channel-wise gate, and reconstructs enhanced time-domain tokens for final pooling. We evaluate FLaG on antimicrobial peptide (AMP) activity prediction with ESM2, image classification with ResNet18 on CIFAR-10 and CIFAR-100, and text classification with RoBERTa on IMDB and GLUE. FLaG achieves its clearest gains on the ESM2-8M antimicrobial peptide tasks and on CIFAR-100, while remaining competitive with strong text baselines on IMDB and GLUE. Then we probe its behavior on the AMP setting with band knockouts, gate summaries, residue perturbations, latent-query readouts, and structure-proxy stratification. We find that low-frequency bands contribute the most overall, and the remaining higher-band pattern is more sample-specific. The gate acts as a broadly shared spectral reweighting stage and the cross-attention patterns are sample-specific with mild query-wise differentiation, and higher-helix peptides exhibit stronger average spectral sensitivity in both bacteria. The supplementary materials, source code and data are released at https://www.healthinformaticslab.org/supp/ and https://github.com/Kewei2023/AMPCliff/tree/FLaG.
Three-dimensional experimental investigation of the interaction between a rising bubble and a vortex ring
arXiv:2606.08193v1 Announce Type: new Abstract: The interaction between turbulent flows and bubbles is a complex phenomenon ubiquitous in natural and industrial settings. In this work, we experimentally investigate, from a fundamental perspective, the interaction between a rising bubble and a vortex ring in counterflow. Using time-resolved three-dimensional Lagrangian Particle Tracking (4D-LPT) coupled with shadowgraphy, we obtain simultaneous measurements of the bubble motion and the surrounding liquid flow. This approach enables detailed observation of bubble dynamics, deformation, and eventual breakup, as well as the fluid motion. We examine several flow configurations by varying the vortex circulation and the Weber number while maintaining a comparable vortex-to-bubble size ratio. Based on these measurements, we classify the interaction events into three categories according to their impact on bubble dynamics and vortex stability over time. Through experiments, we address for the first time the three-dimensional effects of these interactions, which had not been considered in previous studies. The analysed experiments comprise: Case I, corresponding to a weak interaction in which neither the bubble nor the vortex is significantly affected; Case II, where the bubble is captured and advected by the vortex, leading to a strong distortion of the vortex due to the presence of the bubble within its core; and Case III, involving a stronger vortex capable of capturing the bubble and breaking it into two fragments without a severe loss of energy in the vortex core. The analysis of these results provides insight into the bubble breakup process and the mechanisms responsible for the destabilisation of the vortex ring.
GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models
arXiv:2606.08194v1 Announce Type: new Abstract: Large Audio-Language Models (LALMs) integrate audio perception and language understanding within a unified framework, enabling a wide range of real-world applications. Despite recent advances, evaluation for LALMs remains heavily underspecified relative to real-world requirements: most lack true linguistic and cultural authenticity, while others fail to capture acoustic realism. To bridge this gap, we propose GlobeAudio, a multilingual and multicultural benchmark designed to evaluate naturalistic audio understanding. GlobeAudio consists of 5,637 multiple-choice questions across six typologically diverse languages, expertly crafted by native speakers grounded on naturally occurring audio. In order to do well, models must possess higher-level auditory reasoning skills and culturally grounded interpretation. We systematically evaluate representative closed-source and open-source LALMs, as well as cascaded ASR-LLM pipelines. Our experiments reveal substantial performance gaps under natural acoustic conditions, particularly for open-source models and low-resource languages. These findings highlight critical limitations of current LALMs and underscore the importance of naturalistic audio evaluation for future audio-language systems. GlobeAudio can be found at https://huggingface.co/datasets/iNLP-Lab/GlobeAudio .
AlignFed: Alignment-Aware Asynchronous Federated Fine-Tuning for Large Language Models in Heterogeneous Edge Environments
arXiv:2606.08197v1 Announce Type: new Abstract: Large Language Models (LLMs) have significantly propelled the advancement of edge intelligence and have been widely deployed across various scenarios, including autonomous driving, industrial inspection, and personalized IoT services. However, the collaborative adaptation of LLMs on edge devices continues to face formidable challenges due to strict data privacy constraints, highly heterogeneous computing and communication resources, and the non-independent and identically distributed (non-IID) nature of local data. Federated Fine-Tuning (FFT) enables the collaborative optimization of distributed models without exposing raw data. Yet, traditional synchronous aggregation suffers from a severe straggler effect, resulting in high system latency and low resource utilization. Existing asynchronous federated learning methods are predominantly designed for small-to-medium-scale models and struggle to address the specific challenges inherent in LLM fine-tuning namely, model drift caused by stale updates, aggravated client drift stemming from data heterogeneity, and aggregation fairness imbalance resulting from the dominance of fast clients. To address these issues, this paper proposes AlignFed, an asynchronous federated fine-tuning framework for LLMs tailored to heterogeneous edge environments. AlignFed employs a lightweight multi-stage semantic alignment mechanism comprising three core modules: version-aware update grouping, cross-version semantic alignment based on a mini-batch calibration set, and fairness-aware aggregation that integrates both update freshness and client participation frequency. This framework effectively mitigates cross-version model drift and client drift while enhancing aggregation fairness, thereby achieving stable and efficient asynchronous federated optimization in scenarios characterized by high heterogeneity and significant update staleness.
Thermal Decoherence and Population Transfer of MeV Channeling Electrons in Diamond
arXiv:2602.16529v2 Announce Type: replace-cross Abstract: Channeling radiation from MeV-regime electrons is governed by transitions between quantized transverse bound states, but experimental spectra are strongly modified by thermal diffuse scattering. To capture these open-system dynamics, a frozen-phonon multislice framework is combined with bound-state projection analysis to construct depth-dependent reduced density matrices in selected transverse manifolds. Beyond reproducing experimental channeling-radiation transition energies, this approach separates thermal population transfer, intra-manifold decoherence, and cross-manifold coherence loss. Applied to 16.9 MeV axial electron channeling in $\langle100\rangle$ diamond, the results show approximately exponential population decay from the initially occupied states, accompanied by strongly channel-dependent feeding among low-lying manifolds. Starting from a coherent superposition within the degenerate 2p manifold, stochastic symmetry breaking by thermal displacements drives the intra-manifold purity toward the maximally mixed limit, indicating rapid phase scrambling. Under 1s initialization, population transferred into the 2p and 3d manifolds remains internally close to maximally mixed, yet a weak residual 2p-3d cross-manifold coherence persists. This framework goes beyond static mean-field thermal broadening and provides a microscopic basis for evaluating population dynamics and coherence lifetimes in strongly quantized channeling-radiation systems.