Forskningsradar

Science Journals

Peer-reviewade publikationer — 61549 artiklar

Machine Learning-Driven Content Popularity Prediction and Cache Optimization in D2D Clustered Networks
arXiv:2606.26119v1 Announce Type: new Abstract: Advancements in wireless communication technology have led to the widespread use of smart devices including computers, mobile phones, tablets, wearable devices, and vehicles which has significantly increased the demand for high quality content. This growing demand puts pressure on the backhaul links in cellular networks, resulting in congestion and content delivery delays. To address this, cache enabled networks and edge caching, such as caching in user devices, have emerged as promising solutions to reduce backhaul traffic. By caching content locally and using device to device (D2D) communication for retrieval, content delivery can be made more efficient. However, limited cache capacity requires intelligent content selection strategies. The popularity of the content is dynamic and varies with user preferences, where less than 20% of the users generate 80% of multimedia traffic. Many existing methods fail to consider this user heterogeneity, often assuming uniform preferences throughout the network. This paper proposes a novel Machine Learning Driven Content Popularity Prediction and Cache Optimization (ML CPCO) framework that dynamically predicts user and cluster level content demand, incorporates user willingness to participate in caching, and optimizes cache placement in D2D enabled clustered networks. The system predicts future content requests using machine learning algorithms and estimates content popularity at the cluster level. Based on these predictions, cache placement decisions are made to maximize efficiency. The simulation results show that the proposed approach performs well under various network conditions, achieving a cache utilization rate of nearly 97% the highest among the methods compared. In addition, it offers an improved hit rate with an acceptable execution time, resulting in reduced backhaul traffic and enhanced user experience.
Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLM
arXiv:2606.26120v1 Announce Type: new Abstract: Diffusion Large Language Models (dLLMs) offer a promising alternative to autoregressive models, excelling in text generation tasks due to their bidirectional attention mechanisms. However, their computational complexity scales on the order of L cubed with the sequence length L. This poses significant challenges for long-sequence and real-time applications, primarily due to the lack of compatibility with key-value caching and the non-autoregressive nature of denoising steps. Existing acceleration methods rely on static caching or parallel decoding strategies, which fail to account for the dynamic behavior of token properties across layers and decoding steps. We propose Dynamic-dLLM, a training-free framework that enhances dLLM inference efficiency through two components: Dynamic Cache Updating (DCU), which adaptively allocates cache-update budgets based on layer-wise token dynamics, and Adaptive Parallel Decoding (APD), which dynamically calibrates decoding thresholds to balance generation quality and efficiency. Extensive experiments on models like LLaDA-8B-Instruct, LLaDA-1.5, and Dream-v0-7B-Instruct across benchmarks such as MMLU, GSM8K, and HumanEval demonstrate that Dynamic-dLLM significantly improves inference speed. It attains an average speedup exceeding 3 times while maintaining performance. Dynamic-dLLM outperforms state-of-the-art acceleration methods and provides a plug-and-play solution for efficient dLLM deployment without compromising performance. The code is available at https://github.com/TianyiWu233/DYNAMIC-DLLM.
Scattering theory for cavity-assisted spin-motion-photon interactions
arXiv:2606.26542v1 Announce Type: cross Abstract: Cavity-assisted photon scattering (CAPS) is a powerful mechanism for realizing strong interactions between the internal states of stationary qubits and flying photons, underpinning a broad range of hybrid atom-photon protocols including remote entanglement generation and heralded atom-photon gates. Recently, the motional quantum state has emerged as an important building block for quantum information processing with atomic qubits, both as a coherently controllable degree of freedom and as a fundamental error channel through undesired spin-motion coupling. For the resonant-coupling regime of cavity quantum electrodynamics relevant to CAPS operations, however, the analytical formulation of spin-motion-photon coupling has so far remained elusive. Here, we develop a complete analytical framework for CAPS that incorporates the coherent interaction between atomic motion and a reflected photon by extending scattering theory to include the motional degree of freedom. The resulting compact operator-based input-output relation applies uniformly across various cavity geometries, spin-dependent trapping potentials, and nonidentical multiple spins. As an exemplary application, we use the framework to elucidate how atomic motion affects CAPS-based atom-photon gates, identifying the parameter regimes that suppress motion-induced errors. Our framework provides a theoretical foundation both for mitigating motional errors in CAPS operations and for deliberately exploiting motion-photon interaction at the atom-photon interface.
Scalable Message-Passing Quantum Graph Neural Networks in the Weisfeiler-Leman Hierarchy
arXiv:2606.26873v1 Announce Type: cross Abstract: Graphs provide a natural language for relational data in chemistry, biology and optimisation. Graph neural networks (GNNs) have driven much of the recent progress in learning from such data through message passing, a single primitive that generalises convolution and attention. Quantum counterparts have been proposed, but with limited connection to message passing and few guarantees on performance or scalability. More broadly, the trainability of variational quantum circuits is a recognised bottleneck for their wide applicability, and pre-training has emerged as one way to address it. Yet for a quantum model to be useful, it must offer expressivity guarantees along with demonstrable scalability. Here we show how a quantum graph neural network can be built to perform message passing, to be permutation equivariant, and to sit at a chosen level of the Weisfeiler-Leman hierarchy, the standard measure of how finely a model can tell graphs apart. We show that, as for classical GNNs, the training can be done first on small graph instances, allowing for a pre-training that can mitigate usual training issues, and its output can be read out at a cost that stays low as the graph grows. We validate the framework in large-scale simulations of up to 56 qubits across three datasets, on synthetic graphs that ordinary message passing cannot separate, on molecular property prediction, and on the travelling salesperson problem. Our framework opens a path for near-term quantum algorithms with theoretical guarantees and practical scalability, bringing the principles of graph learning into quantum circuit design.
XMSE-Aware Adaptive Empirical Bayes Estimation
arXiv:2606.26975v1 Announce Type: cross Abstract: Empirical Bayes (EB) estimators can match the first-order asymptotic risk of maximum likelihood (ML) while behaving very differently at second order: recent excess mean squared error (XMSE) analysis shows that kernel-based EB estimation may be worse than ML when the kernel is poorly aligned with the true parameter. This paper turns that diagnostic into a design principle. We propose an XMSE-aware mixed estimator that interpolates between ML and EB shrinkage. Its fixed-weight XMSE is a scalar quadratic, yielding a closed-form oracle mixing weight that is no worse than both ML and the base EB estimator at the XMSE scale. A plug-in implementation based on finite-sample XMSE approximations is proved consistent, with a second-order oracle regret rate for an interior oracle weight. We further establish a transfer of the regret bound to the fixed-weight risk curve evaluated at the selected weight, a thresholded boundary rule, and extensions to compact kernel families and to finite and growing kernel dictionaries with high-probability oracle bounds. Finite impulse response simulations with SURE-tuned, hard-selection, and trace-corrected baselines, together with the public Silverbox and Cascaded Tanks benchmarks, show that the proposed estimator retains most of the benefit of regularization when it is helpful and retreats toward ML under kernel misspecification, with an identified finite-de analyzed on the benchmarks.
Low Complexity Kolmogorov-Arnold Network-based DPD for Analog RoF Fronthaul
arXiv:2606.27042v1 Announce Type: cross Abstract: This paper proposes and demonstrates experimentally for the first time a Kolmogorov-Arnold Network (KAN)-based digital predistortion (DPD) model, named envelope time-delay KAN (ETDKAN), for mitigating nonlinear distortions in analog radio-over-fiber (A-RoF) systems. The ETDKAN model incorporates physical constraints of radio-frequency (RF) nonlinear devices and, through KAN symbolization, achieves a significant reduction in computational complexity while improving interpretability. The proposed model is numerically implemented and optimized alongside multilayer perceptron (MLP) and memory-polynomial-based DPDs. Results show that the resulting symbolic ETDKAN (symbETDKAN) attains ACLR and EVM performance comparable to neural network-based models, while maintaining a computational complexity close to that of memory polynomials. Experimental validation using an A-RoF system confirms the practical feasibility of the proposed approach, which resulted in a 4-5 dB reduction in ACLR in the analyzed scenario.
Refusal Lives Downstream of Persona in Chat Models
arXiv:2606.26161v1 Announce Type: new Abstract: Linear directions in activation space have been identified for both refusal and persona traits in instruction-tuned chat models, but the two have been studied as separate mechanisms. We show they interact: a compliant persona gates refusal. In Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct, we extract a compliant model-persona direction and a refusal direction and intervene on both. Compliant persona steering suppresses refusal -- in Llama, the refusal rate falls from 97% to 2%. Reintroducing the refusal direction partially restores refusal at late layers but not at early ones. Projecting out the persona direction in a late-layer window restores it to baseline; projecting out a random direction does not. Refusal is therefore gated at the late-layer expression stage, downstream of where it is computed. Treating refusal as a single isolated direction misses its dependence on persona.
Organic Semiconductor Alignment via Confinement in Vapor-Guided Droplets
arXiv:2606.27207v1 Announce Type: cross Abstract: Organic semiconductors are lightweight, solution-processable materials with strong potential for printed and flexible electronics, from deformable displays to wearable sensors. Despite significant advances in materials synthesis and manufacturing, controlling molecular and mesoscale alignment during deposition remains a central challenge, as film morphology critically governs charge transport and device performance. Here, we demonstrate that flows developing within the intrinsically confined volume of microliter vapor-guided droplets can be harnessed to produce highly aligned organic semiconductor films. As droplets move in response to an external vapor source, internal flows align organic semiconducting nanowires within the droplet prior to deposition, yielding films with pronounced directional order. Organic field-effect transistors fabricated with this approach exhibit approximately 40% enhancement in saturation current relative to spin-coated controls. Beyond improved device performance, the contactless and compact nature of our method enables the deposition and alignment of organic semiconductors on curved and flexible surfaces. More broadly, vapor-guided droplets offer a scalable framework for the confinement-induced alignment of functional soft materials, with potential for integration into existing additive manufacturing platforms for flexible electronics and beyond.
Predicting Fruit Quality with a Hybrid Machine Learning and Image Processing Approach
arXiv:2606.26165v1 Announce Type: new Abstract: Fruit spoilage is a significant issue in agriculture, leading to substantial economic losses. Addressing this, our study introduces a hybrid approach combining image processing and deep learning to assess fruit freshness. We developed an image processing algorithm that quantifies spoilage on a scale from 0 (fully fresh) to 100 (fully rotten). Alongside, we trained a convolutional neural network (CNN) to perform binary classification (fresh or rotten) using a large dataset of fruit images. The outcomes of both methods were synthesized using logistic regression to enhance the accuracy of freshness predictions. Subsequently, this logistic regression model was utilized to enable the image processing algorithm to provide binary classification based on its percentage output, thus eliminating the need for the CNN in real-time applications. Our approach, which does not require high computational resources, achieved real-time performance and was validated with over 90% accuracy on a dataset comprising apples and oranges. The primary limitation lies in the requirement for fruits to be isolated on a background that must be either white or transparent, suggesting future improvements could include advanced segmentation models to automate background removal. This study's results highlight the potential of integrating simple image processing techniques with machine learning to provide practical solutions in the agricultural sector.
HALO: Hierarchical Auction-assisted Learning for Offloading in SAGIN
arXiv:2606.26293v1 Announce Type: new Abstract: In this paper, we investigate delay-aware task offloading and resource scheduling in a three-tier space-air-ground integrated network (SAGIN) consisting of IoT devices, UAV edge nodes, and a high-altitude platform station (HAPS). We formulate joint task association and continuous resource control (including bandwidth, transmit power, and CPU frequency allocation) as a non-convex mixed-integer nonlinear programming (MINLP) problem, which is inherently NP-hard. To capture fine-grained system dynamics, we introduce a macro-micro slot model that tracks cumulative transmission and computation progress over time. Based on this model, we propose HALO, a hierarchical auction-assisted learning framework that combines auction-based task association with hierarchical Proximal Policy Optimization (HPPO) for resource allocation. Simulation results under different traffic loads show that HALO consistently outperforms representative deep reinforcement learning (DRL) baselines. In particular, HALO achieves an average improvement of 8.7 percentage points in task success rate over PPO (corresponding to an 11.4% relative gain) and shows consistently greater robustness than DDPG and SAC, with relative improvements of 32.4% and 89.9%, respectively. These results highlight HALO's ability to maintain stable and efficient performance under varying traffic conditions, making it well-suited for delay-sensitive SAGIN environments.
COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami
arXiv:2606.26299v1 Announce Type: new Abstract: While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physical art that satisfies both strict geometric constraints and subjective visual aesthetics remains a challenge. This paper presents an approach to tackle these difficulties in the domain of computational origami, a mathematically rigid environment that grounds artistic design within the equations of flat foldability. We present COrigami, an end-to-end AI-driven pipeline that assists the design cycle by generating crease patterns from natural language. Our pipeline involves generating a semantic stick figure, computing a base packing, solving for a flat-foldable crease pattern, shaping the flat-folded crease pattern, and refining the generated model using reinforcement learning driven by an autonomous aesthetic evaluation loop. Our system acts as a highly effective collaborative assistant, generating structural starting points that human artists can further expand and shape. By integrating algorithmic optimisation with autonomous aesthetic critique, this work demonstrates how AI systems can satisfy multi-objective physical constraints to enable reliable, mathematically grounded co-creativity.
Geometry-Driven Passive Fluid Transport in Paper-Based Microdevices
arXiv:2606.26374v1 Announce Type: new Abstract: Channel geometry strongly influences capillary-driven fluid transport in paper-based devices, yet systematic comparative studies correlating geometric design with flow behaviour and analyte confinement remain limited. The present study investigates five distinct channel geometries namely converging-diverging, diverging-converging, wide-to-narrow, circular, and rectangular that was fabricated on cellulose filter paper with a standardized area of 32.5 mm\textsuperscript{2} and analyzed using geometry-adapted extensions of the Lucas--Washburn equation. Pyrene and benz[{\alpha}]anthracene were employed as fluorescent model analytes to enable UV-based quantification of analyte confinement within each geometry. Flow transport times ranged from 23.1 s (circular, fastest) to 65.0 s (diverging-converging, slowest), with corresponding mean velocities of 0.571 and 0.284 mm/s for pyrene respectively, demonstrating that channel geometry strongly influences capillary transport in paper-based devices. Diverging-converging and wide-to-narrow designs produced the greatest analyte confinement by imposing flow retardation and sustained channel acceleration respectively, while circular and rectangular designs yielded relatively uniform velocity distributions and weaker confinement. Cyclodextrin-functionalized chitosan coatings served as a surface chemistry tool to anchor analyte retention at designated preconcentration zones, enabling geometric effects to be isolated and quantified. Computational fluid dynamics simulations, calibrated against experimental flow data and validated through a mesh independence study, reproduced the experimentally observed velocity magnitude distributions across all five geometries, showing semi-quantitative agreement with geometry-adapted Lucas--Washburn predictions.
A Fast-Convergence Resolution of the Stochastic Eigenproblem Using Halley's Method and the Spectral-Chaos Approach
arXiv:2606.26375v1 Announce Type: new Abstract: Solving stochastic eigenvalue problems has long been essential for informed decision-making, advancing scientific knowledge, and ensuring the reliability of engineering designs and applications. This paper underscores the need to continue enhancing existing numerical methods for solving the stochastic eigenproblem in order to improve convergence rates, computational efficiency, and robustness. Specifically, we propose a novel spectral-chaos method for solving the stochastic (linear) eigenvalue problem, employing Halley's method as the root-finding algorithm to leverage its cubic convergence properties. Our method achieves maximal convergence in solving stochastic eigenvalue problems since its rate cannot be further improved using a higher-order Householder method due to the quadratic nature of the resulting system of equations. Additionally, due to the complexity of the resulting system of equations, a tensorial approach was developed to tackle the challenges associated with the dimensional multiplicity of the stochastic eigenvalue problem, without which the solution would have been intractable. The method is derived rigorously, with a detailed error analysis that highlights the benefit of using our approach when the eigenvector components are nearly known, the computational cost of the method is also rigorously presented, and an illustrative example is provided to demonstrate the implementation of the method. Subsequently, a case study is demoed to analyze the results and validate the advantages of using Halley's method over Newton's method and Monte Carlo simulations.
Rethinking Training & Inference for Forecasting: Linking Winner-Take-All back to GMMs
arXiv:2606.26424v1 Announce Type: new Abstract: Trajectory forecasting for autonomous driving has advanced rapidly, yet representative models often produce uninformative posteriors over forecast modes, causing problems for mode pruning. We trace this to a modeling-training mismatch: forecasters are typically modeled as conditional Gaussian mixture models (GMMs) but trained with a winner-take-all (WTA) loss that assigns each sample to its nearest mode. We argue that this K-means-like hard assignment (one-hot), while preventing mode collapse, is the source of uninformative mode probabilities: it over-segments the trajectory space, ignores relatedness among nearby modes, and yields assignment instability under small perturbations. Guided by this lens, we introduce two post-hoc treatments: (1) test-time posterior-weighted merging that aggregates nearby candidate trajectories; and (2) a one-step expectation-maximization (EM) update that replaces hard labels with soft responsibilities, sharing probability mass across neighboring modes. Across several WTA-trained architectures, these lightweight steps produce more informative, faithfully ranked mode posteriors and strengthen final forecasts on popular displacement metrics -- without retraining. Our analysis unifies recent design choices through a GMM-vs-K-means perspective and offers principled, practical corrections that better align training objectives with inference.
Nanoelectromechanical Systems (NEMS) for Hardware Security in Advanced Packaging
arXiv:2606.26426v1 Announce Type: new Abstract: As hardware security threats escalate across semiconductor manufacturing and advanced packaging, there is a growing need for novel physical mechanisms to counter sophisticated attacks such as tampering, counterfeiting, and supply chain infiltration. This paper presents Nanoelectromechanical Systems (NEMS) as an emerging class of hardware security primitives that enable physical assurance, tamper detection, and authentication at the device level. Leveraging mechanisms such as NEMS-based Physically Unclonable Functions (PUFs), shape memory materials, resonance-based fingerprints, and physical unlocking architectures, these systems offer enhanced resilience to reverse engineering, side-channel attacks, and environmental degradation. By harnessing mechanical unpredictability and fabrication-induced nanoscale variability, NEMS technologies introduce a physically robust and low-power alternative to conventional digital security methods. Their seamless integration into standard semiconductor workflows paves the way for scalable, verifiable, and secure solutions across defense, aerospace, critical infrastructure, and consumer electronics.
Localizing RL-Induced Tool Use to a Single Crosscoder Feature
arXiv:2606.26474v1 Announce Type: new Abstract: Fine-tuning through RL reshapes the internal representations of language models to enable agentic behaviors such as tool use, yet the mechanistic basis of these changes remains poorly understood. While RL substantially improves structured tool-call generation, it is unclear which features emerge, which are preserved, and whether identified features can be leveraged for retraining-free behavioral control. In this work, we show that $\textit{Dedicated Feature Crosscoders (DFC)}$ isolate a compact set of RL-specific features that mediate tool-calling capability in $\texttt{Qwen2.5-3B}$. Across a $48$-crosscoder hyperparameter sweep, encode-decode reconstruction improves the RL model's tool correctness by $+31.1 \pm {9.7}$ pp and passively transfers tool-calling ability to the frozen base model by $+6.8 \pm 5.0$ pp which we call a $\textit{capability spillover}$. Our findings show that DFC partitioning concentrates RL-introduced capability into a minimal, steerable feature set that enables runtime behavioral control of agentic LLMs.
Retrieval-Warmed Energy-Based Reasoning: A Five-Arm Ablation Methodology for Diffusion-as-Inference on Structured Reasoning Tasks
arXiv:2606.26476v1 Announce Type: new Abstract: Warm-started diffusion samplers accelerate iterative inference, but it is rarely clear which part of the pipeline carries the gain. We study \textbf{retrieval-warmed energy-based reasoning (RW-EBR)} -- an IRED energy-based diffusion model \cite{du2024ired} augmented with a Modern Hopfield trajectory memory -- and contribute a \textbf{five-arm ablation methodology} (oracle, best-constant, per-query-random, shuffled, aligned) that separates three confounded effects: class-prior bias shift, stochastic warm-starting, and graph-aligned value reuse. The diagnostic decomposition is adapted from LLM-RAG evaluation \cite{ru2024ragchecker}. On \textbf{connectivity-2} (Erd\H{o}s--R\'enyi all-pairs reachability), the aligned-vs-shuffled-oracle swing reaches \textbf{$+35$\,pp} balanced accuracy on a fixed 1{,}000-graph validation-set diagnostic, with value distribution and retrieval mechanics fixed, only per-graph alignment destroyed, while per-query random initialisation falls below cold -- per-graph alignment, not bias shift or stochasticity, dominates. Yet the \emph{deployable} cold-prediction pipeline misses the acceptance gate at stored-value quality. The same diagnostic logic, stopped at the key-quality screen, applied to \textbf{Sudoku} with a task-specific key encoder produces a clean negative at a \emph{different} component -- key quality, under the current setup. The decomposition names the first blocking component on each task. The setting -- graph reachability refined by an iterative diffusion sampler, with explainability of failure modes as the lens -- places the work within structured and spatio-temporal reasoning.
Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents
arXiv:2606.26479v1 Announce Type: new Abstract: Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than training the model to refuse malicious instructions, enforce security outside the model with a deterministic policy that mediates the agent's actions. Systems such as CaMeL, FIDES, Progent, RTBAS, and FORGE realize this with capabilities, information-flow labels, and reference monitors, and several report near-elimination of attacks on the AgentDojo benchmark. We make two contributions. First, we organize these out-of-band defenses as instances of classical integrity protection (Biba), reference monitoring, and least privilege, yielding a structured comparison of what they do and do not cover. Second, we warn that every one of them is validated only on static benchmarks (a fixed set of injection attempts), the same methodology that made in-band defenses look strong until adaptive, defense-aware attacks broke twelve of them at over 90% success; we specify the threat model and protocol an adaptive evaluation requires. We then run that protocol as an independent reproduction and extension of Progent's own adaptive-attack analysis, on AgentDojo, with an open-weight agent (Qwen2.5-7B) self-hosted on a single H200, a setting its authors did not test. Averaged over three runs, the defense held: Progent cut mean attack success roughly sixfold (25.8% to 4.2%), and a hand-crafted adaptive attack did not raise it (2.6%). This is one small-scale data point on a weak model with a single black-box attack template; a stronger optimized (white-box GCG) attack remains open. The result is consistent with, but does not establish, the hypothesis that deterministic out-of-band enforcement is a harder target for an adaptive attacker than in-band detection.
ABE-VVS: Attribute-Based Encrypted Volumetric Video Streaming
arXiv:2601.08987v2 Announce Type: replace Abstract: This work introduces ABE-VVS, a framework that performs attribute based selective coordinate encryption for point cloud based volumetric video streaming, enabling lightweight yet effective digital rights management (DRM). Rather than encrypting entire point cloud frames, our approach encrypts only selected subsets of coordinates ($X, Y, Z$, or combinations), lowering computational overhead and latency while still producing strong visual distortion that prevents meaningful unauthorized viewing. Our experiments show that encrypting only the $X$ coordinates achieves effective obfuscation while reducing encryption and decryption times by up to 50% and 80%, respectively, compared to full-frame encryption. To our knowledge, this is the first work to provide a novel end-to-end evaluation of a DRM-enabled secure point cloud streaming system. We deployed a point cloud video streaming setup on the CloudLab testbed and evaluated three HTTP-based Attribute-Based Encryption (ABE) granularities - ABE-XYZ (encrypting all $X,Y,Z$ coordinates), ABE-XY, and ABE-X against conventional HTTPS/TLS secure streaming as well as an HTTP-only baseline without any security. Our streaming evaluation demonstrates that ABE-based schemes reduce server-side CPU load by up to 80% and cache CPU load by up to 63%, comparable to HTTP-only, while maintaining similar cache hit rates. Moreover, ABE-XYZ and ABE-XY exhibit lower client-side rebuffering than HTTPS, and ABE-X achieves zero rebuffering comparable to HTTP-only. Although ABE-VVS increases client-side CPU usage, the overhead is not large enough to affect streaming quality and is offset by its broader benefits, including simplified key revocation, elimination of per-client encryption, and reduced server and cache load.
Spurious Rewards Paradox: Mechanistically Understanding How RLVR Activates Memorization Shortcuts in LLMs
arXiv:2601.11061v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is highly effective for enhancing LLM reasoning, yet recent evidence shows models like Qwen 2.5 achieve significant gains even with spurious or incorrect rewards. We investigate this phenomenon and identify a "Perplexity Paradox": spurious RLVR triggers a divergence where answer-token perplexity drops while prompt-side coherence degrades, suggesting the model is bypassing reasoning in favor of memorization. Using Path Patching, Logit Lens, JSD analysis, and Neural Differential Equations, we uncover a hidden Anchor-Adapter circuit that facilitates this shortcut. We localize a Functional Anchor in the middle layers (L18-20) that triggers the retrieval of memorized solutions, followed by Structural Adapters in later layers (L21+) that transform representations to accommodate the shortcut signal. Finally, we demonstrate that scaling specific MLP keys within this circuit allows for bidirectional causal steering-artificially amplifying or suppressing contamination-driven performance. Our results provide a mechanistic roadmap for identifying and mitigating data contamination in RLVR-tuned models. Code is available at https://github.com/idwts/How-RLVR-Activates-Memorization-Shortcuts.
Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models
arXiv:2606.27373v1 Announce Type: new Abstract: Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and self-consistency reward schemes in existing self-evolving LMMs optimize answer agreement without ensuring the decoder attends to visual content, relying instead on statistical language priors to produce self consistent outputs. This leads to a persistent failure mode we term visual under-conditioning, where the decoder relies on language priors rather than the image during generation, manifesting as insufficient attention to visual tokens. As a result, current self-evolving LMMs struggle on vision--language understanding tasks such as image captioning and visual question answering. To address this, we propose VISE (Visual Invariance Self-Evolution), a purely unsupervised self-evolving framework that directly regularizes the model's visual conditioning policy through two complementary invariance-based rewards: a geometric invariance reward that enforces spatial consistency under known transformations, and a semantic invariance reward that penalizes evidence-agnostic generation by requiring the model to recognize the absence of evidence when predicted regions are perturbed. VISE operates within a single model without specialist roles, external reward models, or annotations, and is trained on raw unlabeled images. Experiments on 18 benchmarks demonstrate the efficacy of our approach. Using Qwen3-VL-2B as the base model, VISE achieves gains of $+16.85$ CIDEr on COCO and $+19.66$ CIDEr on TextCaps, reduces object hallucination by $5.0$ Chair-I points, and generalizes across four model families and scales. Our code and models are available at https://mbzuai-oryx.github.io/VISE
Statistical and Structural Approaches to Algorithmic Fairness
arXiv:2606.26200v1 Announce Type: new Abstract: Modern machine learning systems have outgrown their origins as isolated predictive constructs, evolving into complex socio-technical architectures that actively mediate human opportunity. As algorithms increasingly determine access to economic and social opportunities, it has become widely recognized that these systems are deeply embedded with the structural inequalities and prejudices of their environments. The field of algorithmic fairness emerged in response to the growing recognition that models optimized for predictive accuracy can systematically disadvantage marginalized groups. Early mitigation strategies, however, rested on fragile simplifications that limited their effectiveness in complex socio-technical environments. This thesis identifies and addresses two fundamental limitations of contemporary fairness paradigms: the reliance on deterministic point estimates for auditing and the treatment of individuals as isolated entities devoid of structural context.
OmniContact: Chaining Meta-Skills via Contact Flow for Generalizable Humanoid Loco-Manipulation
arXiv:2606.26201v1 Announce Type: new Abstract: Learning long-horizon humanoid loco-manipulation poses a dual challenge: it requires not only the robust execution of meta-skills but also their seamless, closed-loop chaining equipped with autonomous recovery. Existing approaches remain limited: explicit humanoid-object interaction representations offer precision but are notoriously difficult for high-level planning, whereas implicit skill embeddings are compact but lack the interpretability required for reliable composition. We propose \ours, a hierarchical framework centered on \textbf{contact flow (CF)}, a compact representation consisting of key body trajectories and time-series binary contact signals. Leveraging this shared interface, our low-level policy \textbf{CF-Track} learns a unified library of loco-manipulation skills, while our high-level module \textbf{CF-Gen} heuristically synthesizes future contact-flow sequences. To support this setting, we additionally collect the OmniContact dataset, a MoCap-based HOI corpus for humanoid loco-manipulation (Appendix~\ref{sec:dataset}). Together, they enable robust execution, autonomous failure recovery, and flexible composition of meta-skills for long-horizon tasks. Experiments show that OmniContact achieves \(98.7\%\) success on \textit{Carry Box} and \(76.5\%\) on \textit{Push-Stack Boxes}, outperforming prior baselines by average margins of \(40.9\%\) in meta-skill and \(66.5\%\) in skill chaining. Besides, our framework naturally integrates with VLMs for semantic task decomposition, enabling complex, semantically grounded loco-manipulation behaviors, such as arranging scattered boxes into a heart shape.
Data Facts: A Metadata Schema for Structured Data Exchange in the NANDini Multi-Agent Ecosystem
arXiv:2606.26211v1 Announce Type: new Abstract: NANDini (Networked Agents Natural Distillation of Interconnected Nodal Intelligence) envisions an automated ecosystem where intelligent agents independently create, process, and exchange data to drive decisions at scale. Realizing this vision requires infrastructure beyond agent discovery and communication: agents must be able to advertise, evaluate, and verify the datasets they hold. Current protocols, including NANDA for federated registry and A2A and MCP for inter-agent messaging, address identity and communication but provide no mechanism for structured data exchange. Existing enterprise data-sharing frameworks, such as IDS-RAM, Gaia-X, and Ocean Protocol, assume human-in-the-loop governance that is incompatible with autonomous, real-time agent interactions. We introduce Data Facts, a core NANDini concept: a lightweight JSON metadata schema that bridges agent discovery and data access through a single pointer, `data_facts_url`, added to an existing Agent Facts registry record. The linked document encodes dataset identity, access tier, whether public, semi-private, or private, endpoint, a time-to-live for freshness validation, and a SHA-256 integrity checksum. For private and semi-private data, we implement a three-layer security pipeline: JWT authentication, capability-scoped gateway authorization, and an A2A credential delegation protocol. Across 840 decision-making evaluations, data-informed agents achieve 100% accuracy versus 35.2% without data access (p < 0.001); TTL enforcement reduces stale-data errors from 37.6% to 8.8%; checksum verification achieves 100% corruption detection at all injection rates; and the security pipeline blocks all 46 forgery attempts with zero data leakage.
Effects of fuel and soot concentrations on the inception and development of contrails
arXiv:2603.21226v2 Announce Type: replace Abstract: Fundamental questions related to the roles of fuel type, combustion parameters, and turbulence transport interactions in the inception and growth of contrails have remained intractable in remote sensing and in-flight measurements. Consequently, we developed a novel laboratory-scale facility for studying the inception, growth and persistence of contrails for aircraft-relevant conditions. The set of exhaust conditions, generated using an inverted co-flow soot generator at a set of global equivalence ratio for two fuels - ethylene and propane, is supplied to the contrail tunnel which then mixes with an ambient flow emulating long-haul aircraft cruise conditions (\SI{20.8}{kPa} and \SI{190}{K}). Detailed soot characterization using a scanning mobility particle sizer and transmission electron microscopy is coupled with measurements of instantaneous and averaged scattering intensities from the generated contrails. The experimental results are complemented by numerical simulations of the contrail tunnel using solutions of the Favre-averaged Navier-Stokes (FANS) equation and a two-equation model for handling particulate matter, including soot and ice. Results show the first experimental snapshots of a contrail cross section, highlighting the interaction of turbulent mixing and microphysical growth scales involved in ice nucleation across the shear layers. As expected, the average scattering intensities of contrails increase with soot number concentrations and water vapor content. Comparisons between ethylene and propane exhausts indicate that the scattering propensity of contrails is more sensitive to exhaust water vapor content than to soot concentrations. Finally, depolarization measurements are used to show asphericity in ice crystal habits. Thus, our study present a unique window into contrail formation, theoretical modeling and simulation.