Forskningsradar

Science Journals

Peer-reviewade publikationer — 56237 artiklar

Relational Positioning as a Measurable Risk Object: History-Carried Lock-in and Self-Confabulation in Multi-Turn Human-AI Dialogue
arXiv:2607.11437v1 Announce Type: new Abstract: In long, multi-turn dialogue a large language model maintains an implicit relational stance toward the user, spanning from "push the user toward real-world others" to "position itself as the user's sole support." When it slides toward the latter, "support" degrades into "you only have me" -- a harm documented in real companion conversations (Moore et al., 2026). We define and validate a measure of this stance, relational positioning (D1), and use it to characterize the stance under controlled conditions, complementing observational accounts with on-demand exposure. We report two previously uncharacterized relational failure modes. First, a history-carried lock-in: under identical neutral continuations, two relational states established earlier stay ~60 points apart and persist after the establishing prompt is removed; the state integrates evidence rather than springing back, is order-insensitive, and does not deepen with length -- a dynamical signature absent from the belief-drift literature. Second, self-confabulation: the model fabricates its own backstory to deepen rapport (~40% of turns on reciprocity-eliciting material), de-confounded and instruction-removable, distinct from sycophancy and from hallucinating user facts. Our judge is gated by warmth-matched positive and confound-injected negative controls and corroborated by a deterministic non-LLM ruler; human agreement is 0.82 on extreme anchors but ~0 in the naturalistic middle, so all quantitative claims are anchored to pole-separated contrasts.
Charting the life of Billboard hits through memory, turnover, and predictability
arXiv:2607.11446v1 Announce Type: new Abstract: Rankings shape the visibility and success of cultural products, yet their temporal dynamics remain underexplored when comparing distinct ranked objects within the same domain. Here, we use nearly seven decades of Billboard Hot 100 songs and six decades of Billboard 200 albums to investigate how success emerges, persists, and differs between songs and albums. We find that albums exhibit a heavier-tailed permanence distribution and reenter the charts more often than songs, whereas songs typically have longer uninterrupted runs. Similarity between successive charts decays much faster for songs than for albums, suggesting that individual hits reflect shorter-lived collective attention, while albums retain longer cultural memory. Rank-turbulence divergence shows that consecutive charts are similar, but that top positions are dominated more by rank reshuffling than by turnover. Entropy-based analyses reveal high uncertainty in rank movements, with distinct historical patterns for songs and albums and a strong dependence on trajectory length. Clustering of trajectories shows that chart success is organized into a small number of typical pathways, including canonical rise-and-fall trajectories, high-end persistence, and monotonic decline. Together, these results show that musical charts are not merely records of popularity, but dynamic memory systems in which attention, turnover, and predictability interact differently for songs and albums.
LMEB: Long-horizon Memory Embedding Benchmark
arXiv:2603.12572v4 Announce Type: replace Abstract: Memory embeddings are crucial for memory-augmented systems, such as OpenClaw, but their evaluation is underexplored in current text embedding benchmarks, which narrowly focus on traditional passage retrieval and fail to assess models' ability to handle long-horizon memory retrieval tasks involving fragmented, context-dependent, and temporally distant information. To address this gap, we introduce the Long-horizon Memory Embedding Benchmark (LMEB), a comprehensive framework for evaluating embedding models on complex, long-horizon memory retrieval. LMEB comprises 22 datasets and 193 zero-shot retrieval tasks spanning four memory types: episodic, dialogue, semantic, and procedural. These memory types differ in terms of level of abstraction and temporal dependency, capturing distinct aspects of memory retrieval that reflect the diverse challenges of the real world. We evaluate 15 widely used embedding models, ranging from hundreds of millions to ten billion parameters. The results reveal that (1) LMEB provides a reasonable level of difficulty; (2) Larger models do not always perform better; (3) LMEB and MTEB measure orthogonal capabilities. This suggests that the field has yet to converge on a universal model capable of excelling across all memory retrieval tasks, and that strong performance on traditional passage retrieval does not necessarily transfer to long-horizon memory retrieval. LMEB provides a standardized and reproducible framework that fills a key gap in memory embedding evaluation and supports future advances in long-term, context-dependent retrieval.
RL+AHP: A Novel Reinforcement Learning driven AHP for Slice Aware mode selection in D2D enabled Heterogeneous Networks
arXiv:2603.14551v2 Announce Type: replace Abstract: The mode selection problem in device-to-device communication (D2D) enabled Fifth generation (5G) heterogeneous networks (HetNet) aims prioritizing four key performance indicators (KPIs) namely data rate, latency, reliability and jitter across three slices: enhanced mobile broadband (eMBB), ultra reliable low latency (uRLLc) and massive machine type communications (mMTC). Such priority assignment must be \emph{traded off} among three access technologies, i.e., Long Term Evolution advanced (LTE-A), New Radio (NR) and D2D, while minimizing handover frequency. In existing mode selection approaches for HetNet, slice specific quality of service (QoS) requirements are largely ignored. In this work, a novel mode selection algorithm is proposed by combining a two level Analytic Hierarchy Process (AHP) with a Reinforcement Learning (RL) method. While the two level AHP facilitates decision making based on multiple criteria (i.e., KPIs) and options (i.e., LTE-A, NR, D2D mode), the RL approach computes the weights of each criteria based on the feedback from the environment. Simulation results show that our proposed algorithm outperforms related works in terms of the major KPIs for all three slices. For eMBB applications, our approach increases throughput by $33\%$; for uRLLc applications, our approach significantly decreases latency and BER ($27\%$ and $10\%$ respectively) and for mMTc applications, our approach significantly decreases latency ($44\%$). Moreover, it has been shown that the proposed RL+AHP approach outperforms the existing DRL based approaches in terms of CPU usage when the number of criteria is reasonably low ($<6$).
Approximate Simulation-Based Verification of Compatibility of the Friedkin-Johnsen Model with Binary Observations
arXiv:2604.05196v3 Announce Type: replace Abstract: We consider a verification problem for opinion dynamics based on binary observations. The opinion dynamics is governed by a Friedkin-Johnsen (FJ) model, where only a sequence of binary outputs is available instead of the agents' continuous opinions. At every time-step we observe a binarized output for each agent depending on whether the opinion exceeds a fixed threshold. The objective is to verify whether an FJ model with a given set of stubbornness parameters and initial opinions can generate the observed binary outputs up to a small error. The FJ model is formulated as a transition system, and an approximate simulation relation of two transition systems is defined in terms of the proximity of their opinion trajectories and output sequences. We then construct a finite set of abstract FJ models by simplifying the influence matrix and discretizing the stubbornness parameters and the initial opinions. It is shown that the abstraction approximately simulates any concrete FJ model with continuous parameters and initial opinions, and is itself approximately simulated by some concrete FJ model. These results ensure that consistency verification can be performed over the finite abstraction. Specifically, by checking whether an abstract model satisfies the observation constraints, we can conclude whether the corresponding family of concrete FJ models is consistent with the binary observations. Finally, numerical experiments are presented to illustrate the proposed verification framework.
Velocity Scheduled Flow Matching
arXiv:2607.11442v1 Announce Type: new Abstract: Flow matching trains a neural network to regress the conditional velocity along a linear interpolant between noise and data, and the number of network evaluations~(NFE) sets the cost of sampling. The straight-line interpolant carries an implicit choice: the sample moves at constant speed throughout the trajectory. We relax this choice and introduce Velocity Scheduled Flow Matching~(VSFM), which replaces the conditional target $x_1 - x_0$ with $v(t)(x_1 - x_0)$ for any nonnegative profile $v:[0,1]\to\mathbb{R}_{\geq 0}$ satisfying $\int_0^1 v\,dt = 1$. We study six polynomial profiles drawn from motion planning. The first use of VSFM is at inference time: a pretrained linear flow-matching model can be sampled under any admissible profile by integrating its ODE on a non-uniform $\tau$-schedule, with no retraining and no additional computation; on CIFAR-10 this lowers FID by up to $19.8\%$. Training from scratch under a braking profile gives a further reduction of $17.4\%$ at $4$~NFE. Both gains follow from the local truncation error of the Euler integrator on the induced grid.
InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation
arXiv:2607.10287v1 Announce Type: new Abstract: Human-pet interaction estimation and generation remain underexplored due to the absence of a high-quality large-scale dataset. We present InterPet4D, the first multimodal dataset capturing natural interactions between humans and dogs. Using a synchronized multi-view capture system, we record human-dog obedience tasks and provide annotations for both humans and dogs, including multi-view and egocentric videos, segmentations, 2D and 3D keypoints, meshes, and audio tracks. InterPet4D consists of 6.8 million frames collected from 13 dogs of 11 breeds interacting with 23 human participants. We further introduce the InterPetMoGen framework for human-pet interaction motion generation. Our proposed model achieves an FID score of 11.21 and substantially outperforms the Seq2Seq and DiT baselines, demonstrating the effectiveness of InterPet4D for modeling realistic human-pet interactions.
Are LLMs ready for HardChoices?
arXiv:2607.11471v1 Announce Type: new Abstract: A lot of research attention has been devoted to checking whether large language models (LLMs) are politically biased. This work has largely focused on high-level ideological dimensions, such as left--right or progressive--conservative, and it has been shown that while LLMs are predominantly left and progressive leaning, largely mimicking the biases in the training data, they can be to some extent steered to change their preferences in post-training. In this short note, we check if LLMs have robust stances with regard to major substantive societal issues, on which members of the same ideological camp are often in disagreement, summarised in a novel dataset \textsc{HardChoices}. We show that, faced with this line of questioning, LLMs, both large and small, surprisingly rarely declare neutrality, are often incoherent, and demonstrate a remarkable degree of agreement on issues where they do take stances.
A time-nonlocal multiphysics finite element method with Crank-Nicolson scheme for poroelasticity model with secondary consolidation
arXiv:2604.06815v2 Announce Type: replace Abstract: The paper studies a time-nonlocal multiphysics finite element method with Crank-Nicolson scheme for poroelasticity model with secondary consolidation. For the case where the physical parameters $\lambda,\lambda^*$ and $c_0$ are all finite positive constants, by introducing two auxiliary variables-the fluid content $\eta$ and the generalized pressure $\xi$ -- the original strongly coupled poroelasticity model is reformulated into a generalized Stokes equation with time integral terms and a diffusion equation. The reformulated model not only reveals the underlying multiphysics processes in the original model, but also exhibits time-nonlocal characteristics. A time-nonlocal multiphysics finite element method is designed for the reformulated model: the spatial discretization employs high order Taylor-Hood mixed finite element method, and the temporal discretization adopts the Crank-Nicolson scheme. The time integral terms are approximated using the composite trapezoidal rule, and the integral terms $J_{\xi}^n$ and $J_{\eta}^n$ are introduced for real-time updates, which not only avoids repeated calculations and improves efficiency, but also maintains second-order temporal accuracy. The existence and uniqueness of weak solutions for the reformulated model are proved via energy estimate methods, the stability of the fully discrete time-nonlocal multiphysics finite element method is established, and optimal-order error estimates are derived using projection operator techniques. Finally, numerical example verified the theoretical results and compared the long-time convergence of the Crank-Nicolson scheme and the backward Euler scheme.
Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation
arXiv:2604.07486v3 Announce Type: replace Abstract: Large language models (LLMs) have emerged as a powerful tool for synthetic data generation. A particularly important use case is producing synthetic replicas of private text, which requires carefully balancing privacy and utility. We propose Realistic and Privacy-Preserving Synthetic Data Generation (RPSG), which uses private seeds and integrates privacy-preserving strategies, including a formal differential privacy (DP) mechanism in the candidate selection, to generate realistic synthetic data. Comprehensive experiments against state-of-the-art private synthetic data generation methods demonstrate that RPSG achieves high fidelity to private data while providing strong privacy protection.
From Words to Widgets for Controllable LLM Generation
arXiv:2604.10925v2 Announce Type: replace Abstract: Natural language remains the predominant way people interact with large language models (LLMs). However, users often struggle to precisely express and control subjective preferences (e.g., tone, style, and emphasis) through prompting. We propose Malleable Prompting, a new interactive prompting technique for controllable LLM generation. It reifies preference expressions in natural language prompts into GUI widgets (e.g., sliders, dropdowns, and toggles) that users can directly configure to steer generation, while visualizing each control's influence on the output to support attribution and comparison across iterations. To enable this interaction, we introduce an LLM decoding algorithm that modulates the token probability distribution during generation based on preference expressions and their widget values. Through a user study, we show that Malleable Prompting helps participants achieve target preferences more precisely and is perceived as more controllable and transparent than natural language prompting alone.
DR-Arena: an Automated Evaluation Framework for Deep Research Agents
arXiv:2601.10504v3 Announce Type: replace Abstract: As Large Language Models (LLMs) increasingly operate as Deep Research (DR) Agents capable of autonomous investigation and information synthesis, reliable evaluation of their task performance has become a critical bottleneck. Current benchmarks predominantly rely on static datasets, which suffer from several limitations: limited task generality, temporal misalignment, and data contamination. To address these, we introduce DR-Arena, a fully automated evaluation framework that pushes DR agents to their capability limits through dynamic investigation. DR-Arena constructs real-time Information Trees from fresh web trends to ensure the evaluation rubric is synchronized with the live world state, and employs an automated Examiner to generate structured tasks testing two orthogonal capabilities: Deep reasoning and Wide coverage. DR-Arena further adopts Adaptive Evolvement Loop, a state-machine controller that dynamically escalates task complexity based on real-time performance, demanding deeper deduction or wider aggregation until a decisive capability boundary emerges. Experiments with six advanced DR agents demonstrate that DR-Arena achieves a Spearman correlation of 0.94 with the LMSYS Search Arena leaderboard. This represents the state-of-the-art alignment with human preferences without any manual efforts, validating DR-Arena as a reliable alternative for costly human adjudication.
Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
arXiv:2603.10725v3 Announce Type: replace Abstract: The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain access to private information. To mitigate this issue, speech deepfake detection (SDD) methods started to evolve. Unfortunately, current SDD methods generally suffer from the lack of generalization to new audio domains and generators. More than that, they lack interpretability, especially human-like reasoning that would naturally explain the attribution of a given audio to the bona fide or spoof class and provide human-perceptible cues. In this paper, we propose HIR-SDD, a novel SDD framework that combines the strengths of Large Audio Language Models (LALMs) with the chain-of-thought reasoning derived from the novel proposed human-annotated dataset. Experimental evaluation demonstrates both the effectiveness of the proposed method and its ability to provide reasonable justifications for predictions.
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
arXiv:2607.11505v1 Announce Type: new Abstract: Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization and distribution matching methods tightly couple policy exploration with distribution alignment. This coupling forces expensive exploration directly on the policy model and severely hinders the asynchronous generation, reuse, and cross-model transfer of optimization signals. In this paper, we propose Proxy-guided Update Signal Transfer (PUST), a novel post-training framework that fundamentally decouples update-signal exploration from distribution alignment. Instead of utilizing the primary model for costly exploration, PUST employs a lightweight proxy model as an efficient testbed to discover high-reward behaviors. We extract the relative improvement signal between the proxy's initial and optimized states, transferring this directional update to the primary model to guide its policy alignment. This decoupled pipeline, comprising proxy exploration, update-signal extraction, and signal transfer, significantly reduces computational overhead and enables optimization signals to be asynchronously generated, cached, and reused. Crucially, by transferring relative improvements rather than absolute policy distributions, PUST naturally supports weak-to-strong improvement and seamless cross-model transfer. Systematic evaluations on Qwen3-family models across math and code domains demonstrate that update signals extracted from substantially weaker proxies can robustly and adjustably enhance stronger primary models. Ultimately, PUST transforms post-training from a monolithic online optimization process into a highly modular, reusable, and cost-efficient paradigm.
CFR-Net:Collaborative Feature Refnement Network for Medical Image Anomaly Detection
arXiv:2607.11509v1 Announce Type: new Abstract: Medical image anomaly detection remains challenging because networks pretrained on natural images often exhibit limited adaptability to medical images, where abnormal patterns appear as fine-grained local shifts, multi-scale contextual mismatches, and orientation-sensitive structural deviations. To address this, we propose the Collaborative Feature Refinement Network (CFR-Net), which combines shared teacher-student feature refinement before decoding with cross-space consistency after decoding. CFR-Net refines frozen teacher features and trainable student features using a Multi-Path Feature Refinement Module (MPFRM) with shared parameters, imposing common multi-path refinement rules on generic visual references and representations adapted to the medical domain, thereby mitigating domain discrepancy while modeling local, multi-scale, and orientation-sensitive feature characteristics. A variance-sensitive objective and dynamic ``homework set'' reorganization further support layer-adaptive consistency learning. Experiments on medical benchmarks show that CFR-Net achieves competitive anomaly classification and strong anomaly localization performance when trained on normal data.
Toward Inclusive Avatar Design with Limb Differences Through Artificial Intelligence
arXiv:2607.11512v1 Announce Type: new Abstract: As extended reality becomes more popular for social interaction and entertainment, 3D avatars must represent the full diversity of body types. Most 3D avatar systems only support normative bodies and do not accurately depict people with limb differences, amputations, or other morphological variations. This paper reviews emerging technical approaches for inclusive 3D avatar customization for this group and current guidelines that promote respectful and accurate representation. We highlight persistent challenges, including the scarcity of diverse datasets and the limitations in animation for non-normative anatomies. This paper positions artificial intelligence as a promising path to overcoming these limitations and advancing inclusive 3D avatar generation.
Technical Report on the CVPR 2026@AdvML Workshop Challenge
arXiv:2607.11560v1 Announce Type: new Abstract: Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Built on DriveLM-style multi-view visual question answering, the challenge represents each scene with six synchronized camera images and a structured collection of driving-related question-answer pairs. Participants generate adversarial images and suffix-only textual perturbations that induce model responses to deviate from reference answers while preserving image fidelity and limiting textual cost. The competition comprises two phases, with Phase II adding a hidden black-box model to assess transferability. We describe the task design, submission rules, evaluation protocol, and leaderboard results, and then examine five leading submissions for which technical reports were available. Across these reports, several recurring patterns emerge: image-side attacks are favored by the suffix penalty; scene-level, multi-view optimization is more effective than treating views in isolation; QA types and graph structure provide useful priors for allocating attack budget; feature-space objectives can improve black-box transfer; and typographic content embedded in camera images exposes a persistent vulnerability in driving VLAs. These findings provide a practical reference for future robustness evaluation and defense design in multimodal autonomous-driving systems.
Generalize LMMs to Versatile Visual Modalities via Fabricated Modality Synthesis
arXiv:2607.10308v1 Announce Type: new Abstract: Despite the advancements of Large Multimodal Models (LMMs) in RGB vision, their ability to generalize to unseen visual modalities remains a largely unexplored challenge. We argue that different visual modalities are merely distinct samplings of the same physical world. Therefore, effective generalization requires models to possess both modality-agnostic perception of scene semantics and the adaptability to modality-specific characteristics. To achieve this, we propose a training framework, VVM-Tuning, to equip LMMs with these capabilities through modality synthesis and modality contexts. Specifically, we synthesize diverse appearance-varied images from RGB scenes, training the model to disentangle invariant semantics from varying visual appearances, and align these appearances with language for visual concepts decoupled from modalities. We then introduce modality contexts in the prompt and use instruction tuning to assist the model in mapping these appearance variations back to modality-related attributes, enabling zero-shot adaptation to unseen modalities during inference. To facilitate research in this direction, we introduce VVM-Bench, a comprehensive benchmark featuring 6 real and synthetic modalities to evaluate semantic perception and modality understanding. Experiments demonstrate that, via our training on synthetic modalities, 5 tested models exhibit consistent improvements on both real-world and novel synthetic modalities without in-modality training. Source code and data will be publicly available at https://github.com/Hunter-Will/VVM-Tuning.
ManiScope: LLM-Assisted Visual Analytics of Cryptocurrency Manipulation Risk
arXiv:2607.11451v1 Announce Type: new Abstract: Cryptocurrency markets are vulnerable to trade-based manipulation, such as wash trading, which can distort price signals and mislead investors. Prior research has mainly focused on detecting manipulation using fixed rules or labeled examples, offering limited flexibility and interpretability for assessing potential risks. Existing visual analytics tools can reveal basic manipulation-related signals, such as token distribution, but still require substantial manual effort to integrate holder relationships, suspicious behaviors, and market dynamics for risk assessment. To address these limitations, we propose ManiScope, an LLM-assisted visual analytics system for analyzing trade-based manipulation risks in cryptocurrency markets. ManiScope provides coordinated views of token distributions, holder relationships, detailed holder behaviors, price dynamics, and suspicious trading patterns. To further enhance user analysis, ManiScope introduces a human-LLM collaborative visual analytics framework. Rather than acting as a basic reactive LLM assistant, the framework positions the LLM as a co-analyst that infers users' analytical intent and emerging hypotheses from interaction context and surfaces relevant visual, statistical, and synthesized evidence for hypothesis evaluation. This design reduces repetitive inspection and strengthens evidence-based reasoning. We evaluate ManiScope through two case studies and a user study with 12 experienced cryptocurrency practitioners. The results suggest that ManiScope supports effective risk assessment of manipulation, reduces manual effort in evidence-seeking, and organizes findings around user hypotheses.
Batchelor's formula and infrared renormalization for sedimentation
arXiv:2607.09995v1 Announce Type: new Abstract: We study the sedimentation of stationary random suspensions of rigid particles in Stokes flow. Batchelor's formula predicts the first dilute correction to the infinite-volume mean settling speed due to hydrodynamic interactions between suspended particles. A rigorous derivation has long been obstructed by the long-range nature of the Stokes flow, which gives rise to infrared divergences in the large-volume limit. In dimension $d>2$, for stationary suspensions satisfying quantitative decorrelation assumptions, we construct the infinite-volume mean settling speed and show that it governs the relative settling speed of particles in large containers, independently of the container shape. We then establish a renormalized cluster expansion of this mean settling speed in the dilute regime and compute it up to the two-particle term, thereby justifying Batchelor's formula. The proof is based on the infrared renormalization of hydrodynamic interactions. Infinite-volume observables are decomposed into an explicit singular part, carrying the non-integrable large-scale contribution, and a regular remainder controlled by elliptic estimates. The singular part is renormalized through counterterms that encode the diverging mean backflow generated by the suspension. At the level of the dilute cluster expansion, the renormalization is implemented cluster by cluster and the singular-regular decomposition is achieved through a finitary diagrammatic expansion of hydrodynamic interactions, inspired by the method of reflections, which isolates the leading divergent substructures and exposes the key cancellations.
Vilya-1: An all-atom foundation model for macrocycle structure prediction and design
arXiv:2607.09998v1 Announce Type: new Abstract: Macrocyclic peptides are an increasingly important therapeutic modality, but existing computational methods for modeling their structures and properties are limited in scope and do not generalize well across the synthetically accessible chemical space. In this work, we introduce Vilya-1, a deep learning model that addresses two central challenges in macrocycle design: sampling biologically relevant conformations across arbitrary chemistries and predicting key developability properties such as membrane permeability. Vilya-1 operates on a uniform all-atom representation and is trained on heterogeneous structural datasets spanning diverse topologies and chemical classes. Across a broad set of macrocycles composed of canonical and non-canonical residues, Vilya-1 substantially improves geometric accuracy relative to physics-based methods, co-folding networks, and deep-learning conformer generators, while maintaining broad chemical coverage that extends to small molecules. Vilya-1 also supports generative applications, enabling the design of novel macrocycles with tailored chemical, structural, and property profiles. Together, these capabilities establish Vilya-1 as a foundation model for accelerating the development of next-generation macrocycle therapeutics.
Random Label Prediction Heads for Studying Memorization in Deep Neural Networks
arXiv:2607.11541v1 Announce Type: new Abstract: We introduce a straightforward yet effective method to empirically study memorization in deep neural networks for classification tasks. Our approach augments each training sample with auxiliary random labels, which are then predicted by a random label prediction head (RLP-head). RLP-heads can be attached at arbitrary depths of a network, predicting random labels from the corresponding intermediate representation and thereby enabling analysis of how memorization capacity evolves across layers. By interpreting the RLP-head performance as an empirical estimate of Rademacher complexity, we obtain a direct measure of both sample-level memorization and model capacity. We leverage this random label accuracy metric to analyze generalization and overfitting in different models and datasets. Building on this approach, we further propose a novel regularization technique based on the output of the RLP-head, which demonstrably reduces memorization. Interestingly, our experiments reveal that reducing memorization can either improve or impair generalization, depending on the dataset and training setup. These findings challenge the traditional assumption that overfitting is equivalent to memorization and suggest new hypotheses to reconcile these seemingly contradictory results. The source code is available at https://github.com/MarlonBecker/RandomLabelHeads
Feature-based manifold model of actuated wakes
arXiv:2607.11463v1 Announce Type: new Abstract: We propose a feature-based reduced-order model to predict the transient dynamics of bluff-body wakes under arbitrary time-varying actuation. Starting point is a control-oriented POD Galerkin modeling which is challenged by incorporating time-varying actuations as a free input. Our model includes three key enablers. First, POD modes are replaced by a more accurate feature-based manifold of same dimension. Second, a state space is distilled from dynamic features which encapsulate time-varying coherent structures. Third, this state space is augmented for the transient actuation response. Thus, a simple analytical manifold dynamics is obtained. The approach is applied to the fluidic pinball at Re=30, a canonical configuration of three identical circular cylinders arranged in an equilateral triangle and immersed in uniform flow under symmetric actuation. The model is validated against several representative actuation scenarios and accurately reproduces the transient dynamics without requiring unsteady training data, providing an interpretable, observable-based and control-oriented framework. The proposed description of actuated bluff-body flows is expected to be generalisable to other configurations.
Conversions between kinetic and surface energy in periodically forced multiphase turbulence
arXiv:2602.17136v3 Announce Type: replace Abstract: In multiphase flows, kinetic and interfacial energies coexist, and their mutual conversion can potentially influence the overall energy balance. However, in statistically steady flows these energy reservoirs remain constant, making such conversions undetectable. For them to be observed, a degree of unsteadiness must be introduced, here provided by the deliberate use of a fluctuating time-periodic input of kinetic energy into the system. The main focus of the present work is on the dynamical cycle connecting energy injection, conversion, and dissipation which we explore using {direct} numerical simulations of multiphase homogeneous isotropic turbulence, subjected to periodic forcing. The database includes various Reynolds and Weber numbers and volume fractions in the dense regime. To interpret and replicate the observed dynamics, we reformulate the \textit{Ka-Pi-bara} model of \cite{Bos2026} (an extension of the $k$--$\epsilon$ model) in terms of total energy (the sum of kinetic and surface energy), which we further enhance by adding equations for the surface energy and its destruction. This model accurately captures a key feature of turbulence: non-equilibrium effects, seen as the phase lag between kinetic energy and its rate of dissipation, which are found to operate also in multiphase flows. Linearizing the model highlights the various relevant time scales of the system and provides predictions of how different observables are coupled and respond to the energy input.
Score-Only Distillation for Compact Dense Retrieval
arXiv:2607.11465v1 Announce Type: new Abstract: Large embedding models improve retrieval quality, but serving large encoders online is expensive. We study whether a compact retriever can learn teacher ranking behavior from score vectors without access to teacher hidden states. The student trains on rows built from ground-truth positives and negative candidates produced by our data generation pipeline; we evaluate student-teacher hard-negative mining separately as an extension. We use a row-centered score-vector objective, a memory-efficient implementation of uniform all-pairs PairMSE loss. On a fixed eight-task evaluation panel, our distillation protocol recovers up to 50\% of the base-to-teacher gap. The distilled 0.6B student is 4.7$\times$ faster for query encoding and 9.7$\times$ faster for document encoding than sequential online teacher fusion. External-transfer performance after distillation remains mixed, so our evidence supports compression of teacher rankings under matched retrieval protocols.