Forskningsradar

Science Journals

Peer-reviewade publikationer — 55347 artiklar

Full Quantum and Mixed Quantum--Classical Dynamics of Hot Exciton Cooling in Semiconductor Nanocrystals
arXiv:2605.27708v1 Announce Type: new Abstract: Hot-exciton relaxation in semiconductor nanocrystals (NCs) is often described using perturbative theories, but their accuracy is difficult to assess for realistic exciton--phonon Hamiltonians. Here, we benchmark the perturbative quantum master equation (QME) and several mixed quantum--classical (MQC) methods against fully quantum mechanical dynamics. Using atomistically parameterized models for CdSe core and CdSe/CdS core--shell NCs, we find that bare CdSe exhibits an ultrafast initial decay followed by slower cooling, whereas the core--shell system is dominated by the slower component. Analysis of reduced models shows that the ultrafast component arises from rapid diabatic state mixing driven by thermal fluctuations of low-frequency phonons, rather than from nuclear-assisted energy relaxation. The QME captures the initial fast decay but can fail for the slower relaxation in the diabatic representation, while the mapping approach to surface hopping (MASH) gives the most consistent agreement with both benchmark dynamics and equilibrium populations. These results establish a benchmark for exciton-cooling dynamics in NCs and clarify the physical regimes in which widely used approximate methods are reliable.
HarmoVid: Relightful Video Portrait Harmonization
arXiv:2605.28811v1 Announce Type: new Abstract: We present a method for harmonizing the lighting of a foreground video to match a target background scene, adjusting shadows, color tone, and illumination intensity (relightful harmonization). Unlike images, acquiring labeled data for videos, where identical motions are recorded under different lighting conditions, is practically infeasible and non-scalable. While one way to create such paired data is to apply existing image-based harmonization models frame by frame to a video, the resulting outputs often suffer from significant temporal jitters. We overcome this problem by introducing a novel lighting deflickering model that can stabilize the global and local lighting flickering artifacts. Our video diffusion model learns from these upgraded deflickered data with a volume of real and synthetic videos to generate high-quality video harmonization results. We further propose an asymmetric alpha mask conditioning technique to learn the clean boundaries from real videos. Experiments demonstrate that our model achieves strong temporal coherence, naturalness, cleaner boundaries, and physically meaningful lighting behavior, while maintaining strong relighting expressiveness compared to prior image-based and video-based harmonization methods.
DeepSciVerify: Verifying Scientific Claim--Citation Alignment via LLM-Driven Evidence Escalation
arXiv:2605.27710v1 Announce Type: new Abstract: Misalignment between claims and their cited evidence is a common failure mode in reports generated by large language models, limiting their reliability in scientific and other high-stakes settings. We present DeepSciVerify, a two-stage pipeline for scientific claim-citation verification that combines abstract-level reasoning with selective escalation to passage-level evidence. The system first verifies claims using the abstract and defers uncertain cases, retrieving and analyzing full-text passages only when necessary. This design leverages complementary behaviors across LLMs, as some models are more conservative while others are more decisive under uncertainty. On the SCitance benchmark, DeepSciVerify achieves 86.7 Micro-F1, outperforming strong abstract-only baselines by +4.5 points while resolving 67% of instances without full-text retrieval. These results suggest that selective evidence escalation improves both accuracy and efficiency in claim-citation verification.
Boundary Suppression Asymmetry in Post-trained Assistants: Over-expansion as a Controllability Cost
arXiv:2605.27969v1 Announce Type: new Abstract: Post-trained language-model assistants are often optimized to avoid under-answering, encouraging complete, helpful, cautious, and proactive responses. We ask whether this optimization creates asymmetric controllability costs: when users explicitly request narrower answers, which assistant behaviors remain suppressible, and which continue to shape the response? We study this problem as boundary-suppression asymmetry. Prompt-side probes across multiple high-level response dimensions suggest a selective cost, concentrated around `too-much assistant' directions such as over-completion, extra help, and anti-underanswering. Using controlled assistant-policy variants derived from a shared base model, we find that anti-underanswering policies are harder to pull back than the baseline under matched boundary-control evaluations, while minimal-boundary variants generally avoid this anti-side upward shift in the direct boundary-control comparisons. Mechanism-oriented probes point beyond longer default outputs, pure EOS failure, uncertainty compensation, and local continuation bias, while robustness checks preserve the main anti-over-baseline ordering under shared-system and larger-scale settings. The evidence supports a mixed planning/stopping account, where content-budget overshoot and continuation persistence jointly make boundary correction harder. Overall, post-training may create direction-specific controllability costs: some helpful assistant tendencies remain easy to invoke, yet harder to locally suppress.
On a Constraint on Invariant Measures of Certain Cellular Automata
arXiv:2604.10124v5 Announce Type: replace-cross Abstract: In [6], a constraint on invariant measures of bi-permutative cellular automata has been observed: fixed values at the positive indices determine almost-surely a uniform conditional probability on the subset of values of positive conditional probability at the zero index. When the alphabet is a finite group and the automaton is multiplication of two neighbors, that set is in fact a coset of some subgroup. In the present paper, we strengthen the formulations in [6] and investigate further the implications of this constraint. In the finite group case mentioned above, relations between some attributes of the group structure and the invariant measures are examined. We also inspect a factor, with respect to the shift, that this constraint induces, and analyze the special case in which it has zero measure-theoretical entropy, thus observing an interplay between existence of zero entropy invariant measures on that factor and existence of positive entropy measures corresponding to them on the original system. Then, we leave the setting of bi-permutative cellular automata and generalize our results to a wider class which we named RLP subshifts. The peculiar situation is that although this class may be much larger than the class of bi-permutative cellular automata, we were able to prove only for essentially one other example - the symbolic coding of the times 2 times 3 system on the circle (and its generalizations) - that it belongs to it.
A cavity-less architecture for high-power integrated frequency combs
arXiv:2605.27528v1 Announce Type: new Abstract: Photonic chip-based frequency combs have emerged as a transformative platform, enabling compact, scalable, and high-performance multiwavelength sources with far-reaching impact across science and technology. Most commonly, these sources leverage the cavity enhancement of the nonlinearities to produce a spectrum of equidistant frequency lines via cascaded four-wave mixing in high-quality microresonators pumped with a continuous wave tone. While the presence of the resonator inherently enables low-power threshold operation, it also brings intrinsic limitations in efficiency, tunability, and power per line. Here, we propose and demonstrate a cavity-less approach for the generation of optical frequency combs on-chip, which relies on non-degenerate cascaded four wave mixing in dispersion engineered integrated photonic waveguides. The results presented here enable previously inaccessible regimes of pump-to-comb conversion efficiency, wide-range continuous line-spacing tunability, and power per line for coherent comb states. This work opens new research opportunities in nonlinear integrated photonics and pathways toward high-capacity optical interconnects, scalable photonic AI accelerators, and other power-constrained integrated systems.
Radiative Response of Atomic Systems Illuminated with Approximate Spherical Vector Waves
arXiv:2605.27558v1 Announce Type: new Abstract: The natural electromagnetic modes spontaneously emitted by an atom in free space are spherical vector waves (SVWs). Each SVW mode is uniquely linked to a specific dynamical--spherical--multipole--moment of the atomic system. In this work, we introduce a general formalism for evaluating spherical multipole transition rates under different boundary conditions, considering the superposition between a given SVW and the modes resulting from the boundary conditions. This formalism is applied to study the radiative properties of an atomic system trapped near the focus of a 4$\pi$ optical array. By appropriately selecting the external light field, the juxtaposed lenses of the optical array allow the atom to be illuminated with approximate spherical vector waves. Explicit expressions for the resulting multipole transition rates are presented as a function of the numerical aperture of the lenses. The feasibility of enhancing and inhibiting electric dipole-forbidden transitions using such an array, under current experimental capabilities, is briefly discussed.
A Factory-Floor Deployment Case Study of VLA Pipelines for Industrial Packaging Task: Workflow, Failures, and Lessons
arXiv:2605.27461v1 Announce Type: new Abstract: Vision-Language-Action (VLA) policies have shown promising manipulation capabilities, yet their practical impact is often limited by the reliability demands of real-world deployment. We present a deployment study of an industrial packaging task at Siemens Factory (GWE, Erlangen, Germany), where a robot must pick a transparent accessory bag from a cluttered pile, insert it into the remaining cavity of a cardboard package, and ensure that the bag and its contents remain below the closing plane. Our goal is to understand the practical effort required to adapt a pretrained Pi0.5 policy to a single factory-floor task through iterative fine-tuning and deployment-driven refinement. The pipeline consists of repeated loops of data collection, curation, fine-tuning, evaluation, and targeted recovery data collection. We have accumulated 2535 episodes (10 hours) from the on-site factory settings. In this paper, we contribute an empirical account of a factory-floor VLA deployment, highlighting recurring failure modes and lessons that inform how to improve the deployment workflow.
Corrected Samplers for Discrete Flow Models
arXiv:2601.22519v3 Announce Type: replace-cross Abstract: Discrete flow models (DFMs) have been proposed to learn the data distribution on finite state space, offering a flexible framework as an alternative to discrete diffusion models. A line of recent work has studied samplers for discrete diffusion models, such as tau-leaping and Euler solver. However, these samplers require a large number of iterations to control discretization error, since the transition rates are frozen in time and evaluated at the initial state within each time interval. Moreover, theoretical results for these samplers often require boundedness conditions of the transition rate or they focus on a specific type of source distributions. To address those limitations, we establish non-asymptotic discretization error bounds for those samplers without any restriction on transition rates and source distributions, under the framework of discrete flow models. Furthermore, by analyzing a one-step lower bound of the Euler sampler, we propose two corrected samplers: \textit{time-corrected sampler} and \textit{location-corrected sampler}, which can reduce the discretization error of tau-leaping and Euler solver with almost no additional computational cost. We rigorously show that the location-corrected sampler has a lower complexity than existing parallel samplers. We validate the effectiveness of the proposed method by achieving better generation quality with reduced inference time on simulations and text-to-image generation tasks. Code can be found in https://github.com/WanZhengyan/Corrected-Samplers-for-Discrete-Flow-Models.
ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversations
arXiv:2605.27908v1 Announce Type: new Abstract: Existing emotional support conversation (ESC) systems mainly rely on end-to-end response generation or coarse strategy supervision, offering limited interpretability and little support for systematic skill improvement. We propose ESC-Skills, a skill-centric framework that discovers and self-evolves executable emotional support skills. We first model localized support interactions as Intervention Units (IUs), which capture state--action--outcome dynamics between seeker states, support interventions, and post-response emotional changes. Based on IUs extracted from both successful and failed ESC dialogues, we construct the ESC-Skills Bank, a repository of executable emotional support skills containing intervention guidance, applicability conditions, expected outcomes, and potential risks. To further improve robustness, we introduce a multi-profile self-evolutionary refinement framework in which an ESC agent interacts with diverse simulated seeker profiles under SAGE evaluation. The resulting interaction traces are analyzed to identify missing skills, unsafe interventions, and profile-specific failure patterns, which are then used to refine the Skills Bank through simulation-based verification. Experimental results demonstrate that ESC-Skills improves both response-level quality and dialogue-level emotional outcomes while providing more interpretable and controllable support behaviors. We will release the code, prompts, and ESC-Skills Bank at https://github.com/aliyun/qwen-dianjin.
MSCGC-KAN: Multi-scale Causal Graph Convolution and Kolmogorov-Arnold Feature Mapping for EEG Emotion Recognition
arXiv:2605.26624v2 Announce Type: replace Abstract: Electroencephalogram (EEG)-based emotion recognition is an important affective computing task, and recent EEG foundation models provide useful generic representations for downstream adaptation. However, under the fine-tuning setting, three limitations remain prominent: insufficient modeling of multi-scale emotional dynamics, inadequate exploitation of inter-channel functional connectivity, and the limited expressive power of simple linear classification heads. To address these issues, this paper proposes a new EEG emotion recognition method, termed MSCGC-KAN, which introduces a structured task head composed of multi-scale causal graph convolution and Kolmogorov--Arnold feature mapping. Built on a pre-trained CBraMod backbone, MSCGC-KAN enhances downstream adaptation by jointly strengthening multi-scale temporal modeling, learnable inter-channel connectivity modeling, and nonlinear discriminative mapping within a compact task-specific head. This design preserves the representation advantage of the foundation model while making the classifier more sensitive to emotion-related spatiotemporal patterns. Extensive experiments are conducted on the public FACED and SEED-VII datasets. The proposed method achieves a balanced accuracy of 60.66\%, a Cohen's Kappa of 0.5525, and a weighted F1-score of 60.40\% on FACED, and obtains 33.27\%, 0.2223, and 33.64\%, respectively, on SEED-VII. Compared with the CBraMod+Linear baseline, the balanced accuracy is improved by 5.91 and 2.03 percentage points on the two datasets, respectively. These results indicate that structured task-head design is an effective way to improve EEG emotion recognition when fine-tuning pre-trained EEG models.
A Fibonacci-Based Ontogenetic Discretization of Body-Mass Growth Trajectories
arXiv:2605.27208v2 Announce Type: replace Abstract: Body-mass growth is usually described by continuous models that represent the gradual increase of mass toward an adult or asymptotic value. In this work, we formulate and test an alternative description based on a Fibonacci-inspired ontogenetic coordinate. The model represents growth through a continuous or discretized developmental index, allowing the trajectory to be organized into fine substages. The main analysis compares the Fibonacci-based model with the ontogenetic growth model of West, Brown, and Enquist using digitized growth data from four species with markedly different body sizes and developmental times. The Fibonacci-based formulation was favored in two of the four cases, while the WBE equation remained favored in the other two. A complementary analysis using rodent growth data showed that the Fibonacci discretized model can also approach the performance of classical continuous models in some datasets, although not uniformly. These results suggest that a Fibonacci-based ontogenetic coordinate can provide a useful alternative representation of body-mass growth within a biological systems framework.
Where LLM Annotators Fail: Label-Free Learning on Graphs with LLMs
arXiv:2605.27913v1 Announce Type: new Abstract: Node classification on graphs often requires labeled nodes, yet obtaining labels at graph scale is expensive. When node attributes contain semantic content, such as paper abstracts, web pages, or product descriptions, large language models (LLMs) can provide low-cost supervision by annotating a small subset of nodes. However, these LLM-generated labels are noisy, and existing label-free graph learning methods usually treat this noise as either global or class-conditional. We find that LLM annotation errors are not only class-dependent but also region-dependent: within the same class, reliability can vary sharply across feature-space clusters. In light of this, we propose Cluster-Aware Noise Estimation (CANE), a label-free learning framework that estimates cluster-conditional LLM reliability without ground truth labels, and uses this estimate to decide which pseudo-labels to trust, and which labels to correct. Across various graph benchmarks and GNN backbones, CANE improves over the strongest label-free baselines, with the largest gains on datasets exhibiting stronger cluster-conditional noise.
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
arXiv:2605.27997v1 Announce Type: new Abstract: Large language models frequently generate toxic, hateful, or harmful content, yet existing mitigation methods rely on costly retraining or output-level filtering with no mechanistic insight into where toxicity originates internally. We introduce Meow2X and TRNE, two complementary retraining-free frameworks that localize toxicity to specific layers and neurons by analyzing activation differentials between toxic and neutral prompts, then suppress them via inference-time scaling or minimal rank-one weight edits -- without any gradient descent. Evaluations across five LMs, two benchmarks, and 90 configurations using dual safety evaluators demonstrate consistent toxicity reduction while preserving language modeling quality. Our analysis reveals that toxicity is disproportionately encoded in early MLP layers, varies across architectures, and is systematically underestimated by single-evaluator setups -- underscoring the need for multi-evaluator safety assessment. By bridging mechanistic interpretability with practical detoxification, our framework offers a principled path toward safer, more transparent language models.
Evaluating the Realism of LLM-powered Social Agents: A Case Study of Reactions to Spanish Online News
arXiv:2605.28598v1 Announce Type: new Abstract: LLM-powered social agents are increasingly used to simulate online social behavior, yet their realism remains difficult to validate. Existing work has largely relied on general-purpose benchmarks, while less attention has been paid to short, reactive discourse such as audience replies to online news. In this paper, we evaluate whether LLM-generated reactions to Spanish online news reproduce measurable properties of real audience discourse. Using the Hatemedia dataset, we pair 5,631 news items with 58,555 real audience reactions, and generate a matched synthetic dataset using five LLMs under a shared experimental setting. We compare real and synthetic reactions across three dimensions: hate speech, sentiment, and semantic alignment, considering both off-the-shelf and fine-tuned generation. Results show that off-the-shelf models are poor proxies for real audience reactions: they strongly underproduce hate speech, introduce model-specific sentiment biases, and remain distributionally distant from human replies. Fine-tuning improves fidelity unevenly. Qwen3 provides the most balanced approximation, while Mistral7B achieves the strongest sentiment and semantic alignment but overshoots hate prevalence. Plausible synthetic replies do not necessarily reproduce the distributional properties of public discourse.
EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design
arXiv:2605.19743v2 Announce Type: replace Abstract: Large Language Model (LLM) agents are increasingly applied to engineering design tasks, yet existing evaluation frameworks do not adequately address multi-agent systems that combine simulation, retrieval, and manufacturing preparation. We introduce a benchmark suite with three evaluation dimensions: (1) a workflow benchmark with seven prompt styles targeting distinct cognitive demands-including direct tool use, semantic disambiguation, conditional branching, and working-memory tasks; (2) a Retrieval-Augmented Generation (RAG) benchmark with gated scoring isolating retrieval contributions to parameter selection; and (3) an High Performance Computing (HPC) benchmark evaluating end-to-end ML training orchestration on a SLURM cluster. Alongside the benchmark we present EngiAI, a Multi-Agent System (MAS) reference implementation built on LangGraph that operationalizes the benchmark by coordinating seven specialized agents through a supervisor architecture, unifying topology optimization, document retrieval, HPC job orchestration, and 3D printer control. Across four LLM backends and two EngiBench problems, proprietary models achieve 96-97% average task completion on Beams2D, while open-source 4B-parameter models reach 55-78%, with clear generational improvement. Conditional branching proves most challenging, with task completion dropping to 20-53% for the conditional style on Photonics2D. RAG gating confirms near-perfect retrieval-augmented scores (about 1.0) versus near-zero without retrieval, validating the evaluation design. On HPC orchestration, one model completes all pipeline steps in 100% of runs while another drops to 50%, revealing that multi-step instruction following degrades over long-running workflows.
C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning
arXiv:2605.27860v1 Announce Type: new Abstract: Retrieval-augmented generation combined with reinforcement learning has shown promise for grounding large language models in trustworthy medical evidence. However, existing methods rely on exact-match binary rewards, which in clinical diagnosis cause two issues: (i) semantically relevant but non-verbatim steps receive zero signal, discarding valuable learning signals; and (ii) uni-dimensional rewards cannot effectively supervise heterogeneous reasoning capabilities. To address these issues, we propose C-MIG, a Multi-view Information Gain-based retrieval-augmented generation framework for Clinical diagnosis. C-MIG estimates information gain under a frozen reference model from two complementary views, retrieved-document and document-refinement, to jointly guide what to retrieve and how to refine, alleviating the issues of valuable reward signal loss and credit assignment. We further design a multi-subquery retrieval augmentation strategy that improves knowledge recall coverage in clinical diagnostic scenarios. Comprehensive experiments on four medical benchmarks demonstrate that C-MIG achieves the best performance among all RAG-RL methods on both in-domain and out-of-domain sets, and outperforms state-of-the-art general-purpose LLMs for clinical diagnosis.
Execution and assessment of agentic influence operations in simulated social networks
arXiv:2605.28725v1 Announce Type: new Abstract: This article evaluates AI-enabled influence operations in synthetic social networks through controlled simulations of narrative release, amplification, and counter-messaging. We measure exposure and belief change in agentic audiences, showing that amplification maximizes reach, counter-messaging shifts opinions most, and narrative release requires larger attacker footprints.
Wigner-Eckart Factorization of the Spectral Boltzmann Collision Operator
arXiv:2605.28475v1 Announce Type: new Abstract: We reduce the eight-dimensional weak form of the bilinear Boltzmann collision operator to a five-dimensional kinematic core by rigidly rotating the laboratory frame to align with the colliding pair and integrating over the $\mathrm{SO}(3)$ rotation group. This reduction yields an exact Wigner--Eckart factorization within a spectral Galerkin framework of associated Laguerre polynomials and spherical harmonics. The decomposition decouples the angular geometry from the scattering physics. The former, represented by Clebsch--Gordan coefficients, is evaluated exactly, while the latter is evaluated to machine precision by a spectrally convergent singular quadrature strategy. By explicitly zeroing specific entries, the macroscopic collision invariants are embedded without approximation. Cache-optimized contractions deliver up to a 37-fold single-core speedup and a 1000-fold memory reduction over standard dense Cartesian formulations. The approach is validated against analytical solutions for Maxwell molecules and infinite-order Chapman--Enskog viscosity coefficients for hard spheres.
CXL-ClusterSim: Modeling CXL-based Disaggregated Memory Cluster for Pooling and Sharing using gem5 and SST
arXiv:2605.27745v1 Announce Type: new Abstract: Large-scale AI training and inference require hundreds of gigabytes to terabytes of DRAM with high peak to average utilization ratios, resulting in overprovisioning. In cloud computing, DRAM constitutes a significant share of the cost. Yet, as shown by recent articles, DRAM is heavily under utilized. Memory disaggregation is a solution to both these problems. With the advent of the CXL protocol, there is renewed interest in designing and optimizing computing systems with disaggregated memory. However, at present, there are limited simulation tools available for exploring the design space and evaluating the performance tradeoffs in computer systems with disaggregated memory. In this paper, we propose CXL-ClusterSim, a full-system modeling and simulation framework by combining the gem5 simulator for fidelity, with the Structural Simulation Toolkit (SST) for parallel simulation. We outline the challenges in creating this simulation infrastructure and present a design that is scalable, flexible, and reasonably fast to help computer architects to explore the design space of CXL-based disaggregated memory and identify new opportunities for hardware/software codesign and performance optimization.
SmartIterator: Visual Analytics Workflows for Supervising Unsupervised Data Grouping
arXiv:2605.28219v1 Announce Type: new Abstract: Unsupervised learning methods -- topic modeling, partition-based and density-based clustering -- produce data groupings without human guidance, yet choosing and evaluating those groupings should not itself be unsupervised. We present \emph{SmartIterator}~(SI), a visual analytics approach that treats the full sequence of grouping results across a parameter sweep as a first-class analytical object. For each method family, SI provides a structured six-phase workflow that guides the analyst through systematic exploration of grouping results -- from quality-metric overview through transition-stability assessment, membership-confidence evaluation, content and context inspection, and recurrent-archetype verification to an informed decision -- building cumulative understanding of data structure along the way. The workflows are operationalized through \emph{IteraScope}~(IS), a coordinated visual display combining quality-metric charts with semantic color encoding, a 1D group embedding with Sankey-style transition flows and violin plots of membership confidence, a 2D group embedding with HDBSCAN-detected recurrent archetypes that highlights iterations capturing all persistent patterns, and domain-specific linked views for contextualized interpretation. We demonstrate the three workflows on: (1)~simulated social-media messages from the VAST Challenge 2011 (density-based clustering, validated against ground truth), (2)~EU population statistics across ${\sim}1\,500$ NUTS-3 regions (partition-based clustering), and (3)~30 years of IEEE VIS papers (NMF topic modeling). The workflows constitute the main contribution: they provide actionable, method-specific guidance for navigating parameter spaces, studying how data structure evolves across configurations, and grounding analytical understanding in domain context -- yielding knowledge about the data that no single ``best'' result can provide.
A Policy-Driven Runtime Layer for Agentic LLM Serving
arXiv:2605.27744v1 Announce Type: new Abstract: Multi-agent LLM systems have become the dominant production workload, but the serving stack was not built for them. The agent framework above knows agent identities, role, schemas, and dispatch structure but never sees an engine-level event; the serving engine below sees every event but knows nothing about agents. A surprising number of cross-cutting policies depend on both: prefix caching, batch shaping, speculative execution, fairness, tool-result memoization, safety enforcement, and more. Each lives in the seam between the two layers and is currently solved by a one-off patch into one neighbor or the other. We argue this seam is best addressed by an architectural change rather than point fixes: insert a third tier, an agent runtime layer, between the framework and the engine, exposing four primitives (observe, score, predict, act) into which any agent-aware policy plugs, with agent identity as the shared coordinate. We map nine concrete policies onto the layer and validate the abstraction in depth on the one with the largest immediate serving-cost lever: KV caching across sessions, instantiated as CacheSage, which learns the per-workload agent transition matrix online and uses it for survival-based eviction and between-step prefetch. Preliminary results on five real multi-agent workloads show +13 to +37 pp cache hit-rate lift, 12% to 29% lower mean TTFT, and 6% to 14% higher throughput over an unmodified serving stack.
Robo-Blocks: Generative Scaffolding in End-User Design and Programming of Social Robots
arXiv:2605.28154v1 Announce Type: new Abstract: Programming social robots is challenging for novice robot programmers due to required expertise in planning, interaction design, and programming. While large language models (LLMs) hold significant promise through code generation from natural-language descriptions, they can obscure critical elements of programming and supplant designer intent, eventually resulting in over-reliance instead of developing programming skills. In this paper, we explore how LLM-based social-robot-programming tools can support novice robot programmers through a Research through Design (RtD) process. We designed and prototyped Robo-Blocks, a block-based programming environment that leverages LLMs to offer novice robot programmers generative scaffolding through structured narratives that connect high-level ideas to executable robot behaviors. Through deployment with novices, we discovered emerging user personas and usage patterns for generative scaffolding and showed how this scaffolding shapes end-user design and programming strategies. We present design insights for the effective use of generative scaffolding and its integration into the practice of social-robot programming.
No Safe Dose: How Training Data Drives Unsafe Image Generation
arXiv:2605.28137v1 Announce Type: new Abstract: Text-to-image models trained on large-scale data often inevitably ingest unsafe content. While some people observe input-output amplifications, it remains unclear whether and how training data composition directly drives model output safety or by other factors. We shed light on this question by isolating this variable: we train the same text-to-image model on datasets that differ \emph{only} in their fraction of unsafe images (0\% to 9.6\%), across several dataset scales (100K to 8M). Then we generate images with the resulting models, and evaluate them with four independent safety classifiers. Output unsafety rises monotonically from 16.6\% at 0\% contamination to 25.5\% at 5\%. A factorial design reveals that the \emph{proportion}, not the absolute count, of unsafe training images is the operative variable. The 16.6\% irreducible baseline at zero contamination implicates the other components, e.g. frozen text encoder, as a residual safety risk -- confirmed by a text encoder ablation showing that SafeCLIP reduces this floor to 9.6\%, while the dose-response effect persists across all three encoders tested. Critically, no quality degradation in terms of FID, CLIPscore and ImageReward accompanies safety filtering. These results establish that data curation and text encoder safety are complementary and independently effective interventions. At the same time, the remaining level of unsafety poses questions for future research about emerging capabilities and compositionality.
The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness
arXiv:2605.28190v1 Announce Type: new Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding robustness is multidimensional, since models respond differently to different types of variation, and requires dynamic evaluation to expose failures hidden by static benchmarks. We introduce the Harder Text Embedding Benchmark (HTEB), a dynamic evaluation framework that challenges model robustness along three practically interpretable axes (Lexical/Stylistic, Length and Language) by stochastically transforming inputs at evaluation time with an LLM. Evaluating 16 open-weight embedding models on 32 datasets covering 42 languages under transformations validated by 4,800 human ratings on an English subsample, we find three patterns: (1) Models exhibit specific, partly decoupled robustness profiles across axes. (2) Across three model families, scale increases absolute scores but does not close the gap between original and transformed evaluations. Here, scaling tends to improve specifically the Language axis. (3) English datasets are more sensitive to HTEB transformations than multilingual datasets. This demonstrates that HTEB identifies strengths and weaknesses of models along deployment-relevant axes, challenging current embedding benchmarks and arguing for multidimensional, dynamic robustness evaluation.