Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Score Accuracy Along the Forward Diffusion Does Not Certify Numerical Stability in Diffusion Sampling
arXiv:2607.08757v1 Announce Type: cross Abstract: Score matching controls average error under the forward marginals, but a discretized reverse-time sampler evaluates the learned score along its own trajectory. We show that small forward-marginal error does not guarantee numerical stability. We construct a single smooth score field with arbitrarily small forward-marginal $L^2$ error. The learned reverse-time process is nonexplosive, has moments of every order, and can be arbitrarily close to the exact reverse-time process in path-space total variation. Yet its Euler--Maruyama discretizations converge in probability while every positive moment diverges. Thus weak convergence can hold even though every Wasserstein distance $W_p$, $p\ge1$, diverges. The same failure can occur within one fixed finite neural architecture. We construct a family of bounded, globally Lipschitz denoisers for which both the forward-marginal error and the path-space total variation distance tend to zero, while their Euler--Maruyama endpoints diverge in every $W_p$. For compactly supported data, we also give a simple positive result. Projecting the learned denoiser onto a known bounded closed convex set containing the support preserves pointwise accuracy, gives grid-uniform moment bounds, and yields Wasserstein convergence under mild local regularity. Experiments with a small fixed DiT-style network show large growth along rare numerical trajectories and its suppression by denoiser projection, while overall trajectory errors remain small.
Bridging Cognitive Neuroscience and Graph Intelligence: Hippocampus-Inspired Multi-View Hypergraph Learning for Web Finance Fraud
arXiv:2601.11073v3 Announce Type: replace Abstract: Online financial services constitute an essential component of contemporary web ecosystems, yet their openness introduces substantial exposure to fraud that harms vulnerable users and weakens trust in digital finance. Such threats have become a significant web harm that erodes societal fairness and affects the well-being of online communities. However, existing detection methods based on graph neural networks (GNNs) struggle with two persistent challenges: (1) long-tailed data distributions, which obscure rare but critical fraudulent cases, and (2) fraud camouflage, where malicious transactions mimic benign behaviors to evade detection. To fill these gaps, we propose HIMVH, a Hippocampus-Inspired Multi-View Hypergraph learning model for web finance fraud detection. Specifically, drawing inspiration from the scene conflict monitoring role of the hippocampus, we design a cross-view inconsistency perception module that captures subtle discrepancies and behavioral heterogeneity across multiple transaction views. This module enables the model to identify subtle cross-view conflicts for detecting online camouflaged fraudulent behaviors. Furthermore, inspired by the match-mismatch novelty detection mechanism of the CA1 region, we introduce a novelty-aware hypergraph learning module that measures feature deviations from neighborhood expectations and adaptively reweights messages, thereby enhancing sensitivity to online rare fraud patterns in the long-tailed settings. Extensive experiments on six web-based financial fraud datasets demonstrate that HIMVH achieves 6.42% improvement in AUC, 9.74% in F1 and 39.14% in AP on average over 15 SOTA models.
Retrieval of Scientific and Technological Resources for Experts and Scholars
arXiv:2204.06142v2 Announce Type: replace Abstract: Institutions of higher learning, research institutes and other scientific research units have abundant scientific and technological resources of experts and scholars, and these talents with great scientific and technological innovation ability are an important force to promote industrial upgrading. The scientific and technological resources of experts and scholars are mainly composed of basic attributes and scientific research achievements. The basic attributes include information such as research interests, institutions, and educational work experience. However, due to information asymmetry and other reasons, the scientific and technological resources of experts and scholars cannot be connected with the society in a timely manner, and social needs cannot be accurately matched with experts and scholars. Therefore, it is very necessary to build an expert and scholar information database and provide relevant expert and scholar retrieval services. This paper sorts out the related research work in this field from four aspects: text relation extraction, text knowledge representation learning, text vector retrieval and visualization system.
From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs
arXiv:2607.08009v1 Announce Type: new Abstract: We introduce a Bloom-aligned framework for measuring educational control in Large Language Models (LLMs): the ability to preserve a task's instructional intent while shifting its cognitive demand toward specified learning objectives. We apply this framework to programming tasks in computer science education to study the gap between solving tasks and adapting them for learners. Using revised Bloom's Taxonomy as an operational scale of cognitive demand, we evaluate two intervention settings: general difficulty control, where models are asked to make tasks harder or easier, and Bloom's control, where models are asked to target higher or lower Bloom's levels. We evaluate a matched Qwen3-Next model pair, comparing Qwen3-Next-80B-A3B-Instruct with Qwen3-Coder-Next across 2,520 tasks from three benchmarks. The framework reveals a robust directional asymmetry: both models reliably increase cognitive demand, but struggle to lower it. We further characterize these outcomes with semantic-delta clustering and layer-wise Fisher's Discriminant Ratio probing. Within this controlled comparison, the general model shows clearer middle-layer separability for both general difficulty and Bloom-control contrasts, whereas the coder model shows weaker separability for general difficulty and a deeper peak for Bloom-control contrasts. These results show that strong execution performance does not automatically entail Bloom-aligned educational control.
Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery
arXiv:2607.08408v1 Announce Type: new Abstract: Gaussian splatting is the current state-of-the-art for dense, deformable 3D anatomy reconstruction in robot-assisted minimally invasive surgery (RAMIS); however, most pipelines are offline and depend on accurate camera trajectory priors (often from robotic kinematics), limiting applicability when priors are missing or noisy. To address these limitations, we propose Track2Map, an online 3D Gaussian Splatting pipeline that jointly optimizes camera trajectory and 3D deformable scene representation directly from surgical video. Track2Map is therefore capable of robust 3D reconstructions when camera trajectory priors are either absent or noisy, and due to its online nature it effectively works as a Simultaneous Localisation and Mapping (SLAM) method. To stabilize optimization in the presence of tissue motion and ambiguous visual cues, we introduce a track-anchored deformation initialization using dense 2D point tracks. Track statistics are further utilized to disentangle camera motion from scene deformation by detecting static camera periods and reducing drift during incremental mapping. Experiments on StereoMIS show improved reconstruction quality and camera trajectory against competing SLAM methods, as well as compared to non-SLAM methods that utilize camera trajectory priors. The code is available at https://track2map.github.io/.
DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation
arXiv:2607.08751v1 Announce Type: new Abstract: Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embodiment coverage, or controllable visual variation, hindering studies of cross-task and cross-embodiment generalization. We present DexVerse, a large-scale and modular benchmark for dexterous manipulation. DexVerse includes 100 tasks spanning a broad range of manipulation skills, including object grasping and relocation, articulated-object interaction, functional tool use, bimanual coordination, non-prehensile control, contact-rich behaviors, multi-goal execution, and long-horizon multi-stage task completion. It supports 3 robot arms and 6 dexterous hands, and is extensible to new tasks, assets, and embodiments. To evaluate visuomotor generalization, DexVerse provides configurable visual variations in textures, background, lighting, and camera viewpoints. We further provide a VR-based teleoperation interface and 3,180 demonstrations with synchronized proprioceptive, RGB, depth, point-cloud, and state observations. We benchmark representative methods, including Diffusion Policy, DP3, OpenVLA, and $\pi_{0.5}$, across 19 tasks. Results reveal substantial challenges in task generalization and visuomotor robustness, establishing DexVerse as a promising testbed for general-purpose dexterous manipulation. Project page: https://ycyao216.github.io/DexVerse.site
Out of Sight: Compression-Aware Content Protection against Agentic Crawlers
arXiv:2607.08180v1 Announce Type: new Abstract: The rise of LLM-based agents with reasoning, summarization, and memory capabilities has created a new threat surface for online content that conventional defenses fail to address. Existing defenses like access controls can be circumvented by agents mimicking ordinary browsers, and injection-based defenses often degrade human readability. In this paper, we revisit the agent pipeline and identify context compression, which agents routinely invoke to fit context budgets, as a critical yet overlooked defense layer. We propose CAPE, a framework that protects high-value textual content by injecting invisible perturbations without changing its human-visible surface form, thereby inducing severe information loss during agent compression. CAPE extracts disruptive seed perturbations from an accessible surrogate compressor, then adapts them to query-only target compressors through prior-guided evolution and preference-calibrated candidate prioritization, achieving effective protection under a low query budget. Experiments on three content types and four compression settings show that CAPE improves information loss by up to 75.8% over the strongest baseline while keeping protected content visually indistinguishable from originals. CAPE also transfers to real-world settings, including the LangGraph agent workflow and GitHub Copilot, highlighting its generality and practical value. This paper aims to reveal context compression as a new defense layer, promoting content protection research in the agent era.
PARA-PV: Physics-Aware Retrieval-Augmented PV Prediction Based on Frozen Foundation Model and Distribution Shift Correction
arXiv:2607.08079v1 Announce Type: new Abstract: Accurate photovoltaic (PV) power forecasting is essential for reliable grid dispatch and renewable energy integration, yet it remains challenging because PV generation is jointly shaped by weather variability, day-night transitions, regime-dependent dynamics, and strict physical constraints. We propose PARA-PV, a Physics-Aware Retrieval-Augmented framework that embeds physical knowledge throughout the forecasting process. The framework first encodes multivariate PV observations into patch-level representations and, through a physics-aware retrieval-augmented learner, retrieves historical patches and analog trajectories that are consistent with the current window in temporal shape, power level, PV operating state, and intra-day period; this yields a physically grounded base forecast. To supplement local memory with broader temporal knowledge, the base forecast is then calibrated against a frozen Chronos time-series foundation-model prior through a lightweight residual adapter, so that general temporal regularities are adapted to PV-specific dynamics without overriding the physically grounded prediction. Because residual conditional distribution shifts persist when weather and diurnal regimes change, a physics-aware distribution shift correction module subsequently adjusts the preliminary forecast using power, weather, timestamp, and day/night conditions, applying gated mean-shift and scale corrections selectively. Finally, a physics-constrained loss function partitions the samples into peak, ramping, night-time, and regular regimes and adaptively reweights their error contributions, preventing the dominant regular regime from suppressing learning of operationally critical states. Our code is available at https://github.com/weican1103/PARA-PV.
Benchmark Evaluation of Feredated Learning on Multi-organ Images
arXiv:2607.08219v1 Announce Type: new Abstract: The privacy requirements of medical data and its substantial variations across organs and modalities hinder the clinical implementation of medical AI. Federated learning (FL) is a feasible approach to overcome these challenges. Due to the continuous emergence of FL algorithms and the highly heterogeneous nature of medical data, objectively evaluating their performance in real-world clinical settings remains difficult. Therefore, a comprehensive federated medical imaging benchmark, serving as a unified evaluation standard, is crucial for advancing the technology toward reliable clinical application. Existing federated medical imaging benchmarks have not yet adequately incorporated state-of-the-art algorithms, are limited to data from single organs or modalities, and overly emphasize model accuracy, making it difficult to comprehensively assess the overall efficacy of FL in real-world medical environments. To address these challenges, we developed the MobenFL benchmark. This benchmark integrates 20 cutting-edge FL algorithms and 22 medical imaging datasets, covering 12 critical organs across the human body, surpassing existing benchmark in breadth. In terms of evaluation dimensions, MobenFL not only assesses performance but also systematically incorporates key metrics such as algorithmic efficiency and privacy protection capabilities. Additionally, it conducts specialized evaluations for complex real-world clinical scenarios involving different diseases, devices, and imaging modalities, thereby providing a comprehensive and in-depth evaluation framework for the clinical application of FL in the medical field.
The Context Access Divide: Interaction-Level Architecture as a Complementary Dimension of Agentic Inequality
arXiv:2607.08495v1 Announce Type: new Abstract: Sharp et al. (2025) introduce "agentic inequality" as a framework for analyzing disparities in access to AI agents across three dimensions: availability, quality, and quantity. These person- and organization-level dimensions characterize who can access agents and at what capability, but do not address a structurally important divide operating at a finer level: the individual interaction. Two users with nominally equivalent agent access may experience qualitatively different AI utility depending on whether the system can autonomously retrieve context from the user's knowledge corpus (Dynamic Context Retrieval) or requires the user to manually identify and attach relevant documents at each query (Manual Attachment). We term this the Context Access Divide (CAD). For knowledge-intensive workers whose intellectual capital spans tens of thousands of files, the CAD constitutes a qualitative threshold in AI usefulness: below it, the cognitive burden of context curation falls on the human, reproducing the inefficiencies AI is meant to eliminate. We propose contextuality -- the degree to which an AI system autonomously accesses a user's accumulated knowledge capital -- as a dimension of AI-mediated inequality that complements, but is not reducible to, the Sharp et al. framework. We formalize the CAD with a probabilistic model grounded in the fan effect literature in cognitive psychology, demonstrating that manual context attachment leads to a combinatorial collapse in task-success probability as corpus size and task conjunctivity grow, while dynamic retrieval architectures are structurally insulated from this collapse. We analyze the technical basis of this divide in the Model Context Protocol (MCP) and retrieval-augmented generation (RAG) architectures, and examine its implications for knowledge-work stratification and AI platform governance.
Decomposition-Based QAOA for Maximum Coverage Location Problem in Satellite Constellation Design
arXiv:2607.08102v1 Announce Type: cross Abstract: An increase in earth observation missions has increased the demand of efficient design and optimization of satellite constellations. Maximizing coverage of the target while effectively utilizing the limited orbital resources is one of the critical design challenges for complex combinatorial optimization problems. The maximal covering location problem (MCLP), serves as a base for orbital coverage modeling, is NP-hard and computationally intractable for large-constellation instances. Using heuristics, metaheuristics, and mixed-integer linear programming, classical solvers have achieved optimal or near-optimal results, yet their scalability is limited as the problem size increases. Quantum computing advancements, including the quantum approximate optimization algorithms, offer a potential solution to NP-hard combinatorial optimization problems. Current quantum hardware limitations, such as low qubit counts and circuit depth, restrict solutions for small-scale instance problems. To address this challenge, this paper proposes a scalable quantum optimization framework for MCLP in satellite constellation design. A decomposition-based quantum methodology is proposed, in which large MCLP instances are partitioned into subgraphs by classical decomposition, optimized independently via quantum optimization circuits, and combined using quantum reconstruction strategies. Computational results across different constellation sizes reveal better scalability in less time while maintaining competitive coverage performance compared to classical solvers.
A Quantized Native Runtime for On-Device Semantic Audio Generation
arXiv:2607.08526v1 Announce Type: new Abstract: Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather than through framework-heavy datacenter stacks. We present \textit{aria}, a dependency-free native runtime that runs the complete text-to-music pipeline of Stable Audio~3 (SA3) on ordinary GPUs, CPU-only machines, and a Raspberry~Pi~5, with no Python or deep-learning framework underneath. Our main contribution is a study of quantization: running the model at lower numerical precision to fit tight memory budgets, saving memory in place rather than adding to it. Because the runtime owns every internal tensor, it also exposes activation steering, a low-cost way to steer what the model generates. We judge the quality cost with three independent measures of the output (prompt adherence, overall audio quality, taste preservation), each compared against the ordinary variation between random seeds. Eight-bit precision shows no measurable quality loss on any measure while sharply cutting memory, and it is the fastest mode on the GPU; four-bit adds a small, bounded cost but shrinks the footprint enough to run the $1.2$-billion-parameter model on an $8$\,GB Pi. Against the official implementation, aria matches or exceeds generation speed and starts about seven times faster. A case study of the steering interface generates music carrying taste associations (\emph{sonic seasoning}), with genuine but bounded control for a subset of attributes. These results make a compact, quantized runtime with built-in control a practical basis for on-device semantic audio in Internet-of-Sounds settings. The \textit{aria} runtime is released at https://github.com/matteospanio/aria.
XOV-Action: Towards Generalizable Open-Vocabulary Action Recognition
arXiv:2403.01560v3 Announce Type: replace Abstract: Inspired by the impressive success of image-text foundation models, recent works have proposed to adapt these foundation models to video data, leading to efficient and effective video models for open-vocabulary action recognition. However, through a comprehensive evaluation, our work finds that state-of-the-art open-vocabulary action recognition models still struggle with generalization to video domains that they have not encountered. To address this limitation, we introduce \textit{generalizable open-vocabulary action recognition}, which aims to develop action recognition models capable of generalizing to both novel action categories and unseen video domains. Our work contributes a novel model named XOV-Action to overcome two critical challenges: (1) understanding novel action concepts of open-set categories, and (2) mitigating the scenario discrepancy between training and test datasets. Specifically, XOV-Action first proposes to capture diverse action-related concepts by learning diversified elaboration representations, which enables better generalization to open-set action categories. Second, XOV-Action learns scene-agnostic video representations to overcome the scene bias, which improves the generalization in unseen video domains. Additionally, to evaluate models in generalizable open-vocabulary action recognition, we contribute a new cross-domain action benchmark named XOVABench, which covers multiple video domains with varying degrees of gaps and consists of both closed-set and open-set action categories. Extensive quantitative and qualitative experiments demonstrate that our proposed XOV-Action can effectively improve action recognition performance for both closed-set and open-set categories across video domains. The benchmark is available at https://github.com/KunyuLin/XOV-Action/.
LLM for the development of FCM
arXiv:2607.04983v2 Announce Type: replace Abstract: This article is about the development of a fuzzy cognitive map using a local large language model. In the light of recent advances it is evident that large language models, and even local large language models are capable of extracting quantities from textual data. In other words, a local LLM like Qwen2.5-32B, or probably larger, can accept entities as prompt input and determine relevant quantitative data as the model output. In turn, this output can be utilized for the construction of a data driven fuzzy cognitive map. Hence, this implementation is achieved and then the model is thoroughly tested; Qwen2.5-32B is used and the data is extracted from hotel reviews from TripAdvisor. Furthermore, the extracted documents pass through the model unfiltered and then a fuzzy cognitive map is trained and evaluated. A case is made about Greek reviews where a star topology FCM is formed that indicates the preferences of the reviewers. Finally, external validation is performed to establish whether the fuzzy cognitive map can correlate the star rating of the review -an outcome outside the model's inference scope -with its predicted satisfaction.
Optimal slit width for high-precision orbital angular momentum measurement using angular double-slit interferometry
arXiv:2607.08549v1 Announce Type: new Abstract: We demonstrate an optimization of angular double-slit interferometry for accurate measurement of orbital angular momentum (OAM) of vortex beams. By scanning the dynamic double slits, the topological charge (TC) magnitude is directly determined from the oscillation frequency of the on-axis intensity. Based on repeated experimental investigations, we establish a critical criterion for slit width selection: to avoid phase truncation or period overlap, the angular width of each slit must exactly match the spiral phase period 2{\pi}/|l|. Under this optimal condition, the interference pattern exhibits the highest visibility and the measurement error is minimized. Experiments for l = 5, 10, and 15 are performed as representative examples, and the universality of this criterion is confirmed. Furthermore, by introducing an additional phase shift, the sign of the TC is unambiguously determined, as demonstrated for l = 10 and l = -15. This simple, robust method provides a high-precision pathway for OAM metrology.
CommuniWave:A Machine Learning Model for Quantifying the Degree of Temporary Informal Behavior in Urban Communities
arXiv:2607.08554v1 Announce Type: new Abstract: For urban managers and designers, improving the functional attributes of urban communities to enhance territorial resilience in the face of complexity and uncertainty is crucial. Currently, community planning often follows a top-down approach and lacks effective metrics to quantify informal behaviors of residents, leading to frequent conflicts with original plans. This study introduces CommuniWave, a machine learning model designed to efficiently detect and quantify the Degree of Informal Behavior (DIB) in urban communities. The model integrates a Behavior Capture Net (BCN) based on mmaction2, a self-developed YOLOv10 model (YLX), and a Behavior Eval Model (BEM) using random forest. Ultimately, by generating DIB fluctuation charts from street videos, the model facilitates dynamic monitoring, supporting urban managers in making refined decisions to enhance the overall resilience of communities.
Adaptive and Robust Watermark for Generative Tabular Data
arXiv:2409.14700v4 Announce Type: replace Abstract: In recent years, watermarking generative tabular data has become a prominent framework to protect against the misuse of synthetic data. However, while most prior work in watermarking methods for tabular data demonstrate a wide variety of desirable properties (e.g., high fidelity, detectability, robustness), the findings often emphasize empirical guarantees against common oblivious and adversarial attacks. In this paper, we study a flexible and robust watermarking algorithm for generative tabular data. Specifically, we demonstrate theoretical guarantees on the performance of the algorithm on metrics like fidelity, detectability, robustness, and hardness of decoding. The proof techniques introduced in this work may be of independent interest and may find applicability in other areas of machine learning. Finally, we validate our theoretical findings on synthetic and real-world tabular datasets.
Empirical Calibration and Conditional-Reliability Diagnostics for Bearing RUL Prediction under Operating-Regime Shift
arXiv:2607.08273v1 Announce Type: new Abstract: Remaining useful life (RUL) estimates support reliability and maintenance decisions only if both point accuracy and prediction intervals remain trustworthy when operating conditions change. Convenient mixed splits can hide that failure. This paper studies the question on a documented 10-bearing PHME subset with time-varying load and speed. Derived load-speed regimes define the held-out evaluation units, while models receive only measured load and speed as context. A calibrated predictive-representation model fuses raw vibration windows, engineered descriptors, and operating context, then forms intervals by empirical residual calibration. Under strict train/validation/calibration/test separation, the model reaches normalized MAE 0.1477, empirical 90% coverage 0.900, and retrospective absolute-step MAE 285.26; a 400-tree random forest reaches 0.1538, 0.871, and 294.57. The results do not show uniform dominance: conditional diagnostics expose non-uniform reliability, including 0.666 coverage in a low-load/high-speed cell, and a post-hoc pooled regime-conditioned residual diagnostic raises that cell to 0.941 only as motivation for future pre-specified conditional calibration. Stress tests further identify raw-channel loss as the largest tested reliability failure mode. The contribution is therefore a bounded reliability-evaluation protocol for the processed 10-bearing subset, with conditional undercoverage and raw-channel loss reported explicitly as failure modes rather than deployment guarantees.
Extracting inter-nuclear distances in the oxygen molecule interacting with an XFEL-pulse: a fundamental system for understanding Coulomb explosion imaging
arXiv:2607.08211v1 Announce Type: new Abstract: We investigate the interaction of O$_{2}$ with an X-ray Free Electron laser (XFEL) pulse of short duration. We consider a photon energy of 570 eV, which allows for the formation of molecular states with up to two core holes. We compute the sum of the final kinetic energies, i.e. the kinetic energy release (KER) of the atomic fragments for different fragmentation channels. We demonstrate that with the knowledge of the potential energy curves the equilibrium inter-nuclear distance of O$_{2}$ is best extracted from the O$^{+}$+O$^{+}$ channel and not from higher-charged channels such as O$^{2+}$+O$^{2+}$. This challenges our current understanding of Coulomb explosion imaging, where from just applying quasi-classical dynamics one expects that the equilibrium inter-nuclear distance of molecules is best extracted from higher-charged fragmentation channels. On the other hand, we show that the KER of the O$^{+}$+O$^{+}$ channel is inconsistent with the one obtained by just calculating the Coulomb repulsion of the atomic fragments resulting from molecular dissociation.
How Analysts Use AI in High-Stakes Crime Linkage: An Industrial Study
arXiv:2607.08274v1 Announce Type: new Abstract: Crime linkage analysis is used in many countries to identify series of offences that may have been committed by the same individual. In practice, specialist analysts manually search for behavioural and situational connections across large crime databases, an effort that is time-consuming, cognitively demanding, and can involve repeated exposure to disturbing material. To support this work, an Artificial Intelligence (AI)-enabled decision-support tool was co-developed with a UK law enforcement agency to assist analysts in identifying likely crime linkages. This paper reports an industrial evaluation of the crime-linkage tool. We conducted a mixed-methods usability study combining direct observation, eye-tracking, mouse-tracking, and surveys to examine how analysts engage with AI predictions and with the model features presented as explanations. Our findings show that analysts used the AI predictions selectively and frequently validated them against behavioural (non-AI) evidence, reflecting partial trust and an ongoing reliance on established analytical practices. We also found that analysts attended to the presented model features and valued their availability, while identifying opportunities to improve how explanations are presented and integrated into the workflow. Overall, our results highlight the need for AI-enabled decision-support tools to better integrate explanations and traditional analytical methods, and demonstrate the importance of in-situ evaluation for engineering usable and trustworthy AI in high-stakes settings.
SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling
arXiv:2607.08565v1 Announce Type: new Abstract: LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by agents instead of humans. This shifts the workload in two ways: (1) agents act only on complete responses, making the cluster's tokens per second (TPS) the primary goal and relaxing--not eliminating--per-token latency requirements; and (2) requests share much of their KV\$-reuse exceeds 80% of request tokens in a production trace from BAILIAN, versus 54-62% in chat. This paper first contributes a systematic study of request scheduling for agents on two real-world traces. We find that to increase KV\$ reuse, existing schedulers overly prioritize routing requests to instances caching their KV\$, overloading a few while leaving the rest idle, capping TPS. We thus present two key insights: (1) load balance need not sacrifice all KV\$ reuse, thanks to the global-tier KV\$ store and (2) by utilizing the workload's intra-session locality, balancing a small fraction of requests--the first request in each agent session--suffices to balance the cluster without sacrificing most KV\$ reuse on local instances. SMETRIC realizes these insights with balanced session-centric scheduling: it routes each session's first request purely for load balance and its follow-up requests in a cache-aware manner, preserving load balance and local reuse while keeping demand on the global tier low. Using the session turn information as the scheduling metric is deliberate: it is derived efficiently and accurately from the user inputs alone, so the scheduler stays clean and stateless. SMETRIC improves cluster TPS by 10-16% under prefill-decode colocation with a global store and prefill TPS by 2-34% under disaggregation over state-of-the-art schedulers, also with a better per-token latency.
Orthogonality of coastal trapped waves
arXiv:2607.07975v1 Announce Type: new Abstract: Coastal trapped wave modes are shown to be orthogonal in the sense that they make independent contributions to the energy of the wave field. The hydrostatic Boussinesq dynamics on an $f$-plane, linearized around a state of rest, are formulated as a (generalized) Schr\"odinger equation, which exposes the Hermitian structure of the wave operator that implies the orthogonality of eigenmodes. This formulation, which becomes particularly simple in weak form, is parlayed into a finite-element discretization that preserves the symmetries of the original problem and therefore the energy conservation and orthogonality of modes. The geostrophic momentum approximation, under which the orthogonality of modes has been recognized previously, is reprised and discussed in the present context to emphasize that low-frequency coastal trapped waves are edge waves. Their dynamics are governed by boundary potential-vorticity anomalies, both Bretherton-type contributions familiar from the quasi-geostrophic equations and lateral contributions that are important for steep slopes and coastal walls.
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend
arXiv:2607.08215v1 Announce Type: new Abstract: Non-GPU AI accelerators are increasingly adopted as alternatives to general-purpose GPUs for large-model inference, but the real engineering cost of migrating demanding workloads beyond CUDA remains poorly documented. We present a field study of deploying two large inference workloads on a 16-device Huawei Ascend 910 system using CANN and vLLM-Ascend: an LLM-as-a-judge safety and alignment evaluation pipeline based on a W8A8 MoE judge model, DeepSeek-V4-Flash, and a multimodal medical vision--language benchmark based on DeepSeek-V4-Flash-Vision for MMMU and MMMU-Pro. Making these workloads reliable required twelve source-level patches to the vendor inference plugin, disabling several high-throughput features to preserve numerical correctness, and adding operational safeguards for recurring device-level failures. We summarize the main platform limitations in eight categories: incomplete operator and feature support, fragile parallelism, numerical faults in low-level kernels, immature graph compilation, unstable advanced features, limited scalability, weak observability, and ecosystem fragmentation. For each category, we report the symptoms, evidence, and likely causes. We also quantify the integration effort, concurrency behavior, and benchmark quality to show that both workloads were served correctly. Our study provides a reproducible reference for teams evaluating or operating non-GPU accelerators for large-model inference.
Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models
arXiv:2607.08282v1 Announce Type: new Abstract: While Large Language Models (LLMs) have become essential productivity tools, their integration into workflows without adequate safeguards creates significant risks. This paper proposes an open-source, privacy-focused, user-facing firewall designed to secure both web-based and programmatic LLM interactions. The architecture combines a browser extension and a proxy for total traffic interception across both HTTP(S) and WebSocket communications. At its core, a flexible multi-agent pipeline delivers data leakage prevention through a hybrid approach combining deterministic detectors with LLM-driven semantic analysis, proprietary code leakage prevention, and extensible components designed for future security enhancements such as prompt injection evasion. The framework's layered architecture enables deployment across heterogeneous environments, allowing organizations to balance computational cost, detection depth and latency. Evaluation results demonstrate it achieves F1 scores of up to 94.93% on optimal configurations.
ArtMine: Discovering and Formalizing Artistic Processes
arXiv:2607.08331v1 Announce Type: new Abstract: Understanding how artworks are created requires reasoning about the iterative decisions, material operations, and contextual influences that shape artistic production. While recent generative AI systems can synthesize artworks with high fidelity, they primarily model distributions over finished artifacts rather than the creative processes underlying their creation. In practice, artistic workflows are only partially documented through fragmented sources such as archival records, preparatory studies, correspondence, etc., making process-level understanding difficult to formalize computationally. In this work, we introduce ArtMine, a framework for discovering and formalizing artistic processes from heterogeneous historical evidence. Our approach synthesizes heterogeneous artwork evidence into a structured repository, from which a Peircean abductive agent infers evidence-grounded production steps. These steps are converted into a compositional graph and rendering prompt, then optimized through self-reflection over deviations between the generated and reference artworks. We provide a preliminary proof-of-concept case study using open-domain historical sources across multiple artists and artistic movements, demonstrating that fragmented documentary evidence can support coherent, interpretable, and auditable representations of artistic workflows. By modeling creative processes rather than only final artifacts, our work moves toward process-centred human-AI co-creativity systems that can support artistic interpretation, creative education, reflective collaboration, and computational studies of cultural production.