Forskningsradar

Science Journals

Peer-reviewade publikationer — 56237 artiklar

SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL
arXiv:2607.11185v1 Announce Type: new Abstract: Computer use agents (CUAs) are emerging as a powerful interface for automating complex digital workflows through visual perception and GUI execution. Online reinforcement learning with verifiable rewards (RLVR) has emerged as a key direction for scaling their capabilities. However, this paradigm is bottlenecked by verifiable data scarcity and online RL inefficiency. To break these barriers, we introduce ScaleCUA, a unified framework that scales online RL for CUAs via verifiable task synthesis and efficient training. At the data level, we design VeriGen, an end-to-end framework for generating verifiable RL tasks through iterative docker interactions and a multi-agent feedback loop. Scaled to 100+ concurrent agent workers via a shared docker interaction probe, this pipeline produces 24K+ verifiable tasks and nearly 3K high-quality RL tasks. To maximize sample efficiency, we propose Frontier Sampling, which tracks per-task capability and allocates rollouts to the current learning frontier. On the training side, we further design Visual Context Segmentation, a sliding window over recent visual context that balances rollout and training-engine pressure, yielding a 2.83x training speedup over step-wise decomposition. Together, ScaleCUA achieves 68.7% on OSWorld and 54.0% on ScienceBoard, establishing new state-of-the-art performance among open-source computer use agents. Code, models, and datasets are available at https://github.com/THUDM/SCALE-CUA.
[AAFLOW+] Stateful Operator Abstraction with Zero-Copy Distributed KV Cache Orchestration for Multi-Agent Workflows
arXiv:2607.10987v1 Announce Type: new Abstract: Multi-agent LLM systems increasingly integrate retrieval, planning, and reasoning, but remain fundamentally text-centric, requiring agents to repeatedly recompute shared context through expensive prefill. Although single-request inference is known to be accelerated by KV-cache management, it is usually restricted to local serving scopes. We introduce AAFLOW+, a stateful extension of agentic workflow operators that makes KV cache a first-class distributed systems object. AAFLOW+ builds processes into communication-aware graphs that concurrently optimize data, prompts, and reusable model state. It also provides operators for KV materialization, transfer, fork, composition, and eviction. Its runtime enables zero-copy, transfer-aware execution, allowing agents to reuse long context without recomputation. AAFLOW+ reduces TTFT by up to 50.2x, achieves up to 7.63x reduced multi-agent compute cost at 16-agent scale, reduces KV memory by 1.72-6.10x, and increases throughput by more than 7.74x, based on an analytical cost model parameterized by empirical hardware microbenchmarks. The results demonstrate that KV transmission outperforms recomputation on networks with moderate to high bandwidth, making sure KV-state sharing greatly increases efficiency in multi-agent LLM systems by replacing text passing.
A variational formulation of the adjoint Kutta condition in potential flow
arXiv:2606.06937v2 Announce Type: replace Abstract: We give a variational formulation of the continuous adjoint Kutta condition for two-dimensional subcritical potential flow, with emphasis on the Kutta condition and the role of the wake. We show that the adjoint Kutta condition can be imposed by a penalty term evaluated at the trailing edge, with the corresponding Lagrange multiplier determined by stationarity of the Lagrangian with respect to circulation, and that a wake treatment is not required. Some of the implications of these results for adjoint consistency are also briefly discussed.
Youtu-Parsing: Perception, Structuring and Recognition via High-Parallelism Decoding
arXiv:2601.20430v2 Announce Type: replace Abstract: This paper presents Youtu-Parsing, an efficient and versatile document parsing model designed for high-performance content extraction. The architecture employs a native Vision Transformer (ViT) featuring a dynamic-resolution visual encoder to extract shared document features, coupled with a prompt-guided Youtu-LLM-2B language model for layout analysis and region-prompted decoding. Leveraging this decoupled and feature-reusable framework, we introduce a high-parallelism decoding strategy comprising two core components: token parallelism and query parallelism. The token parallelism strategy concurrently generates up to 64 candidate tokens per inference step, which are subsequently validated through a verification mechanism. This approach yields a 5--11x speedup over traditional autoregressive decoding and is particularly well-suited for highly structured scenarios, such as table recognition. To further exploit the advantages of region-prompted decoding, the query parallelism strategy enables simultaneous content prediction for multiple bounding boxes (up to five), providing an additional 2x acceleration while maintaining output quality equivalent to standard decoding. Youtu-Parsing encompasses a diverse range of document elements, including text, formulas, tables, charts, seals, and hierarchical structures. Furthermore, the model exhibits strong robustness when handling rare characters, multilingual text, and handwritten content. Extensive evaluations demonstrate that Youtu-Parsing achieves state-of-the-art (SOTA) performance on both the OmniDocBench and olmOCR-bench benchmarks. Overall, Youtu-Parsing demonstrates significant experimental value and practical utility for large-scale document intelligence applications.
Non-Abelian holonomic transformations in digitally coupled acoustic waveguides guided by the global adiabatic criterion
arXiv:2607.11000v1 Announce Type: new Abstract: An acoustic platform is validated for implementing compact non-Abelian holonomic transformations (NHTs) guided by a global adiabatic criterion (GAC). A tripod model is mapped onto a digitally coupled four-waveguide structure, where designed coupling envelopes and an acoustically-induced-transparency phase-control module implement a two-stage phase-stitched holonomic evolution. Compared with a reference Gaussian envelope, the GAC-guided power-law profile flattens the spatial distribution of the global nonadiabatic burden, thereby providing a quantitative basis for compact acoustic implementation. Full-wave simulations show Pauli-$X$ and Hadamard-type target transformations, with excellent agreement between the extracted normalized intensities and analytical coupled-mode predictions. These target responses are obtained with half the coupling length required by the reference Gaussian implementations. More uniquely, the same phase-stitched structure also supports unidirectional acoustic mode conversion, which is closely related to a reduced two-mode non-Hermitian picture associated with an encircled exceptional point (EP). These results validate acoustic NHTs as a robust geometric route for compact wave control, establish the GAC as a powerful guideline for fast adiabatic transport in digitally coupled systems, and further demonstrate that the same phase-stitched architecture supports unidirectional mode conversion through EP-assisted branch selection.
A Novel Graph Fraud Detector via Grouped Attribute Completion and Confidence-Aware Contrastive Learning
arXiv:2607.11107v1 Announce Type: new Abstract: Graph fraud detection plays a pivotal role in safeguarding the security and integrity of modern digital ecosystems. Graph Neural Networks (GNNs) are commonly adopted for graph fraud detection. However, the practical performance of existing GNN-based detectors is severely hindered by incomplete node attributes and extreme class imbalance within graphs. To mitigate these limitations, this paper proposes a novel framework for Graph Fraud Detection with Grouped attribute completion and Confidence-aware Contrastive learning, named GFD-GC. Specifically, it first imitates heterogeneous neighborhood structures to implement group-wise aggregation, which obtains informative complete node features by capturing fine-grained graph contextual patterns. Further, it introduces a confidence-aware supervised contrastive learning strategy to augment scarce labeled fraud nodes with high confidence pseudo-fraud nodes, which enhances the compactness of fraud representations and their separability from non-fraud nodes. Extensive experiments demonstrate the superiority of the proposed GFD-GC over state-of-the-art baselines on the graph fraud detection task, thereby providing an effective solution for real-world fraud scenarios.
Secondary Flows and Near-Wall Turbulence in Channel Flow with Longitudinal Ribs
arXiv:2607.10699v1 Announce Type: new Abstract: This study employs the Large Eddy Simulation (LES) to investigate secondary flows and near-wall turbulence induced by two types of surface-mounted longitudinal ribs, namely rectangular and triangular, in a channel flow. The friction Reynolds number, based on friction velocity and channel height H, is set at 220. The rib aspect ratio W/h, where W and h represent the width and height of the rib, is 2, and the rib spacing, S is 0.6H. The results indicate formation of two counter-rotating vortices between the adjacent ribs for both the cases considered. The roughness function is higher with the rectangular rib as compared to that of the triangular rib. At the location of the mid-plane on the rectangular ribs, the wall shear stress is relatively lower as compared to that of the location of mid-plane between the ribs. Conversely, for the case of triangular rib, the opposite pattern is observed. Normal Reynolds stress exhibits strong anisotropic behaviour near the wall for both the cases, overlapping above 0.4H. Between 0.4H and 0.8H, the variation in normal Reynolds stresses is linear. Near the wall, higher production of turbulent kinetic energy (TKE) and normal Reynolds stresses are observed with the triangular rib as compared to those of the rectangular rib. The ratio of production to dissipation is unity in the log-law region.
Where Experts Disagree, Models Fail: Detecting Implicit Legal Citations in French Court Decisions
arXiv:2603.22973v2 Announce Type: replace Abstract: Applying computational methods to law at scale requires separating genuine legal reasoning from surface similarity. We study this through a concrete task: detecting implicit citations of the French Civil Code, where a court applies a statutory rule without naming it: a post-hoc question about the reasoning a court actually used. We release a benchmark of 1,015 passage-article pairs annotated by three legal experts. Our central finding is that their disagreement is itself informative: the third of cases the experts dispute are where models fail. Our best ensemble reaches an F1 score of 0.70 overall. Yet, two-thirds of its false positives fall on those disputed cases, a concentration that holds across all ten models we evaluate. Disagreement is a signal of intrinsic difficulty, not annotation noise. This should not block useful tools, however: reframed as top-$k$ ranking with multi-model consensus, the same signals reach 76% precision for the top-200 candidates without supervision.
Near-Optimal Parallel Approximate Counting via Sampling
arXiv:2604.01263v2 Announce Type: replace Abstract: The computational equivalence between approximate counting and sampling is well established for polynomial-time algorithms. The most efficient general reduction from counting to sampling is achieved via simulated annealing, where the counting problem is formulated in terms of estimating the ratio $Q={Z(\beta_{\max})}/{Z(\beta_{\min})}$ between partition functions $Z(\beta)=\sum_{x\in \Omega} \exp(\beta H(x))$ of Gibbs distributions $\mu_\beta$ over $\Omega$ with Hamiltonian $H$, given access to a sampling oracle that produces samples from $\mu_\beta$ for $\beta \in [\beta_{\min}, \beta_{\max}]$. The best bound achieved by known annealing algorithms with relative error $\varepsilon$ is $O(q \log h / \varepsilon^2)$, where $q, h$ are parameters which respectively bound $\ln Q$ and $H$. However, all known algorithms attaining this near-optimal complexity are inherently sequential, or *adaptive*: the queried parameters $\beta$ depend on previous samples. We develop a simple non-adaptive algorithm for approximate counting using $O(q \log^2 h / \varepsilon^2)$ samples, as well as an algorithm that achieves $O(q \log h / \varepsilon^2)$ samples with just two rounds of adaptivity, matching the best sample complexity of sequential algorithms. These algorithms naturally give rise to work-efficient parallel (RNC) counting algorithms. We discuss applications to RNC counting algorithms for several classic models, including the anti-ferromagnetic 2-spin, monomer-dimer and ferromagnetic Ising models.
Satellite-Free Training for Drone-View Geo-Localization
arXiv:2604.01581v3 Announce Type: replace Abstract: Drone-view geo-localization (DVGL) aims to determine the location of drones in GPS-denied environments by retrieving the corresponding geotagged satellite tile from a reference gallery given UAV observations of a location. In many existing formulations, these observations are represented by a single oblique UAV image. In contrast, our satellite-free setting is designed for multi-view UAV sequences, which are used to construct a geometry-normalized UAV-side location representation before cross-view retrieval. Existing approaches rely on satellite imagery during training, either through paired supervision or unsupervised alignment, which limits practical deployment when satellite data are unavailable or restricted. In this paper, we propose a satellite-free training (SFT) framework that converts drone imagery into cross-view compatible representations through three main stages: drone-side 3D scene reconstruction, geometry-based pseudo-orthophoto generation, and satellite-free feature aggregation for retrieval. Specifically, we first reconstruct dense 3D scenes from multi-view drone images using 3D Gaussian splatting and project the reconstructed geometry into pseudo-orthophotos via PCA-guided orthographic projection. This rendering stage operates directly on reconstructed scene geometry without requiring camera parameters at rendering time. Next, we refine these orthophotos with lightweight geometry-guided inpainting to obtain texture-complete drone-side views. Finally, we extract DINOv3 patch features from the generated orthophotos, learn a Fisher vector aggregation model solely from drone data, and reuse it at test time to encode satellite tiles for cross-view retrieval. Experimental results on University-1652 and SUES-200 show that our SFT framework substantially outperforms satellite-free generalization baselines and narrows the gap to methods trained with satellite imagery.
Neural Discovery of Memory and Nonlocal Kernels in Integro-Differential Equations with Constrained Kolmogorov--Arnold Networks
arXiv:2607.11110v1 Announce Type: new Abstract: Discovering the memory or nonlocal kernel governing an integro-differential equation (IDE) from sparse and noisy observations is an ill-posed inverse problem. Existing identification methods often rely on problem-specific analytical derivations, specialized observation requirements, or restrictive assumptions about the kernel, limiting their applicability across different classes of IDEs. In this work, we propose a differentiable-solver-based framework for discovering memory and nonlocal kernels directly from spatiotemporal observations. Within the solver, the unknown kernel is represented using a constrained Kolmogorov--Arnold Network (KAN) parameterization, with the physical constraints imposed through two different approaches: a Bernstein-polynomial-based Monotone--Convex KAN (MC-KAN), whose coefficient constraints enforce positivity, monotonic decrease, and convexity by construction, and a Chebyshev-based KAN (Cheb-KAN), in which the same properties are encouraged through soft penalty terms. After training, symbolic regression is applied to the learned kernels to obtain interpretable closed-form representations. We evaluate both methods on benchmarks spanning a one-dimensional Volterra equation, a one-dimensional viscoelastic wave partial integro-differential equation, and a two-dimensional nonlocal reaction-diffusion equation with an anisotropic coupled kernel. For the 1D problems, both methods recover the correct kernel functional form and achieve comparable solution-reconstruction accuracy. In contrast, for the sparse and noisy 2D nonlocal problem, the hard-constrained MC-KAN consistently achieves lower kernel reconstruction errors than the soft-constrained Cheb-KAN. Our results demonstrate that enforcing physically motivated shape constraints by construction provides greater robustness than soft penalties for multidimensional kernel discovery from sparse and noisy observations.
Compact and Stable Representation of Real-Frequency Spectral Functions for Machine Learning
arXiv:2607.11190v1 Announce Type: new Abstract: We introduce a compact and stable moment representation for real-frequency Green's functions, hybridization functions, and self-energies for machine-learning applications, avoiding the inefficiency of dense frequency grids as well as the ill-posed analytic continuation of Matsubara approaches. The representation is constructed from Cayley-mapped trigonometric moments with the Jacobian included, which preserve spectral-weight normalization, tie the moment sequence to a positive matrix-valued spectral measure, and admit a systematic route to a pole representation via ESPRIT. This provides a fixed-dimensional learning target in which physical constraints such as normalization and positivity can be imposed directly. Using a graph-attention neural network with FiLM conditioning, we benchmark the representation on single-orbital DMFT, antiferromagnetic DMFT, and a two-orbital impurity model. The results demonstrate accuracy matching or exceeding that of direct frequency-domain learning, reliable reproduction of the density and staggered magnetization, stable self-energy reconstruction through Dyson equation inversion, and accurate recovery of matrix-valued spectra with orbital mixing.
HandFlow: Fully Generative 4D Hand Recovery with Flow Matching
arXiv:2607.11221v1 Announce Type: new Abstract: Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions. Temporal models improve consistency by aggregating information across frames, but they are typically deterministic regressors, making them vulnerable to ambiguous observations caused by occlusion and motion blur. Generative modeling offers a natural alternative by learning a prior over plausible hand motion sequences, enabling coherent hand-state recovery when visual evidence is incomplete or unreliable. Motivated by this observation, we present HandFlow, a fully generative flow-matching framework for temporally coherent 3D hand pose and shape estimation from monocular video. Given visual and skeletal observations, HandFlow denoises an entire temporal window of MANO parameters through a single ODE integration. To support this, we use a Flux-style dual-stream transformer that attends across the full sequence to capture long-range dependencies without autoregressive decoding, and a confidence-aware continuous masking mechanism that blends observed features with learnable mask tokens to handle noisy or missing observations. Experiments on DexYCB and HOT3D show that HandFlow achieves state-of-the-art performance, with particularly large gains in world-space accuracy and temporal smoothness. It reduces world-space pose error by over 30% compared with the strongest baseline and achieves the lowest acceleration error among all evaluated methods, while remaining competitive in per-frame pose accuracy. Moreover, on a single GPU HandFlow reconstructs a 150-frame sequence at 47 fps, about 12x faster than the fastest prior video-based method, with reconstruction itself accounting for only a small fraction of the end-to-end latency.
Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution
arXiv:2607.11111v1 Announce Type: new Abstract: LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone to factual errors caused by insufficient repository understanding. Recent methods attempt to mitigate this limitation through pre-repair repository exploration; however, their fix-driven strategies explore repositories without identifying the agent's knowledge gaps, often yielding imprecise context that fails to bridge the underlying understanding deficit. In this paper, we propose ACQUIRE, a QA-driven framework for software issue resolution. Mirroring how experienced developers first comprehend unfamiliar code before attempting a fix, ACQUIRE explicitly acquires repository knowledge prior to repair. The framework decouples knowledge acquisition from patch generation through two stages: in the first stage, a Questioner and an Answerer collaborate to acquire structured repository knowledge, where the Questioner poses targeted questions and the Answerer produces evidence-grounded answers through autonomous exploration; in the second stage, the Resolver leverages the resulting QA knowledge to generate informed patches. By transforming implicit knowledge gaps into explicit, factually reliable understanding, ACQUIRE accelerates knowledge-intensive repair stages and enables more accurate resolution. Experiments on SWE-bench Verified demonstrate that ACQUIRE consistently outperforms representative pre-repair methods, raising Pass@1 by up to 4.4 percentage points with modest additional cost and time.
Stop to Decide: Latency-Aware Proprioceptive Navigation Primitives for Mapping-Free Quadruped Inspection
arXiv:2607.11204v1 Announce Type: new Abstract: Compute-constrained quadrupeds often run their navigation loop far below the controller's design rate: sharing the onboard Jetson Orin with the vision pipeline slows our stair loop to about 15 Hz. This latency breaks a standard proprioceptive pattern: declaring stair-summit arrival from the body-pitch signal while still climbing. On a stepped platform whose 50 cm top is shorter than the robot (Unitree Go2, about 75 cm), in-motion detection overshoots the top edge with probability rising with the per-period advance v/f (the slowest about 15 Hz cell partly diluted by a separate non-arrival mode), whereas a climb-settle cadence holds overshoot near zero at every loop rate (pooled 22/45 vs 1/45 over about 30/20/15 Hz; Fisher p about 2.4e-7; 7/15 vs 0/15 at the deployed about 15 Hz). A logistic dose-response model in v/f captures the failure; a pre-specified 40 Hz out-of-sample test favours the protocol-clean fit (33% observed vs 43%/22% predicted), giving a deployment rule (critical loop rate about 19 Hz at 0.30 m/s). The detector sits in a fully onboard, mapping-free and learning-free stack: built-in inertial measurement unit, four foot-force channels, three 1-D ranges, one line camera, chaining line-following, a three-segment maneuver for 90-degree corners in a 55 cm corridor (20/20 contact-free vs 14/20 with 12 wall contacts for in-place yaw; exit-heading error 1.56 degrees vs 5.64 degrees), and stair traversal, completing the inspection course in 18/20 trials (90%). Results are from a single course geometry, platform, and operator.
LLM-Generated Design Problems for Assessing Higher-Order Thinking in Project-Based Learning
arXiv:2607.11032v1 Announce Type: new Abstract: Project-based learning (PjBL) is common in computing education, but traditional assessments of PjBL often fail to capture higher-order thinking (HOT), especially in transfer contexts. This study introduces "design problems" (DPs): concise, scenario-based prompts that require applying project concepts in new situations, to address this gap. We examined instructor perceptions, the ability of large language models (LLMs) to generate DPs, and student experiences. Surveys of 31 instructors, evaluation of 80 LLM-generated DPs, and student performance data showed that while instructors value DPs, creation effort is a barrier. LLMs helped by producing high-quality prompts with strong expert agreement. Students rated DPs from different LLMs similarly, and their performance on DP tasks showed negligible correlation with traditional project grades, suggesting DPs may capture distinct aspects of HOT. Keystroke data also suggested deeper cognitive engagement of students through planning and revision behaviors. Overall, DPs appear to be a useful complement to traditional assessments, especially in situations where AI use or collaboration may undermine individual learning.
FastTPS: An Optimized Method for LLM Token Phase for AI accelerators
arXiv:2607.11211v1 Announce Type: new Abstract: The popularity of large language models (LLMs) escalates an ongoing demand for effective inference. However, due to the sequential processing of tokens during the token phase in decoder-only LLMs inference, the inherent low parallelism leads to reduced throughput and suboptimal utilization of the computing units on artificial intelligence (AI) accelerators, particularly when handling long-sequence inputs that impose significant memory overhead. Recently, many reported methods have been developed as potential solutions, since they emerge with numeric deviation. This paper presents FastTPS, a high performance and low-precision loss method for accelerating the token-phase in LLM inference on general AI accelerators which includes three key components: (1) AI accelerator-enabled reloading-free KV Cache concatenation which decreases memory access overhead as well as enables full fusion of Attention, (2) high-efficiency and high-accuracy 'RoPE' attention based on the tiling optimized FLAT, and (3) highly-fused MLP with fine-grain pipeline scheduling. Our results confirm that FastTPS significantly alleviates memory bottlenecks in the token phase, delivering a 6x speed improvement (compared to none-fusion) on an AMD Ryzen AI 300 series NPU with BF16 precision while sustaining 93% peak memory bandwidth utilization during Phi3-mini-4k-instruct inference.
SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception
arXiv:2607.11008v1 Announce Type: new Abstract: Open-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remarkable progress of recent CLIP-based approaches, we identify a critical limitation: synonym-induced grounding inconsistency, where semantically equivalent expressions yield disparate spatial attention patterns. This inconsistency undermines the robustness and performance of existing methods in real-world OVDP applications. To address this issue, we propose SynCLIP, a Synonym-Coherent Language-Image Pretraining framework that enhances synonym-robust grounding for OVDP. SynCLIP introduces a Semantic-consistent Spatial Attention alignment (SSA) module to enhance spatial attention consistency by minimizing discrepancies between attention maps of original and synonymous expressions. Furthermore, a Spatial Attention Refinement (SAR) module selectively strengthens the most semantically relevant spatial regions within aligned maps for more precise and stable grounding. To support synonym-coherent pretraining, we also construct a Synonym-Enriched Visual Corpus (SEViC), which augments each category with multiple synonyms and textual definitions. Extensive experiments on multiple benchmarks demonstrate that SynCLIP substantially improves grounding consistency under diverse linguistic variants and achieves state-of-the-art performance among CLIP-based OVDP methods. Code is available at https://github.com/Justlovesmile/SynCLIP.
PREF-Gate: Provenance-Constrained Relational Evidence Fusion with Validation-Gated Selection for Graph Fraud Detection
arXiv:2607.11212v1 Announce Type: new Abstract: Relational fraud detection can exploit both label-free graph context and label-derived neighborhood evidence, but these two information sources obey different validity conditions. In particular, neighborhood risk becomes invalid when a queried node's own label, or any validation or test label, enters its construction. We formulate this issue as provenance-constrained relational evidence use and present PREF-Gate, an auditable decision framework with two fixed experts and a finite validation gate. The context expert uses attributes, one-hop means, feature residuals, and degree descriptors without labels. The evidence expert adds self-excluded, training-label-only neighborhood risk and empirical-Bayes summaries that expose support, uncertainty, availability, and shrinkage. Before test inference, the gate selects either expert or one of three pre-specified probability mixtures and fixes the decision threshold. On Amazon, YelpChi, and TFinance, using five identical stratified splits and 14 same-protocol methods, PREF-Gate obtains mean AUPRC values of 0.9085, 0.8104, and 0.8913. It selects the label-free expert on all Amazon and YelpChi splits and an evidence mixture on all TFinance splits. Thus, the main result is conditional rather than universal: label-derived relational evidence is useful only where held-out validation supports it. The framework couples competitive ranking performance with an explicit label-provenance contract, finite selection policy, failure accounting, and review-budget evaluation, providing an auditable knowledge-based decision pipeline for graph fraud detection.
When the Target Domain Changes: AI-Mediated Construct Drift in High-Stakes English Language AssessmenW
arXiv:2607.11213v1 Announce Type: new Abstract: High-stakes English proficiency tests treat standardized, unaided performance as evidence for score interpretations about academic English proficiency. This interpretation remains meaningful, but as target language use domains increasingly involve generative AI, the extrapolation from unaided test performance to academic communicative readiness becomes less self-evident. This conceptual validity argument reframes AI as a score-interpretation problem in high-stakes language testing, not only an operational issue of scoring, feedback, security, or misconduct. Synthesizing current literature in three uneven layers, the paper shows that most work treats AI as assessment infrastructure, while far less theorizes its implications for construct validity and extrapolation warrants. It defines AI-mediated construct drift as the misalignment that arises when communicative abilities required in the target domain change through AI mediation while test constructs remain anchored to an unaided-performance model. It proposes bounded AI mediation as a validity-oriented design principle: a standardized condition in which all test takers access the same institutionally controlled AI assistant, with predefined assistance boundaries, logged interactions, and tasks that distinguish comprehension support from answer generation. The paper argues that score interpretations should be narrowed and supplemented when used to support claims about AI-mediated academic communication.
A Novel Method to Evaluate Models on Unreliable, Noisy and Inconsistent Labels: Adaptive Resolution Label Aggregation (ARLA)
arXiv:2607.11214v1 Announce Type: new Abstract: Labels are critical for both training and evaluating deep learning segmentation models, but are often inconsistent, noisy, or ambiguous at class boundaries. Many approaches have been developed to support training models on weak labels, but few to none currently exist to facilitate evaluating models on unreliable labels. We therefore introduce a method called "Adaptive Resolution Label Aggregation", or "ARLA", which dynamically adapts the resolution of both the label and the model prediction at inference time before the evaluation metrics are computed. We demonstrate how ARLA can be used to better analyse model behaviour with a practical application to a real flood prediction model, where ARLA was able to overcome issues with inconsistent labelling of forested areas and errors in labels within regions of heavy cloud cover. Our work presents a new approach to evaluating segmentation models, with adjustable parameters to adapt the aggregated resolution to the precision of the label or the level of label noise. Fundamentally, ARLA exploits the information encapsulated by a label but minimises the label error, extracting from the noise a clearer signal of a model's true performance.
The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium Reasoning
arXiv:2607.11116v1 Announce Type: new Abstract: Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved equilibrium should encode the result of genuine iterative inference. We report a cautionary study of a port-Hamiltonian DEQ with a learned initialization on two reasoning tasks -- ProofWriter entailment over frozen DeBERTa embeddings and a BFS-verified graph-reachability benchmark -- in which the implicit computation is a silent no-op. Across tasks, seeds, and controlled ablation arms, the solved equilibrium equals the solver's start point to numerical precision, and bypassing the solver entirely changes test accuracy by +0.00 percentage points in 18 of 19 training runs. Controlled interventions falsify the tempting explanation: removing the anchoring term reproduces every result, and retraining with noise-decoupled starts yields a solver that converges to the noisy start while the decoder learns to ignore it. The single escaping run diverges instead ($\|h^{*}-z_0\|=171$), producing a co-adapted noise channel whose removal improves accuracy. Iteration counts are uncorrelated with ground-truth difficulty ($r=0.009$), and the full apparatus never outperforms a two-layer MLP on either task. We trace the mechanism to gradient starvation along two distinct routes, show that the standard zeroing ablation is confounded and gives wildly seed-dependent answers where the correct substitution test gives a stable zero, and distill a four-test diagnostic protocol for auditing claimed implicit computation. All experiments run on a single free Colab GPU; code, raw logs, and analysis scripts are released.
DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs
arXiv:2607.11228v1 Announce Type: new Abstract: While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. Existing bias evaluation protocols predominantly rely on static datasets, which provide only a superficial assessment, as their fixed test cases cannot adaptively evolve to measure the true depth and limits of model vulnerabilities. We introduce DeepBias, an adaptive framework for the in-depth probing of social biases in LVLMs with carefully designed agents. Our approach operates through a dynamic ''generation-evolution-probing'' loop. First, a generative ProposerAgent synthesizes test data and is iteratively updated via Direct Preference Optimization (DPO) based on the target LVLM's responses, exploring model-specific failure modes. Second, an autonomous skill-driven DiggerAgent rewrites each test data across multiple probing turns, adaptively selecting from a curated skill library of deepening and rewriting strategies. At each turn, this process is conditioned on the model's previous response, enabling progressively deeper biases to be exposed. Furthermore, we build a benchmark named DeepBiasBench using our framework. By employing an ensemble of five diverse state-of-the-art LVLMs as anchors, the benchmark captures vulnerabilities shared across architectures. Comprehensive experiments demonstrate the effectiveness of our framework and show that DeepBias provides a challenging benchmark for in-depth bias evaluation, establishing an evolutionary paradigm for LVLM safety assessment.
Discrete Wavelet Transform for Serial X-ray Crystallography Image Segmentation
arXiv:2605.19199v2 Announce Type: replace Abstract: Upcoming LCLS-II/II-HE operation at repetition rates approaching 1MHz demands on-detector data reduction to manage the resulting data volumes. We present a 2D discrete wavelet transform (DWT) pre-processing algorithm that segments background scatter from crystal diffraction in serial crystallography images, enabling early data analysis and, when combined with peak finding, lossy compression by transmitting only the identified diffraction peaks. The method zeroes the approximation (LL) coefficients of a multi-level Haar wavelet decomposition and reconstructs from detail subbands only, exploiting the natural separation of smooth background and sharp Bragg peaks in the wavelet domain. Evaluated on 100 simulated nanoBragg frames with known ground truth, the pipeline achieves $F1 \approx 0.96$ at four decomposition levels ($J = 4$), substantially outperforming the established peakfinder8 algorithm ($F1 \approx 0.37$) in both precision ($P \approx 1.00$ vs.\ $0.94$) and recall ($R \approx 0.92$ vs.\ $0.24$). A comparison of 12 wavelet families confirms that Haar is optimal for Bragg-peak detection due to its minimal filter support. Downstream crystallographic analysis performed on real ePix10kA data shows that CC* and $R_\mathrm{split}$ converge at $J = 4$ and track the unprocessed baseline through the practical resolution limit. Under added noise exceeding $\sim$50 ADU, the current pipeline's precision degrades significantly more than that of the pf8 algorithm, exposing a limitation of the proposed strategy. We also demonstrate an FPGA implementation of the DWT filters on an Alveo U200 at 200MHz, with a projected resource footprint compatible with integration into the upcoming ePixUHR firmware and a path to on-detector ASIC implementation in SparkPix detector family.
IEnSF: Iterative Ensemble Score Filter for Reducing Error in Posterior Score Estimation in Nonlinear Data Assimilation
arXiv:2510.20159v2 Announce Type: replace Abstract: The Ensemble Score Filter (EnSF) is a score-based diffusion model approach for solving high-dimensional and nonlinear data assimilation problems. While initial applications of EnSF to the Lorenz-96 model and the quasi-geostrophic system showed potential, the current method employs a heuristic weighted sum to combine the prior and the likelihood score functions. This introduces a structural error into the estimation of the posterior score function in the nonlinear setting. This work addresses this challenge by developing an iterative ensemble score filter (IEnSF) that applies an iterative algorithm as an outer loop around the reverse-time stochastic differential equation solver. When the state dynamics or the observation operator is nonlinear, the iterative algorithm can gradually reduce the posterior score estimation error by improving the accuracy of approximating the conditional expectation of the likelihood score function. The number of iterations required depends on the distance between the prior and posterior distributions. Numerical experiments demonstrate that the IEnSF algorithm substantially reduces the error in posterior score estimation in the nonlinear setting and thus improves the accuracy of tracking high-dimensional dynamical systems.