Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

ERP Data Provisioning Financial Control Testing
arXiv:2607.09712v1 Announce Type: new Abstract: Financial control testing increasingly depends on representative enterprise resource planning (ERP) data in quality environments, yet direct production copies expose personal, supplier, banking, and commercially sensitive records. This work presents Secure ERP Quality Provisioning for Financial Control Testing (SEQ-FCT), a governed data-provisioning framework that combines deterministic masking, synthetic scenario expansion, referential tokenization, policy-based release approval, and automated validation for reconciliation, fraud-rule testing, and audit analytics. A single synthetic dataset is used for evaluation. It contains 186,000 finance-process records from six subsidiaries over 2022-2025, including accounts payable invoices, payments, general-ledger journals, accounts receivable receipts, and bank-statement lines. The dataset includes entity relationships, monetary values, approval paths, tax attributes, banking markers, exception labels, fraud-rule triggers, and control-failure outcomes. Because the dataset is synthetic, reported results demonstrate controlled internal consistency rather than production validation. Against a production-clone upper bound, static masking, rules-only synthesis, conditional tabular generative synthesis, and a hybrid baseline, SEQ-FCT achieved 0.932 reconciliation F1, 0.887 fraud-trigger recall, 0.914 control-failure F1, and an estimated leakage-risk score of 0.018. The analysis indicates that financial process behavior can be preserved more reliably when masking, synthetic data, and governance checks are evaluated as a single release pipeline instead of independent utilities.
EvidentialRAG: Quantifying and Mitigating Information Conflict in Multi-Source Retrieval-Augmented Generation via Evidential Deep Learning
arXiv:2607.10491v1 Announce Type: new Abstract: Retrieval-augmented generation grounds large language models in external evidence, but most pipelines still treat retrieved passages as deterministic and mutually consistent context. In open information environments, retrieved sources may disagree because of temporal drift, source error, ambiguity, or genuine uncertainty. This paper introduces ERAG, an uncertainty-aware RAG framework that converts retrieved chunks into probabilistic evidence before generation. A lightweight evaluator extracts candidate claims and maps chunk-level support to Dirichlet evidence. A conflict-preserving Dempster-Shafer fusion rule then transfers unresolved disagreement into epistemic uncertainty rather than normalizing it away. The generator is routed to direct answering, conflict-aware answering, or abstention according to the fused uncertainty score. Experiments on CRAG, ConflictQA, and MuSiQue show that ERAG remains competitive with the strongest matched baseline on standard question answering while improving behavior under conflict. On the CRAG ambiguous subset, hallucination decreases from 45.3% for Corrective RAG to a human-calibrated estimate of 34.8%, conflict resolution increases from 35.2% to 51.2%, and expected calibration error improves to 0.122. These results suggest that evidential modeling is a practical mechanism for trustworthy information processing in foundation-model-based retrieval systems.
Narrix: Remixing Narrative Strategies from Examples for Story Writing
arXiv:2604.07643v1 Announce Type: cross Abstract: Experienced storytellers decompose stories into local narrative strategies and how these strategies shape higher-level arcs. This decomposition helps writers recognize patterns in others' work and adapt those patterns to tell new stories. Novices, however, struggle to identify these strategies or to reuse them effectively. We present Narrix, a novel writing tool that helps novice writers recognize narrative strategies in example stories and repurpose these strategies in their own writing. Narrix analyzes strategies in example stories, highlights them with color-coded lexical cues and explanations, and situates them on an interactive story arc for exploration by emotional shifts and turning points. Writers then drag strategies onto multi-dimensional tracks and apply block-scoped edits to revise or continue their drafts through controlled generation steered by specified strategies. Through a within-subjects study (N=12), Narrix showed improved participants' retention, confidence, and creative adaptation of narrative strategies compared to a baseline chat-based writing interface.
The Theory of Strategic Evolution: Games with Endogenous Players and Strategic Replicators
arXiv:2512.07901v3 Announce Type: replace Abstract: Von Neumann founded both game theory and the theory of self-reproducing automata, but the two programs never merged. This paper provides the synthesis. The Theory of Strategic Evolution analyzes strategic replicators: entities that optimize under resource constraints and spawn copies of themselves. We introduce Games with Endogenous Players (GEPs), where lineages (not instances) are the fundamental strategic units, and define Evolutionarily Stable Distributions of Intelligence (ESDIs) as the resulting equilibrium concept. The central mathematical object is a hierarchy of strategic layers linked by cross-level gain matrices. Under a small-gain condition (spectral radius less than one), the system admits a global Lyapunov function at every finite depth. We prove closure under meta-selection: adding governance levels, innovation, or constitutional evolution preserves the dynamical structure. The Alignment Impossibility Theorem shows that unrestricted self-modification destroys this structure; stable alignment requires bounded modification classes. Applications include AI deployment dynamics, market concentration, and institutional design. The framework shows why personality engineering fails under selection pressure and identifies constitutional constraints necessary for stable multi-agent systems.
Frequency Locking to Environmental Forcing Suppresses Oscillatory Extinction in Phage-Bacteria Interactions
arXiv:2512.08224v2 Announce Type: replace Abstract: Bacteriophage-bacteria interactions are central to microbial ecology, influencing evolution, biogeochemical cycles, and pathogen behavior. Most theoretical models assume static environments and passive bacterial hosts, neglecting the joint effects of bacterial traits and environmental fluctuations on coexistence dynamics. This limitation hinders the prediction of microbial persistence in dynamic ecosystems such as soils and oceans. Using a minimal ordinary differential equation framework, we demonstrate that environmental fluctuations can suppress destructive oscillations through resonance, promoting coexistence where static models otherwise predict collapse. Counterintuitively, we find that lower bacterial growth rates are helpful in enhancing survival under high infection pressure, elucidating the observed post-infection growth reduction. Our studies highlight bacterial hosts as active builders of ecological dynamics and environmental variation as a potential stabilizing force. Our findings thus bridge a theory-experiment gap and provide a framework for predicting microbial responses to environmental stress, which might have potential implications for phage therapy, microbiome management, and climate-impacted community resilience as well.
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
arXiv:2512.10719v3 Announce Type: replace Abstract: End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning capabilities obtained from the large-scale pretraining. However, we find that current VLMs struggle to understand fine-grained 3D spatial relationships which is a fundamental requirement for systems interacting with the physical world. To address this issue, we propose SpaceDrive, a spatial-aware VLM-based driving framework that treats spatial information as explicit positional encodings (PEs) instead of textual digit tokens, enabling joint reasoning over semantic and spatial representations. SpaceDrive employs a universal positional encoder to all 3D coordinates derived from multi-view depth estimation, historical ego-states, and text prompts. These 3D PEs are first superimposed to augment the corresponding 2D visual tokens. Meanwhile, they serve as a task-agnostic coordinate representation, replacing the digit-wise numerical tokens as both inputs and outputs for the VLM. This mechanism enables the model to better index specific visual semantics in spatial reasoning and directly regress trajectory coordinates rather than generating digit-by-digit, thereby enhancing planning accuracy. Extensive experiments validate that SpaceDrive achieves state-of-the-art open-loop performance on the nuScenes dataset and the second-best Driving Score of 78.02 on the Bench2Drive closed-loop benchmark over existing VLM-based methods. Code is available at: https://github.com/zhenghao2519/SpaceDrive.
Sharp convergence bounds for sums of POD and SPOD weights
arXiv:2512.13068v2 Announce Type: replace Abstract: This work analyzes the convergence of sums of the form $S_{\boldsymbol{\gamma}}(m)=\sum_{v\subseteq \mathbb{N}}\gamma_v m^{|v|}$ with product and order dependent (POD) weights $\gamma_v$. We establish that for a nonnegative sequence $\{\Upsilon_j\mid j\in \mathbb{N}\}$, $$\sum_{v\subseteq \mathbb{N}} |v|! m^{|v|}\prod_{j\in v} \Upsilon_j<\infty \text{ for all } m>0 \text{ if and only if } \sum_{j=1}^\infty \Upsilon_j<\infty.$$ We further characterize the growth of $S_{\boldsymbol{\gamma}}(m)$ when $\gamma_v=(|v|!)^{\sigma}\prod_{j\in v}j^{-\rho}$ and prove that $\log S_{\boldsymbol{\gamma}}(m)$ is of asymptotic order $m^{1/(\rho-\sigma)}$ when $\rho>\sigma\geq 0$. We subsequently generalize both the convergence criterion and the asymptotic order of $\log S_{\boldsymbol{\gamma}}(m)$ to smoothness-driven product and order dependent (SPOD) weights, while noting that a full necessary-and-sufficient analogue remains open. Finally, we apply our theory to quasi-Monte Carlo (QMC) integration, showing that interlaced polynomial lattice rules achieve a dimension-independent convergence rate without a commonly imposed assumption in the QMC literature.
MoLingo: Motion-Language Alignment for Text-to-Human Motion Generation
arXiv:2512.13840v3 Announce Type: replace Abstract: We introduce MoLingo, a text-to-motion (T2M) model that generates realistic, lifelike human motion by denoising in a continuous latent space. Recent works perform latent space diffusion, either on the whole latent at once or auto-regressively over multiple latents. In this paper, we study how to make diffusion on continuous motion latents work best. We focus on two questions: (1) how to build a semantically aligned latent space so diffusion becomes more effective, and (2) how to best inject text conditioning so the motion follows the description closely. We propose a semantic-aligned motion encoder trained with frame-level text labels so that latents with similar text meaning stay close, which makes the latent space more diffusion-friendly. We also compare single-token conditioning with a multi-token cross-attention scheme and find that cross-attention gives better motion realism and text-motion alignment. With semantically aligned latents, auto-regressive generation, and cross-attention text conditioning, our model sets a new state of the art in human motion generation on standard metrics and in a user study. We will release our code and models for further research and downstream usage.
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
arXiv:2512.17776v5 Announce Type: replace Abstract: Recent advances in large language models have enabled deep research systems that generate expert-level reports through multi-step reasoning and evidence-based synthesis. However, evaluating such reports remains challenging: report quality is multifaceted, making it difficult to determine what to assess and which criteria to use; LLM-based judges may miss errors that require domain expertise to identify; and because deep research relies on retrieved evidence, report-wide claim verification is also necessary. To address these issues, we propose DEER, a benchmark for evaluating expert-level deep research reports. DEER systematizes evaluation criteria with an expert-developed taxonomy (7 dimensions, 25 subdimensions) operationalized as 101 fine-grained rubric items. We also provide task-specific Expert Evaluation Guidance to support LLM-based judging. In addition to rubric-based assessment, we propose a claim verification architecture that verifies both cited and uncited claims and quantifies evidence quality. Experiments show that current systems produce structurally plausible, evidence-citing reports, but still struggle to fully satisfy expert-level user requests and achieve logical completeness. Beyond performance comparisons, DEER makes system strengths and limitations interpretable and provides diagnostic signals for improvement.
MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
arXiv:2512.19311v2 Announce Type: replace Abstract: This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the input of a prediction network at one training timestep is the corresponding ground-truth noisy data that is an interpolation of the noise and the data, and during testing, the input is the generated noisy data. We present a novel training approach, named MixFlow, for improving the performance. Our approach is motivated by the Slow Flow phenomenon: the ground-truth interpolation that is the nearest to the generated noisy data at a given sampling timestep is observed to correspond to a higher-noise timestep (termed slowed timestep), i.e., the corresponding ground-truth timestep is slower than the sampling timestep. MixFlow leverages the interpolations at the slowed timesteps, named slowed interpolation mixture, for post-training the prediction network for each training timestep. Experiments over class-conditional image generation (including SiT, REPA, and RAE) and text-to-image generation validate the effectiveness of our approach. Our approach MixFlow over the RAE models achieve strong generation results on ImageNet: 1.43 FID (without guidance) and 1.10 (with guidance) at 256 x 256, and 1.55 FID (without guidance) and 1.10 (with guidance) at 512 x 512.
Integral modelling of weakly evaporating 3D liquid film with variable substrate heating
arXiv:2512.21299v2 Announce Type: replace Abstract: Analysing the dynamics of phase-changing liquid films is essential for enhancing the performance of thermal management systems. Still, direct simulation of the full governing equations is computationally expensive. To circumvent this limitation, I derived a weighted-integral boundary-layer (WIBL) model under long-wave assumptions, weak evaporation, and strong surface tension, also accounting for variable substrate heating. In the linear regime, the WIBL reproduces growth rates and the cutoff wavenumber of unstable modes with significantly higher accuracy than commonly used Benney-type models for Re<40, as compared to the Orr-Sommerfeld equations. The linear analysis further reveals a threshold separating streamwise- and spanwise-dominated instabilities in hanging films, arising from the competition between Kapitza and Rayleigh-Taylor mechanisms; the WIBL predicts this threshold accurately for small Re and inclination angles. In the nonlinear regime, with substrate heating that varies in both space and time, the WIBL model captures the evolution of free-surface thickness and temperature within approximately 6% of the original Navier-Stokes equations. Three-dimensional simulations show that a condensing film undergoes dry-out due to Kapitza instability, whereas unsteady substrate heating promotes spanwise momentum spreading, modifies wave dynamics, and prevents dry-out. The WIBL model provides a good level of accuracy at a low computational cost, enabling extensive parametric studies, nonlinear stability analyses, and the design of optimal substrate-heating control strategies.
Implicit Fine-tuning via Context Engineering: A Curriculum Learning Framework for Multimodal Entity Alignment
arXiv:2607.10532v1 Announce Type: new Abstract: Multimodal Entity Alignment (MMEA) aims to identify equivalent entities across different modalities. While existing methods enhance MMEA performance through black-box context engineering strategies, their reliance on LLM parameter capacity and lack of theoretical interpretability remain unresolved. To this end, we first theoretically validate the mathematical equivalence between context engineering and model fine-tuning in MMEA tasks, demonstrating that prompt components simulate contrastive learning-based sequential fine-tuning in MMEA. Building on this foundation, we then propose PTFEA, a curriculum-learning-inspired framework that translates fine-tuning strategies into interpretable context engineering. Specifically, adaptive difficulty modulation dynamically adjusts information injection stages using confidence thresholds, establishing mathematical equivalence between curriculum learning weights and context sample selection; and three-stage progressive inference incorporates entity information from simple to complex cases, mirroring the gradient descent process in fine-tuning. Experiments on five public datasets demonstrate that PTFEA consistently outperforms strong baselines. In particular, on the ICWIKI dataset, PTFEA narrows the H@1 gap between Qwen2.5-72B and 14B to 0.6%. Moreover, compared with the representative context-engineering-based MMEA method MM-ChatAlign, PTFEA reduces the runtime of Qwen2.5-72B from 21 hours to 1 hour and lowers token consumption from 2200-3000 to 200-400, achieving over 80% reduction on the ICWIKI dataset. This work provides the first theoretical framework unifying context engineering and fine-tuning in MMEA, paving the way for future research that seeks to translate additional fine-tuning strategies into context engineering paradigms. Our code is available at https://github.com/DMiC-Lab-HFUT/PTFEA.
Tool-Adaptive LLM Reranker
arXiv:2607.10555v1 Announce Type: new Abstract: Generative Large Language Models (LLMs) have revolutionized information retrieval, yet their strictly parametric nature frequently leads to severe factual hallucinations when confronted with complex queries beyond their epistemic boundaries. While external tool-calling can mitigate this, indiscriminately invoking search tools for every document during reranking incurs prohibitive latency overheads, creating an intractable accuracy-efficiency dilemma. To address this challenge, we propose TALRanker, a novel framework that formalizes pointwise relevance scoring as an agentic Markov decision process. We optimize it via a two-stage training paradigm. An initial warm-up utilizes a language-preserving hybrid loss to prevent the catastrophic forgetting of native generative capacities. Subsequently, an asymmetric cost-aware reward equipped in reinforcement learning forces the policy to autonomously bypass tools for maximum efficiency when confident, while selectively retrieving external evidence to avert severe hallucination penalties when uncertain. Extensive evaluations demonstrate that TALRanker achieves state-of-the-art performance across standard and reasoning-intensive retrieval benchmarks, matching throughput with pointwise rerankers while outperforming parameter-heavy reasoning models.
UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp
arXiv:2607.10557v1 Announce Type: new Abstract: Multimodal BrowseComp tasks require agents to combine perception, tool use, and long-horizon reasoning over dynamic web content, challenging their ability to handle compositional structure, open-world uncertainty, and multimodal integration across extended interactions. Crucially, real-world multimodal browsing involves three distinct information-flow patterns: text-only, image-to-text, and text-to-image, yet existing data construction methods cover only the text-only and image-to-text patterns, leaving text-to-image largely unaddressed and limiting agent generality and robustness. We introduce UNIBROWSE, a unified data pipeline that for the first time simultaneously generates training data covering all three patterns, augments curated knowledge graphs with live web retrieval for improved fidelity, and introduces a novel metric of exploration degree to filter low-signal instances for efficient reinforcement learning. Through this pipeline, we produce high-quality cold-start tool-use trajectories and exploration-rich QA pairs, and train a 35B-scale agent via supervised fine-tuning and exploration-aware RL.The resulting UNIBROWSE agent achieves state-of-the-art performance on multimodal BrowseComp benchmarks, attaining an average accuracy of 54.4 across five diverse benchmarks -- an improvement of 10.5 points over its base model Qwen3.5-35B-A3B -- and surpassing serveral closed-source agent workflows such as GPT-5 (42.9), Gemini-2.5 Pro (44.8), and Gemini-2.5 Flash (41.3).
Student Perspectives on Traditional Pedagogy Used in Graduate Physics Coursework
arXiv:2607.10955v1 Announce Type: new Abstract: Graduate programs in STEM disciplines are central to preparing future researchers and professors. Program requirements for students often include taking several graduate-level courses. Anecdotal evidence suggests that graduate coursework in physics in particular features outdated and ineffective pedagogical methods, with high emphasis on mathematical rigor in place of conceptual learning or connections made with authentic research. Prior physics education research indicates that students can leave graduate courses with shortcomings in conceptual understanding of covered topics. This pilot study is designed to document student perspectives on their graduate coursework in a single U.S. R1 physics Ph.D. program. A total of 14 semi-structured interviews were conducted with students enrolled in the program, and thematic analysis was conducted on five of these interviews for this paper. The resulting themes are discussed, including the prevalence of traditional passive lecture pedagogy and students placing high value on course content relevant to research.
Dynamical dark energy in the Bianchi Type-V Universe with DESI DR2 BAO, SNIa compilation and RSD measurements
arXiv:2607.09718v1 Announce Type: new Abstract: We investigate the cosmological implications of dynamical dark energy (DDE) models within an anisotropic, spatially homogeneous Bianchi Type-V spacetime framework using a $1+3$ covariant thermodynamics approach. By implementing both constant ($w$) and time-varying ($w_0, w_a$) parameterized equations of state, we evaluate the background expansion history and track linear matter perturbations via the quasi-static approximation. We confront these scenarios with the latest cosmological datasets, including the Dark Energy Spectroscopic Instrument (DESI) DR2 Baryon Acoustic Oscillations (BAO), the Union3 and Dark Energy Survey 5-year (DESY5) Type Ia Supernovae compilations, Cosmic Chronometers (CC), and Redshift-Space Distortion (RSD) measurements. Our joint statistical analyses reveal that the introduction of spatial anisotropy coupled with DDE efficiently accommodates recent late-time measurements and provides a viable mechanism to mitigate the persistent $H_0$ and $S_8$ cosmological tensions. Model selection metrics show that while Akaike criteria strongly support the extended Bianchi Type-V scenarios across most joint data combinations, Bayesian criteria continue to favor the simpler standard $\Lambda$CDM baseline due to its lower dimensionality. Finally, we establish tight constraints on the current matter density parameter $\Omega_{m,0}$, the shear parameter $\Omega_{\sigma,0}$, and the dark energy evolution parameters, confirming that anisotropic extensions remain viable and testable frameworks for modern precision cosmology.
The critical radius of compressible capillary drops: viscosity, thermodynamics, and diffuse-interface scales
arXiv:2607.10609v1 Announce Type: new Abstract: A compressible capillary drop differs from an incompressible one in that its radius is a dynamical degree of freedom. When the drop is sufficiently small, this spherical mode can become unstable, defining a critical radius set by the competition between surface tension and compressibility. This paper examines the robustness and physical meaning of that critical radius. For a non-relativistic viscous fluid, starting from the viscous compressible equations and the free-surface stress condition, we derive the radial dispersion relation and show that shear and bulk viscosities change the eigenvalues but not the onset radius. A thermodynamic argument identifies the same radius as the point where the energy of a uniformly compressed drop changes from locally stable to unstable, explaining why the threshold is not set by viscous dissipation. This energetic interpretation can also be applied to a special-relativistic fluid, where the non-relativistic mass-density factor is replaced by the corresponding enthalpy-density factor. We then compare the critical radius with diffuse-interface length scales in two Cahn--Hilliard free-energy models to determine whether the instability can occur within the range of validity of the sharp-interface description. In a symmetric quartic model the critical radius is much smaller than the interface thickness, so the instability is absent throughout that range. In a shallow-well model, however, the critical radius can become parametrically larger than the interface thickness, leaving a range of sharp-interface drops that are unstable. Whether such a range exists therefore depends on the diffuse-interface free-energy model.
WALA Learning Executable Latent Actions from Action-Labeled Demonstrations and Action-Free Videos
arXiv:2607.11397v1 Announce Type: new Abstract: Generalizable robot policies typically rely on action-labeled robot demonstrations, which are expensive to collect and difficult to scale. In contrast, large-scale human and robot videos contain rich physical interactions but often lack executable robot action labels. We present WALA, a framework for learning executable latent actions from both action-labeled demonstrations and action-free videos. WALA first pretrains a semantic-geometric latent action model from videos by modeling the evolution between current observations and sparsely sampled future observations. Instead of reconstructing raw pixels, WALA predicts future deltas in the DINOv3 feature space and dense depth space, preserving task-relevant semantic and geometric structure while reducing sensitivity to appearance details. During policy training, the pretrained encoder provides stable latent action targets, and the decoder serves as a trainable latent world model. The latent actions generated by the vision-language backbone are jointly supervised by robot action prediction, latent action target matching, and future dynamics prediction. This enables action-labeled demonstrations to provide executable control supervision, while action-free videos contribute dynamics supervision without requiring robot action annotations. Experiments show that WALA achieves strong performance on RoboTwin, sets a new state-of-the-art result on RoboCasa with 75.2% average success, and improves both policy performance and generalization in real-world manipulation tasks.
Representation Learning for Semiparametric Causal Mediation Analysis under No Essential Heterogeneity
arXiv:2607.10540v1 Announce Type: cross Abstract: We propose a two-stage estimator for structural mediation parameters that combines deep representation learning with G-estimation under the "no essential heterogeneity" (NEH) assumption. We call the method UNIT. In the first stage,TARNet estimates the heterogeneous effect of a randomized treatment on a mediator by learning a shared covariate representation across treatment arms.The resulting conditional average treatment effect (CATE) estimate provides a plug-in approximation to the heterogeneity-dependent component of the weight function entering the G-estimating equation of Zheng and Zhou (2015), which identifies the structural parameters even in the presence of unmeasured mediator-outcome confounding. We show that more accurate first-stage representation learning can yield a more informative plug-in weight and thereby improve the precision of the structural parameter estimator. In simulations with non-Gaussian covariates and nonlinear mediator effects, TARNet weights reduce the Stage-2 standard error of the mediation coefficient by a factor of $1.45$ to $1.51$ (median across replications, $n \ge 2000$) relative to the classical approach, at no cost to bias or coverage.
MC-RAG System: A Structure-Driven RAG System for Multi-Constraint Queries
arXiv:2607.10151v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems are widely adopted in question answering, yet they often fail to satisfy complex multi-constraint queries, leading to constraint violations, factual inconsistencies, or hallucinations. We present Structure-Driven RAG System for Multi-Constraint Queries(MC-RAG), a structure-driven RAG system that reformulates retrieval as a subgraph matching problem over a knowledge graph. By integrating semantic and structural embeddings with path-level indexing, MC-RAG performs interpretable, structure-aware, and constraint-consistent retrieval and generation. During the demonstration, participants can input medical or encyclopedic multi-constraint queries, visualize how the system parses constraints, performs structural matching, and generates answers, thereby experiencing an end-to-end, interactive, and explainable RAG pipeline. A demo video is available at https://youtu.be/J8kahzmAnu0.
The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese
arXiv:2607.10745v1 Announce Type: new Abstract: This paper describes the first ChineseBabyLM challenge, which will be held in the 2026 NLPCC conference. The challenge calls for researchers to train language models from scratch with 100 million Chinese tokens and evaluates the models on 3 tracks of tasks: NLU, cognitive alignment and Hanzi knowledge. There is no restriction on tokenizer, model architecture and the number of training epochs. Details of the challenge can be found in https://chinese-babylm.github.io/.
DiffLens: A Visualization System to Explore Local Differences in Graph Sampling
arXiv:2607.11424v1 Announce Type: new Abstract: Graph sampling techniques have been widely used to simplify network computation and visualization, which also results in inevitable differences between the sampled networks and the original networks in terms of nodes, edges and structures. Investigating such differences can inform graph sampling technique users of the pros and cons of different techniques and select the appropriate one, and can also help graph sampling developers evaluate their own technique. However, there are still no systematic ways to achieve such a goal. This paper fills this research gap by first proposing systematic and generic quantitative measures to quantify three categories of graph differences (i.e., neighbor-based, path-based, and structure-based). Built upon this, we further propose DiffLens, a novel visualization system to help graph sampling developers and users intuitively explore local differences at different regions of their interest within a sampled graph, where three new lens-based visual designs are presented to display the neighbor-based, path-based, and structure-based differences respectively. We conducted two case studies and a user study using real-world network datasets to evaluate DiffLens. The results confirmed its effectiveness and usability in helping users explore local differences and compare different graph sampling strategies.
Completeness of Logical Atomicity for Linearizability in Concurrent Separation Logic
arXiv:2607.11435v1 Announce Type: new Abstract: Linearizability is a standard correctness condition for concurrent data structures. It guarantees that operations behave as if they took effect at some atomic instant between their call and return points. Despite the central role linearizability plays, prior work has argued for instead using a style of specification that internalizes the atomicity of operations in terms of the logic's reasoning rules, known as logical atomicity. These logically atomic specifications are intended to be easier to compose inside of the logic than linearizability. Prior work has shown that in the Iris separation logic framework, a certain form of logically atomic specifications implies that a data structure is linearizable. However, the converse remained an open question: for every linearizable data structure, is it always possible to derive a corresponding logically atomic specification? This paper resolves this question in the affirmative. We prove a completeness theorem for Iris that derives a logically atomic specification for any linearizable data structure. As a consequence, we are able to embed a variety of linearizability proof techniques into Iris and use them to derive logically atomic specifications. We apply this to three linearizability proof methods: aspect-oriented linearizability proofs, forward simulations with commit points, and meta-configuration tracking. Using these embeddings, we derive logically atomic specifications for the Herlihy-Wing queue and the Baskets Queue. We furthermore establish a connection between logical atomicity and an encoding of refinement in Iris that has been used in prior logical relations models. This result allows us to transport logically atomic specifications across refinements, which we apply to the Folly MPMC queue implementation. All of the results in this paper have been mechanized in the Rocq Prover.
Event-based Neural Decoding for Neuroprosthetic Motor Control
arXiv:2607.11445v1 Announce Type: new Abstract: A substantial number of patients experience diminished mobility due to disabilities, diseases, or accidents. Although modern prostheses, powered by deep neural networks, hold the promise of significantly enhancing the quality of life for these individuals, their widespread adoption is hindered by significant latency, energy consumption, and spatial requirements. Wired connections to external high-performance processors restrict patient mobility, while wireless connections limit the volume of information that can be transmitted to these processors. Spiking neural networks offer the potential for compressed communication and low-power inference, yet they often lag behind state-of-the-art deep learning models in various applications. In this study, we propose a high-performance neural decoding method that effectively balances task performance and efficiency. An eventbased gated recurrent unit generates a sparse communication pattern with graded spikes, surpassing classical spiking neural networks in terms of task performance. Utilising an efficient training method and sparse inference, our model presents new opportunities for on-device neural decoding.
A Multimodal Dataset for Large Language Model Applications in the Energy Domain
arXiv:2607.11459v1 Announce Type: new Abstract: This paper presents the mAIEnergy dataset, an open-access, multimodal corpus developed to support Large Language Model (LLM) applications in the energy sector. The dataset integrates approximately 50,000 textual documents, 20,000 images, 25 million numerical time series records, and 2 million geospatial and relational data entries. It includes policy and regulatory texts, scientific articles and news articles, satellite and contextual imagery, electricity system measurements, weather observations, statistical indicators, and geospatial representations of energy infrastructure and related entities. All data have been harmonized into structured, ready-to-use formats, accompanied by consistent metadata and reproducible data retrieval and preparation workflows. The dataset can serve as a foundational energy knowledge base, allowing energy stakeholders to integrate additional open-source or proprietary data. The mAIEnergy dataset adheres to Findable, Accessible, Interoperable, and Reusable (FAIR) principles, enhancing its applicability for AI-driven energy research, modeling, and decision-making.