Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

IG-GAN: A Generative Adversarial Network for Aerodynamic Data Generation Based on Intrinsic Geometry
arXiv:2607.11497v1 Announce Type: new Abstract: Existing generative models learn data distributions in flat Euclidean space. However, most data in our real world are manifolds embedded in high dimensional Euclidean space. Therefore, we propose an intrinsic-geometry-based generative adversarial network (IG-GAN) for data generation in the field of aerodynamics. The generator of the IG-GAN represents aerodynamic data as a piecewise smooth manifold constructed by B\'ezier surfaces, and the generator tries to learn the coefficients of each B\'ezier surface to further combine multiple B\'ezier surfaces into a smooth manifold automatically. The discriminator in the IG-GAN is a radial-basis-function based discriminator (RBF-D). Experimental results show that IG-GAN achieves lower predicted Mean Squared Errors (MSEs) than those of three baselines. Specifically, on the Burgers' equation dataset, IG-GAN reduces the predicted MSE of velocity u by 97.41% compared with state of the art SSL-Transformer. Additionally, on the ONERA M6 aircraft dataset, IG-GAN reduces the overall MSE of nine aerodynamic coefficients by 82.95% compared with SSL-Transformer.
TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging
arXiv:2607.09818v1 Announce Type: new Abstract: Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute corresponding actions as embodied agents. Recently, autoregressive token-based action generation has driven the development of many representative VLA models. However, this paradigm often reduces action generation to next-token prediction, thereby lacking explicit modeling of the spatiotemporal structure of action sequences and the disentanglement between vision-language representations and actions, which can limit performance in long-horizon and complex scenarios. In this paper, we propose TS-Mask VLA, a vision-language-action framework for robot manipulation. TS-Mask VLA is built upon two key designs: (1) a Discrete Diffusion Action Expert equipped with a Bridge Attention conditioning bridge, which enables multi-layer conditioning from the VLM and facilitates more accurate and stable action generation; and (2) a temporal-spatial 2D masking strategy for discrete action tokens that strengthens the model's understanding of cross-time dependencies and inter-dimensional coupling, leading to more structurally consistent action sequences. We conduct extensive experiments on simulation benchmarks and real-world tasks. On LIBERO, TS-Mask VLA achieves a 95.7 percent average success rate with only 0.5B parameters, outperforming significantly larger models. On CALVIN, it attains the best average sequence length of 4.19 and strong long-horizon performance. Comprehensive analyses and ablations further validate the effectiveness of our design.
Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization
arXiv:2607.10191v1 Announce Type: new Abstract: Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely. We reveal that this trade-off arises not from the constraints of streaming architectures, but from an inappropriate choice of optimization anchor. Directly optimizing against audio quality metrics induces catastrophic reward hacking, where content critical to pronunciation and intelligibility is systematically erased to maximize a proxy score. To break this bottleneck, we propose two complementary improvements: an enlarged Conformer convolution kernel for richer local spectro-temporal modeling, and WavLM-anchored Direct Preference Optimization (DPO) fine-tuning strategy. DPO preference pairs are ranked by WavLM cosine similarity, a deep acoustic feature encoding both phonetic structure and speaker identity, providing an optimization anchor that resists hacking. Under a 560 ms streaming chunk size, the proposed method achieves a 10.9% relative intelligibility improvement (word error rate: 0.138 to 0.123), with marginal simultaneous gains in audio quality and speaker similarity.
UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters
arXiv:2512.21095v2 Announce Type: replace Abstract: Text and formulas constitute the core informational components of many documents. Accurately and efficiently recognizing both is crucial for developing robust and generalizable document parsing systems. Recently, vision-language models (VLMs) have achieved impressive unified recognition of text and formulas. However, they are large-sized and computationally demanding, restricting their usage in many applications. In this paper, we propose UniRec-0.1B, a unified recognition model with only 0.1B parameters. It is capable of performing text and formula recognition at multiple levels, including characters, words, lines, paragraphs, and documents. To implement this task, we first establish UniRec40M, a large-scale dataset comprises 40 million text, formula and mixed samples, enabling the training of a powerful yet lightweight model. Secondly, we identify two challenges when building such a lightweight but unified expert model. They are: structural variability across levels and semantic entanglement between textual and formulaic content. To tackle these, we introduce a hierarchical supervision training that explicitly guides structural comprehension, and a semantic-decoupled tokenizer that separates text and formula representations. Finally, we develop a comprehensive evaluation benchmark covering Chinese and English documents from multiple domains and with multiple levels. Experimental results on this and public benchmarks demonstrate that UniRec-0.1B outperforms both general-purpose VLMs and leading document parsing expert models, while achieving 2-9x speedup, validating its effectiveness and efficiency. Codebase and Dataset: https://github.com/Topdu/OpenOCR.
Mitigating LLM Sycophancy in Code Smell Detection Using Evidence-Guided Reasoning Prompts
arXiv:2607.10411v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used for code smell detection tasks due to their ability to interpret program semantics. However, their reliability in this context remains poorly explored, particularly under varying prompt conditions where model predictions may be influenced by external cues rather than code characteristics. One such limitation is sycophancy bias, where models tend to align their outputs with user-provided assumptions instead of performing objective analysis. In this paper, we present the first systematic empirical study of sycophancy bias in LLM-based code smell detection. Using the MLCQ dataset, we evaluate how different prompt framings like confirmation bias, contradictory hints, and false premises affect model predictions. Our results show that LLMs are highly sensitive to prompt variations, with Decision Flip Rates reaching up to 72% and False Alignment Rates exceeding 90%, indicating substantial instability and agreement with misleading prompts. To address this issue, we propose Evidence-Guided Debiasing Prompting (EGDP), a structured prompting strategy that enforces evidence-first reasoning. EGDP reduces decision instability and improves robustness, lowering Decision Flip Rates to as low as 12% and False Alignment Rates to as low as 21%, while increasing reliance on structurally grounded evidence. Our findings demonstrate that sycophancy bias poses a critical threat to the reliability of LLM-based code smell detection, and that evidence-guided reasoning provides an effective and generalizable mitigation approach.
HermesHFL: Incentive-Compatible Hierarchical Federated Unlearning for Dynamic LLM Fine-Tuning
arXiv:2607.11528v1 Announce Type: new Abstract: Hierarchical federated unlearning (HFUL) for large language model (LLM) fine-tuning faces significant challenges due to hierarchical aggregation, dynamic client participation, and strong parameter coupling in LLM adaptation. Selectively removing client contributions is particularly difficult because model updates propagate across multiple aggregation stages while unlearning requests may coincide with client departures and rejoining. To address these issues, we propose **HermesHFL**, a hierarchical federated learning framework that supports selective unlearning, dynamic client participation, and client reintegration for scalable LLM fine-tuning via parameter-efficient fine-tuning (PEFT) with LoRA. We formulate a unified optimization problem that jointly models client participation, edge association, incentive allocation, and unlearning under heterogeneous client behaviors. To solve this problem efficiently, we develop **Neogen**, a neural-guided bilevel evolutionary optimization framework that combines CMA-ES for continuous incentive optimization with a CHC-based evolutionary mechanism for discrete participation and association decisions. A neural surrogate further accelerates optimization and improves search efficiency. Extensive experiments on LLM fine-tuning tasks demonstrate that HermesHFL consistently outperforms state-of-the-art baselines in model utility, unlearning effectiveness, convergence stability, and resource efficiency.
Uncovering Students' Mental Models of Generative Artificial Intelligence
arXiv:2607.11692v1 Announce Type: new Abstract: In this paper we present a study of students' mental models of generative AI (GenAI). A student's mental model of GenAI influences not only how they perceive the technology's capabilities and limitations but also how they choose to integrate it into their academic work. Whether they view it as a collaborative partner, a shortcut to complete tasks, or something in between, depends on how they conceptualize its use. This study addresses the following questions: (I) What mental models do undergraduate students hold about GenAI? and (II) What aspects of conceptual knowledge - declarative, procedural, and conditional - are present in these mental models? Sixty-four concept maps were collected from students enrolled in a course on technology ethics. Students were asked to construct concept maps representing their understanding of GenAI use. The concept maps were analyzed using a structured codebook and the analysis revealed five categories of mental models: technical process based, educational tool based, transition model, consequence aware model, integrated model. Declarative knowledge was most dominant across maps, suggesting that students largely understood GenAI primarily at a surface level - knowing its names, tools, and applications but demonstrate limited procedural understanding of how it works and limited conditional knowledge about when and why it should or should not be used. By identifying students' mental models, we can improve students' AI literacy by designing curriculum and guidelines that improve cognition while ensuring responsible and ethical use.
Adaptive Routing for Efficient Diffusion Transformer-Based PNI Prediction
arXiv:2607.11533v1 Announce Type: new Abstract: Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. However, its preoperative prediction from magnetic resonance imaging (MRI) remains challenging due to subtle imaging features that extend beyond tumor boundaries into surrounding regions. Conventional convolutional neural networks are limited in capturing long-range spatial dependencies. Transformer-based architectures improve global modeling of volumetric MRI by aggregating spatially distributed contextual cues, yet capturing subtle and noise-sensitive patterns in peritumoral regions remains challenging. Diffusion-based classifiers offer an alternative formulation by leveraging denoising-based class scoring to better capture such subtle patterns. However, these approaches introduce substantial computational overhead due to the combination of transformer-based modeling and iterative denoising processes. To address these challenges, we formulate PNI prediction as a diffusion-based classification problem and implement the denoising network using a transformer-based representation. To improve computational efficiency, we introduce adaptive routing across attention heads, spatial tokens, and MLP width. Experimental results demonstrate that the proposed approach achieves an AUC of 0.731 with 257.57 GFLOPs.
VecHeart: Holistic Four-Chamber Cardiac Anatomy Modeling via Hybrid VecSets
arXiv:2604.19403v2 Announce Type: replace Abstract: Accurate cardiac anatomy modeling requires the model to be able to handle intricate interrelations among structures. In this paper, we propose VecHeart, a unified framework for holistic reconstruction and generation of four-chamber cardiac structures. To overcome the limitations of current feed-forward implicit methods, specifically their restriction to single-object modeling and their neglect of inter-part correlations, we introduce Hybrid Part Transformer, which leverages part-specific learnable queries and interleaved attention to capture complex inter-chamber dependencies. Furthermore, we propose Anatomical Completion Masking and Modality Alignment strategies, enabling the model to infer complete four-chamber structures from partial, sparse, or noisy observations, even when certain anatomical parts are entirely missing. VecHeart also seamlessly extends to 3D+t dynamic mesh sequence generation, demonstrating exceptional versatility. Experiments show that our method achieves state-of-the-art performance, maintaining high-fidelity reconstruction across diverse challenging scenarios. Code is available at https://github.com/Scalsol/VecHeart.
An efficient method based on the evolutionary center algorithm for optimizing chemical-diffusive models for flame acceleration and DDT
arXiv:2604.19812v2 Announce Type: replace Abstract: This paper presents an efficient method based on Evolutionary Center Algorithm (ECA) for accurately and efficiently determining the optimal reaction and diffusion parameters for Chemical-Diffusive Models (CDM) to simulate flame acceleration (FA) and deflagration-to-detonation transition (DDT). The proposed method leverages the global search capability of the ECA and the local optimization strength of the Nelder-Mead (NM) algorithm. The hybrid approach (ECA-NM) can efficiently optimize CDM parameters that are capable of accurately reproducing the major properties of combustion waves. The CDMs for premixed flames and detonations of hydrogen in air or oxygen were developed using the present ECA-NM method and validated against canonical tests of combustion waves and previous experiments of FA and DDT. The results show that the major flame and detonation properties calculated using the developed CDMs match those obtained from detailed chemical reaction mechanisms over a wide range of equivalence ratio. The simulated FA and DDT in a channel also agree qualitatively and quantitatively with experiments in terms of complex flame instabilities (e.g., tulip and distorted tulip flames), flame displacement speed, and detonation occurrence. In addition, detailed comparisons to the traditional genetic algorithm demonstrate that the developed ECA-NM method diminishes the global error by four orders of magnitude while reducing the computational cost by two orders of magnitude. This work provides a significantly efficient method for developing chemical-diffusive models that allows quantitative multi-scale simulations of transient flames and detonations in complex scenarios.
Optical Lineshape Models and the Generalized Einstein Relation between Absorption and Stimulated Emission
arXiv:2604.22173v2 Announce Type: replace Abstract: Recently, Ryu et al. generalized Einstein's three coefficients for absorption, stimulated emission, and spontaneous emission between two quantum levels to a set of four spectra between two broadened bands. The spectra obey generalized Einstein relationships at thermal equilibrium; Einstein's relations are obtained as an approximation for line spectra. Here, the generalized Einstein relation between absorption and stimulated emission dipole-strength spectra is applied to investigate optical lineshape models. Lineshapes for the Bloch model, the stochastic model, and the semi-classical Brownian oscillator model do not obey the generalized Einstein relation and therefore fail to satisfy detailed balance with Planck blackbody radiation. The quantum Brownian oscillator model treats a harmonic quantum vibration that is bi-linearly coupled to a thermal bath of quantum harmonic oscillators which generate damping and a random force. The two-state quantum Brownian oscillator lineshape model provides lineshapes for transitions between two displaced, but otherwise identical, harmonic potential energy surfaces on which the same quantum vibration is coupled to the same thermal bath of quantum harmonic oscillators. The absorption and stimulated emission lineshapes were calculated using the quantum Brownian oscillator model in under-damped, critically damped, and over-damped cases. The thermal and reorganization energy were each varied from much less to greater than the vibrational quantum of energy. All quantum Brownian oscillator lineshapes obey the generalized Einstein relation within the numerical precision of the calculation (14 to 30 digits), suggesting this lineshape model is compatible with detailed balance. The formula giving the electric-dipole transition cross-section in terms of these lineshapes is presented.
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
arXiv:2604.22226v2 Announce Type: replace Abstract: Sports videos are a challenging domain for multimodal understanding because they involve complex and dynamic human activities. Despite rapid progress in Multimodal Large Language Models (MLLMs), long-horizon reasoning in sports videos remains difficult, as answering questions requires both locating temporally sparse evidence and integrating it into reasoning. We attribute this limitation to two closely coupled factors: insufficient supervision over temporally dispersed evidence, and the lack of methods that require models to identify, localize, and justify temporal evidence. To address these gaps, we introduce SportsTime, a large-scale benchmark for long-form sports video understanding, comprising 14K+ open-ended QA pairs and 50K+ step-wise temporal evidence annotations. Building on SportsTime, we propose Chain-of-Time Reasoning (CoTR), which treats reasoning as a process of temporally grounded evidence composition. Specifically, during training, CoTR introduces a temporal-reward GRPO to encourage temporally grounded reasoning. During inference, it employs an anchor-observe-infer evidence-seeking loop to iteratively localize, verify, and compose temporal evidence before producing the final answer. Experiments demonstrate the usefulness of SportsTime as a benchmark and the effectiveness of CoTR, which consistently improves temporal compositional reasoning and step-wise grounding quality over strong MLLM baselines.
First-Order Softmax Weighted Switching Gradient Method for Distributed Stochastic Minimax Optimization with Stochastic Constraints
arXiv:2603.05774v2 Announce Type: replace Abstract: This paper addresses the distributed stochastic minimax optimization problem subject to stochastic constraints. We propose a novel first-order Softmax-Weighted Switching Gradient method tailored for federated learning. Under full client participation, our algorithm achieves the standard $\tilde{\mathcal{O}}(\epsilon^{-4})$ oracle complexity to satisfy a unified bound $\epsilon$ for both the optimality gap and feasibility tolerance. We extend our theoretical analysis to the practical partial participation regime by quantifying client sampling noise through a stochastic superiority assumption. Furthermore, by relaxing standard boundedness assumptions on the objective functions, we establish a strictly tighter lower bound for the softmax hyperparameter. We provide a unified error decomposition and establish a sharp $\mathcal{O}(\log\frac{1}{\delta})$ high-probability convergence guarantee. Ultimately, our framework demonstrates that a single-loop primal-only switching mechanism provides a stable alternative for optimizing worst-case client performance, effectively bypassing the hyperparameter sensitivity and convergence oscillations often encountered in traditional primal-dual or penalty-based approaches. We verify the efficacy of our algorithm via experiment on the Neyman-Pearson (NP) classification, fair classification, and federated safe reinforcement learning tasks.
Can we Trust Unreliable Voxels? Exploring 3D Semantic Occupancy Prediction under Label Noise
arXiv:2603.06279v2 Announce Type: replace Abstract: 3D semantic occupancy prediction is a cornerstone of robotic perception, yet real-world voxel annotations are inherently corrupted by structural artifacts and dynamic trailing effects. This raises a critical but underexplored question: can autonomous systems safely rely on such unreliable occupancy supervision? To systematically investigate this issue, we establish OccNL, the first benchmark dedicated to 3D occupancy under occupancy-asymmetric and dynamic trailing noise. Our analysis reveals a fundamental domain gap: state-of-the-art 2D label noise learning strategies collapse catastrophically in sparse 3D voxel spaces, exposing a critical vulnerability in existing paradigms. To address this challenge, we propose DPR-Occ, a principled label-noise-robust framework that constructs reliable supervision through dual-source partial label reasoning. By synergizing temporal model memory with representation-level structural affinity, DPR-Occ dynamically expands and prunes candidate label sets to preserve true semantics while suppressing noise propagation. Extensive experiments on SemanticKITTI demonstrate that DPR-Occ prevents geometric and semantic collapse under extreme corruption. Notably, even at 90% label noise, our method achieves significant performance gains (up to 2.57% mIoU and 13.91% IoU) over existing label noise learning baselines adapted to the 3D occupancy prediction task. By bridging label noise learning and 3D perception, OccNL and DPR-Occ provide a reliable foundation for safety-critical robotic perception in dynamic environments. The benchmark and source code will be made publicly available at https://github.com/mylwx/OccNL.
Local Message-Passing for Discrete Graph Generation
arXiv:2603.08825v2 Announce Type: replace Abstract: Discrete graph generation has emerged as a powerful paradigm for modeling graph-structured data, yet state of the art models often rely on Graph Transformers or higher order architectures. We revisit this design assumption by introducing GenGNN, a modular message passing backbone for graph generation. GenGNN enables powerful generation by persisting edge fields through latent refinement of coupled node edge graph states, all without requiring global attention. Diffusion models integrating GenGNN achieve over 90 percent validity on standard benchmark datasets, performing within margins of Graph Transformer backbones and achieving up to 2x or even 5x faster inference. Systematic ablations isolate how GenGNN is resilient to oversmoothing during generative denoising, indicating each GenGNN component is necessary for downstream generation quality. Finally, representation-space analysis suggests GenGNN learns functionally similar representations to more theoretically-expressive architectures; even at deeper layers. As such, GenGNN uplifts local message-passing to challenge prevailing assumptions that performant discrete graph generation requires global attention or higher-order representations. Source Code Available Here
Philosopher and Prophet Inequalities for Divisible Items
arXiv:2607.11742v1 Announce Type: new Abstract: We study online welfare maximization with divisible resources. A sequence of $n$ players arrive one by one; upon arrival, each player draws a valuation function over $m$ divisible items from a known distribution, reveals this valuation, and must be allocated an irrevocable fractional bundle subject to unit supply constraints. While online welfare maximization has been extensively studied for indivisible items and combinatorial valuations, much less is known when the resources are divisible and players have multi-dimensional concave valuations. We give approximation algorithms for monotone concave valuations satisfying diminishing returns. Our main result is a $2/3$-approximation to the optimal online policy, also known as the philosopher benchmark. The algorithm is guided by a low-dimensional concave relaxation of the online benchmark and rounds it via a new single-item capped online contention resolution scheme. This Capped-OCRS problem allocates to each realized type no more than its prescribed fractional bundle while preserving a $2/3$-fraction of that bundle in expectation. Its analysis uses a submartingale potential for the remaining side, we show that computing the optimal online policy is #P-hard even for a single divisible item. We also obtain a tight prophet inequality against the offline hindsight optimum. We show that a fixed-price auction with one linear per-unit price for each original divisible item achieves a $1/2$-approximation to the offline/prophet benchmark. The prices are obtained by aggregating Aumann--Shapley supporting prices, a continuous analogue of supporting prices for submodular/XOS set functions, and yield simple item prices rather than copy-dependent prices arising from discretization. The factor $1/2$ for the prophet benchmark is information-theoretically tight even for one item with linear valuations.
HiFi-LLP: High-Fidelity, Low-Cost Latency Predictors with Confidence for Robust HW-NAS
arXiv:2607.11746v1 Announce Type: new Abstract: With deep neural networks (DNNs) increasingly deployed on edge devices, hardware (HW)-aware optimization techniques--such as HW-aware compression and HW-aware neural architecture search (HW-NAS)--have become essential. These methods rely on real feedback from the target hardware to tailor DNN architectures for efficient deployment. While the search can be parallelized, latency measurements via hardware-in-the-loop (HIL) remain a bottleneck due to their sequential nature. Recent approaches use latency predictors to replace costly HIL feedback, but challenges persist: (1) platform-specific predictors often require tens of thousands of samples, and (2) inaccurate predictions can mislead the NAS process. To address this, we introduce HiFi-LLP, a high-fidelity, low-cost latency predictor based on graph attention networks, augmented with a confidence metric. HiFi-LLP outperforms prior platform-specific predictors by up to 9 percentage points (p.p.) in the 10% accuracy bound and achieves a Spearman's rank correlation of up to 0.996 across six devices in the LatBench dataset. We further propose a hybrid NAS framework that routes low-confidence predictions to HIL, achieving up to 8.6$\times$ speedup compared to typical NAS while maintaining a competitive Pareto front.
A Fokker-Planck approach to a stochastic multiplicative wealth model with taxation and redistribution
arXiv:2607.11755v1 Announce Type: cross Abstract: We develop a Fokker-Planck description of the dynamics of wealth distribution in a stochastic multiplicative economic growth model with taxation and redistribution, as introduced by P.M.C. de Oliveira. Extending the original formulation, our theoretical framework includes general redistribution protocols, encompassing a broad class of state-dependent transfer mechanisms. As a particular case, we investigate a two-state protocol designed to emulate conditional cash transfer programs. Analytical expressions for the stationary wealth distributions are derived, revealing how the interplay between multiplicative noise, taxation, and redistribution shapes the emergence of inequality. The theoretical results are corroborated by agent-based simulations. To quantify and compare the impact of the different protocols, we employ the Gini index as a measure of inequality. Our analysis highlights how specific nonuniform redistribution schemes can significantly mitigate wealth disparities.
Sensor operating point calibration and monitoring of the ALICE Inner Tracking System during LHC Run 3
arXiv:2510.27592v3 Announce Type: replace Abstract: The new Inner Tracking System (ITS2) of the ALICE experiment began operation in 2021 with the start of LHC Run 3. Compared to its predecessor, ITS2 offers substantial improvements in pointing resolution, tracking efficiency at low transverse momenta, and readout-rate capabilities. The detector employs silicon Monolithic Active Pixel Sensors (MAPS) featuring a pixel size of 26.88$\times$29.24 $\mu$m$^2$ and an intrinsic spatial resolution of approximately 5 $\mu$m. With a remarkably low material budget of 0.36% of radiation length ($X_{0}$) per layer in the three innermost layers and a total sensitive area of about 10 m$^2$, the ITS2 constitutes the largest-scale application of MAPS technology in a high-energy physics experiment and the first of its kind operated at the LHC. For stable data taking, it is crucial to calibrate different parameters of the detector, such as in-pixel charge thresholds and the masking of noisy pixels. The calibration of 24120 monolithic sensors, comprising a total of 12.6$\times$10$^{9}$ pixels, represents a major operational challenge. This paper presents the methods developed for the calibration of the ITS2 and outlines the strategies for monitoring and dynamically adjusting the detector's key performance parameters over time.
A Polynomial-Time Algorithm for the Next-to-Shortest Path Problem on Positively Weighted Directed Graphs
arXiv:2511.04345v2 Announce Type: replace Abstract: Given a graph and a pair of terminals $s$, $t$, the next-to-shortest path problem asks for an $s\!\to \!t$ (simple) path that is shortest among all not shortest $s\!\to \!t$ paths (if one exists). This problem was introduced in 1996, and soon after was shown to be NP-complete for directed graphs with non-negative edge weights, leaving open the case of positive edge weights. Subsequent work investigated this open question, and developed polynomial-time algorithms for the cases of undirected graphs and planar directed graphs. In this work, we resolve this nearly 30-year-old open problem by providing an algorithm for the next-to-shortest path problem on directed graphs with positive edge weights.
Learning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning
arXiv:2604.22770v2 Announce Type: replace Abstract: Most digital language learning curricula rely on discrete-item quizzes that test recall rather than applied conversational proficiency. When progression is driven by quiz performance, learners can advance despite persistent gaps in using grammar and vocabulary during interaction. Recent work on LLM-based judging suggests a path toward scoring open-ended conversations, but using interaction evidence to drive progression and review requires scoring protocols that are reliable and validated. We introduce Learning in Blocks, a framework that grounds progression in demonstrated conversational competence evaluated using CEFR-aligned rubrics. The framework employs heterogeneous multi-agent debate (HeteroMAD) in two stages: a scoring stage where role-specialized agents independently evaluate Grammar, Vocabulary, and Interactive Communication, engage in debate to address conflicting judgments, and a judge synthesizes consensus scores; and a recommendation stage that identifies specific grammar skills and vocabulary topics for targeted review. Progression requires demonstrating 70% mastery, and spaced review targets identified weaknesses to counter skill decay. We benchmark four scoring and recommendation methods on CEFR A2 conversations annotated by ESL experts. HeteroMAD achieves a superior score agreement with a 0.23 degree of variation and recommendation acceptability of 90.91%. An 8-week study with 180 CEFR A2 learners demonstrates that combining rubric-aligned scoring and recommendation with spaced review and mastery-based progression produces better learning outcomes than feedback alone.
PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging
arXiv:2604.22823v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementary cross-modal alignment capabilities. Model merging provides a cost-effective mechanism for integrating multiple expert MLLMs with complementary strengths into a unified model. However, existing model merging research mainly focuses on post-finetuning scenarios, leaving the pre-training stage largely unexplored. We argue that the core of MLLM pre-training lies in establishing effective cross-modal alignment, which bridges visual and textual representations into a unified semantic space. Motivated by this insight, we introduce the post-alignment merging task, which aims to integrate cross-modal alignment capabilities learned from heterogeneous multimodal pre-training. This setting introduces two key challenges: cross-domain parameter interference, where parameter updates learned from different data distributions conflict during merging, and layer-wise alignment contribution disparity, where different layers and projectors contribute unevenly to cross-modal alignment. To address them, we propose \textbf{PivotMerge}, a post-alignment merging framework for cross-modal projectors. PivotMerge incorporates two key components: Shared-space Decomposition and Filtering, which disentangles shared alignment patterns from domain-specific variations and suppresses conflicting directions, and Alignment-guided Layer-wise Merging, which assigns layer-specific merging weights based on differing alignment contributions. We construct systematic CC12M-based post-alignment merging scenarios for evaluation. Extensive experiments on multiple multimodal benchmarks show that PivotMerge consistently outperforms existing baselines, demonstrating its effectiveness and generalization ability.
Remotely programming the weights of a spintronic neural network by a radiofrequency broadcast signal
arXiv:2604.24561v4 Announce Type: replace Abstract: Selectively programming large number of non-volatile synaptic weights without compromising scalability is a key challenge for in-memory computing. Here, we demonstrate remote programming of synaptic weights in series-connected chains of 11 vortex-based magnetic tunnel junctions using broadcast radiofrequency signals applied through a shared strip line. The programming relies on frequency-selective reversal of the vortex-core polarity and therefore does not require individual access lines or selector devices. By reconfiguring the binary states of these chains, we reshape the weighted sums they perform on frequency-multiplexed RF inputs. Using a 22-synapse network composed of two such chains, we remotely reconfigure the same hardware to perform two distinct tasks: handwritten-digit classification and drone RF-signature identification. The digit-optimized configuration reaches 94.91 +/- 0.26% accuracy on handwritten digits but only 13.17 +/- 0.47% on drone RF signatures, whereas the drone-optimized configuration reaches 97.33 +/- 0.62% on drones but only 47.59 +/- 1.5% on digits. Broadcast RF programming thus provides a compact and scalable route to rapidly reconfigurable spintronic neuromorphic hardware.
Research Team Identification Based on Representation Learning of Academic Heterogeneous Information Network
arXiv:2311.00922v2 Announce Type: replace Abstract: Academic networks in the real world can usually be described by heterogeneous information networks composed of multiple types of nodes and relationships. Existing representation-learning research for homogeneous information networks lacks the ability to explore the heterogeneity of such networks and therefore cannot be directly applied to heterogeneous information networks. To meet the practical need to identify and discover scientific research teams from academic heterogeneous information networks composed of massive and complex scientific and technological data, this paper proposes a research-team identification method based on representation learning. Node-level and meta-path-level attention mechanisms learn low-dimensional, dense, real-valued vector representations while retaining rich topological information and meta-path semantics. Scientific research teams and important team members are then identified by maximizing node influence. Experimental results show that the proposed method outperforms the comparison methods.
HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA
arXiv:2402.01767v3 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) has rapidly advanced the language model field, particularly in question-answering (QA) systems. By integrating external documents during the response generation phase, RAG significantly enhances the accuracy and reliability of language models. This method elevates the quality of responses and reduces the frequency of hallucinations, where the model generates incorrect or misleading information. However, these methods exhibit limited retrieval accuracy when faced with numerous indistinguishable documents, presenting notable challenges in their practical application. In response to these emerging challenges, we present HiQA, an advanced multi-document question-answering (MDQA) framework that integrates cascading metadata into content and a multi-route retrieval mechanism. We also release a benchmark called MasQA to evaluate and research in MDQA. Finally, HiQA demonstrates the state-of-the-art performance in multi-document environments.