Forskningsradar

Science Journals

Peer-reviewade publikationer — 56237 artiklar

Lasing from a Quantum-Dot-Like Buried Heterostructure in an InP Nanobeam Cavity
arXiv:2603.16636v2 Announce Type: replace Abstract: We report lasing from a lithographically defined buried heterostructure with an estimated lateral footprint of (107 nm)^2, embedded in an InP photonic-crystal nanobeam cavity. This represents the smallest laterally confined buried heterostructure gain region from which lasing has been observed. Despite etching of the active region during cavity definition and the associated risk of surface-related nonradiative recombination, optically pumped devices exhibit a clear lasing threshold and a narrow linewidth. By systematically varying the BH size, we investigate how the lasing threshold depends on the active volume under optical pumping. The estimated intrinsic threshold under ideal carrier injection is 57 nW, comparable to values reported for single quantum-dot nanolasers, highlighting the potential of quantum-dot-scale buried heterostructures as deterministic, scalable gain media for nanophotonic lasers.
Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens
arXiv:2606.16847v3 Announce Type: replace Abstract: Diffusion Large Language Models (dLLMs) offer a promising avenue for parallel generation but face a trade-off between decoding speed and quality. While revocable decoding strategies attempt to mitigate errors by verifying and remasking tokens, they typically operate within a mixed-quality context. This leads to two critical failures: \textit{Error Propagation}, where new tokens absorb toxic information from erroneous context, and \textit{Local Error Reinforcement}, where errors mutually reinforce each other to evade detection. To alleviate these challenges, we propose ASRD (Anchor Supervised Revocable Decoding), a training-free framework that operates within the embedding space. ASRD explicitly decouples the decoding context into trusted \textit{Anchor Tokens}, which are identified via temporal consistency, and uncertain candidates. Leveraging a dynamic Anchor Tokens Cache, we introduce two complementary mechanisms: (1) Anchor-Guided Generation, which injects entropy-weighted anchor signals into masked positions to implicitly rectify attention toward the reliable global skeleton; and (2) Anchor-Perturbed Verification, which applies orthogonal perturbations to uncertain candidate tokens, destabilizing and remasking errors driven by fragile local consensus. Extensive experiments on math and coding benchmarks demonstrate that ASRD outperforms recent remasking baselines, achieving accuracy improvements of up to 6.4\% while accelerating inference throughput by up to 7.2$\times$.
LapSurgie: Humanoid Robots Performing Surgery via Teleoperated Handheld Laparoscopy
arXiv:2510.03529v3 Announce Type: replace Abstract: Robotic laparoscopic surgery has gained increasing attention in recent years for its potential to deliver more efficient and precise minimally invasive procedures. However, adoption of surgical robotic platforms remains largely confined to high-resource medical centers, exacerbating healthcare disparities in rural and low-resource regions. To close this gap, a range of solutions has been explored, from remote mentorship to fully remote telesurgery. Yet, the practical deployment of surgical robotic systems to underserved communities remains an unsolved challenge. Humanoid systems offer a promising path toward deployability, as they can directly operate in environments designed for humans without extensive infrastructure modifications -- including operating rooms. In this work, we introduce LapSurgie, the first humanoid-robot-based laparoscopic teleoperation framework. The system leverages an inverse-mapping strategy for manual-wristed laparoscopic instruments that abides to remote center-of-motion constraints, enabling precise hand-to-tool control of off-the-shelf surgical laparoscopic tools without additional setup requirements. A control console equipped with a stereo vision system provides real-time visual feedback. Finally, a comprehensive user study across platforms demonstrates the effectiveness of the proposed framework and provides initial evidence for the feasibility of deploying humanoid robots in laparoscopic procedures.
Stochastic Finite Volume Approximation with Clustering in the Parameter Space for the Forward Uncertainty Quantification of Differential Equations with Random Parameters
arXiv:2510.12109v2 Announce Type: replace Abstract: The uncertainty quantification (UQ) for mathematical models with random parameters is important for many science and engineering problems. Forward UQ quantifies the impact of random parameters on the output of system. In the current study, we propose a new stochastic finite volume (SFV) scheme by combining SFV with clustering algorithm in the parameter space such that each cluster can be regarded as a finite volume with implicit boundaries. The advantage of SFV is that no specific form of the random variable is required such that discontinuous solutions and sharp interfaces can be accurately approximated. Compared to classic SFV based on structured grids, the new SFV-cluster scheme extends SFV to parameter spaces of higher dimensions. For demonstration and validation, we present the construction of SFV schemes for the Kraichnan-Orszag three-mode problem and the Buckley-Leverett equation. The error analysis of SFV is extended and the new algorithm is validated by numerical experiments.
SWARM+: Scalable and Resilient Multi-Agent Consensus for Decentralized Data-Aware Workload Management
arXiv:2603.19431v4 Announce Type: replace Abstract: Distributed scientific workflows are increasingly executed across heterogeneous and geo-distributed computing environments, where centralized workload orchestration becomes a scalability and resilience bottleneck. This paper presents SWARM+, a decentralized workload management system that coordinates workload placement through hierarchical multi-agent consensus, reducing coordination overhead and dramatically improving scalability, while tolerating failures and dynamic membership changes. SWARM+ enables data-aware scheduling policies that incorporate resource availability, data transfer node (DTN) connectivity, and data locality into workload placement decisions. We evaluate SWARM+ on the distributed FABRIC testbed using heterogeneous scientific workloads derived from production workflow traces obtained from the Pegasus Workflow Management System (WMS). Experimental results show that SWARM+ scales coordination to 990 distributed agents with approximately 1\,s per-job selection time at 110 agents. SWARM+ demonstrates balanced workload distribution, maintains over 97% job completion under distributed failures with graceful degradation (mean ~95% job completion) during correlated site outages, tolerates coordinator agent failures gracefully, improves schedule quality by employing data-aware policies, and reduces both selection time and scheduling latency by 97-98% when compared to the prior SWARM system.
ParaTutor: Coordinating Parent and Child Math Tutoring through Role Separated LLM Scaffolding
arXiv:2606.18030v2 Announce Type: replace Abstract: Parent and child tutoring is a collaborative learning setting with asymmetric roles. Parents guide children s problem solving, while children are expected to remain actively engaged in understanding and reasoning. However, most LLM based learning systems are designed for single users or relatively symmetric collaboration, leaving parent and child tutoring with distinct instructional roles underexplored. Through a formative study, we found that parent and child math tutoring was often disrupted by cognitive misalignment, emotional escalation, and method mismatch. To address these challenges, we present ParaTutor, a multiple agents LLM based scaffolding system for home math word problem tutoring. ParaTutor distributes support across user roles by providing parents with strategy, language, repair, and phase scaffolds, while providing children with visual grounding for problem interpretation. We evaluated ParaTutor with 23 parent and child dyads (children aged 10 to 12) across four tutoring conditions that varied how LLM assistance was delivered. Results show that generic LLM assistance often provided useful explanations but did not consistently support parent led tutoring or children s active reasoning. In contrast, ParaTutor helped redistribute tutoring work across parents and children, increased children s engagement with word problems, supported shared understanding through visual grounding, and helped parents translate LLM generated methods into child facing tutoring moves. These findings suggest that in family learning, the value of LLM support depends not only on model capability, but also on how support is coordinated across users with different roles. Our work contributes design implications for LLM systems that support role sensitive scaffolding in parent and child learning.
Investigating students' gender expression and its relation to sense of belonging in introductory physics courses
arXiv:2603.17966v2 Announce Type: replace Abstract: Despite nation efforts to promote diversity and inclusion, women and gender minorities remain underrepresented in physics. One common approach to studying gender in physics contexts treats gender as a categorical identity variable (e.g. "man," "woman"). In contrast, approaches that center gender expression focus on the nuanced and context-dependent ways in which gender is socially enacted and interpreted. They are therefore well-suited for exploring how gender permeates the small-scale interactions that ultimately shape students' persistence and perceptions of inclusivity. In the present study, we utilized gender expression as a lens to investigate gendered patterns in introductory undergraduate physics students' sense of belonging in the discipline. Specifically, this qualitative investigation expands on our previous quantitative work to investigate why students may feel misperceived by their peers in physics and how that experience influences their belonging. Results indicated that students' sense of belonging may be impacted by perceived pressures to alter their gender expression in physics contexts. Many interviewees expressed a felt need to present themselves more "masculinely" to fit in. Contrastingly, pressure to present "femininely" was most often associated with standing out. Implications for supporting students' authentic self-ex in physics contexts are discussed.
On splitting strategies for the numerical solution of stochastic delay differential equations with correlated noises
arXiv:2603.21858v2 Announce Type: replace Abstract: In this article we investigate the numerical solution of a scalar semilinear stochastic delay differential equation (SDDE) where the linear instantaneous feedback and nonlinear delayed feedback terms are perturbed by a pair of standard Brownian motions with correlation $\rho$. Such SDDEs may be naturally decomposed into two subsystems: a linear stochastic differential equation (SDE) without delay, and a nonlinear SDDE. Splitting methods work by solving each subsystem separately and composing the results over a single step. Our main theoretical result provides a bound on the mean-square error of a particular strategy for doing this, known as Lie-Trotter splitting. This bound implies that the method is mean-square strongly convergent with order $1/2$ when $\rho=0$, so that the noises are uncorrelated, but assurances of convergence are lost when $\rho\neq 0$. Indeed we develop an upper bound on the global mean-square error with a term that depends linearly on the magnitude of the correlation, and is independent of the stepsize. While our theoretical error bound is an estimate from above, we conduct numerical experiments that confirm the order of mean-square strong convergence of Lie-Trotter splitting in the $\rho=0$ case, and demonstrate a rapid fall-off to effectively zero as $|\rho|$ increases. Similar numerical results are observed for an alternative commonly used strategy known as Strang splitting. Nonetheless, by carefully reorganising the subsystems into which we split the SDDE, we can improve the range of values of $\rho$ over which a nonzero order of convergence is observed numerically.
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
arXiv:2603.25112v2 Announce Type: replace Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy) with how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity). We apply Signal Detection Theory (SDT) to decompose these capacities, treating token-level normalised log-probability as a graded confidence variable and answer correctness as the state to be discriminated. We characterise the Type-2 ROC of this signal, including its unequal-variance structure via z-ROC analysis, and -- because the meta-d' efficiency ratio is not well defined for open-ended QA, which lacks a two-alternative Type-1 decision -- quantify metacognitive efficiency with a model-free information measure, normalised metacognitive information (meta-I_2r). Applied to four LLMs (Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.3, Llama-3-8B-Base, Gemma-2-9B-Instruct) across 224,000 factual QA trials, we find: (1) metacognitive information varies more than two-fold across models and co-varies inversely with accuracy -- the least accurate model has the most informative confidence -- though with four models this ordering cannot be separated from an error-difficulty confound, so we report it as coupling, not decoupling; (2) the confidence signal has model-specific unequal-variance structure (z-ROC slopes 0.81 to 1.18) invisible to calibration metrics; (3) metacognitive information is domain-specific, strongest in Arts & Literature for every model; (4) temperature dissociates Type-1 accuracy from metacognitive information, which stays stable while accuracy shifts. All estimates carry permutation nulls and bootstrap confidence intervals. Pre-registered; code and data public.
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
arXiv:2603.27706v2 Announce Type: replace Abstract: Reference Audio-Visual Segmentation (Ref-AVS) aims to segment objects in audible videos based on multimodal cues in reference expressions. Previous methods overlook the explicit recognition of expression difficulty and dominant modality in multimodal cues, over-rely on the quality of the instruction-tuning dataset for object reasoning, and lack reflective validation of segmentation results, leading to erroneous mask predictions. To address these issues, in this paper, we propose a novel training-free Multi-Agent Recognition, Reasoning, and Reflection framework to achieve high-quality Reference Audio-Visual Segmentation, termed MAR3. Incorporating the sociological Delphi theory to achieve robust analysis, a Consensus Multimodal Recognition mechanism is proposed that enables LLM agents to explicitly recognize the difficulty of reference expressions and the dominant modality of multimodal cues. Based on our modality-dominant difficulty rule, we propose an adaptive Collaborative Object Reasoning strategy to reliably reason about the referred object. To further ensure precise mask prediction, we develop a Reflective Learning Segmentation mechanism, in which a check agent examines intermediate segmentation results and iteratively corrects the object text prompt of the segment agent. Experiments demonstrate that MAR3 achieves superior performance (69.2% in J&F) on the Ref-AVSBench dataset, outperforming SOTA by 3.4% absolutely.
Continuous Policy and Value Iteration for Stochastic Control Problems and Its Convergence
arXiv:2506.08121v2 Announce Type: replace-cross Abstract: We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This framework applies to both the entropy-regularized relaxed control problems and the classical control problems, with infinite horizon. We establish policy improvement and demonstrate convergence to the optimal control under the monotonicity condition of the Hamiltonian. By utilizing Langevin-type stochastic differential equations for continuous updates along the policy iteration direction, our approach enables the use of distribution sampling and non-convex learning techniques in machine learning to optimize the value function and identify the optimal control simultaneously.
Beyond Logit Adjustment: A Residual Decomposition Framework for Long-Tailed Reranking
arXiv:2604.01506v2 Announce Type: replace Abstract: Long-tailed classification, where a small number of frequent classes dominate many rare ones, remains challenging because models systematically favor frequent classes at inference time. Existing post-hoc methods such as logit adjustment address this by adding a fixed classwise offset to the base-model logits. However, the correction required to restore the relative ranking of two classes need not be constant across inputs, and a fixed offset cannot adapt to such variation. We study this problem through Bayes-optimal reranking on a base-model top-k shortlist. The gap between the optimal score and the base score, the residual correction, decomposes into a classwise component that is constant within each class, and a pairwise component that depends on the input and competing labels. When the residual is purely classwise, a fixed offset suffices to recover the Bayes-optimal ordering. We further show that when the same label pair induces incompatible ordering constraints across contexts, no fixed offset can achieve this recovery. This decomposition leads to testable predictions regarding when pairwise correction can improve performance and when cannot. We develop REPAIR (Reranking via Pairwise residual correction), a lightweight post-hoc reranker that combines a shrinkage-stabilized classwise term with a linear pairwise term driven by competition features on the shortlist. Experiments on nine benchmarks confirm that the decomposition explains where pairwise correction helps and where classwise correction alone suffices. These span text classification, visual recognition, and multimodal rare-disease diagnosis.
The Darkside-20k Data Acquisition System
arXiv:2604.03059v3 Announce Type: replace Abstract: DarkSide-20k is a WIMP search experiment using liquid argon as a target, designed to perform a background-free search for dark matter with unprecedented sensitivity, and is currently under construction at INFN Laboratori Nazionali del Gran Sasso, Italy. The detector comprises a dual-phase Time Projection Chamber complemented with external veto systems and is equipped with a total of 2720 SiPM-based readout channels. This work presents the DAQ system designed for DarkSide-20k. The system is capable of continuous, triggerless digitisation of the waveforms with high single-photoelectron detection efficiency and online processing, ensuring data reduction for long-term storage. The DarkSide-20k DAQ system employs commercial CAEN VX2745 digitisers with custom FPGA firmware implementation. Timing and synchronisation across all 48 digitisers are provided by custom Global and Crate Data Manager boards distributing a phase-aligned clock derived from a disciplined rubidium standard. Waveform segments are processed in real time by Front End Processor machines. Data are organised into collections containing whole detector information and distributed across a farm of Time Slice Processors for event reconstruction, classification, and further reduction before storage and offline analysis. A full "Quadrant" of the system, corresponding to one quarter of the final DAQ, has been assembled and validated at TRIUMF laboratory in Canada. The Quadrant has been stress-tested with simultaneous pulses and demonstrated sustained digitizer readout exceeding expected physics rates and stable long-term performance.
Understanding LoRA as Knowledge Memory: An Empirical Analysis
arXiv:2603.01097v4 Announce Type: replace Abstract: Continuous knowledge updating for pre-trained large language models (LLMs) is increasingly necessary yet remains challenging. Although inference-time methods like In-Context Learning (ICL) and Retrieval-Augmented Generation (RAG) are popular, they face constraints in context budgets, costs, and retrieval fragmentation. Departing from these context-dependent paradigms, this work investigates a parametric approach using Low-Rank Adaptation (LoRA) as a modular knowledge memory. Although few recent works examine this concept, the fundamental mechanics governing its capacity and composability remain largely unexplored. We bridge this gap through the first systematic empirical study mapping the design space of LoRA-based memory, ranging from characterizing storage capacity and optimizing internalization to scaling multi-module systems and evaluating long-context reasoning. Rather than proposing a single architecture, we provide practical guidance on the operational boundaries of LoRA memory. Overall, our findings position LoRA as the complementary axis of memory alongside RAG and ICL, offering distinct advantages.
LogiDroid: Individual Functional Test Generation via Business Logic Extraction and Adaptation
arXiv:2602.24108v2 Announce Type: replace Abstract: Functional testing is essential for verifying that the business logic of mobile applications aligns with user requirements. Despite its importance, functional testing remains heavily dependent on manual effort due to two core challenges. First, acquiring and reusing business logic from unstructured requirements remains difficult, which hinders the understanding of specific functionalities. Second, a significant semantic gap exists when adapting business logic to the diverse GUI environments, which hinders the generation of test cases for specific mobile applications. To address the preceding challenges, we propose LogiDroid, a two-stage approach that generates individual functional test cases by extracting business logic and adapting it to target applications. First, in the Knowledge Retrieval and Fusion stage, two LLM-based agents are employed to construct a functional test dataset, retrieve relevant test cases, and extract structured business logic for the target functionality. Second, in the Context-Aware Test Generation stage, two other LLM-based agents jointly analyze the extracted business logic and the real time GUI environment to incrementally generate context adaptive functional test cases. This design allows LogiDroid to accurately understand application semantics and use domain expertise to generate complete test cases with verification assertions. We assess the effectiveness of LogiDroid using two widely-used datasets that cover 28 real-world applications and 190 functional requirements. Experimental results show that LogiDroid successfully tested 40% of functional requirements on the FrUITeR dataset (an improvement of over 25% compared to the state-of-the-art approaches) and 65% on the Lin dataset (an improvement of over 55% compared to the state-of-the-art approaches). These results demonstrate the significant effectiveness of LogiDroid in functional test generation.
Mistake gating leads to energy and memory efficient continual learning
arXiv:2604.14336v2 Announce Type: replace Abstract: Synaptic plasticity is metabolically expensive, yet animals continuously update their internal models without exhausting energy reserves. However, when artificial neural networks are trained, the network parameters are typically updated on every sample that is presented, even if the sample was classified correctly. Inspired by the human negativity bias and error-related negativity, we propose 'memorized mistake-gated learning' -- a biologically plausible plasticity rule where synaptic updates are strictly gated by current and past classification errors. This reduces the number of updates the network needs to make by $50\%\sim80\%$. Mistake gating is particularly well suited in two cases: 1) For incremental learning where new knowledge is acquired on a background of pre-existing knowledge, 2) For online learning scenarios when data needs to be stored for later replay, as mistake-gating reduces storage buffer requirements. The algorithm can be implemented in a few lines of code, adds no hyper-parameters, and comes at negligible computational overhead. Learning on mistakes is an energy efficient and biologically relevant modification to commonly used learning rules that is well suited for continual learning.
A Practical Guide to PID Controller Implementation
arXiv:2604.15918v3 Announce Type: replace Abstract: How difficult can it be to implement a PID controller? The answer is twofold. Implementing the PID control law is simple and computationally inexpensive. However, this basic form will not work in practical applications. The primary reason for this is the various physical limitations of the actuator. Measurement noise, different implementations depending on the various structures (P, PI, PD or PID), bumpless transfer, and varying sampling interval also result in problems rendering the basic form inoperable. PID implementation is therefore more difficult than meets the eye. This paper introduces a reference implementation of the PID controller which considers these practical issues. It includes pseudo-code, discussion of the implementation choices and simulation of carefully selected, important test cases.
MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models
arXiv:2607.01117v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) have shown strong progress in video understanding, yet they still suffer from hallucinations that are inconsistent with visual evidence. Existing benchmarks mainly focus on object hallucination or coarse action perception, leaving a key video-specific problem underexplored: motion hallucination, in which models infer human motions that are absent from the video. We present MoHallBench, a benchmark for diagnosing motion hallucination in VideoLLMs. MoHallBench systematically evaluates three major sources of hallucination: co-occurrence priors, sequential inference, and similarity confusion. It contains 11,306 video clips and 40,493 question-answer pairs, covering binary-choice, multiple-choice, and generative settings. We further introduce a bi-directional questioning protocol with bias-aware metrics to reduce affirmation bias in binary evaluation. Experiments on ten recent open-source VideoLLMs reveal a clear decoupling between action recognition and hallucination resistance, as models that perform well on positive action recognition often fail on adversarial negatives. Among all settings, sequential inference hallucination is the most severe, showing that current models tend to over-infer expected outcomes from partial motion cues. Our analyses further confirm that stronger priors and finer-grained similarity substantially amplify hallucination. We hope MoHallBench can facilitate future evaluation and mitigation of motion hallucination in VideoLLMs.
DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models
arXiv:2607.02374v2 Announce Type: replace Abstract: Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior context, then injecting this information into future prompts. We study whether such memory reshapes reasoning on open-ended questions where no single ground-truth answer exists. To quantify this effect, we introduce DRIFTLENS, a ground-truth-free framework that maps each expressed reasoning step to a value category and measures divergence between a question's no-memory trajectory and its trajectory under injected user-attribute memory. We first validate that DRIFTLENS distinguishes content-free pragmatic noise from substantive reasoning changes. Across four LLMs and 10 user-attribute categories, including age, occupation, and disability, user-attribute memory induces medium-to-large reasoning drift above each model's pragmatic-noise floor, even when final answers remain fluent, on-topic, and plausible. We then evaluate GRPO- and DPO-based post-training methods for reducing drift. Both reduce drift, but neither uniformly dominates; effects on downstream capability, helpfulness, and instruction following are model-and reward-dependent. These results suggest that memory-induced reasoning drift is a measurable and only partly mitigated failure mode of personalized language models.
SPECTRA: Context-Conditioned Spectral Movement Primitives for Robot Skill Generalization
arXiv:2607.06978v2 Announce Type: replace Abstract: Robot imitation learning for manipulation should preserve demonstrated task geometry while producing dynamically admissible robot motions. Existing pipelines often learn task-dependent trajectories and impose execution limits afterward through filtering, smoothing, clipping, or time scaling, which may distort task-critical end-effector paths. We propose the Spectral Movement Primitive (SMP), a frequency-domain imitation learning framework that couples task-space skill generation with joint-space execution regulation. Demonstrations are represented by truncated finite-horizon Fourier coefficients. An empirically selected low-frequency task band captures the dominant motion geometry, while higher harmonics contribute disproportionately to derivative growth. A frame-aware context-conditioned GMM/GMR prior predicts the task-band coefficients in a canonical task frame, and the resulting Cartesian trajectory is mapped to joint space through sequential inverse kinematics. A phase-coupled regulator then limits the requested phase progression without modifying the spectral coefficients, thereby enforcing joint velocity and acceleration limits while preserving the represented path. Experiments evaluate task-band reconstruction, robustness to composite demonstration corruption, out-of-distribution cross-board generalization, joint-space dynamic admissibility, end-effector path preservation, and deployment on a Franka Panda robot. Results show compact geometric reconstruction, consistent transfer across unseen task frames, substantial reductions in dynamic violations and jerk, and preservation of the intended end-effector path during phase regulation.
MLPTR-CC: Multi-label Pathology Test Recommendation using Classifier Chains and SHAP
arXiv:2607.08299v2 Announce Type: replace Abstract: Diagnostic decision making often relies on a sequence of pathology tests that bridge patient symptoms and final disease diagnosis. Existing clinical decision-support systems typically focus on predicting single diseases and do not explicitly recommend sets of intermediate tests or model dependencies among them. In this paper, we formulate pathology test recommendation as a multi-label classification problem where each case is associated with multiple, interdependent tests. We propose an AI-based framework that applies classifier chains with logistic regression, decision trees, random forests, and their ensemble to capture label dependencies between tests. Experiments on an expert-curated dataset from a private pathology laboratory show that classifier-chain models outperform their independent counterparts, improving F1-score and reducing Hamming loss while maintaining high accuracy across common and rare tests. To enhance trust and transparency, we integrate SHAP-based explainable AI, providing symptom-level attributions that align with established clinical reasoning in most cases. The results demonstrate that classifier chains combined with SHAP offer an effective and interpretable approach for multi-label pathology test recommendation, with potential to support clinicians in selecting appropriate diagnostic tests at an early stage.
Machine Learning Based Mesh Movement for Non-Hydrostatic Tsunami Simulation
arXiv:2603.06152v2 Announce Type: replace Abstract: This study investigates the use of machine learning based mesh movement method, specifically the Universal Mesh Movement Network (UM2N), with depth integrated non-hydrostatic shallow water models. Motivation for this comes from the need for models which balance efficiency and accuracy for use in probabilistic coastal hazard assessment. Implementations are built on the discontinuous Galerkin finite-element (DG-FE) based software, Thetis, which leverages the partial differential equation (PDE) framework Firedrake for automated code generation. Verification on benchmark test cases and validation against laboratory measurements of coastal hazards, focusing on tsunami propagation, run-up, and inundation is performed. In these tests, the UM2N-driven meshes help resolve key non-hydrostatic dynamics including wave refraction over a conical shoal, run-up with wetting-drying on a conical island, and tsunami inundation in the Monai Valley laboratory benchmark, and yield numerical solutions in close agreement with reference fine-mesh computations and measured data. Notably, in the Monai Valley case, UM2N achieves a ~91% reduction in wave-peak error at the nearshore gauge compared with ~74% for the conventional Monge--Amp\`ere (MA) mesh movement, both relative to the coarse fixed mesh. The UM2N surrogate based approach accelerates the conventional mesh movement step, achieving a ~32% reduction in total runtime and ~2 times speed-up in mesh movement step time over the MA solver on GPU, while offering a significant improvement in robustness over long integration periods and under strongly nonlinear wave conditions.
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
arXiv:2603.06592v2 Announce Type: replace Abstract: Contemporary studies in mechanistic interpretability have uncovered many puzzling phenomena in the neural information processing of Transformer-based language models, such as induction heads, function vectors, and the Hydra effect. Some of these individual phenomena have been independently tied to different data distributional properties, while some have been loosely associated with model architecture and how Transformers process information. However, a unified understanding of the relationship between data, model architecture, and optimization remains lacking, failing to answer the fundamental question: why do these three phenomena appear universally across different model families and scales, despite their seeming disconnect? In this work, we answer this question by unifying these three phenomena as consequences of hierarchical latent structures in the data generation process, coupled with decorrelated gradients across additive model components and directional concavity in the representation geometry. We validate our theoretical results in a toy model regime and in a large-scale synthetic data regime, comparing them with language models trained on natural language data.
Git-Assistant: Planning-Based Support for Updating Git Repositories
arXiv:2607.09224v2 Announce Type: replace Abstract: Version control systems are essential for collaborative software development, yet tools like git remain challenging for many practitioners. Recent advances in Large Language Models (LLMs) offer promising capabilities for interpreting developer intent, but their effectiveness in repository management tasks is limited by the need for formal reasoning. This work introduces Git-Assistant, an AI-based assistant that combines LLMs with automated planning to support developers in executing non-trivial git operations. The assistant analyzes repository context, translates natural language requests into actionable command sequences, and incorporates planning techniques to ensure correctness and safety. We present a systematic evaluation methodology using synthetic and randomized git environments, comparing the performance of LLM-only and planning-augmented variants across multiple metrics. Experimental results demonstrate that integrating formal reasoning with LLMs improves reliability and reduces errors in repository management, highlighting the potential of hybrid AI approaches for intelligent developer assistance.
Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination
arXiv:2605.05686v3 Announce Type: replace Abstract: Language models draw on two knowledge sources: facts baked into weights (parametric memory, PM) and information in context (working memory, WM). We study two mechanistically distinct failure modes--conflict, when PM and WM disagree and interfere; and hallucination, when the queried fact was never learned. Both produce confident output regardless, making output-based monitoring blind by design. We show both failures share a unified geometric account. In the hidden-state space of autoregressive generation, learned facts form attractor basins. Conflict is basin competition: WM disrupts convergence to the correct basin without raising output entropy. Hallucination is basin absence: the hidden state drifts freely when no memorized basin exists. The frozen LM head, designed for next-token prediction, cannot distinguish these cases and fires confidently either way. We verify this account in a controlled synthetic task-entity identifiers mapped to unique codes with PM installed via LoRA adapters--where ground truth is exact and component roles can be causally isolated through targeted adapter placement. Geometric margin--the hidden state's distance to the nearest memorized basin--reads this geometry directly and separates correct recall from hallucination far more cleanly than output entropy, with zero false refusals where entropy-based detection cannot avoid rejecting the vast majority of correct outputs. The separation holds on natural-language factual queries from the pretrained model with no adaptation, confirming attractor geometry is structural rather than a fine-tuning artifact. The fraction of confident hallucinations follows a scaling law $C = \exp(-c/\bar\Delta)$, growing with scale even as overall error rates fall. Hidden states reliably encode epistemic state; the frozen output head systematically erases it--and this erasure worsens with scale.