Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning
arXiv:2601.19624v3 Announce Type: replace Abstract: Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entropy, causing over-exploration during stable periods and under-exploration after drift, and leaving unanswered the principled question of how exploration intensity should scale with drift magnitude. We show that, under standard assumptions, entropy scheduling in non-stationary maximum-entropy RL can be cast as the dynamic-regret trade-off between tracking a drifting comparator and stabilizing updates, yielding a square-root scaling rule for the entropy weight in terms of a online non-stationarity proxy. Building on this, we propose AES--Adaptive Entropy Scheduling--which adaptively adjusts the entropy coefficient/temperature online using observable drift proxies during training, requiring almost no structural changes and incurring minimal overhead. Across 4 algorithm variants, 12 tasks, and 4 drift modes, AES significantly reduces the fraction of performance degradation caused by drift and accelerates recovery after abrupt changes.
PINN-based short-term forecasting of fault slip evolution during the 2010 slow slip event in the Bungo Channel, Japan
arXiv:2601.21516v2 Announce Type: replace Abstract: Monitoring and forecasting fault slip evolution are fundamental for understanding earthquake cycles and assessing future seismic hazards. This study proposes a physics-based data assimilation framework that integrates geodetic observations with fault mechanics introducing spatial heterogeneity in frictional properties, with a particular focus on short-term fault slip forecasting. The proposed method employs physics-informed neural networks (PINNs) to calculate fault slip evolutions and to optimize the spatial distribution of frictional properties and is applied to the 2010 slow slip event beneath the Bungo Channel, southwest Japan, by changing the data period to be assimilated. When only the initial phase of slip acceleration is assimilated, a velocity-weakening frictional region is inferred beneath southwest Shikoku, corresponding to the initial nucleation are of the slow slip event. Out results demonstrate that the PINN-based data assimilation framework successfully forecasts slow transient slip even when only slip acceleration data are assimilated, whereas forecasts based on frictionally homogeneous models result in unstable fast slip. This difference can be interpreted as a consequence of introducing frictional heterogeneity, which allows both the characteristic size of the slipping region and the critical nucleation size to be variable, leading to stable slip evolution consistent with observations. When longer observation periods are assimilated, a velocity-strengthening region emerges around the slip-weakening patch, progressively restricting the direction of slip propagation. This velocity-strengthening region is interpreted as a mechanical constraint imposed by fault physics, linking the slip regions required to reproduce the observed geodetic time series. The results highlight the capability of PINN-based data assimilation incorporating geodetic observations and fault mechanics.
Decision-Making Under Complete Uncertainty: You Will Not Regret Being Greedy
arXiv:2502.07593v2 Announce Type: replace Abstract: In this paper, we propose a game-theoretic model to study the properties of the worst-case regret of the greedy strategy under complete (Knightian) uncertainty. In a game between a decision-maker (DM) and an adversarial agent (Nature), Nature chooses an unknown state determining the distribution of ratings for each product. The DM observes a realization of product ratings and then chooses a product according to a strategy. For arbitrary numbers of products and ratings, we first study the equal-observations case in which every product has the same number of observations. In this benchmark, we establish matching upper and lower bounds on the worst-case regret, showing that the regret vanishes as the number of observations increases and that the greedy strategy is rate-optimal up to universal constants. In the special case with two products and two ratings, we show that with one observation per product the greedy strategy is minimax-optimal with respect to worst-case regret. We then allow products to have different numbers of observations. Greedy remains robust in a conservative sense: its worst-case regret is controlled by the least-reviewed product. However, unequal numbers of observations can also change greedy's exact worst-case behavior. In particular, adding observations for only one product can increase greedy's worst-case regret. Finally, we test the model on data collected from Google reviews for restaurants, showing that the greedy strategy's empirical performance closely aligns with the theoretical findings.
Influence of bending parameters on crystalline undulator radiation peak stability for 530 MeV positron channelling
arXiv:2601.06921v2 Announce Type: replace Abstract: We investigate the stability of crystalline undulator radiation (CUR) peaks emitted by 530 MeV positron channelling in periodically bent C(110) crystals with varying bending amplitudes and bending periods. Relativistic molecular dynamics simulations were performed to quantify how these parameters affect the intensity and position of the CUR peak. The continuous potential approximation was used to identify isolines of constant peak energy, providing a reference for regions of spectral stability. MD results show that increasing the bending amplitude shifts the CUR peak to lower photon energies, while decreasing the period shifts it to higher energies, with both trends accompanied by enhanced dechannelling. For crystal parameters similar to recent experiments conducted at the MAinz MIkrotron (MAMI), the simulated CUR peak appears near 0.515 MeV. These results demonstrate that the CUR peak remains stable across a broad range of bending amplitudes and periods, providing quantitative estimates of the sensitivity of the emitted radiation to variations in the crystal bending parameters.
Structure-Semantic Co-optimized Latent Diffusion Model for Fast Visual Anagram Synthesis
arXiv:2606.16241v3 Announce Type: replace Abstract: Visual anagram is an intriguing form of art creation wherein a single image presents different conceptual interpretations under transformations such as flipping or rotation. Recent work has achieved visual anagram synthesis by leveraging pretrained text-to-image (T2I) diffusion models, yet still suffers from several key limitations including computational inefficiency, suboptimal aesthetic quality, and weak semantic fidelity and expressiveness. This work focuses on generating visual anagrams with substantially improved visual quality at minimal computational cost, thereby advancing intelligent creation of illusionary digital art. To increase image resolution while reducing time overhead, we adapt the cutting-edge parallel denoising algorithm from pixel-based T2I model to the adversarially distilled latent-based one, and accordingly propose a structure-semantic co-optimization (S2CO) framework to counteract the consequent visual degradation. As the core of our approach, S2CO framework comprises three key innovations: (\romannumeral1) null-text structure alignment optimization; (\romannumeral2) semantic enhancement optimization; (\romannumeral3) attention-guided noise fusion. Building upon these components, our method dubbed \textbf{S2CO-Anagram} is able to generate higher-resolution anagram images with noticeably superior visual harmony and semantic faithfulness than related SOTA approaches, all while achieving substantially faster inference speed. Code will be publicly available.
Lasing from a Quantum-Dot-Like Buried Heterostructure in an InP Nanobeam Cavity
arXiv:2603.16636v2 Announce Type: replace Abstract: We report lasing from a lithographically defined buried heterostructure with an estimated lateral footprint of (107 nm)^2, embedded in an InP photonic-crystal nanobeam cavity. This represents the smallest laterally confined buried heterostructure gain region from which lasing has been observed. Despite etching of the active region during cavity definition and the associated risk of surface-related nonradiative recombination, optically pumped devices exhibit a clear lasing threshold and a narrow linewidth. By systematically varying the BH size, we investigate how the lasing threshold depends on the active volume under optical pumping. The estimated intrinsic threshold under ideal carrier injection is 57 nW, comparable to values reported for single quantum-dot nanolasers, highlighting the potential of quantum-dot-scale buried heterostructures as deterministic, scalable gain media for nanophotonic lasers.
Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens
arXiv:2606.16847v3 Announce Type: replace Abstract: Diffusion Large Language Models (dLLMs) offer a promising avenue for parallel generation but face a trade-off between decoding speed and quality. While revocable decoding strategies attempt to mitigate errors by verifying and remasking tokens, they typically operate within a mixed-quality context. This leads to two critical failures: \textit{Error Propagation}, where new tokens absorb toxic information from erroneous context, and \textit{Local Error Reinforcement}, where errors mutually reinforce each other to evade detection. To alleviate these challenges, we propose ASRD (Anchor Supervised Revocable Decoding), a training-free framework that operates within the embedding space. ASRD explicitly decouples the decoding context into trusted \textit{Anchor Tokens}, which are identified via temporal consistency, and uncertain candidates. Leveraging a dynamic Anchor Tokens Cache, we introduce two complementary mechanisms: (1) Anchor-Guided Generation, which injects entropy-weighted anchor signals into masked positions to implicitly rectify attention toward the reliable global skeleton; and (2) Anchor-Perturbed Verification, which applies orthogonal perturbations to uncertain candidate tokens, destabilizing and remasking errors driven by fragile local consensus. Extensive experiments on math and coding benchmarks demonstrate that ASRD outperforms recent remasking baselines, achieving accuracy improvements of up to 6.4\% while accelerating inference throughput by up to 7.2$\times$.
LapSurgie: Humanoid Robots Performing Surgery via Teleoperated Handheld Laparoscopy
arXiv:2510.03529v3 Announce Type: replace Abstract: Robotic laparoscopic surgery has gained increasing attention in recent years for its potential to deliver more efficient and precise minimally invasive procedures. However, adoption of surgical robotic platforms remains largely confined to high-resource medical centers, exacerbating healthcare disparities in rural and low-resource regions. To close this gap, a range of solutions has been explored, from remote mentorship to fully remote telesurgery. Yet, the practical deployment of surgical robotic systems to underserved communities remains an unsolved challenge. Humanoid systems offer a promising path toward deployability, as they can directly operate in environments designed for humans without extensive infrastructure modifications -- including operating rooms. In this work, we introduce LapSurgie, the first humanoid-robot-based laparoscopic teleoperation framework. The system leverages an inverse-mapping strategy for manual-wristed laparoscopic instruments that abides to remote center-of-motion constraints, enabling precise hand-to-tool control of off-the-shelf surgical laparoscopic tools without additional setup requirements. A control console equipped with a stereo vision system provides real-time visual feedback. Finally, a comprehensive user study across platforms demonstrates the effectiveness of the proposed framework and provides initial evidence for the feasibility of deploying humanoid robots in laparoscopic procedures.
Stochastic Finite Volume Approximation with Clustering in the Parameter Space for the Forward Uncertainty Quantification of Differential Equations with Random Parameters
arXiv:2510.12109v2 Announce Type: replace Abstract: The uncertainty quantification (UQ) for mathematical models with random parameters is important for many science and engineering problems. Forward UQ quantifies the impact of random parameters on the output of system. In the current study, we propose a new stochastic finite volume (SFV) scheme by combining SFV with clustering algorithm in the parameter space such that each cluster can be regarded as a finite volume with implicit boundaries. The advantage of SFV is that no specific form of the random variable is required such that discontinuous solutions and sharp interfaces can be accurately approximated. Compared to classic SFV based on structured grids, the new SFV-cluster scheme extends SFV to parameter spaces of higher dimensions. For demonstration and validation, we present the construction of SFV schemes for the Kraichnan-Orszag three-mode problem and the Buckley-Leverett equation. The error analysis of SFV is extended and the new algorithm is validated by numerical experiments.
SWARM+: Scalable and Resilient Multi-Agent Consensus for Decentralized Data-Aware Workload Management
arXiv:2603.19431v4 Announce Type: replace Abstract: Distributed scientific workflows are increasingly executed across heterogeneous and geo-distributed computing environments, where centralized workload orchestration becomes a scalability and resilience bottleneck. This paper presents SWARM+, a decentralized workload management system that coordinates workload placement through hierarchical multi-agent consensus, reducing coordination overhead and dramatically improving scalability, while tolerating failures and dynamic membership changes. SWARM+ enables data-aware scheduling policies that incorporate resource availability, data transfer node (DTN) connectivity, and data locality into workload placement decisions. We evaluate SWARM+ on the distributed FABRIC testbed using heterogeneous scientific workloads derived from production workflow traces obtained from the Pegasus Workflow Management System (WMS). Experimental results show that SWARM+ scales coordination to 990 distributed agents with approximately 1\,s per-job selection time at 110 agents. SWARM+ demonstrates balanced workload distribution, maintains over 97% job completion under distributed failures with graceful degradation (mean ~95% job completion) during correlated site outages, tolerates coordinator agent failures gracefully, improves schedule quality by employing data-aware policies, and reduces both selection time and scheduling latency by 97-98% when compared to the prior SWARM system.
ParaTutor: Coordinating Parent and Child Math Tutoring through Role Separated LLM Scaffolding
arXiv:2606.18030v2 Announce Type: replace Abstract: Parent and child tutoring is a collaborative learning setting with asymmetric roles. Parents guide children s problem solving, while children are expected to remain actively engaged in understanding and reasoning. However, most LLM based learning systems are designed for single users or relatively symmetric collaboration, leaving parent and child tutoring with distinct instructional roles underexplored. Through a formative study, we found that parent and child math tutoring was often disrupted by cognitive misalignment, emotional escalation, and method mismatch. To address these challenges, we present ParaTutor, a multiple agents LLM based scaffolding system for home math word problem tutoring. ParaTutor distributes support across user roles by providing parents with strategy, language, repair, and phase scaffolds, while providing children with visual grounding for problem interpretation. We evaluated ParaTutor with 23 parent and child dyads (children aged 10 to 12) across four tutoring conditions that varied how LLM assistance was delivered. Results show that generic LLM assistance often provided useful explanations but did not consistently support parent led tutoring or children s active reasoning. In contrast, ParaTutor helped redistribute tutoring work across parents and children, increased children s engagement with word problems, supported shared understanding through visual grounding, and helped parents translate LLM generated methods into child facing tutoring moves. These findings suggest that in family learning, the value of LLM support depends not only on model capability, but also on how support is coordinated across users with different roles. Our work contributes design implications for LLM systems that support role sensitive scaffolding in parent and child learning.
Investigating students' gender expression and its relation to sense of belonging in introductory physics courses
arXiv:2603.17966v2 Announce Type: replace Abstract: Despite nation efforts to promote diversity and inclusion, women and gender minorities remain underrepresented in physics. One common approach to studying gender in physics contexts treats gender as a categorical identity variable (e.g. "man," "woman"). In contrast, approaches that center gender expression focus on the nuanced and context-dependent ways in which gender is socially enacted and interpreted. They are therefore well-suited for exploring how gender permeates the small-scale interactions that ultimately shape students' persistence and perceptions of inclusivity. In the present study, we utilized gender expression as a lens to investigate gendered patterns in introductory undergraduate physics students' sense of belonging in the discipline. Specifically, this qualitative investigation expands on our previous quantitative work to investigate why students may feel misperceived by their peers in physics and how that experience influences their belonging. Results indicated that students' sense of belonging may be impacted by perceived pressures to alter their gender expression in physics contexts. Many interviewees expressed a felt need to present themselves more "masculinely" to fit in. Contrastingly, pressure to present "femininely" was most often associated with standing out. Implications for supporting students' authentic self-ex in physics contexts are discussed.
On splitting strategies for the numerical solution of stochastic delay differential equations with correlated noises
arXiv:2603.21858v2 Announce Type: replace Abstract: In this article we investigate the numerical solution of a scalar semilinear stochastic delay differential equation (SDDE) where the linear instantaneous feedback and nonlinear delayed feedback terms are perturbed by a pair of standard Brownian motions with correlation $\rho$. Such SDDEs may be naturally decomposed into two subsystems: a linear stochastic differential equation (SDE) without delay, and a nonlinear SDDE. Splitting methods work by solving each subsystem separately and composing the results over a single step. Our main theoretical result provides a bound on the mean-square error of a particular strategy for doing this, known as Lie-Trotter splitting. This bound implies that the method is mean-square strongly convergent with order $1/2$ when $\rho=0$, so that the noises are uncorrelated, but assurances of convergence are lost when $\rho\neq 0$. Indeed we develop an upper bound on the global mean-square error with a term that depends linearly on the magnitude of the correlation, and is independent of the stepsize. While our theoretical error bound is an estimate from above, we conduct numerical experiments that confirm the order of mean-square strong convergence of Lie-Trotter splitting in the $\rho=0$ case, and demonstrate a rapid fall-off to effectively zero as $|\rho|$ increases. Similar numerical results are observed for an alternative commonly used strategy known as Strang splitting. Nonetheless, by carefully reorganising the subsystems into which we split the SDDE, we can improve the range of values of $\rho$ over which a nonzero order of convergence is observed numerically.
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
arXiv:2603.25112v2 Announce Type: replace Abstract: Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accuracy) with how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity). We apply Signal Detection Theory (SDT) to decompose these capacities, treating token-level normalised log-probability as a graded confidence variable and answer correctness as the state to be discriminated. We characterise the Type-2 ROC of this signal, including its unequal-variance structure via z-ROC analysis, and -- because the meta-d' efficiency ratio is not well defined for open-ended QA, which lacks a two-alternative Type-1 decision -- quantify metacognitive efficiency with a model-free information measure, normalised metacognitive information (meta-I_2r). Applied to four LLMs (Llama-3-8B-Instruct, Mistral-7B-Instruct-v0.3, Llama-3-8B-Base, Gemma-2-9B-Instruct) across 224,000 factual QA trials, we find: (1) metacognitive information varies more than two-fold across models and co-varies inversely with accuracy -- the least accurate model has the most informative confidence -- though with four models this ordering cannot be separated from an error-difficulty confound, so we report it as coupling, not decoupling; (2) the confidence signal has model-specific unequal-variance structure (z-ROC slopes 0.81 to 1.18) invisible to calibration metrics; (3) metacognitive information is domain-specific, strongest in Arts & Literature for every model; (4) temperature dissociates Type-1 accuracy from metacognitive information, which stays stable while accuracy shifts. All estimates carry permutation nulls and bootstrap confidence intervals. Pre-registered; code and data public.
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
arXiv:2603.27706v2 Announce Type: replace Abstract: Reference Audio-Visual Segmentation (Ref-AVS) aims to segment objects in audible videos based on multimodal cues in reference expressions. Previous methods overlook the explicit recognition of expression difficulty and dominant modality in multimodal cues, over-rely on the quality of the instruction-tuning dataset for object reasoning, and lack reflective validation of segmentation results, leading to erroneous mask predictions. To address these issues, in this paper, we propose a novel training-free Multi-Agent Recognition, Reasoning, and Reflection framework to achieve high-quality Reference Audio-Visual Segmentation, termed MAR3. Incorporating the sociological Delphi theory to achieve robust analysis, a Consensus Multimodal Recognition mechanism is proposed that enables LLM agents to explicitly recognize the difficulty of reference expressions and the dominant modality of multimodal cues. Based on our modality-dominant difficulty rule, we propose an adaptive Collaborative Object Reasoning strategy to reliably reason about the referred object. To further ensure precise mask prediction, we develop a Reflective Learning Segmentation mechanism, in which a check agent examines intermediate segmentation results and iteratively corrects the object text prompt of the segment agent. Experiments demonstrate that MAR3 achieves superior performance (69.2% in J&F) on the Ref-AVSBench dataset, outperforming SOTA by 3.4% absolutely.
Continuous Policy and Value Iteration for Stochastic Control Problems and Its Convergence
arXiv:2506.08121v2 Announce Type: replace-cross Abstract: We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This framework applies to both the entropy-regularized relaxed control problems and the classical control problems, with infinite horizon. We establish policy improvement and demonstrate convergence to the optimal control under the monotonicity condition of the Hamiltonian. By utilizing Langevin-type stochastic differential equations for continuous updates along the policy iteration direction, our approach enables the use of distribution sampling and non-convex learning techniques in machine learning to optimize the value function and identify the optimal control simultaneously.
Beyond Logit Adjustment: A Residual Decomposition Framework for Long-Tailed Reranking
arXiv:2604.01506v2 Announce Type: replace Abstract: Long-tailed classification, where a small number of frequent classes dominate many rare ones, remains challenging because models systematically favor frequent classes at inference time. Existing post-hoc methods such as logit adjustment address this by adding a fixed classwise offset to the base-model logits. However, the correction required to restore the relative ranking of two classes need not be constant across inputs, and a fixed offset cannot adapt to such variation. We study this problem through Bayes-optimal reranking on a base-model top-k shortlist. The gap between the optimal score and the base score, the residual correction, decomposes into a classwise component that is constant within each class, and a pairwise component that depends on the input and competing labels. When the residual is purely classwise, a fixed offset suffices to recover the Bayes-optimal ordering. We further show that when the same label pair induces incompatible ordering constraints across contexts, no fixed offset can achieve this recovery. This decomposition leads to testable predictions regarding when pairwise correction can improve performance and when cannot. We develop REPAIR (Reranking via Pairwise residual correction), a lightweight post-hoc reranker that combines a shrinkage-stabilized classwise term with a linear pairwise term driven by competition features on the shortlist. Experiments on nine benchmarks confirm that the decomposition explains where pairwise correction helps and where classwise correction alone suffices. These span text classification, visual recognition, and multimodal rare-disease diagnosis.
The Darkside-20k Data Acquisition System
arXiv:2604.03059v3 Announce Type: replace Abstract: DarkSide-20k is a WIMP search experiment using liquid argon as a target, designed to perform a background-free search for dark matter with unprecedented sensitivity, and is currently under construction at INFN Laboratori Nazionali del Gran Sasso, Italy. The detector comprises a dual-phase Time Projection Chamber complemented with external veto systems and is equipped with a total of 2720 SiPM-based readout channels. This work presents the DAQ system designed for DarkSide-20k. The system is capable of continuous, triggerless digitisation of the waveforms with high single-photoelectron detection efficiency and online processing, ensuring data reduction for long-term storage. The DarkSide-20k DAQ system employs commercial CAEN VX2745 digitisers with custom FPGA firmware implementation. Timing and synchronisation across all 48 digitisers are provided by custom Global and Crate Data Manager boards distributing a phase-aligned clock derived from a disciplined rubidium standard. Waveform segments are processed in real time by Front End Processor machines. Data are organised into collections containing whole detector information and distributed across a farm of Time Slice Processors for event reconstruction, classification, and further reduction before storage and offline analysis. A full "Quadrant" of the system, corresponding to one quarter of the final DAQ, has been assembled and validated at TRIUMF laboratory in Canada. The Quadrant has been stress-tested with simultaneous pulses and demonstrated sustained digitizer readout exceeding expected physics rates and stable long-term performance.
Understanding LoRA as Knowledge Memory: An Empirical Analysis
arXiv:2603.01097v4 Announce Type: replace Abstract: Continuous knowledge updating for pre-trained large language models (LLMs) is increasingly necessary yet remains challenging. Although inference-time methods like In-Context Learning (ICL) and Retrieval-Augmented Generation (RAG) are popular, they face constraints in context budgets, costs, and retrieval fragmentation. Departing from these context-dependent paradigms, this work investigates a parametric approach using Low-Rank Adaptation (LoRA) as a modular knowledge memory. Although few recent works examine this concept, the fundamental mechanics governing its capacity and composability remain largely unexplored. We bridge this gap through the first systematic empirical study mapping the design space of LoRA-based memory, ranging from characterizing storage capacity and optimizing internalization to scaling multi-module systems and evaluating long-context reasoning. Rather than proposing a single architecture, we provide practical guidance on the operational boundaries of LoRA memory. Overall, our findings position LoRA as the complementary axis of memory alongside RAG and ICL, offering distinct advantages.
LogiDroid: Individual Functional Test Generation via Business Logic Extraction and Adaptation
arXiv:2602.24108v2 Announce Type: replace Abstract: Functional testing is essential for verifying that the business logic of mobile applications aligns with user requirements. Despite its importance, functional testing remains heavily dependent on manual effort due to two core challenges. First, acquiring and reusing business logic from unstructured requirements remains difficult, which hinders the understanding of specific functionalities. Second, a significant semantic gap exists when adapting business logic to the diverse GUI environments, which hinders the generation of test cases for specific mobile applications. To address the preceding challenges, we propose LogiDroid, a two-stage approach that generates individual functional test cases by extracting business logic and adapting it to target applications. First, in the Knowledge Retrieval and Fusion stage, two LLM-based agents are employed to construct a functional test dataset, retrieve relevant test cases, and extract structured business logic for the target functionality. Second, in the Context-Aware Test Generation stage, two other LLM-based agents jointly analyze the extracted business logic and the real time GUI environment to incrementally generate context adaptive functional test cases. This design allows LogiDroid to accurately understand application semantics and use domain expertise to generate complete test cases with verification assertions. We assess the effectiveness of LogiDroid using two widely-used datasets that cover 28 real-world applications and 190 functional requirements. Experimental results show that LogiDroid successfully tested 40% of functional requirements on the FrUITeR dataset (an improvement of over 25% compared to the state-of-the-art approaches) and 65% on the Lin dataset (an improvement of over 55% compared to the state-of-the-art approaches). These results demonstrate the significant effectiveness of LogiDroid in functional test generation.
Mistake gating leads to energy and memory efficient continual learning
arXiv:2604.14336v2 Announce Type: replace Abstract: Synaptic plasticity is metabolically expensive, yet animals continuously update their internal models without exhausting energy reserves. However, when artificial neural networks are trained, the network parameters are typically updated on every sample that is presented, even if the sample was classified correctly. Inspired by the human negativity bias and error-related negativity, we propose 'memorized mistake-gated learning' -- a biologically plausible plasticity rule where synaptic updates are strictly gated by current and past classification errors. This reduces the number of updates the network needs to make by $50\%\sim80\%$. Mistake gating is particularly well suited in two cases: 1) For incremental learning where new knowledge is acquired on a background of pre-existing knowledge, 2) For online learning scenarios when data needs to be stored for later replay, as mistake-gating reduces storage buffer requirements. The algorithm can be implemented in a few lines of code, adds no hyper-parameters, and comes at negligible computational overhead. Learning on mistakes is an energy efficient and biologically relevant modification to commonly used learning rules that is well suited for continual learning.
A Practical Guide to PID Controller Implementation
arXiv:2604.15918v3 Announce Type: replace Abstract: How difficult can it be to implement a PID controller? The answer is twofold. Implementing the PID control law is simple and computationally inexpensive. However, this basic form will not work in practical applications. The primary reason for this is the various physical limitations of the actuator. Measurement noise, different implementations depending on the various structures (P, PI, PD or PID), bumpless transfer, and varying sampling interval also result in problems rendering the basic form inoperable. PID implementation is therefore more difficult than meets the eye. This paper introduces a reference implementation of the PID controller which considers these practical issues. It includes pseudo-code, discussion of the implementation choices and simulation of carefully selected, important test cases.
MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models
arXiv:2607.01117v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) have shown strong progress in video understanding, yet they still suffer from hallucinations that are inconsistent with visual evidence. Existing benchmarks mainly focus on object hallucination or coarse action perception, leaving a key video-specific problem underexplored: motion hallucination, in which models infer human motions that are absent from the video. We present MoHallBench, a benchmark for diagnosing motion hallucination in VideoLLMs. MoHallBench systematically evaluates three major sources of hallucination: co-occurrence priors, sequential inference, and similarity confusion. It contains 11,306 video clips and 40,493 question-answer pairs, covering binary-choice, multiple-choice, and generative settings. We further introduce a bi-directional questioning protocol with bias-aware metrics to reduce affirmation bias in binary evaluation. Experiments on ten recent open-source VideoLLMs reveal a clear decoupling between action recognition and hallucination resistance, as models that perform well on positive action recognition often fail on adversarial negatives. Among all settings, sequential inference hallucination is the most severe, showing that current models tend to over-infer expected outcomes from partial motion cues. Our analyses further confirm that stronger priors and finer-grained similarity substantially amplify hallucination. We hope MoHallBench can facilitate future evaluation and mitigation of motion hallucination in VideoLLMs.
DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models
arXiv:2607.02374v2 Announce Type: replace Abstract: Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response. Modern LLMs personalize interactions by storing user attributes, preferences, and prior context, then injecting this information into future prompts. We study whether such memory reshapes reasoning on open-ended questions where no single ground-truth answer exists. To quantify this effect, we introduce DRIFTLENS, a ground-truth-free framework that maps each expressed reasoning step to a value category and measures divergence between a question's no-memory trajectory and its trajectory under injected user-attribute memory. We first validate that DRIFTLENS distinguishes content-free pragmatic noise from substantive reasoning changes. Across four LLMs and 10 user-attribute categories, including age, occupation, and disability, user-attribute memory induces medium-to-large reasoning drift above each model's pragmatic-noise floor, even when final answers remain fluent, on-topic, and plausible. We then evaluate GRPO- and DPO-based post-training methods for reducing drift. Both reduce drift, but neither uniformly dominates; effects on downstream capability, helpfulness, and instruction following are model-and reward-dependent. These results suggest that memory-induced reasoning drift is a measurable and only partly mitigated failure mode of personalized language models.
SPECTRA: Context-Conditioned Spectral Movement Primitives for Robot Skill Generalization
arXiv:2607.06978v2 Announce Type: replace Abstract: Robot imitation learning for manipulation should preserve demonstrated task geometry while producing dynamically admissible robot motions. Existing pipelines often learn task-dependent trajectories and impose execution limits afterward through filtering, smoothing, clipping, or time scaling, which may distort task-critical end-effector paths. We propose the Spectral Movement Primitive (SMP), a frequency-domain imitation learning framework that couples task-space skill generation with joint-space execution regulation. Demonstrations are represented by truncated finite-horizon Fourier coefficients. An empirically selected low-frequency task band captures the dominant motion geometry, while higher harmonics contribute disproportionately to derivative growth. A frame-aware context-conditioned GMM/GMR prior predicts the task-band coefficients in a canonical task frame, and the resulting Cartesian trajectory is mapped to joint space through sequential inverse kinematics. A phase-coupled regulator then limits the requested phase progression without modifying the spectral coefficients, thereby enforcing joint velocity and acceleration limits while preserving the represented path. Experiments evaluate task-band reconstruction, robustness to composite demonstration corruption, out-of-distribution cross-board generalization, joint-space dynamic admissibility, end-effector path preservation, and deployment on a Franka Panda robot. Results show compact geometric reconstruction, consistent transfer across unseen task frames, substantial reductions in dynamic violations and jerk, and preservation of the intended end-effector path during phase regulation.