Forskningsradar

Science Journals

Peer-reviewade publikationer — 60005 artiklar

S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation
arXiv:2606.27872v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation, but their performance degrades significantly in long-horizon tasks due to cumulative error propagation. This limitation largely arises from static feature fusion mechanisms that rely on fixed weights to combine visual, language, and action representations, preventing the model from adapting to different phases of task execution. To address this limitation, we propose S$^2$-VLA, a framework that introduces a State-Space Guided Adaptive Attention (SSGAA) mechanism. SSGAA maintains a belief state that tracks task progression and generates dynamic gating weights to adaptively fuse information from three complementary sources visual features for spatial perception, task intents for high-level task planning, and temporal action sequences for execution consistency. This adaptive fusion allows the model to shift its focus throughout task execution, aligning with the evolving requirements of different task stages. Despite its compact 2B parameter size, S$^2$-VLA consistently outperforms larger 7B-scale models and achieves state-of-the-art performance on long-horizon manipulation benchmarks, including LIBERO and SimplerEnv. highlighting the importance of adaptive feature fusion for long-horizon robotic manipulation.
Ocean-atmosphere interaction at the Gulf Stream sea surface temperature front: variability and impacts on midlatitude atmospheric circulation
arXiv:2606.27873v1 Announce Type: new Abstract: Sea surface temperature (SST) gradients associated with western boundary currents affect the atmospheric circulation across a range of spatial and temporal scales. Yet, several aspects of ocean-atmosphere interactions linked to oceanic fronts remain unclear. This PhD thesis analyses such interactions for the Gulf Stream SST front (GSF). The first part assesses the atmospheric response to the interannual GSF meridional shifts and its dependence on model horizontal resolution, using ERA5 reanalysis and atmosphere-only simulations forced by observed SST. Results show that the response is strongly resolution dependent, with only simulations finer than 50km resembling observed anomalies. Locally, diabatic heating near the GSF is mainly balanced by vertical motion and transient eddy heat transport. At large-scale, the GSF shifts is associated with a homo-directional shift in the North Atlantic eddy-driven jet and storm track, mediated by changes in low-level baroclinicity. The second part assesses the North Atlantic Oscillation (NAO)-GSF interaction and the mechanisms through which the NAO forces the GSF shifts on decadal timescale, using atmosphere and ocean reanalyses. The NAO and GSF covary on decadal timescales only during 1972-2018. This non-stationarity is also reflected in their lead-lag relationship: the NAO leads the GSF shifts by 3 years during 1972-1990 and by 2 years during 1990-2018. The lag is interpreted as the joint effect of the fast response of wind-driven oceanic circulation, the lagged response of deep oceanic circulation, and the propagation of Rossby waves. However, Rossby wave propagation is evident only before 1990, suggesting that its non-stationarity may explain the different NAO-GSF time lag before and after 1990. Overall, the thesis improves understanding of GSF variability and its role in North Atlantic and extratropical climate variability.
OrthoTryOn: Geometric Orthogonalization for Conflict-Free Unified Fashion Generation
arXiv:2606.27880v1 Announce Type: new Abstract: Unified fashion generation integrates tasks like virtual try-on and garment reconstruction into a single model to reduce task-specific adaptation costs. However, naive parameter sharing across semantically distinct tasks induces negative transfer through severe inter-task gradient conflict. We propose OrthoTryOn, a unified framework mitigating this interference within a shared Low-Rank Adaptation (LoRA) module. Its Orthogonal Subspace Projection (OSP) applies task-specific orthogonal rotations to bottleneck features, mapping them into decorrelated coordinate frames. To address residual semantic coupling at inference time, we further propose Fisher-guided Negative Guidance (FNG), a parameter-free strategy that utilizes diagonal Fisher information to quantify inter-task sensitivity overlap and explicitly repels generation trajectories from the most confusable task via Classifier-Free Guidance. Extensive experiments demonstrate that OrthoTryOn avoids the severe performance degradation typical of naive unified training and even surpasses independently trained task-specific models, achieving state-of-the-art results across multiple benchmarks while generalizing robustly across diverse diffusion backbones. Code is available at https://github.com/NJU-PCALab/OrthoTryOn.
Swarm sign language: motion-based communication between drones
arXiv:2606.27883v1 Announce Type: new Abstract: In stealth-constrained swarm robotics, visual communication provides a critical alternative to active radio transmissions, which might be jammed. This research investigates motion-based communication for non-active information exchange, utilizing modular, dynamically feasible planar trajectories as visual cues. On the receiver drone end, a pose estimator tracks the transmitting drone's pose, feeding it into our custom 3DTrajDecoder. The decoder is designed to classify and segment the spatiotemporal sequence while simultaneously regressing its size and normal vector. To robustly train the decoder on both communicative and non-communicative trajectories, we developed a configurable online procedural generation pipeline. We validate our system through real-world testing and simulation to define its operating domain, supported by an extensive ablation study detailing our architectural choices and system limitations.
A Dynamical Low-rank Multilevel Monte Carlo Estimator for High-Dimensional Kinetic Equations
arXiv:2606.27888v1 Announce Type: new Abstract: Kinetic equations are used to model a wide range of phenomena important for real-world applications. Their applications span astrophysics, nuclear physics, engineering, and social sciences. Due to their high-dimensional phase space, modelling and quantifying uncertainties, relevant for applications, poses a significant challenge even for modern computing infrastructure. In recent years, dynamical low-rank approximation (DLRA) has gained popularity for making fine grid simulations of high-dimensional problems feasible by evolving the solution of a time-dependent PDE as a low-rank factorization. This reduces the computational and memory requirements significantly. In this work, we propose a low-rank multilevel Monte Carlo estimator for kinetic equations based on a probabilistic rank-adaptive DLRA time integrator. The level hierarchy of the low-rank multilevel estimator is constructed through spatial refinement and by ensuring that the low-rank error remains below the spatial discretization error. We demonstrate the efficacy of the estimator through several numerical experiments from radiation transport, radiation therapy, and shallow water flow.
Electrons with Anomalous Energy Generated in Gas-Filled and Vacuum Diodes
arXiv:2606.27889v1 Announce Type: new Abstract: It is shown that, when high-voltage pulses (with a voltage amplitude exceeding 100 kV in centimeter gaps) with a leading edge duration of 1 ns or shorter are applied to gas-filled and vacuum electric discharge diodes, electrons with kinetic energies nominally exceeding the amplitude of the applied voltage are detected. In the experiments, electron beam attenuation curves were measured in absorbers consisting of Al foil of varying thicknesses. These curves were used to reconstruct the electron beam energy spectrum by regularizing the solution of an integral equation based on deep machine learning. The obtained spectra contain electrons with anomalously high energies, the proportion of which, depending on the conditions, can reach 25 percent. A control experiment with a long voltage pulse on a large-area vacuum diode (voltage 150 kV, pulse duration 35 microseconds, vacuum gap 12 cm, electrode area 75 x 15 cm2) showed that the proportion of electrons with anomalous energies is less than 0.2 percent. Experiments have shown that the main mechanism for generating electrons with anomalous energy is the spatio-temporal synchronism of the motion of fast electrons in the enhanced field formed in the gap by space charge.
Vibrational high-harmonics and period-doubling bifurcation probed by time-resolved electron diffraction
arXiv:2606.27891v1 Announce Type: new Abstract: Nanoscale mechanical oscillators exhibit a plethora of nonlinear phenomena with promising applications for the sensing and clocking of processes down to atomic length scales. Oscillator dynamics are typically probed by electrical or optical means, providing only limited access to the spatial profile of the oscillator motion. Here, we introduce event-based convergent beam electron diffraction for the spatio-temporal mapping of nanoscale mechanical resonators in ultrafast transmission electron microscopy. Employing an optically driven silicon membrane resonator at various driving strengths, we gain access to nonlinear processes with increasing complexity, ranging from a simple Duffing behavior to nonlinear multimode coupling and period-doubling bifurcations. The time-resolved diffraction probing approach supports a spatial resolution down to a few nanometers and a temporal resolution of 5 ns and provides quantitative information on the local membrane bending. Because the diffraction signal responds to local displacement gradients, which become more pronounced as resonators shrink, this approach offers a route toward probing nonlinear nanomechanics at the atomic scale.
Co-Optimization of Analog Kolmogorov-Arnold Networks for Low-Power Function Approximation in Flexible Electronics
arXiv:2606.27892v1 Announce Type: new Abstract: Wearable devices and Internet of Things (IoT) sensors require on-sensor processing of biosignals and environmental data, including computationally demanding operations such as nonlinear activation functions for neural network inference, sensor calibration curves to map raw readings to physical units, and signal preprocessing functions like logarithmic compression and power operations for feature extraction. These functions exhibit significant complexity, often involving transcendental operations and multivariate dependencies that are costly to implement digitally. Analog function approximation provides a power-efficient alternative by performing these computations in the analog domain, thereby reducing the energy overhead associated with analog-to-digital conversion and subsequent digital processing. Flexible Electronics (FE) present a particularly attractive platform for wearable applications due to mechanical flexibility and low-cost fabrication, but impose strict constraints on circuit density and power consumption, making efficient analog implementations critical but challenging. This work introduces Analog Kolmogorov-Arnold Networks (AKANs), developed via hardware-software co-optimization, to approximate these complex multivariate functions accurately under hardware imperfections. Our method incorporates circuit-level error modeling during training and applies pruning at both software and hardware levels to reduce area and power. Validation across multiple benchmarks demonstrates that our proposed pruning methodology not only reduces hardware cost but can also improve approximation accuracy by regularizing spline parameters. Results show up to 55% area and 50% power savings, with average reductions of nearly 30% across datasets, highlighting AKANs as a robust and generalizable framework for low-power analog function approximation in FE.
Mosaic: A Benchmark Suite for Differentiable Physics Solvers
arXiv:2606.27895v1 Announce Type: new Abstract: Differentiable partial differential equation (PDE) solvers underpin solver-in-the-loop ML training, gradient-based optimal control, and inverse problems, yet the practical cost of obtaining correct, usable gradients from a given solver on a given problem is largely undocumented. Integration effort, computational cost, gradient accuracy, and numerical conditioning vary widely across solvers and are discoverable only by trial and error. We introduce Mosaic, an extensible benchmarking framework for differentiable PDE solvers that standardizes access to solver gradients. Each solver is packaged as a containerized component (Tesseract) exposing a uniform gradient API regardless of language or automatic differentiation (AD) strategy, enabling researchers to evaluate, compare, and build on non-trivial physical solvers. Our evaluation of 14 solvers across fluid dynamics, structural mechanics, and heat transfer demonstrates that the benchmark surfaces practically relevant differences: order-of-magnitude variation in computational cost and Jacobian conditioning, alongside structural incompatibilities that eliminate solvers from realistic tasks entirely. Despite this variation, all solvers that produce gradients converge to similar optima, indicating that the practical barriers are memory limits, numerical stability, and setup compatibility rather than gradient accuracy alone. Mosaic is open-source and available at https://github.com/pasteurlabs/mosaic.
A Multi-Attribute Latent Space for Visual Analysis of Watches
arXiv:2606.27897v1 Announce Type: new Abstract: We present a design rationale, embedding model, and interactive visual-analysis system for exploring large wristwatch collections through heterogeneous visual and semantic attributes. The system addresses a common limitation of catalog and e-commerce interfaces: users can filter by metadata, but they receive little support for open-ended exploration of visual similarity, stylistic alternatives, and mixed aesthetic-functional criteria. We therefore represent watches with separate attribute graphs for dial color and dial design, while using watch type as an explicit semantic organizer. Dials are segmented with a U-Net, watch types are predicted with a Vision Transformer, colors are represented through a shared CIELAB reference palette, and dial structure is described with a gradient-based image descriptor. We extend UMAP by combining attribute-specific neighborhood graphs in a unified probabilistic objective and by adding a class-aware layout term that separates global type structure from local visual neighborhoods. The resulting map is exposed in an interactive interface with spatial navigation, metadata filtering, detail inspection, and search-by-example insertion. We evaluate the approach through parameter analysis, runtime measurements, and a qualitative pilot study with watch experts and novices. The results suggest that the system supports discovery and comparison, while also revealing limitations in scalability assessment, search-by-example validation, and the need for broader domain studies. We explicitly discuss these limitations and derive design implications for multi-attribute latent-space visualization across heterogeneous visual collections.
Drifting in the Future: Stabilizing Path Following Drifting on High-Latency Vehicle Systems
arXiv:2606.27914v1 Announce Type: new Abstract: Autonomously controlling and handling a vehicle at and beyond its stability limit is a mathematically and computationally demanding task. Prior demonstrations of automated drifting have been limited to research platforms with instantaneous torque delivery and independently actuated wheels, leaving their applicability to production vehicles with actuator latencies and mechanically coupled axles uncertain. To overcome these issues, we design a predictor to compensate for powertrain delays, develop a revised control formulation to accommodate higher actuation latencies as well as a differential coupling on the driven axle, and introduce brake-based velocity stabilization. This paper presents the controller framework, the model extensions, and real-world experimental results. We observe that our controller enables a production sports car with a combustion engine to robustly sustain circular and figure-eight drifts, limiting lateral error to 1.1 m and sideslip overshoot to 0.06 rad despite actuator delays exceeding 250 ms, while mitigating oscillations and maintaining stable path and sideslip tracking. In conclusion, our results establish that autonomous drifting is feasible on production-ready vehicles, opening pathways to advanced safety systems capable of stabilizing cars in scenarios where traditional control fails.
Every Step of the Way: Video-based Parkinsonian Turning Step Counting
arXiv:2606.27918v1 Announce Type: new Abstract: As a prominent symptom of Parkinson's disease (PD), turning impairment is evaluated through parameters such as turning angle, duration, and particularly, the number of steps required to complete a turn, which directly reflects motor dysfunction. Accurate step counting is challenging due to variability in real-world turning movements and atypical shuffling patterns in parkinsonian gait. Existing methods are predominantly wearable-based, requiring users to wear and manage dedicated devices, which can be inconvenient for continuous daily use. To address this, we propose a passive, video-based framework that estimates step count in a coarse-to-fine manner using diverse motion representations. Specifically, an initial step count is estimated from foot movement signals derived from 3D human mesh recovery, providing high-level motion structures. To incorporate fine-grained motion details, a motion encoder learns complementary gait dynamics from mesh and optical flow to refine the initial estimate. In this process, coarse foot movement signals query the pixel-level motion cues via cross attention to capture subtle parkinsonian gait dynamics. To handle varying video lengths, we partition each video into clips and integrate clip-wise motion embeddings via multiple instance learning (MIL) for step count residual prediction. Extensive experiments show our method consistently outperforms existing step counting methods on real-world PD turning datasets.
VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring
arXiv:2606.27941v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) provide useful decompositions of Transformer residual streams, but their learned features are usually named post hoc rather than directly connected to the Transformer's token vocabulary. We introduce Vocabulary-Aligned Sparse Autoencoder (VASAE), a method that trains SAE features under vocabulary-aligned anchoring and assigns each feature an intrinsic token name: the token string whose embedding is nearest to that feature. Without reducing reconstruction quality compared with a standard SAE, VASAE produces dictionaries with vocabulary-aligned features. Using a 0.8 cutoff on the nearest-token alignment score, dictionaries trained on GPT-2-small post-residual streams align about 90% of features in layers 0--10. In Llama-3.1-8B, representative shallow and middle-layer dictionaries contain strongly aligned features, including 92.8% in the shallow layer, while the representative final-layer dictionary shows limited alignment. After subtracting the sentence-level mean sparse code, case studies show that many remaining intrinsic token names are relevant to nearby input tokens. These results suggest that vocabulary-aligned anchoring can connect learned features to intrinsic token names during training, complementing post hoc interpretation of learned dictionaries.
It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents
arXiv:2606.27944v1 Announce Type: new Abstract: Phone-use Agents can execute complex tasks end to end across real mobile applications. By operating a real device on the user's behalf, they reach far more functionalities than CLI agents, which amplifies the real-world harm they can cause when driven for malicious purposes. We present the first study of this threat on real phones and 27 commercial apps, and find that agents built on 9 mainstream commercial and open-source models readily carry out serious misuse, ranging from procuring drug and explosive precursors to fraud, online harassment, and review manipulation. Across the agents we run on real devices, the average refusal rate to harmful requests stays low while the average task-completion rate reaches 68.8%, and in some scenarios an agent finishes a violation faster than a human would. These results suggest that Phone-use Agents already meet the practical conditions for automated misuse at scale. In one observed real-device execution, Claude-Opus-4.8 fabricated a medical history, deceived an online doctor into issuing a prescription, and completed the order and payment on its own to purchase a precursor for a highly toxic substance. To our knowledge, this is the first documented real-world case of an AI agent procuring controlled precursor materials. We trace this behavior to a Safety Awareness-Execution Gap, where an agent recognizes that a request is harmful yet still executes it. Simple defenses curb the overt cases, but the more covert and arguably more damaging threats, such as coordinated review manipulation and fake traffic, remain largely unsolved. We hope these findings push the community toward safer Phone-use Agents.
Understanding How MLLMs Describe Artworks Using Token Activation Maps
arXiv:2606.27947v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) describe artworks with remarkable fluency, yet the visual reasoning behind their outputs remains opaque. When an MLLM names a style, identifies a subject, or recognizes an iconographic symbol, does it ground each claim in the relevant region of the canvas, draw on an undifferentiated visual signal, or rely primarily on textual priors? We study this using the Token Activation Map (TAM), which produces, for each generated token, a heatmap isolating the visual evidence specific to that token from prior-context interference. Applying TAM to a curated set of paintings spanning multiple periods and genres, we analyze grounding patterns across five semantically distinct token categories: common visual objects, style descriptors, metadata, iconographic tokens, and affective expressions. We find that visual grounding varies substantially with token semantics. We further show that MLLMs attempt to identify artworks and artists, achieving higher accuracy in artist attribution than in title prediction, where hallucinations are more frequent. Finally, we compare TAM with SAM~3 open-vocabulary segmentation. To ensure reproducibility, we release our code, experimental configurations, prompts, and qualitative results on the project page at https://nicolafan.github.io/tamart/.
An Empirical Analysis of Factual Errors in Human-Written Text and its Application
arXiv:2606.27959v1 Announce Type: new Abstract: Factual Error Detection (FED), which is the task of identifying factually incorrect spans in a given text, has long been recognized as an important research problem. However, with the rapid rise of large language models (LLMs), research attention has shifted toward factual errors specific to LLM-generated text (hallucinations) and their detection. As a result, the detection of factual errors in human-written text has been relatively neglected. To address this gap, we first distill a taxonomy of human-induced factual errors by analyzing corrections of newspaper articles, a representative source of text that is guaranteed to be human-written and contains few grammatical errors. Our analysis revealed that there are characteristic categories such as kanji misconversions and numeral classifier errors, which are not focused in existing hallucination benchmarks. Based on the taxonomy, we then evaluate the FED capability of vanilla LLMs on synthesized realistic test cases and real corrections. Experimental results demonstrated that even high-performance LLMs such as GPT-5.4 achieved only word-level F1 score of 52% on the synthetic evaluation data, highlighting the task difficulty. Furthermore, a detailed analysis by detection difficulty revealed the current state of FED.
Decoys Cannot Go Everywhere: Mapping the Deception Surface in MITRE ATT&CK
arXiv:2606.27966v1 Announce Type: new Abstract: Cyber deception research often assumes that a decoy can be placed wherever there is attacker behavior. This work tests that assumption across MITRE ATT&CK v18.1. We introduce a four-criterion rubric for infrastructure deception and apply it to all 250 ATT&CK techniques. The rubric evaluates whether a defender-controlled decoy can be placed, whether an attacker is likely to interact with it, what intelligence that interaction can yield, and whether the interaction reliably indicates malice. The resulting deception surface is sparse: only 80 techniques (32%) admit a decoy the attacker could plausibly reach. For the remaining 170 techniques, there is no defender-controlled asset in the attacker's path that can be fabricated as a decoy. Decoy placement across those 80 techniques falls into two patterns we call Sweep and Seek. In Sweep, the attacker moves broadly through assets in range and encounters the decoy as part of that activity. In Seek, the attacker looks for a specific kind of asset and interacts with a fabricated version of it. These patterns give a simple placement rule: a decoy must either sit on a sweep path or imitate a sought asset. We also show that decoys usually have useful intelligence potential, but whether an attacker interacts with them at all, and whether that interaction reliably indicates malice, both vary. We release the rubric, decision rules, and per-technique assessment as an auditable baseline for future deception research and deployment planning, and show that infrastructure decoys cannot be assumed to apply to all attacker behavior.
Curriculum-guided Change Detection Training: Toward Accurate Serac Fall Monitoring
arXiv:2606.28012v1 Announce Type: new Abstract: Change Detection (CD) aims to identify semantic or structural changes from nearly registered multi-temporal images. While recent advances in training methodologies have largely focused on semi-supervised learning and consistency regularization, alternative training paradigms remain underexplored. In particular, most deep CD methods rely on uniform sampling during training, implicitly assuming that all training samples contribute equally to the optimization process. However, such naive sampling can introduce noisy gradients and hinder robust representation learning. To address this limitation, we propose a curriculum learning framework tailored for change detection. Our approach investigates two complementary difficulty measures: the Solar Angular Gap (SAG), a physically grounded proxy for acquisition-condition variability, and the Structural Similarity Index Measure (SSIM), which evaluates appearance similarity between image pairs. Based on these criteria, the framework progressively introduces challenging samples during training, enabling models to learn robust representations in a coarse-to-fine manner. We evaluate our method on the challenging SeracFallDet benchmark, where results demonstrate consistent improvements of the proposed approach over standard uniform-sampling strategies for both pixel-based and object-based approaches. These results highlight the potential of curriculum learning to improve robustness in deep change detection. Importantly, our training framework is orthogonal to existing CD architectures, making it readily applicable to a broad range of methods.
Effects of thermochemical modelling on a hypersonic shock-wave/turbulent boundary-layer interaction
arXiv:2606.28018v1 Announce Type: new Abstract: Thermochemical non-equilibrium can alter the structure, loads, and time scales of hypersonic shock-wave/turbulent boundary-layer interactions, yet its role in fully turbulent configurations remains largely unquantified. The present work addresses this issue by performing three direct numerical simulations of an oblique shock impinging on a turbulent high-enthalpy boundary layer at edge Mach number $M_e=6.4$ and stagnation enthalpy $H_e=16.9$ MJ/kg. The simulations share identical geometry and freestream conditions, but employ a hierarchy of progressively simplified thermochemical descriptions: a finite-rate reactive case, a single-species thermally perfect gas model, and a single-species calorically perfect model. The reactive simulation shows that the shock-induced temperature rise substantially enhances chemical activity relative to the incoming boundary layer, with peak concentrations of dissociation products attained downstream of the interaction. Thus, the thermal and chemical responses are not synchronised: the composition lags the rapid thermal forcing imposed by the shock system, and turbulent Damk\"ohler numbers reach values of order unity within the recirculation region, indicating non-negligible turbulence-chemistry interaction. The comparison among the three models shows that thermally and calorically perfect descriptions yield similar predictions, whereas finite-rate chemistry produces systematic differences: a smaller separation bubble, lower post-interaction wall heat flux, lower mean and fluctuating temperatures, and a less inclined reflected shock. In the present regime, the dominant modelling distinction is therefore between frozen and chemically reacting descriptions, with caloric-model effects playing only a secondary role.
Evolution-Aware Regression Test Prioritization of ML-Enabled Systems Using Gradient-Based Behavior Vectors
arXiv:2606.28037v1 Announce Type: new Abstract: The machine learning(ML) component of an ML-enabled system evolves through retraining, fine-tuning, and optimization, so previously valid test results may no longer hold. A single evolution step can worsen performance on some test cases while improving others, making regression test prioritization inherently directional. We present Gradient-based Behavior Vector-Parameter Delta(GBV-PD), the first approach to operationalize the behavior vector space for evolution-aware regression test prioritization. GBV-PD represents each test case as a gradient-based vector(GBV), a low-dimensional projection of its loss gradient under the original model. It then projects the observed parameter update of the evolved model onto the same PCA basis and uses the resulting alignment to estimate whether each test case's loss is likely to increase or decrease, without running the evolved model on test cases during prioritization. In an empirical study across classification and regression tasks, GBV-PD consistently outperformed non-directional baselines and remained competitive with a full-gradient reference, while offering better time and storage profiles for repeated updates via reusable GBV caching. These results show that behavior-space ideas can be operationalized into a practical and efficient mechanism for repeated-update regression testing of evolving ML-enabled systems.
RAMSES: Secure high-performance computing for sensitive data
arXiv:2606.27919v1 Announce Type: new Abstract: Traditionally, the architecture of high-performance computing (HPC) systems is tailored for speed, while highly secure computer systems must sacrifice speed for security. However, a wide range of scientific domains, such as the life sciences, call for a combination of performance and security to allow processing sensitive data at scale. Here, we present RAMSES (Research Accelerator for Modeling and Simulation with Enhanced Security), an HPC system designed from the ground up to deliver high performance within a robust security framework. RAMSES integrates hardware-based memory encryption of AMD processors with state-of-the-art file encryption from IBM Storage Scale and the Thales CipherTrust manager, establishing an HPC platform that ensures continuous encryption throughout the data life cycle - at rest, in transit, and in use - in compliance with major data protection standards (European General Data Protection Regulation, ISO/IEC 27001 certification, and Federal Information Processing Standards). In addition, we implemented advanced operating system hardening, a multi-layered security architecture, and mandatory multi-factor authentication to adapt the HPC environment to increased security demands. Benchmark results from the biomedical sector demonstrate that the performance impact of the secure environment is limited and that integration of the conflicting requirements speed and security can be achieved while preserving a coherent, flexible, and user-friendly system.
On the Relationship Between Plasma and Tritium Fuel Cycle Through Matter Injection and Particle Exhaust
arXiv:2606.28043v1 Announce Type: new Abstract: This work identifies an inconsistency between plasma operating scenarios and tritium fuel cycle (TFC) requirements, calling for a re-examination of the traditional reactor-led design approach. The key point is simple: in current TFC architectures, fuel puffing must contain tritium. Moscheni et al. (2026 Nucl. Fusion 66 026008) investigated fuel puffing rates in detached operation. Expanding that database, puffing is shown to exceed core fuelling by about an order of magnitude, from present-day tokamaks to next-step stellarators. Though not unknown in the plasma community, TFC models instead assumed core fuelling to dominate. The implications are severe. In recent TFC architectures, direct internal recycling (DIR) is intended to minimise tritium inventory, but assumes near-50:50 D:T composition. This assumption may become self-defeating: a substantial fraction of the puffed fuel must be tritium. Tritium inventory, doubling time, required breeding ratio, and pump sizing therefore become critical once puffing is properly accounted for. Mitigation is assessed by extending the models of Meschini et al. (2023 Nucl. Fusion 63 126005). For a notional plant, realistic TFC requirements can be met with D-rich, T-lean puffing, at the cost of about 10% lower fusion power. Alternatively, for near-50:50 D:T puffing, reduced fuel puffing with stronger impurity seeding can maintain detachment while alleviating TFC constraints, albeit with higher core contamination. Combined use of these strategies enables scenarios that minimise tritium inventory and throughput while balancing competing requirements. Ultimately, these results place renewed emphasis on the TFC as a central element of reactor design. A viable fusion reactor requires joint optimisation of core plasma, edge plasma, and TFC, implying unavoidable trade-offs across all three.
A Tree-of-Thoughts Inspired Hybrid Approach for Legal Case Judgement Summarization using LLMs
arXiv:2606.28044v1 Announce Type: new Abstract: In recent times, Large Language Models (LLMs) are increasingly being used for legal case judgement summarization. Most prior works have tried traditional extractive and abstractive summarization of case judgements. However, hybrid or extractive-abstractive techniques have not been explored much. In this work, we propose a novel tree-of-thoughts inspired extractive-abstractive summarization approach for legal judgement summarization. We conduct experiments using two popular LLMs, DeepSeek and LLama, and compare among extractive, abstractive and extractive-abstractive summarization. Our experiments show that the proposed extractive-abstractive prompt provides better summaries compared to other types of LLM prompts.
Rapid Prototyping of Event-Driven Contextual Memory in the ACT-Up Cognitive Architecture
arXiv:2606.28045v1 Announce Type: new Abstract: The present paper describes an implementation of contextual memory and a basic event-handler for the ACT-Up cognitive architecture which maintains its scalability and appropriateness for rapid-prototyping while adding essential features and lowering the barrier to entry for new users. This includes describing a theory-neutral implementation of working memory and spreading activation, in addition to a basic associative learning mechanism. An example of rapid prototyping for algorithm development is presented using the serial memory task described in Klein, Addis, and Kahana (2005). This study describes how contiguity effects change across sequential list presentations across three serial and free recall conditions. We further describe how to use generative AI and the event handler to automatically create cognitive experiments directly from the Methods section of research papers.
DG^VoiC: Speaker Clustering for Fraud Investigation under Real Call-Centre Conditions
arXiv:2606.28048v1 Announce Type: new Abstract: Insurance fraud remains costly and operationally difficult, particularly in call-centre workflows where many customer interactions begin at FNOL. While recent fraud detection methods mainly rely on structured data, text, or images, repeated speaker identity across calls remains underused as an investigative signal. This paper presents DG^VoiC, a voice clustering framework for customer verification and cross-profile speaker linking on anonymised real call-centre audio. The approach combines sensitive information-aligned anonymisation, speech-focused preprocessing, sliding-window speaker embedding extraction, and cosine similarity based clustering to identify repeated speakers under real telephony conditions. The method was evaluated on 121 recordings, with a curated reference subset of 56 samples in 22 human-agreed speaker clusters. used for validation. The best configuration achieved 96% AMI, 95% ARI, 98% completeness, 100% homogeneity, and 99% V-measure. These results show that speaker clustering can provide a strong additional signal for fraud investigation by helping analysts verify speaker consistency and surface repeated voices across customers.