Forskningsradar

Science Journals

Peer-reviewade publikationer — 60005 artiklar

The Signal-Coverage Matrix: Stratifying Type and Semantic Errors in Statement Autoformalization
arXiv:2606.28013v1 Announce Type: new Abstract: Headline type-correctness (TC\%) of LLM autoformalization has climbed from $\sim$53\% to $\sim$76\% in two years, yet this scalar conceals which errors each method resolves. We propose a signal-coverage matrix that crosses the Lean elaborator (pass/fail) with a semantic-equivalence judgment (equivalent/not), sorting every output into one of four cells: true success (TS), type-only (TO), semantic-only (SO), or both fail (BF). On ProofNet\# and MiniF2F-test with DeepSeek V4-Pro across Vanilla, Lean-Retry, Sample-Filter, and Stratified Autoformalization (SAF): (1) the +34 to +36 TS gain across the three elab-feedback methods is $\sim$64\% type-stratum recovery, with SO flat on net (87.5\% of original semantic errors rescued, 8 newly created). (2) The TO-to-TS rate is 23/61 for each method (Wilson 95\% CI [26.6\%, 50.3\%]), and this stratum-level recovery rate predicts $\Delta$TS on held-out methods to within 2/186 and renders $\Delta$TC linear in the Vanilla elab-fail rate across six (model, dataset) cells ($R^2=0.96$). (3) The two judges disagree by 26 to 37 pp on elab-feedback outputs (vs. 7 pp on Vanilla), with 30 to 56\% of symbolic-judge false negatives traceable to elaborator-forced rewrites. The persistent residual reduces to two gold-formalization errors. TC\% gains should be credited by which cell moved, not by the scalar alone.
Indecision and accuracy under social information across groups sizes
arXiv:2606.28025v1 Announce Type: new Abstract: Observing the decisions and actions of others provides social information that can inform decisions such as whether to follow. We consider a model where agents simultaneously gather stochastic private information, each deciding once sufficiently confident. Observed decisions and indecision provide social information that triggers discrete waves of collective response: a first decision causes others to update and potentially follow, whose decisions in turn provide further social information, generating successive waves. We explore this model across a range of group sizes and report three main findings. First, social information leads to faster and more accurate decisions than individual decision-making, but agent-level accuracy is maximised at a finite optimal group size. This contrasts with the accuracy of the majority choice, which increases monotonically with the number of agents. Second, waves frequently fail to resolve collective indecision, particularly for smaller groups and when the first decision is incorrect, leaving a subgroup of agents unconvinced. Third, these remaining undecided agents are systematically biased and make less accurate subsequent decisions, with this inaccuracy growing with group size.
EMOSH: Expressive Motion and Shape Disentanglement for Human Animation
arXiv:2606.28026v1 Announce Type: new Abstract: High-fidelity and expressive controllable human animation is essential for content creation and digital avatar applications. However, existing methods face a dilemma between expressiveness and disentanglement. Mainstream 2D pose-conditioned approaches suffer from "motion-shape entanglement", leading to the leakage of the driving subject's body shape. Conversely, methods relying on 3D priors (e.g., SMPL) achieve geometric disentanglement but struggle to capture facial expressions and complex gestures, resulting in rigid animations. To this end, we propose EMOSH, a novel framework for high-fidelity controllable human video generation. First, an Expressive Human Model (EHM) is introduced as the core control representation. By explicitly disentangling shape and pose parameters, we fundamentally resolve the body shape leakage issue. Alongside this, a robust motion tracker is designed to accurately estimate EHM parameters from video. Second, we propose a Coarse-to-Fine Hybrid Motion Injection strategy, enabling more fine-grained control over expressions and gestures. Furthermore, we introduce a Spatially-Aligned Conditioning mechanism to bridge the domain gap between training and inference, improving identity consistency. Extensive experiments demonstrate that EMOSH outperforms previous methods in both self-driven and cross-driven scenarios, producing high-fidelity videos with vivid expressions while maintaining shape disentanglement.
A Structure-Preserving Neural-Spectral Method for Reconstructing Controls of Wave Equations
arXiv:2606.28029v1 Announce Type: new Abstract: The numerical reconstruction of controls for partial differential equations remains comparatively underdeveloped, despite the extensive analytical literature on controllability. This difficulty is particularly pronounced for wave equations, whose conservative structure, oscillatory dynamics, and high-frequency behavior make direct discretization and optimization challenging. In this work, we introduce a Neural-Spectral method for approximating controls of wave equations. The method represents both the state and the control in a Dirichlet spectral basis and parameterizes the time-dependent modal coefficients using shallow neural networks. In this way, the spatial oscillatory structure of the wave equation is built into the approximation, and the learning task is reduced to reconstructing temporal coefficients. We prove approximation results showing that, under the standing assumption that an exact control exists in the relevant energy framework, the control-state pairs found can approximate exact controlled trajectories uniformly in time in the energy norm, while also approximating the corresponding controls in \(L^2\). We also state a conditional computable error estimate that separates spectral truncation, neural-network approximation, quadrature, and optimization errors. In addition, we discuss structural obstructions faced by standard time-stepping schemes for conservative wave dynamics: explicit Euler amplifies high frequencies, implicit Euler introduces artificial dissipation, and Crank--Nicolson preserves amplitudes but compresses high-frequency phases. Numerical experiments in one, two, and three space dimensions illustrate the method on nonlinear, linear-reference, and high-dimensional control benchmarks.
Performance Analysis and Optimal Design of ORB-Type GRAND Algorithms
arXiv:2606.28030v1 Announce Type: new Abstract: Guessing Random Additive Noise Decoding (GRAND) performs decoding by sequentially guessing channel error patterns (EPs). Ordered Reliability Bits GRAND (ORBGRAND) is a notable instance suitable for efficient implementation, as it schedules EPs solely according to the ranking of soft channel outputs. In this paper, we generalize this principle to a broader class of GRAND algorithms whose testing order depends only on reliability ranking, referred to as ORB-type GRAND. We develop a unified analytical framework based on a key quantity termed the average guessing posterior (AGP), which captures the effectiveness of each EP and reduces decoding into an ordering problem over the EP space. For random code ensembles, we derive exact expressions for the block error rate (BLER), stopping-time distribution, and average number of tests under a fixed test budget. The analysis separates target-miss and target-preemption errors and shows that ordering EPs by non-increasing AGP is optimal over the EP set under consideration. For fixed linear block codes, we derive the BLER expression that isolates the code-dependent target-preemption term and characterize this term through higher-order weight relationships of codeword tuples, with a computable first-order upper bound as a useful special case. Guided by these insights, we formulate ReShuffled-ORBGRAND (RS-ORBGRAND) as an offline AGP-based reshuffling scheme. Numerical results for the Bose--Chaudhuri--Hocquenghem (BCH)$(127,113)$ code show that RS-ORBGRAND consistently improves existing ORB-type GRAND algorithms and lies within $0.1$~dB of a maximum-likelihood decoding lower-bound benchmark at a BLER of $10^{-6}$.
A robust mixed finite element formulation for third medium contact
arXiv:2606.28036v1 Announce Type: new Abstract: Third medium contact provides a smooth continuum alternative to classical contact algorithms by replacing explicit contact constraints with a highly compliant fictitious medium. In this work, an auxiliary-field stabilization is introduced in which a deformation-gradient-like field is treated as an independent unknown in the third medium and coupled to the physical deformation gradient by a penalty term. A gradient contribution acting on the auxiliary field provides the regularization mechanism without requiring a direct evaluation of higher displacement derivatives. Linear and quadratic interpolation spaces are investigated, including continuous and element-wise discontinuous auxiliary-field approximations. The numerical results show that continuous low-order auxiliary fields provide an effective gradient-type stabilization of the third medium, even when the displacement field is approximated by first-order finite elements. For element-wise discontinuous auxiliary fields, the additional unknowns remain local to each element and can be eliminated locally by static condensation, so that the global system does not necessarily contain additional auxiliary degrees of freedom. Benchmark problems involving large deformation, progressive self-contact and severe third-medium compression are used to assess the formulation.
Context-Aware Explanations for Spatialized Document Layouts
arXiv:2606.28081v1 Announce Type: new Abstract: Spatialized document layouts are widely used for exploratory analysis of text corpora, but interpreting the spatial organization of documents and the relationships between regions remains challenging. Existing approaches primarily summarize document content or explain how layouts are generated, providing limited support for understanding spatial relationships within the layout itself. We present CAPE, a context-aware explanation framework that generates natural-language explanations grounded in both document semantics and layout-derived spatial context. CAPE identifies salient spatial patterns (e.g., clusters, subgroups, outliers, and bridging documents) and constructs multi-level contextual representations to guide LLM-based explanation generation. It supports both AI-guided overview and user-driven exploration, with explanations available at multiple levels of detail. We demonstrate CAPE on news and scholarly document layouts and evaluate it in a controlled user study against keyword-based and content-only LLM baselines. Our results suggest that spatially grounded explanations are perceived as more helpful than content-only baselines for interpreting the spatial organization of document layouts.
AB-Sync: Attention-Based Slot-Level Clock Synchronization Method for UWB-TDOA Localization Networks
arXiv:2606.28087v1 Announce Type: new Abstract: Ultra-wideband (UWB) time-difference-of-arrival (TDOA) localization networks provide high-update-rate indoor location services for IoT and cyber-physical applications, but their accuracy depends on nanosecond-level clock synchronization among anchors. Existing wireless clock synchronization (WCS) methods typically estimate clock states at the synchronization-stage or interval level, whereas TDMA-based UWB-TDOA systems localize tags from blinks transmitted in discrete short slots inside each synchronization stage. We identify this granularity mismatch as a source of residual TDOA error and present AB-Sync, an attention-based slot-level clock synchronization method. AB-Sync models the relationship between the slot-specific clock-speed ratio required by a target tag blink and neighboring clock-fluctuation observations, thereby enabling tag-slot-level timestamp mapping without adding extra UWB synchronization messages. On a real UWB-TDOA testbed, AB-Sync reduces the multi-anchor average TDOA ranging STD.V by 9.4% and improves representative static localization accuracy by 18.6% compared with Deferred+3S-KF, the leading low-overhead baseline in our evaluation. In a five-slot multi-tag experiment, AB-Sync consistently improves localization stability across all TDMA slots, reducing STD.V by 5.3% on average and up to 16.2% per slot with no extra UWB synchronization overhead.
Diffusion Model Attribution via Spectral Coupling of Denoiser Responses
arXiv:2606.28092v1 Announce Type: new Abstract: Attributing a generated image to its source diffusion model is a fundamental challenge in provenance verification and intellectual property protection. This problem is particularly difficult because diffusion models trained on different datasets can converge to similar score functions and thus similar output distributions, making the generated images themselves unreliable as attribution evidence. Existing non-invasive methods either fail on architecturally similar variants or rely on signals that vanish when models share the same autoencoder. We propose Spectral Denoising Signatures (SDS), a non-invasive attribution method that identifies the source model by fingerprinting each candidate model's denoising behavior. Our key insight is that a model's denoising score function exhibits a distinctive spectral geometry, reflected in how it redistributes energy across spatial frequency bands during denoising. By probing this behavior with frequency-controlled perturbations, SDS extracts a stable signature that is intrinsic to the model, requiring only standard forward passes with no inversion, optimization, or generation-time enrollment. Our results demonstrate that SDS achieves approximately 99.9% accuracy across eight diverse diffusion models and 96.2% under cross-domain prompt shift, outperforming non-invasive baselines across variations in training data, architecture, and training procedure, establishing spectral geometry as a principled and practical basis for diffusion model attribution. Code is available at: https://github.com/Pragati-Meshram/SGS
Fair Classification with Efficient and Post-hoc Controllable Fairness-Accuracy Trade-off
arXiv:2606.28097v1 Announce Type: new Abstract: Post-hoc controllability of fair machine learning models, the ability to control the trade-off between fairness and accuracy after training, is valuable for practical deployment. Existing post-processing methods provide such post-hoc controllability but often suffer from significant accuracy degradation, whereas in-processing methods achieve efficient trade-offs but require computationally expensive retraining for each change in trade-off ratio. To achieve both post-hoc controllability and efficient trade-offs, we propose a novel fair classification algorithm that learns effective feature representations to improve the trade-off efficiency of post-processing fair classifiers, by a gradient-based optimization approach. Experimental results on real-world datasets demonstrate that our method achieves trade-off efficiency comparable to, or even surpassing, in-processing methods, without requiring any retraining.
Cross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation Training
arXiv:2606.28104v1 Announce Type: new Abstract: Vision-based assessment can provide convenient and cost-effective evaluation in Traditional Chinese Medicine (TCM) rehabilitation training, where action quality assessment (AQA) from computer vision offers a promising solution. Existing automatic AQA frameworks for physical therapy typically rely on skeletal data captured from a single viewpoint, which is inefficient for TCM techniques such as acupuncture or Tuina that involve dense hand self-occlusion and complex hand-object interactions. To address these challenges, we propose CME-AQA, a cross-view, multimodal vision-based assessment framework that integrates visual-pose fusion to enhance understanding of environmental context and leverages both first-person and third-person videos during training to improve inference robustness. We collected two dual-view datasets, TCM-AQA61-A (Acupuncture) and TCM-AQA61-T (Tuina), each containing synchronized first-person and third-person recordings of 61 subjects with expert annotations. Experimental results show that our approach achieves superior or comparable mean performance against competitive baselines, achieving over 10% relative improvement in weighted F1 over the best competing method on key rating tasks such as Needle Depth and Quick Needle Insertion, while also reducing mean absolute error in quantitative measures such as insertion time and manipulation frequency. Testing on a CPR dataset further demonstrates comparable performance on several posture-based criteria, suggesting applicability to related structured simulated clinical skill assessments where participant motion is central to evaluation. Overall, CME-AQA enhances assessment accuracy for structured TCM rehabilitation training and facilitates more convenient and effective training-oriented skill evaluation.
BiDeMem: Bidirectional Degradation Memory for Explainable Image Restoration
arXiv:2606.28112v1 Announce Type: new Abstract: Degradation-aware prompts, conditions, and latent priors are increasingly used in image restoration, yet they are usually judged by a single endpoint: whether the restored image obtains higher PSNR. This is a weak test of semantics. A condition can help by adding capacity, acting as a global correction bias, or exploiting dataset shortcuts, without becoming an interpretable degradation prior. We propose BiDeMem, a bidirectional degradation memory for explainable image restoration. A query built from restoration features and input statistics retrieves a compact top-k subset of memory slots. The same selected slot identity supports the restoration path at inference time and a training-only forward-degradation explanation path. The study centers on verifiability in a controlled multi-degradation NAFNet setting. New controls separate the gain from a correction head alone, a dense query prior, and a static global prior: these variants are 0.2588, 0.2586, and 0.2839 dB below BiRank, respectively. Strong residual supervision and a wider degradation head also remain below the full bidirectional memory model. Intervention probes show that BiRank preserves restoration quality while increasing wrong-prior and native-prior sensitivity, framing degradation memory as both a restoration module and a falsifiable explanation mechanism.
Higher-Order Fourier Neural Operator: Explicit Mode Mixer for Nonlinear PDEs
arXiv:2606.28122v1 Announce Type: new Abstract: Neural operators provide deep neural networks for learning mappings between function spaces. Among them, the Fourier Neural Operator (FNO) is particularly effective: its spectral convolution relies on low-dimensional Fourier-domain representations and can handle inputs at different resolutions. This design aligns well with settings where the Fourier basis diagonalizes the underlying operator, such as linear, constant-coefficient PDEs on periodic domains, in which Fourier modes evolve independently. However, nonlinear PDEs may benefit from an additional inductive bias, as they exhibit structured interactions between modes, governed by polynomial nonlinearities. To capture this inductive bias, we introduce the Higher-Order Spectral Convolution, a spectral mixer that extends FNO from diagonal modulation to explicit n-linear mode mixing, aligned with the dynamics of nonlinear PDEs. Our experiments on standard benchmarks show that the proposed Higher-Order FNO (HO-FNO) retains the efficiency of FNO-based architectures and consistently improves over other spectral neural operators. HO-FNO also performs on par with or better than state-of-the-art transformers and state-space models on several datasets, with stronger gains in highly nonlinear regimes, such as the Poisson equation with polynomial forcing, where a single HO-FNO layer outperforms FNO models with up to 16 layers. We open-source our code for reproducibility at: https://github.com/AlexColagrande/HO-FNO.
PA-BiCoop: A Primary-Auxiliary Cooperative Framework for General Bimanual Manipulation
arXiv:2606.28192v1 Announce Type: new Abstract: Bimanual manipulation is essential for advanced robotic systems because it offers higher efficiency and flexibility compared to single-arm configurations. However, existing approaches either lack inter-arm interaction or ignore the need for a dynamic division of labor, treating the arms as functionally equivalent. To address these limitations, this paper draws inspiration from human bimanual manipulation where one arm handles core operations and the other provides auxiliary support, and proposes PA-BiCoop, a new single-model bimanual cooperation framework with dynamic primary-auxiliary arm differentiation. PA-BiCoop categorizes robotic arms into primary and auxiliary arms with adaptively adjustable roles across task stages, employs two specialized decoders that share a global feature encoder: the primary decoder generates the primary arm's base-coordinate pose and core-task affordance heatmaps, and the auxiliary decoder outputs the auxiliary arm's relative pose in the primary arm's coordinate system. Moreover, we design a dynamic role assignment module to automatically map roles to left/right arms without manual pre-definition. This design facilitates inter-arm knowledge sharing and coordinated manipulation. Extensive experiments demonstrate that our PA-BiCoop achieves superior performance: it outperforms state-of-the-art baselines by 48% on average in RLBench2 simulation tasks and by over 50% on average in real world tasks, thereby verifying its effectiveness and advancement in bimanual manipulation.
DexCompose: Reusing Dexterous Policies for Multi-Task Manipulation with a Single Hand
arXiv:2606.28323v1 Announce Type: new Abstract: Dexterous manipulation policies can solve individual skills, but composing them to perform multiple tasks with a single hand remains challenging. Adding a new task on top of an existing manipulation skill often imposes conflicting demands on overlapping fingers and contact modes, causing destructive interference between preserving an existing manipulation outcome and executing a new one. We propose DexCompose, a role-aware residual composition framework that reuses pretrained dexterous policies for multi-task manipulation through explicit finger-level action ownership. Given two pretrained full-hand policies, DexCompose first collects successful post-task states from the first skill and performs release tests over candidate finger masks to identify which fingers are necessary for maintaining the established skill state. It then trains two asymmetric residual modules: a bounded residual stabilizer for task preservation, and a context-aware residual that adapts the frozen downstream policy only within the action subspace assigned to the new task. We evaluate the framework on 16 composite dexterous manipulation tasks spanning four object-retention skills and four downstream interactions. DexCompose achieves a 77.4% average composite success rate, demonstrating that structural action ownership with dual residuals offers a promising direction for composing dexterous skills beyond conventional policy chaining.
DFM: Difference Feature Modeling with Text-Guided Gated Contrastive Loss for Remote Sensing Image Change Captioning
arXiv:2606.27410v1 Announce Type: cross Abstract: The primary goal of Remote Sensing Image Change Captioning (RSICC) is to automatically generate descriptions of changes between remote sensing images captured at different time points. Existing models still rely on a single autoregressive generation paradigm, which tends to prioritize learning easily generated vocabulary over capturing discriminative differences between images. To address this, we reframe the training paradigm and propose a novel Difference Feature Modeling (DFM) framework. Specifically, we introduce a Text-guided Gated Contrastive Loss (TGCL) to guide the vision encoder to extract critical features from a text-modal perspective. Additionally, we incorporate a pre-trained Change Detection model to transfer stable change detection knowledge. In order to further enhance the representation, we design a Joint Feature Modeling (JFM) module to achieve the fusion of multi-scale difference representations, thereby capturing comprehensive spatiotemporal variations between multi-temporal images. Extensive experiments on multiple datasets demonstrate the effectiveness of our approach.
Quantum Multi-Party Threshold Private Set Intersection with Explicit Cardinality Testing
arXiv:2606.27996v1 Announce Type: cross Abstract: Threshold private set intersection (TPSI) allows parties to reveal their intersection only when its cardinality reaches a prescribed threshold. Existing quantum TPSI protocols typically rely on a third party (TP) to interpret the final results, which deviates from the cardinality-testing paradigm of TPSI. In this paper, we propose a quantum multiparty TPSI protocol with explicit cardinality testing. Our protocol develops a rotation-based quantum construction in which single-photon sequences are sequentially processed through participant-side data rotations, TP--participant masking rotations, and correlated aggregate rotations. This design produces hidden-label measurement vectors: TP can complete the final measurement, but cannot interpret the semantic meaning of the outcomes. Based on these hidden measurements, we further realize the threshold decision through an oblivious linear evaluation (OLE)-based inner product procedure and a lightweight garbled circuit, revealing only \(\mathbf 1[|\bigcap_i X_i|\ge \tau]\) before conditional intersection reconstruction. We prove the correctness and security of the proposed protocol, and further validate its feasibility through quantum-circuit simulations implemented on the IBM \textsf{Qiskit} platform.
In-flight calibration of the Wide-field X-ray Telescope on board the Einstein Probe
arXiv:2606.28101v1 Announce Type: cross Abstract: By utilizing novel lobster-eye optics, the Wide-field X-ray Telescope (WXT) onboard the Einstein Probe (EP) satellite achieves an unprecedented combination of a large instantaneous field-of-view (FoV) and high sensitivity for monitoring the dynamic X-ray sky. In this paper, we present the in-orbit calibration results of the WXT during its first two and a half years of operations. By conducting observations of standard celestial sources--including the Crab Nebula, Scorpius X-1, and Cassiopeia A--we systematically characterized key instrumental properties. Our analysis demonstrates that the in-orbit performance of the WXT agrees with prelaunch ground calibrations well. The spatial resolution, denoted by the full width at half maximum (FWHM) of the focal spot, typically ranges from $3'$ to $6'$ across $\sim$90% of the FoV, with a median of $\sim 4.3'$. The post-calibration source positioning accuracy achieves $1.3'$ (at the 90% confidence level). The in-orbit effective area is consistent with model predictions and ground measurements, exhibiting an overall systematic uncertainty of $\lesssim 10\%$ (90% C.L.) in the 0.5-4 keV band. While the vast majority of the detectors remain highly stable, a noticeable long-term degradation at the low-energy end ($\sim30\%$-$40\%$, 0.4-0.6 keV) is observed in a few specific modules. Furthermore, spectral evaluations using Cas A confirm the stability of the energy scale and spectral resolution of the focal-plane Complementary Metal-Oxide Semiconductor (CMOS) detectors. All derived calibration products have been incorporated into the WXT calibration database (CALDB). These results comprehensively verify the instrumental capabilities of the WXT, providing a solid foundation for the reliable analysis of scientific observations.
Coexisting Regular and Chaotic Dynamics in the Dysprosium Feshbach Spectrum
arXiv:2606.28233v1 Announce Type: cross Abstract: Strongly dipolar gases, such as dysprosium, erbium and thulium, exhibit dense Feshbach spectra whose level statistics have been associated with quantum chaos arising from couplings among many molecular channels. Here, we combine a precise calibration of the Feshbach spectrum of $^{162}$Dy with spectroscopic measurements of the differential magnetic moments of bound states associated with more than 80 resonances between 0 and 30 G. These magnetic moments provide an eigenstate-sensitive probe of the molecular states underlying the resonance spectrum. We find that the level statistics are not uniform: resonances associated with states near the center of the magnetic-moment distribution display enhanced level repulsion, whereas those near the lower edge remain close to Poisson statistics. Our results reveal hidden structure within the chaotic dysprosium Feshbach spectrum and identify molecular-state composition as a key ingredient in the emergence of quantum chaos in strongly dipolar scattering.
Composing Quantum Instruments
arXiv:2606.28291v1 Announce Type: cross Abstract: We study the composition of classically-controlled quantum instruments--the natural quantum analogue of Markov kernels. Classically, Markov kernels compose by integrating one kernel against another. Defining this composition for quantum instruments with continuous outcomes requires an integral of quantum channel-valued functions with respect to a quantum instrument. We construct this integral in the Heisenberg picture using the Okamura-Ozawa normal extension to a von Neumann tensor product. This integral recovers the expected finite formula, preserves normal complete positivity and subunitality, and provides the multiplication for a monad governing the composition of quantum instruments. As an immediate consequence, we identify the category of quantum Markov kernels as the Kleisli category of this monad.
Continual Memorization of Factoids in Language Models
arXiv:2411.07175v3 Announce Type: replace Abstract: As new knowledge rapidly accumulates, language models (LMs) with pretrained knowledge quickly become obsolete. A common approach to updating LMs is fine-tuning them directly on new knowledge. However, recent studies have shown that fine-tuning for memorization may be ineffective in storing knowledge or may exacerbate hallucinations. In this work, we introduce a setting we call continual memorization, where a model must memorize and retain a set of factoids through multiple stages of fine-tuning on subsequent datasets. We characterized the forgetting patterns through extensive experiments and show that LMs widely suffer from forgetting, especially when needing to memorize factoids in the second stage. We posit that forgetting can be alleviated by modifying training dynamics: (1) protecting the memorization process when learning factoids or (2) reducing interference from subsequent training stages. Intriguingly, we find that mixing randomly generated word sequences or generic data sampled from pretraining corpora at different training stages effectively mitigates forgetting REMIX: Random and Generic Data Mixing). REMIX can recover performance from severe forgetting, outperforming replay methods and other continual learning baselines. We analyze how REMIX influences the learning process and find that robust memorization follows a distinct pattern: the model stores factoids in earlier layers than usual and diversifies the layers that retain them, which results in easier recall and manipulate of the learned factoids.
Major Space Weather Risks Identified via Coupled Physics-Engineering-Economic Modeling
arXiv:2412.18032v3 Announce Type: replace Abstract: Space weather poses an important but under-quantified threat to society. While severe geomagnetic storms are recognized as potential global catastrophes, their socio-economic impacts remain poorly quantified. We present a novel physics-engineering-economic framework that links geophysical drivers to power grid geoelectric fields, transformer vulnerability, and macroeconomic consequences. Using the United States as an example, we estimate daily U.S. economic losses for a 250-year geomagnetic storm from transformer thermal heating of 2.04 billion USD (95 percent confidence interval: 1.86 to 2.22 billion USD), disrupting power for approximately 5.7 million people and 150,000 businesses. These estimates are conservative lower bounds, reflecting only transformer thermal heating effects and excluding voltage collapse, cascading failures, and restoration costs. The true societal risk is likely substantially higher. Nonetheless, the contribution is in providing the first nationwide end-to-end coupling from space physics to potential macroeconomic loss, with quantified uncertainties. Our results demonstrate that coupled socio-economic modeling of space weather is both feasible and essential, and the framework is scalable and transferable, offering a template for assessing space weather risk to critical infrastructure in other countries.
CO-DEFEND: Continuous Decentralized Federated Learning for Secure DoH-Based Threat Detection
arXiv:2504.01882v2 Announce Type: replace Abstract: The use of DNS over HTTPS (DoH) tunneling by an attacker to hide malicious activity within encrypted DNS traffic poses a serious threat to network security, as it allows malicious actors to bypass traditional monitoring and intrusion detection systems while evading detection by conventional traffic analysis techniques. ML techniques can be used to detect DoH tunnels; however, their effectiveness relies on large datasets containing both benign and malicious traffic. Sharing such datasets across entities is challenging due to privacy concerns. In this work, we propose CO-DEFEND framework that enables multiple entities to collaboratively train a classification machine learning model for DoH threat detection while preserving data privacy, enhancing scalability and resilience against single points of failure. The proposed DFL framework provides a realistic implementation for DoH threat detection, enabling multiple entities to train their local models online with incoming DoH flows in real-time batches as they are processed - an approach that fits naturally within modern Internet architectures. This framework adapts four classical machine learning algorithms, Support Vector Machines, Logistic Regression, Decision Trees, and Random Forest, for federated scenarios and efficient training. In addition, a key methodological feature of CO-DEFEND is the use of DT and RF as model selection rather than aggregation mechanisms, allowing each participant to retain interpretable and locally optimal decision structures while benefiting from collective updates. We compare our proposed method by using the dataset CIRA-CIC-DoHBrw-2020 with existing machine learning approaches, including more computationally complex alternatives such as neural networks, to demonstrate its effectiveness in detecting malicious DoH tunnels while improving scalability and computational efficiency.
Usability Testing of an Explainable AI-enhanced Tool for Clinical Decision Support: Insights from the Reflexive Thematic Analysis
arXiv:2504.04703v2 Announce Type: replace Abstract: Artificial intelligence-augmented technology represents a considerable opportunity for improving healthcare delivery. Significant progress has been made to demonstrate the value of complex models to enhance clinicians` efficiency in decision-making. However, the clinical adoption of such models is scarce due to multifaceted implementation issues, with the explainability of AI models being among them. One of the substantially documented areas of concern is the unclear AI explainability that negatively influences clinicians` considerations for accepting the complex model. With a usability study engaging 20 U.S.-based clinicians and following the qualitative reflexive thematic analysis, this study develops and presents a concrete framework and an operational definition of explainability. The framework can inform the required customizations and feature developments in AI tools to support clinicians` preferences and enhance their acceptance.
VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization
arXiv:2507.17455v2 Announce Type: replace Abstract: Geo-localization from a single image at planet scale (essentially an advanced or extreme version of the kidnapped robot problem) is a fundamental and challenging task in applications such as navigation, autonomous driving and disaster response due to the vast diversity of locations, environmental conditions, and scene variations. Traditional retrieval-based methods for geo-localization struggle with scalability and perceptual aliasing, while classification-based approaches lack generalization and require extensive training data. Recent advances in vision-language models (VLMs) offer a promising alternative by leveraging contextual understanding and reasoning. However, while VLMs achieve high accuracy, they are often prone to hallucinations and lack interpretability, making them unreliable as standalone solutions. In this work, we propose a novel hybrid geo-localization framework that combines the strengths of VLMs with retrieval-based visual place recognition (VPR) methods. Our approach first leverages a VLM to generate a prior, effectively guiding and constraining the retrieval search space. We then employ a retrieval step, followed by a re-ranking mechanism that selects the most geographically plausible matches based on feature similarity and proximity to the initially estimated coordinates. We evaluate our approach on multiple geo-localization benchmarks and show that it consistently outperforms prior state-of-the-art methods, particularly at street (up to 4.51%) and city level (up to 13.52%). Our results demonstrate that VLM-generated geographic priors in combination with VPR lead to scalable, robust, and accurate geo-localization systems.