Forskningsradar

Science Journals

Peer-reviewade publikationer — 55483 artiklar

Beyond Memorization: Distinguishing Between Pattern-Based and Epistemic Reasoning in LLMs Using Epistemic Puzzles
arXiv:2603.21350v3 Announce Type: replace Abstract: Epistemic reasoning requires agents to infer the state of the world from partial observations and information about other agents' knowledge. Prior work evaluating LLMs on epistemic puzzles often frames failures as memorization rather than reasoning. We argue that this dichotomy is too coarse for newer models: memorization is a limiting case of pattern-based reasoning, where a model matches a task to a familiar template and applies the corresponding solution. We introduce a two-dimensional benchmark over DEL-style puzzles, separating narrative familiarity from inference complexity, allowing us to distinguish pattern-based from epistemic reasoning. We find that models are substantially more robust to surface form changes than prior work suggested, yet consistently struggle in asymmetric settings where familiar patterns no longer apply and success requires tracking fragmented epistemic states.
SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices
arXiv:2601.08303v3 Announce Type: replace Abstract: Recent advances in diffusion transformers (DiTs) have set new standards in image generation, yet remain impractical for on-device deployment due to their high computational and memory costs. In this work, we present an efficient DiT framework tailored for mobile and edge devices that achieves transformer-level generation quality under strict resource constraints. Our design combines three key components. First, we propose a compact DiT architecture with an adaptive global-local sparse attention mechanism that balances global context modeling and local detail preservation. Second, we propose an elastic training framework that jointly optimizes sub-DiTs of varying capacities within a unified supernetwork, allowing a single model to dynamically adjust for efficient inference across different hardware. Finally, we develop Knowledge-Guided Distribution Matching Distillation, a step-distillation pipeline that integrates the DMD objective with knowledge transfer from few-step teacher models, producing high-fidelity and low-latency generation (e.g., 4-step) suitable for real-time on-device use. Together, these contributions enable scalable, efficient, and high-quality diffusion models for deployment on diverse hardware.
VesselTok: Tokenizing Vessel-like 3D Biomedical Graph Representations for Reconstruction and Generation
arXiv:2603.18797v2 Announce Type: replace Abstract: Spatial graphs provide a lightweight and elegant representation of curvilinear anatomical structures such as blood vessels, lung airways, and neuronal networks. Accurately modeling these graphs is crucial in clinical and (bio-)medical research. However, the high spatial resolution of large networks drastically increases their complexity, resulting in significant computational challenges. In this work, we aim to tackle these challenges by proposing VesselTok, a framework that approaches spatially dense graphs from a parametric shape perspective to learn latent representations (tokens). VesselTok leverages centerline points with a pseudo radius to effectively encode tubular geometry. Specifically, we learn a novel latent representation conditioned on centerline points to encode neural implicit representations of vessel-like, tubular structures. We demonstrate VesselTok's performance across diverse anatomies, including lung airways, lung vessels, and brain vessels, highlighting its ability to robustly encode complex topologies. To prove the effectiveness of VesselTok's learnt latent representations, we show that they (i) generalize to unseen anatomies, (ii) support generative modeling of plausible anatomical graphs, and (iii) transfer effectively to downstream inverse problems, such as link prediction.
MOSAIC: Modular Scalable Autonomy for Intelligent Coordination of Heterogeneous Robotic Teams
arXiv:2601.23038v3 Announce Type: replace Abstract: Mobile robots have become indispensable for exploring hostile environments, such as in space or disaster relief scenarios, but often remain limited to teleoperation by a human operator. This restricts the deployment scale and requires near-continuous low-latency communication between the operator and the robot. We present MOSAIC: a scalable autonomy framework for multi-robot scientific exploration using a unified mission abstraction based on Points of Interest (POIs) and multiple layers of autonomy, enabling supervision by a single operator. The framework dynamically allocates exploration and measurement tasks based on each robot's capabilities, leveraging team-level redundancy and specialization to enable continuous operation. We validated the framework in a space-analog field experiment emulating a lunar prospecting scenario, involving a heterogeneous team of five robots and a single operator. Despite the complete failure of one robot during the mission, the team completed 82.3% of assigned tasks at an Autonomy Ratio of 86%, while the operator workload remained at only 78.2%. These results demonstrate that the proposed framework enables robust, scalable multi-robot scientific exploration with limited operator intervention. We further derive practical lessons learned in robot interoperability, networking architecture, team composition, and operator workload management to inform future multi-robot exploration missions.
Metric Differential Privacy at the User-Level Via the Earth Mover's Distance
arXiv:2405.02665v3 Announce Type: replace Abstract: Metric differential privacy (DP) provides heterogeneous privacy guarantees based on a distance between the pair of inputs. It is a widely popular notion of privacy since it captures the natural privacy semantics for many applications (such as, for location data) and results in better utility than standard DP. However, prior work in metric DP has primarily focused on the item-level setting where every user only reports a single data item. A more realistic setting is that of user-level DP where each user contributes multiple items and privacy is then desired at the granularity of the user's entire contribution. In this paper, we initiate the study of one natural definition of metric DP at the user-level. Specifically, we use the earth-mover's distance ($d_\textsf{EM}$) as our metric to obtain a notion of privacy as it captures both the magnitude and spatial aspects of changes in a user's data. We make three main technical contributions. First, we design two novel mechanisms under $d_\textsf{EM}$-DP to answer linear queries and item-wise queries. Specifically, our analysis for the latter involves a generalization of the privacy amplification by shuffling result which may be of independent interest. Second, we provide a black-box reduction from the general unbounded to bounded $d_\textsf{EM}$-DP (size of the dataset is fixed and public) with a novel sampling based mechanism. Third, we show that our proposed mechanisms can provably provide improved utility over user-level DP, for certain types of linear queries and frequency estimation.
Hyperparameter Transfer in Graph Neural Networks
arXiv:2607.05017v1 Announce Type: new Abstract: The performance of deep learning models crucially depends on the settings of hyperparameters like learning rate, initialization scale, and weight decay. Hyperparameter transfer aims to make near-optimal hyperparameter settings consistent across model scale, so that large models can be optimized by proxy tuning their smaller, cheaper-to-optimize counterparts. While transfer principles are well-studied in the context of dense neural networks in language and vision tasks, they remain comparatively under-explored for graph neural networks (GNNs). We develop and validate a transfer parameterization for GNNs trained with SGD, Adam, and AdamW. Through theoretical scaling analyses and controlled experiments, we show that the proposed parameterization yields stable feature updates, learning rate transfer, and improved performance as width and depth increase. For SGD, we identify graph-dependent first-layer correction factors and show that their use can accelerate early training in graphs with sparse bag-of-words inputs. For Adam, we explore how different message passing normalizations affect early- and late-training transfer behavior, illustrating the importance of message passing normalization and advocating for an associated hyperparameter. For AdamW, we adapt a parameterization that allows for the joint transfer of weight decay and learning rate. Together, these results provide a practical recipe for scaling GNNs across a variety of learning tasks and training scenarios.
A Simulation Framework for Electromagnetic Signal Injection Attacks on Image Sensors
arXiv:2408.05124v2 Announce Type: replace Abstract: Image sensors are fundamental to many intelligent systems, allowing visual perception and AI-driven decision-making. However, their integrity can be compromised by electromagnetic signal injection attacks (ESIA), which manipulate captured images without modifying sensor hardware or software. Despite the growing threat, system-level understanding of the attacks, as well as the development of defenses, remains limited, in part because collecting adversarial data is often complex and requires specialized attack setups. To address this challenge, we model ESIA and develop a simulation framework for generating synthetic adversarial images. Our analysis shows that these synthetic images are statistically indistinguishable from those produced by real attacks. The proposed framework enables faster vulnerability evaluation of computer vision (CV) algorithms, without the need for dedicated attack hardware. We also present a pilot study showing that the robustness of the algorithms can be improved by adversarial training, demonstrating a practical and scalable path toward mitigating ESIA threats.
Proofdoors and Efficiency of CDCL Solvers
arXiv:2603.26286v3 Announce Type: replace Abstract: We propose a new parameter called proofdoor in an attempt to explain the efficiency of CDCL SAT solvers over a certain class of formulas derived from circuit (esp., arithmetic) verification applications. Informally, given an unsatisfiable CNF formula F over n variables, a proofdoor decomposition consists of a chunking of the clauses into A_1, ..., A_k together with a sequence of interpolants connecting these chunks. Intuitively, a proofdoor captures the idea that an unsatisfiable formula can be refuted by reasoning chunk by chunk, while maintaining only a summary of the information (i.e., interpolants) gained so far for subsequent reasoning steps. We prove several theorems in support of the proposition that proofdoors can explain the efficiency of CDCL solvers for some class of circuit verification problems. First, we show that formulas with small proofdoors (i.e., where each interpolant is O(n) sized, each chunk A_i has small pathwidth, and each interpolant clause has at most O(log n) backward dependency on the previous interpolant) have short resolution (Res) proofs, and a certain configuration of CDCL solvers can compute such proofs in time polynomial in n. Second, we show that commutativity (miter) formulas over floating-point addition have small proofdoors and hence short Res proofs, even though they have large pathwidth. Third, we identify limits of the proofdoor framework: we show that a poor decomposition of arithmetic miter instances can force exponentially large interpolants, and hence our framework derives exponentially large Res refutations from such a decomposition, even when a different decomposition (i.e., a small proofdoor) yields short proofs. As a byproduct, these interpolant lower bounds imply new lower bounds for the partially ordered resolution proof system.
Modeling Cell-Cycle-Aware Single-Cell Drug Perturbation Responses
arXiv:2606.30695v2 Announce Type: replace-cross Abstract: Single-cell drug perturbation models should capture transcriptional response magnitude and whether a treatment changes the proliferative state of the cell. This is difficult because cell-cycle variation is often treated as a nuisance factor, and benchmark processing rarely makes drug-induced phase changes a primary prediction target. We introduce scCycleMol, a cell-cycle-aware perturbation prediction framework built on a curated 24-hour SciPlex3 benchmark with standardized molecule identities, dose and cell-line metadata, modeled genes, and expression-derived cell-cycle supervision. scCycleMol derives cell-cycle supervision from the treated state and applies it to predicted treated expression without using phase as an input covariate. The model includes a learnable full-expression cell-cycle head with circular G1/S/G2M targets, and we evaluate readout-only supervision (with stop-gradient) versus closed-loop supervision (backpropagating through decoder, dose-response module, and drug representation). We also compare molecular representations and pretraining sources to isolate the effect of the cell-cycle objective. On a processed 24-hour SciPlex3 benchmark (635,541 cells, 186 perturbations, 188 compound embeddings, 3 cell lines, 4 doses plus DMSO, 5,080 genes), the best LINCS-pretrained circular variant reaches 0.9093 mean all-gene R-squared and 0.6843 mean DE-gene R-squared. Under matched preprocessing, closed-loop cell-cycle supervision improves phase accuracy by 0.54-0.62 points while keeping mean all-gene R-squared within 0.003 of matched chemCPA no-cell-cycle models; Tahoe-pretrained readout-only circular supervision achieves the strongest phase accuracy at 0.9609.
A rechargeable AA battery supporting Qi wireless charging
arXiv:2408.09852v3 Announce Type: replace Abstract: Wireless power transfer is one of the key drivers in modern consumer electronics, as it allows one to enhance the convenience and usability of many devices. However, in most cases, wireless charging is accessible only to devices with incorporated receivers or at least to gadgets with standard charging connectors, such as USB Type-C, that allow to attach an external receiver. We propose a rechargeable battery that has the size and output voltage of a standard AA battery but supports wireless power transfer from charging stations of the widely used Qi standard. The proposed design uses a series resonant circuit with a curved receiving coil, as well as load modulation using detuning capacitors switched by a microcontroller unit to implement a receiver compatible with the Qi Baseline protocol. It also utilizes a number of DC-DC converters to store energy in a Li-ion cell and convert it to the 1.5 V voltage level. Our design is supported by numerical simulations of magnetic field distributions and scattering parameters of the introduced battery coupled to a planar transmitting coil. The performance of the proposed battery has been studied experimentally, including measurements of the maximal distance between the battery and a charging station that allows wireless charging at various rotation angles and the charge curve. The developed battery design facilitates the addition of wireless charging functionality to a wide range of electronic devices in a universal way.
AnyDesign: Versatile Area Fashion Editing via Mask-Free Diffusion
arXiv:2408.11553v5 Announce Type: replace Abstract: Fashion image editing aims to modify a person's appearance based on a given instruction. Existing methods require auxiliary tools like segmenters and keypoint extractors, lacking a flexible and unified framework. Moreover, these methods are limited in the variety of clothing types they can handle, as most datasets focus on people in clean backgrounds and only include generic garments such as tops, pants, and dresses. These limitations restrict their applicability in real-world scenarios. In this paper, we first extend an existing dataset for human generation to include a wider range of apparel and more complex backgrounds. This extended dataset features people wearing diverse items such as tops, pants, dresses, skirts, headwear, scarves, shoes, socks, and bags. Additionally, we propose AnyDesign, a diffusion-based method that enables mask-free editing on versatile areas. Users can simply input a human image along with a corresponding prompt in either text or image format. Our approach incorporates Fashion DiT, equipped with a Fashion-Guidance Attention (FGA) module designed to fuse explicit apparel types and CLIP-encoded apparel features. Both Qualitative and quantitative experiments demonstrate that our method delivers high-quality fashion editing and outperforms contemporary text-guided fashion editing methods.
On sampling diluted Spin-Glasses with unbounded interactions
arXiv:2603.22432v2 Announce Type: replace Abstract: Spin-glasses are natural Gibbs distributions that have been studied in Theoretical CS for many decades. Recently, they have been gaining attention from the community as they emerge naturally in neural computation and learning, network inference, optimisation and other areas. We study the problem of efficiently sampling from spin-glass distributions when the underlying graph is a typical instance of $G(n,d/n)$, i.e., the random graph on $n$ vertices such that each edge appears independently with probability $d/n$, and $d=\Theta(1)$. Our focus is on the 2-spin model at inverse temperature $\beta$. We consider this distribution to be one of the most interesting case of spin-glasses, and one of the most challenging to analyse, since its Gaussian couplings give rise to unbounded interaction. We employ the well-known Glauber dynamics to sample from the aforementioned distribution. We show that for the typical instances of the 2-spin model on $G(n,d/n)$, the mixing time of Glauber dynamics is $O\left(n^{1+\Theta(\frac{1}{\sqrt{d}})}\right)$, for any $\beta\leq \frac{1}{4\sqrt{d}}$. Our results can also be adapted for the case of spin-glass distributions with bounded interactions. In that respect, we obtain rapid mixing of Glauber dynamics for the Viana-Bray model on $G(n,d/n)$ when $\beta\leq \frac{1}{4\sqrt{d}}$. This improves on the current best bound which is $\beta<\frac{0.18}{\sqrt{d}}$. We utilise stochastic localisation, and in particular, we build and improve on the scheme introduced in [Liu, Mohanty, Rajaraman and Wu: FOCS 2024]. This is the first time that stochastic localisation is used for diluted spin-glasses, where both degrees and interactions can be unbounded.
LivingWorld: Interactive 4D World Generation with Environmental Dynamics
arXiv:2604.01641v2 Announce Type: replace Abstract: We introduce LivingWorld, an interactive framework for generating 4D worlds with environmental dynamics from a single image. While recent advances in 3D scene generation enable large-scale environment creation, most approaches focus primarily on reconstructing static geometry, leaving scene-scale environmental dynamics such as clouds, water, or smoke largely unexplored. Modeling such dynamics is challenging because motion must remain coherent across an expanding scene while supporting low-latency user feedback. LivingWorld addresses this challenge by progressively constructing a globally coherent motion field as the scene expands. To maintain global consistency during expansion, we introduce a geometry-aware alignment module that resolves directional and scale ambiguities across views. We further represent motion using a compact hash-based motion field, enabling efficient querying and stable propagation of dynamics throughout the scene. This representation also supports bidirectional motion propagation during rendering, producing long and temporally coherent 4D sequences without relying on expensive video-based refinement. On a single RTX 5090 GPU, generating each new scene expansion step requires 9 seconds, followed by 3 seconds for motion alignment and motion field updates, enabling interactive 4D world generation with globally coherent environmental dynamics. Video demonstrations are available at paper.pnu-cvsp.com/LivingWorld.
SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
arXiv:2602.06358v3 Announce Type: replace Abstract: We propose SHINE (Scalable Hyper In-context NEtwork), a scalable hypernetwork that can map diverse meaningful contexts into high-quality LoRA adapters for large language models (LLMs). By reusing the frozen LLM's own parameters in an in-context hypernetwork design and introducing architectural innovations, SHINE overcomes key limitations of prior hypernetworks and achieves strong expressive power with a relatively small number of parameters. We introduce a pretraining and instruction fine-tuning pipeline, and train our hypernetwork to generate high quality LoRA adapters from diverse meaningful contexts in a single forward pass. It updates LLM parameters without any fine-tuning, and immediately enables complex question answering tasks related to the context without directly accessing the context, effectively transforming in-context knowledge to in-parameter knowledge in one pass. Our work achieves outstanding results on various tasks, greatly saves time, computation and memory costs compared to SFT-based LLM adaptation, and shows great potential for scaling. Our code is available at https://github.com/MuLabPKU/SHINE
Human Thinking under Plural LLM Assistance: Mathematical Problem Solving and Open-Ended Writing
arXiv:2604.02677v2 Announce Type: replace Abstract: Large language models are changing not only the kind of assistance people receive, but also how that assistance is organized. Instead of working with a single general-purpose chatbot, people can now receive help from systems arranged as peers, specialists, or multiple agents with distinct roles. However, it remains unclear how these forms of plural LLM assistance affect human performance, confidence, and diversity of thought. We conducted two controlled experiments involving 562 participants to examine the effects of using multiple LLMs on mathematical problem-solving and writing. In a math task, participants worked with no LLM, an expert assistant, peer-like agents that surfaced common errors, or both an expert and a peer-like assistant. The expert-plus-peer condition produced the strongest unassisted post-task performance. In a writing task, participants wrote with no LLM, a single generalist assistant, or a pair of role-specialized assistants. LLM assistance improved essay quality, but the role-specialized pair preserved greater idea diversity than the single assistant. Together, these findings identify the arrangement of LLM assistance as a consequential design variable for human-AI collaboration.
Toward Efficient Agents: Memory, Tool learning, and Planning
arXiv:2601.14192v2 Announce Type: replace Abstract: Recent years have witnessed increasing interest in extending large language models into agentic systems. While the effectiveness of agents has continued to improve, efficiency, which is crucial for real-world deployment, has often been overlooked. This paper therefore investigates efficiency from three core components of agents: memory, tool learning, and planning, considering costs such as latency, tokens, steps, etc. Aimed at conducting comprehensive research addressing the efficiency of the agentic system itself, we review a broad range of recent approaches that differ in implementation yet frequently converge on shared high-level principles including but not limited to bounding context via compression and management, designing reinforcement learning rewards to minimize tool invocation, and employing controlled search mechanisms to enhance efficiency, which we discuss in detail. Accordingly, we characterize efficiency in two complementary ways: comparing effectiveness under a fixed cost budget, and comparing cost at a comparable level of effectiveness. This trade-off can also be viewed through the Pareto frontier between effectiveness and cost. From this perspective, we also examine efficiency oriented benchmarks by summarizing evaluation protocols for these components and consolidating commonly reported efficiency metrics from both benchmark and methodological studies. Moreover, we discuss the key challenges and future directions, with the goal of providing promising insights.
Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation
arXiv:2601.14788v2 Announce Type: replace Abstract: Diffusion models have seen widespread adoption for text-driven human motion generation and related tasks due to their impressive generative capabilities and flexibility. However, current motion diffusion models face two major limitations: a representational gap caused by pre-trained text encoders that lack motion-specific information, and error propagation during the iterative denoising process. This paper introduces Reconstruction-Anchored Diffusion Model (RAM) to address these challenges. First, RAM leverages a motion latent space as intermediate supervision for text-to-motion generation. To this end, RAM co-trains a motion reconstruction branch with two key objective functions: self-regularization to enhance the discrimination of the motion space and motion-centric latent alignment to enable accurate mapping from text to the motion latent space. Second, we propose Reconstructive Error Guidance (REG), a testing-stage guidance mechanism that exploits the motion diffusion model's inherent self-correction ability to mitigate error propagation. At each denoising step, REG uses the motion reconstruction branch to reconstruct the previous estimate, reproducing the prior error patterns. By amplifying the residual between the current prediction and the reconstructed estimate, REG highlights the improvements in the current prediction. Extensive experiments demonstrate that RAM achieves significant improvements and state-of-the-art performance. Our code will be released.
Complexity of Normalized Persistence Problems for Topological Data Analysis and Local Hamiltonians
arXiv:2607.03278v1 Announce Type: cross Abstract: Topological data analysis (TDA) is a machine learning technique that uses topology to extract patterns from data and has shown the potential to exhibit quantum advantage. A key concept in TDA is persistent homology, which measures the robustness of topological information at different lengthscales. In this paper, we introduce and study the problem of normalized persistence, a practically motivated and easily interpretable version of persistent homology that counts the fraction of holes that persist at different lengthscales. We prove that a variant of normalized persistence is $\mathsf{DQC}_1$-hard and contained in $\mathsf{BQP}$, giving evidence of an exponential quantum speedup for TDA under the standard assumption that $\mathsf{DQC}_1 \not\subseteq \mathsf{BPP}$. These are the first $\mathsf{DQC}_1$-hardness results that are directly applicable to TDA instances. We also find a close connection between normalized persistence and the complexity of estimating spectral quantities in the low-energy subspace of local Hamiltonians. We study a family of such problems, including a low-energy normalized subtrace and spectral density. We show that these are $\mathsf{DQC}_1$-hard for $O(1)$-local Hamiltonians, strengthening previous results that required log-local interactions. We also introduce a variant of $\mathsf{DQC}_1$ with perfect completeness ($\mathsf{SDQC}_1$) to characterize the hardness of problems normalized by an exact kernel. This includes normalized persistence for $O(1)$-local Hamiltonians, which we show is $\mathsf{SDQC}_1$-hard.
Information mechanics: conservation and assimilation
arXiv:2601.15028v2 Announce Type: replace Abstract: Inference and learning are commonly cast in terms of optimisation, yet the invariant constraints governing uncertainty reduction remain unclear. This work presents information mechanics (infomechanics), a first-principles framework that describes informational structure in two canonical state coordinates. Starting from the pointwise identity implied by Bayes' rule, minimal requirements of additivity, symmetry, and finite-resolution robustness select only two distinct additive projections, yielding conservation identities for Shannon entropy, governing global uncertainty, and Fisher information, encoding complementary local geometry. Because part of Fisher information is fixed by entropy, the residual structure is captured by a non-additive, coordinate-scale-invariant state function, the information potential $\Phi$. This yields a two-coordinate description that separates the entropic baseline from residual geometric complexity. $\Phi$ vanishes uniquely for isotropic Gaussian distributions and decreases under Gaussian coarse-graining. In finite-resolution multimodal landscapes, $\Phi$ asymptotically scales with the logarithm of the effective number of local optima, linking information geometry to inference difficulty. The same two-coordinate formalism extends to the Markov chain linking hidden states, observations, and internal representations, yielding assimilation inequalities that constrain faithful external-state inference. Together, these results identify invariant constraints underlying inference, learning, and computation across biological and artificial systems.
Tracing 3D Anatomy in 2D Strokes: A Multi-Stage Projection Driven Approach to Cervical Spine Fracture Identification
arXiv:2601.15235v4 Announce Type: replace Abstract: Cervical spine fractures require rapid and accurate diagnosis, yet automatic CT interpretation remains challenging as subtle injuries must be assessed across large 3D volumes. We ask whether full 3D vertebra segmentation is necessary for automated fracture recognition, or whether vertebra masks approximated from 2D projections can preserve sufficient diagnostic context. We propose an end-to-end pipeline that localizes the cervical spine, estimates C1-C7 vertebra masks from optimized 2D projections, and uses the resulting vertebra-level volumes for downstream fracture classification. A YOLOv8 detector first localizes spine regions of interest from multi-view variance projections, achieving a 3D mean Intersection over Union of 94.45%. Multi-label vertebra segmentation is then performed with a DenseNet121-Unet on energy-based sagittal and coronal projections, attaining a mean Dice score of 87.86%. The predicted 2D masks are back-projected and fused into approximate 3D masks for each vertebra to extract volumes of interest from the original CT. These volumes are analyzed by an ensemble of 2.5D spatio-sequential CNN-Transformer models, yielding vertebra-level and patient-level F1 scores of 68.15 and 82.26, area under the receiver operating characteristic curve of 91.62 and 90.95, and area under the precision-recall curve of 75.60 and 92.00, respectively. The projection-derived volumes achieved fracture-recognition performance comparable to a full 3D-segmentation baseline, while shifting the vertebra segmentation stage into a lower-dimensional domain. Saliency-based explainability and interobserver variability analysis further examine interpretability and reliability. Overall, the results indicate that projection-based mask approximation is a viable proxy for full 3D vertebra segmentation in cervical fracture recognition.
Scalable Dexterous Robot Learning with AR-based Remote Human-Robot Interactions
arXiv:2602.07341v2 Announce Type: replace Abstract: This paper focuses on the scalable robot learning for manipulation in the dexterous robot arm-hand systems, where the remote human-robot interactions via augmented reality (AR) are established to collect the expert demonstration data for improving efficiency. In such a system, we present a novel method to address the general manipulation task problem. Specifically, the proposed method consists of two phases: i) In the first phase for pretraining, the policy is created in a behavior cloning (BC) manner, through leveraging the learning data from our AR-based remote human-robot interaction system; ii) In the second phase, a contrastive learning empowered reinforcement learning (RL) method is developed to obtain more efficient and robust policy than the BC, and thus a projection head is designed to accelerate the learning progress. An event-driven augmented reward is adopted for enhancing the safety. To validate the proposed method, both the physics simulations via PyBullet and real-world experiments are carried out. The results demonstrate that compared to the baselines, our method not only significantly speeds up the training process, but also achieves much better performance in terms of the success rate for fulfilling the manipulation tasks. By conducting the ablation study, it is confirmed that the proposed RL with contrastive learning overcomes policy collapse. Supplementary demonstrations are available at https://cyberyyc.github.io/.
Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models
arXiv:2409.16663v5 Announce Type: replace Abstract: We propose the use of latent space generative world models to address the covariate shift problem in autonomous driving. A world model is a neural network capable of predicting an agent's next state given past states and actions. By leveraging a world model during training, the driving policy effectively mitigates covariate shift without requiring an excessive amount of training data. During end-to-end training, our policy learns how to recover from errors by aligning with states observed in human demonstrations, so that at runtime it can recover from perturbations outside the training distribution. Additionally, we introduce a novel transformer-based perception encoder that employs multi-view cross-attention and a learned scene query. We present qualitative and quantitative results, demonstrating significant improvements upon prior state of the art in closed-loop testing in the CARLA simulator, as well as showing the ability to handle perturbations in both CARLA and NVIDIA's DRIVE Sim.
OnePath: Efficient and Privacy-Preserving Decision Tree Inference in the Cloud
arXiv:2409.19334v3 Announce Type: replace Abstract: The vast storage capacity and computational power of cloud servers have led to the widespread outsourcing of machine learning inference services. While offering significant operational benefits, this practice also introduces privacy risks, such as the exposure of proprietary models and sensitive user data. In this paper, we present OnePath, a framework for secure and efficient decision tree inference in cloud environments. Unlike existing methods that traverse all internal nodes of a decision tree, our traversal protocol processes only the nodes on the prediction path, significantly improving inference efficiency while preserving privacy. To further optimize privacy and performance, OnePath is the first to employ functional encryption for evaluating decision tree nodes. Notably, our protocol enables both model providers and users to remain offline during the inference phase, offering a crucial advantage for practical deployment. We provide formal security analysis to demonstrate that OnePath provides comprehensive privacy protections during the model inference process. Extensive experimental results show that our approach processes query data in microseconds, highlighting its efficiency. OnePath offers a practical solution that strikes a balance between security and performance, making it a promising option for a wide range of cloud-based decision tree inference applications.
VLMGuard: Bootstrapping Malicious Prompt Detectors from Unlabeled Vision-Language Prompts in the Wild
arXiv:2410.00296v2 Announce Type: replace Abstract: Vision-language Models (VLMs) are essential for contextual understanding of both visual and textual information. However, their vulnerability to adversarially manipulated inputs presents significant risks, leading to compromised outputs and raising concerns about the reliability in VLM-integrated applications. Detecting these malicious prompts is thus crucial for maintaining trust in VLM generations. A major challenge in developing a safeguarding prompt classifier is the lack of a large amount of labeled benign and malicious data. To address the issue, we introduce VLMGuard, a novel learning framework that leverages the unlabeled user prompts in the wild for malicious prompt detection. These unlabeled prompts, which naturally arise when VLMs are deployed in the open world, consist of both benign and malicious information. To harness the unlabeled data, we present an automated maliciousness estimation score for distinguishing between benign and malicious samples within this unlabeled mixture, thereby enabling the training of a binary prompt classifier on top. Notably, our framework does not require extra human annotations and is robust to realistic prompt variations, offering strong flexibility and practicality for real-world applications. Extensive experiments show that VLMGuard achieves superior detection results, improving AUROC by 5.39% on average over the state-of-the-art method. Disclaimer: This paper may contain offensive examples; reader discretion is advised. Code is available at: https://github.com/radiolab-ntu/vlmguard.
Variance of the $SIS$ Epidemic on Networks: A Diffusion Approximation
arXiv:2607.03300v1 Announce Type: cross Abstract: Functional laws of large numbers (FLLNs) describe the mean-field trajectory of epidemics on networks, but say nothing about the fluctuations around it. These fluctuations are governed by moments of the degree distribution not relevant at the level of the mean. A rigorous functional central limit theorem (FCLT) exists for the susceptible--infected ($SI$) process on configuration-model graphs, but no analogue exists for $SIS$, where recovery reintroduces vertices into the susceptible pool with partially known neighborhoods, breaking the clean neighborhood distribution the $SI$ derivation relies on. We develop a tractable variance approximation for Markovian $SIS$ on configuration-model graphs, combining Gleeson's approximate master equation (AME) framework with a van Kampen system-size expansion in the spirit of the $SI$ FCLT. We derive a closed drift and diffusion matrix for a reduced susceptible/$SI$-edge/$SS$-edge count vector and obtain the time-dependent covariance via the associated Langevin/Lyapunov equation. Validation against Gillespie simulation across Poisson, regular, and power-law networks shows close agreement, with deviations near the epidemic threshold and in strongly heterogeneous networks.