Forskningsradar

Science Journals

Peer-reviewade publikationer — 58997 artiklar

Mitigating Scoring Errors and Compensating for Nonverbal Subtests in Speech-Based Dementia Assessment
arXiv:2606.18979v1 Announce Type: cross Abstract: Early detection of cognitive impairment relies on neuropsychological tests to minimize subjectivity by assessing multiple cognitive domains. Speech-based evaluation can support diagnostics and improve accessibility, but transcription errors and the omission of nonverbal subtests (e.g., motor skills) limit accuracy. Beyond conventional test scores, speech-derived features can provide additional insights into cognitive status. This study investigates the speech-based evaluation of the German "Syndrom-Kurz-Test," a standardized dementia screening test comprising verbal and motor subtests. We train models that integrate transcript-derived scores and Whisper embeddings per verbal subtest to reduce scoring errors. To compensate for missing motor subtests, we then leverage these fused representations to approximate expert overall ratings. Despite omitting subtests, our models strongly correlate with expert ratings and efficiently and accurately discriminate between cognitive status groups.
Optimizing Incomplete, Large-Scale and Sparse Multi-Graph Matching in Bioimaging
arXiv:2406.18215v2 Announce Type: replace Abstract: Multi-graph matching is a fundamental problem in computer vision. Our work is motivated by a challenging application in bioimaging, where dozens or even hundreds of 3D microscopy images of worms must be brought into correspondence. Existing datasets do not cover this large-scale regime, and virtually all existing methods are inapplicable because they assume a complete or dense problem setting. To support further research, our first contribution is a new large-scale dataset based on problem instances from bioimaging. Our second contribution is a comprehensive analysis of the two main multi-graph matching paradigms: direct and permutation synchronization-based formulations. We argue, in part by proof, that practical large-scale methods must explicitly address problem sparsity and incompleteness. Since standard permutation synchronization approaches fail in this setting, we further introduce a sparse permutation synchronization paradigm. Our final contribution is GREEDA, a general method for sparse and incomplete problems that can be instantiated across cost orders and paradigms. While our paper focuses on objective functions up to quadratic order, GREEDA is inherently generalizable to arbitrary orders. On larger, sparse instances, GREEDA outperforms competing methods in both objective value and runtime. For example, for moderately-sized problems based on 30 worm images GREEDA produces a high-quality solution within 2 minutes, whereas competitors require at least half an hour and yield far worse results. On smaller dense problems, GREEDA remains on par with leading methods while being an order of magnitude faster.
Co-evolution of the global research collaboration network and the performance of nations in science and technology
arXiv:2606.18549v1 Announce Type: new Abstract: Researchers have long suspected that international research collaboration (IRC) and scientific and technological (S\&T) performance are subject to reciprocal causality, yet the endogenous co-evolution of these twin phenomena has yet to be tested by large-scale empirical analysis. This study tests IRC network effects on national research performance and vice versa simultaneously using a longitudinal co-evolution model on three decades of global network and national performance data. Stochastic actor oriented models (SAOM) are used to analyze data on 166 countries from 1993 to 2022. Yearly IRC networks are constructed from Web of Science's XML database, and performance data are gathered from Elsevier's fractional field-weighted citation index (FWCI). The models also account for geographic, economic, demographic, and political factors, as well as endogenous network processes. The results provide support for reciprocal co-evolution. However, notably, geographic distance appears to moderate the interaction between research performance and network dynamics, suggesting researchers may rely more on visible performance metrics when selecting geographically distant collaborators. This finding points to the role of citation based performance metrics as a signaling mechanism for collaborator selection.
Hierarchical Multi-Modal Retrieval for Knowledge-Grounded News Image Captioning
arXiv:2606.18553v1 Announce Type: new Abstract: Traditional image captioning methods often struggle to generate comprehensive, context-rich descriptions, especially for details not directly observable from visual cues. To overcome this, we propose a novel retrieval-augmented image captioning framework that generates captions with deeper insights, such as object attributes, event context, and underlying significance, by leveraging external knowledge. Our approach features a hierarchical multi-modal article retrieval mechanism that moves beyond monolithic text entities. This retrieval considers article structure-aware features, including weighted textual components (e.g., headlines, body sections) and visual placement patterns, alongside multi-faceted similarity computations (content--visual, visual--visual, and discourse positioning). A subsequent contextual relevance refinement stage further enhances the retrieved information. The retrieved articles then serve as the knowledge base for caption generation: first, a VLM generates a concise image description; second, we segment relevant information from the retrieved articles based on this description; and finally, an LLM utilizes both the description and extracted knowledge to generate a comprehensive, contextually detailed caption. We participated in the ACM Multimedia EVENTA 2025 Challenge and achieved 5th place with an overall score of 0.2824 on the private test set of the OpenEvent-V1 dataset. Source code is publicly released at https://github.com/mf0212/EVENTA-Challange.
Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
arXiv:2606.17846v2 Announce Type: replace Abstract: Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale. In this report, we investigate whether this scaling recipe can be applied to robotic manipulation to achieve genuine generalization. This is challenging because, unlike text, manipulation data is heterogeneous by nature, expensive to collect, and narrow in diversity, making alignment and scale simultaneously difficult. We present Qwen-RobotManip, a generalizable Vision-Language-Action foundation model built on Qwen-VL. Qwen-RobotManip introduces a unified alignment framework across the representation, motion, and behavioral dimensions of manipulation, making large-scale multi-source training coherent rather than conflicting. This alignment capability in turn enables Qwen-RobotManip to absorb manipulation data at a scale that prior training regimes could not sustain. A human-to-robot synthesis pipeline converts egocentric hand demonstrations into robot trajectories across 15 platforms, and a rigorous curation pipeline harmonizes heterogeneous datasets. Using only open-source datasets and human videos without proprietary data collection, Qwen-RobotManip constructs a ~38,100-hour pretraining corpus and exhibits emergent generalization capabilities, including zero-shot instruction following, robustness to perturbations, reactive error recovery, and cross-embodiment transfer. We find that standard benchmarks fail to capture pretraining quality and instead adopt OOD settings including RoboCasa365, LIBERO-Plus, EBench, RoboTwin-Clean2Rand, RoboTwin-IF, and RoboTwin-XE. Qwen-RobotManip substantially outperforms prior state-of-the-art models, including $\pi$0.5, across all OOD settings, ranks 1st in RoboChallenge with a 20% relative improvement, and is validated on real-robot platforms including AgileX ALOHA, Franka, UR, and ARX.
Riemann invariant-based alternative WENO scheme for a two-layer thin film model
arXiv:2606.17862v2 Announce Type: replace Abstract: In this article, we develop a multi-dimensional two-layer thin film model extending the thin film model proposed in \cite{barthwal2025hyperbolic}. The model considered in \cite{barthwal2025hyperbolic} considered a very specific Marangoni scale by choosing Marangoni numbers in both layers to be $1$. We relax this condition here and prove that the obtained system possesses a full set of Riemann invariants. Based on these findings, we develop a Riemann Invariant-based Local Characteristic Decomposition WENO (RI-WENO) method for the two-layer thin film model in one and two dimensions. The method is built upon a specially designed variable transformation constructed from the derived Riemann invariants of the system. This transformation partially diagonalizes the governing equations and yields a sparse structure in the transformed eigenvector matrices. As a result, the proposed RI-WENO framework significantly reduces the computational cost of the standard Local Characteristic Decomposition WENO approach while retaining its strong capability to suppress spurious oscillations. Numerical experiments, including new benchmark test cases, demonstrate that the RI-WENO method achieves an effective balance between accuracy and computational efficiency, making it a promising and practical choice for solving the two-layer thin film model.
Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding
arXiv:2606.18101v2 Announce Type: replace Abstract: Graphical user interface (GUI) grounding requires vision-language models (VLMs) to identify small target elements in high-resolution screenshots and predict precise screen coordinates. On-policy self-distillation (OPSD) is a promising post-training approach for this coordinate-sensitive task, since it provides dense token-level teacher signals beyond hard coordinate labels. However, naive OPSD is not well suited to GUI grounding: OPSD evaluates the teacher on student-generated prefixes, the quality of coordinate-token teacher signals can degrade when the prefix has already deviated from the target coordinate, leading to unreliable teacher signal. To mitigate this, We propose quality-aware self-distillation for VLM-based GUI grounding, which improves coordinate-token teacher-signal quality through soft correctness-aware gating and teacher-probability scaling. The soft correctness-aware gate checks whether the teacher's current coordinate-token prediction can still be completed into the ground-truth box under the student-generated prefix. If not, the corresponding teacher signal is down-weighted. Teacher-probability scaling then uses the teacher's confidence as a lightweight factor to further calibrate the strength of the gated supervision. A key empirical finding is that neither component alone improves overall performance, whereas combining them consistently improves performance. This suggests that the two mechanisms play complementary roles: correctness-aware gating suppresses unreliable coordinate-token supervision, while teacher-probability scaling calibrates the strength of the remaining signals. Experiments across six GUI grounding benchmarks show that our method consistently improves the base model and outperforms strong baselines.
FluidViews: Adaptive Drag-and-Drop Token Filters for Heterogeneous Multi-View Visual Analytics
arXiv:2606.18260v1 Announce Type: new Abstract: Interactive visual analytics workflows are often disrupted by rigid filter panels and context switches that break analysts' cognitive flow. We introduce FluidViews, a web-based framework that elevates filters to first-class, manipulable objects through two novel direct-manipulation interactions. Copy-as-Highlight enables users to duplicate any visual mark into a persistent highlight token for rapid, transient cross-view comparison, while Drag-as-Filter allows analysts to pick up a mark and drop it onto another view to apply context-sensitive filters in place no menus, panels, or modal dialogs required. An optional pop-out micro-view provides on-demand, spatially independent subviews for detailed inspection without disrupting the primary workspace. By embedding these lightweight gestures into coordinated multi-view environments, FluidViews preserves analytic momentum, reduces cognitive overhead, and supports fluid, multi-step exploration across heterogeneous datasets. We describe the system's design and implementation, illustrate its application in exploratory workflows, and discuss how tangible filter objects can transform interactive data exploration.
Simulating Hate Speech Cascades with Multi-LLM Agents: Empirical Grounding, Modeling Fidelity, and Intervention Strategies
arXiv:2606.18264v1 Announce Type: new Abstract: Faithful modeling of hateful content propagation on online platforms remains an open problem for moderation research. Classical cascade models that do not explicitly represent the profile, community, and content factors associated with hateful-content propagation may yield moderation strategies that behave less effectively when deployed in real-world scenarios. Multi-agent large language model (LLM) systems can, in principle, make each reshare decision depend on the user's profile, the surrounding community, and the post's content, but it remains unclear whether this added flexibility actually reproduces real hateful cascades more faithfully than classical baselines. We study three hateful Bluesky cascades and a size-matched benign control. In the empirical Bluesky data, we found that: 97.4--99.7\% of reposters take a hostile stance; toxicity-engagement homophily is higher on the diffusion tree than on the follower graph for hateful cascades; topology is star-like for the hateful cascades (most reposts come directly from the root) versus tree-like for the benign cascade (reposts propagate through multi-hop chains). In simulation, a multi-LLM-agent simulator reproduces the stance monoculture and the toxicity-delta direction. A structured ablation identifies agent heterogeneity as the leading fidelity factor, and amplifier targeting on dense networks yields 7.5--12.9\% reduction at 5.7\% benign collateral.
Graph Instance Landscapes: When Structural Similarity Does (Not) Reflect Shortest-Path Performance
arXiv:2606.18267v1 Announce Type: new Abstract: Benchmarking shortest-path algorithms is commonly based on aggregate performance over heterogeneous graph sets, which limits insight into how different search paradigms react to instance structure. We adopt an instance-landscape view of graph benchmarking by embedding graphs into a low-cost structural feature space and clustering them into regions of similar structure. Three benchmark suites are studied: weighted Erd\H{o}s--R\'enyi graphs, random geometric (wireless) graphs, and real-world road networks. We evaluate four representative shortest-path solvers spanning uninformed exact search (Dijkstra), bidirectional exact search (bidirectional Dijkstra), heuristic-guided exact search (A$^{*}$), and deque-based strategies (DEQ). Clustering robustness is analyzed under multiple feature-selection schemes, and runtime distributions are compared across landscape regions using non-parametric tests. While generator parameters induce stable structural regions, we find that feature-space similarity does not necessarily imply performance similarity: significant runtime shifts are frequently observed even within the same landscape region. A merged-suite analysis further shows that different benchmark families occupy largely disjoint regions. These results highlight both the potential and the limits of structural landscapes for the structure-aware benchmarking of shortest-path algorithms.
Optimizing Lithium Production Decisions under Geological, Demand, and Pricing Uncertainties: A POMDP Framework for Multi-Objective Decision Making
arXiv:2606.18598v1 Announce Type: new Abstract: Decision making in lithium production is challenging, whether from an investor's perspective or a strategic production standpoint. Determining which mines to open and when to open them involves not only geological and price uncertainties, but also complexities around the choice of extraction method, from direct lithium extraction to hard rock mining. Prior work explored models of this problem and different methods to optimize mining decisions; these models did not account for uncertainty in pricing, uncertainty in demand, or different mining technologies to extract lithium. Incorporating different pricing models and extraction technology into these models enables more robust strategies for determining not only when and where to open a mine, but also which method of production to pursue. We frame the problem as a partially observable Markov decision process (POMDP) and solve using belief state planning methods to get optimal decision making. In our study, we show that POMDP solvers outperform human inspired heuristics by dynamically adapting to shifting lithium price regimes (static, linear, exponential, and stochastic) through belief state planning and explicit uncertainty management. By optimally sequencing exploration, production, and technology choice, the framework achieves higher demand fulfillment and more balanced economic environmental outcomes over the projects lifetime in all different pricing and deposit scenarios.
OmniPlan: An Adaptive Framework for Timely and Near-Optimal Network Planning Optimization
arXiv:2606.18105v2 Announce Type: replace Abstract: Network planning optimization is a fundamental problem across diverse domains, including transportation systems, communication networks, and power grids. It requires simultaneous optimization of multiple competing objectives under complex constraints. Existing network planning optimization frameworks rely on mixed integer programming (MIP) solvers, heuristics, and deep reinforcement learning (DRL) models to compute planning decisions. However, they lack effective adaptability to diverse and dynamic user intents, thus leading to the trade-off between execution time and optimality. In this paper, we propose OmniPlan, an adaptive framework that achieves both timeliness and near-optimality in network planning optimization. To achieve the adaptability lacking in existing solutions, OmniPlan employs a large language model (LLM)-based interpreter to convert heterogeneous natural-language intents into a unified and quantifiable user-preference vector. Then it employs a mixture-of-experts architecture that integrates MIP solvers, heuristics, and DRL models as specialized experts, where OmniPlan adapts to diverse intents by dynamically selecting timely and near-optimal experts. Finally, it incorporates a DRL-based expert configuration module that fine-tunes optimization objective weights to align planning decisions with user-specific preferences. We evaluate OmniPlan with a representative real-world workload, i.e., distributed machine learning (ML), where we leverage OmniPlan to offload a wide spectrum of ML inference tasks, e.g., decision trees, SVM, naive Bayes, XGBoost, and random forests, onto a network of hardware devices. Our experiments on a real-world testbed indicate that OmniPlan achieves near-optimal and low-execution-time offloading for real-world ML inference tasks, reducing latency by up to 97.8\% and network device resource consumption by up to 11.5\%.
PEPSKit.jl: A Julia package for projected entangled-pair state simulations
arXiv:2605.19960v2 Announce Type: replace-cross Abstract: We present PEPSKit$.$jl, a Julia package for simulating two-dimensional quantum many-body systems with infinite projected entangled-pair states (iPEPS). PEPSKit$.$jl builds on the TensorKit$.$jl package for tensor computations and provides high-level algorithms for iPEPS simulations that support both Abelian and non-Abelian symmetries, as well as fermionic systems. This work gives an overview of the main package features, which include support for ground-state, time-evolution, and finite-temperature simulations in systems with different physical symmetries and lattice geometries. These capabilities are illustrated through various examples and technical benchmarks.
Why SWAVE May Not Be All You Need:A Concept-Evolution Retrospective on Complex-Valued Recurrent Language Models
arXiv:2606.18324v1 Announce Type: new Abstract: SWave is a complex-valued recurrent language model (169.26M parameters, D=384, L=16, T=2048) trained on FineWeb-Edu using 2xH100 NVL. It was designed around three founding premises: that representing language as complex waves rather than real-valued numbers enables richer information encoding; that a Cayley-parameterised unitary transition provides a mathematical guarantee against state decay or explosion; and that a hidden state which rotates rather than shrinks preserves signal integrity over arbitrarily long contexts. The core of SWave evolved substantially across three development phases. The Resonance Head was found to structurally admit imaginary-channel collapse as a global loss minimum (a failure mode we term cos-domination collapse) and was superseded by an untied head with independent real and imaginary embedding tables from the Phase-Associative Memory (PAM) architecture. This resolved the degenerate minimum and enabled stable 200,000-step training (best-step PPL 22.0 at step 89,861). ComplexNorm and the Wave Propagation Scan proved load-bearing throughout all three phases and were retained to the final architecture. ProtectGatedScan was reframed as a structural prior rather than a learned behaviour. The four multi-scale retention concepts showed no measurable improvement under controlled evaluation and were found non-load-bearing. The ComplexGatedUnit was superseded by a real-valued squared-ReLU channel mixer with fewer parameters. The auxiliary training objectives showed no benefit once structural constraints were resolved. The investigation yields a formal characterisation of cos-domination collapse, a parallel scan with a log-space backward pass for numerical stability, six transferable engineering principles for complex-valued recurrent training, and a plan-to-code traceability methodology for catching structural divergences that conventional test suites miss.
Time Entangled Quantum Blockchain with Phase Encoding for Classical Data
arXiv:2507.14839v5 Announce Type: replace-cross Abstract: With rapid advancements in quantum computing, it is widely anticipated that scalable quantum hardware may threaten classical cryptography and hence, the internet and the current information security infrastructure in the coming decade. This is mainly due to the operational realizations of quantum algorithms such as Grover and Shor, to which the current classical encryption protocols are vulnerable. Blockchains, i.e., blockchain data structures and their data, rely heavily on classical cryptography. One approach to secure blockchains is to attempt to achieve conceptual information-theoretic security under certain assumptions by defining blockchains on quantum technologies. There have been two major conceptualizations of blockchains data structures on quantum registers: the time-entangled Greenberger-Horne-Zeilinger (GHZ) state blockchain and the quantum hypergraph blockchain. We conceptualize a new quantum blockchain framework combining features of both these schemes to achieve the conceptual information-theoretic protection against undetected measurement attack (physics-based disturbance detectability) of the time-entangled GHZ blockchain and the scalability and efficiency of the quantum hypergraph blockchain in the proposed quantum blockchain data structure and framework. In this work, we propose a novel quantum blockchain architecture that integrates temporal GHZ entanglement with phase encoding inspired by the quantum hypergraph blockchain. The proposed design combines the conceptual information-theoretic tamper sensitivity/resistance of temporal entanglement with improved encoding efficiency, offering a unified conceptual framework for scalable and secure quantum blockchains.
Experimental measurement of quantum first-passage-time distributions
arXiv:2508.21790v2 Announce Type: replace-cross Abstract: Classical First-Passage-Time Distributions (FPTDs) have been extensively studied both theoretically and experimentally. Their quantum counterparts-Quantum First-Passage-Time Distributions (QFPTDs)-remain largely unexplored and have deep implications for both fundamental physics and the development of emerging quantum technologies. We measure the first QFPTDs using a motional mode of a single trapped ion. We develop a novel composite-phase laser pulse sequence to perform tunable stroboscopic single-shot projective measurements of the motional state of a trapped ion. We measure QFPTDs of the ion energy when coupled to electric-field noise. The measurement protocol developed here is broadly applicable to other quantum systems and provides a powerful method for exploring a broad range of QFPTD phenomena. With these results we open a new field of experimental investigations of QFPT processes with potential future relevance to quantum search algorithms, unraveling connections between classical and quantum dynamics, and study of the quantum measurement problem.
SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding
arXiv:2601.12805v4 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. However, their ability to reliably reason from gene-level knowledge to functional understanding, a core requirement for knowledge-enhanced cell atlas interpretation, remains largely underexplored. To address this gap, we introduce SciHorizon-GENE, a large-scale gene-centric benchmark constructed from authoritative biological databases. The benchmark integrates curated knowledge for over 190K human genes and comprises more than 540K questions covering diverse gene-to-function reasoning scenarios relevant to cell type annotation, functional interpretation, and mechanism-oriented analysis. Motivated by behavioral patterns observed in preliminary examinations, SciHorizon-GENE evaluates LLMs along four biologically critical perspectives: research attention sensitivity, hallucination tendency, answer completeness, and literature influence, explicitly targeting failure modes that limit the safe adoption of LLMs in biological interpretation pipelines. We systematically evaluate a wide range of state-of-the-art general-purpose and biomedical LLMs, revealing substantial heterogeneity in gene-level reasoning capabilities and persistent challenges in generating faithful, complete, and literature-grounded functional interpretations. Our benchmark establishes a systematic foundation for analyzing LLM behavior at the gene scale and offers insights for model selection and development, with direct relevance to knowledge-enhanced biological interpretation.
Differential Equation Inductive Robustness Axiomatization
arXiv:2606.18685v1 Announce Type: new Abstract: This article establishes the completeness of an axiomatization for the robust safety of dynamical systems with polynomial differential equations on bounded time horizons. Safety properties of robust systems are uniformly reduced to a sound axiomatization of polynomial invariants, resulting in reliable logical proofs of correctness. Approximate decidability results are also established: there is a computable algorithm such that, given any perturbation parameter $\delta$, it either produces a symbolic proof of robust safety (hence correctly decides the dynamical system to be robustly safe), or correctly decides that the system is not robustly safe under a perturbation of level $\delta$. In contrast to earlier works, this article crucially leverages results from subanalytic geometry to retain a level of exactness, thereby establishing positive results of provability/decidability allowing for arbitrary bounded (semialgebraic) initial/post conditions even without positive separation at their (topological) boundaries. This enables the generation of proofs of inductive safety beyond finite time horizons for general hybrid dynamical systems.
Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment
arXiv:2606.18703v1 Announce Type: new Abstract: Pretrained biological language models expose per-token probability distributions through masked-token prediction, providing the likelihood interface central to sequence design, variant scoring, and mechanistic interpretation. Yet these distributions are learned from broad unlabeled corpora and are not naturally conditioned on task-specific biological contexts such as interaction partners, cellular environments, or therapeutic interventions. Existing contextual matching methods often distort this interface through pooled embeddings, contrastive latent spaces, or task-specific prediction heads. We introduce LOGICA (Logit-space Contrastive Alignment), a framework for context-conditioned prediction that performs contrastive learning directly in output-logit space. Using gated cross-modal adapters compatible with each model's native token head, LOGICA preserves the pretrained likelihood interface and converts contextualized token log-likelihoods into matching scores. Alignment is defined through context-sensitive token probabilities rather than proximity in a shared embedding space, enabling learning from sparse paired data across models with distinct vocabularies, without a shared tokenizer or decoder. LOGICA is particularly effective for mutation-local variant ranking, where comparisons reduce to context-conditioned likelihoods of mutant tokens at perturbed sites. Across protein--ligand binding, TCR--peptide activity, and drug-conditioned resistance prediction, LOGICA improves over prior state-of-the-art methods, including matched latent-contrastive and conditional MLM baselines, while retaining a token-level interface for interpretation and generation. On held-out-gene single-mutation drug-resistance prediction, LOGICA improves AUC from near-random latent-space baselines of $\sim$0.55 to $\sim$0.65.
Image Prompt Reconstruction Attacks on Distributed MLLM Inference Frameworks
arXiv:2606.18710v1 Announce Type: new Abstract: Distributed large language model (LLM) inference frameworks connect isolated consumer-grade devices for large-scale model inference, substantially reducing hardware constraints. However, recent studies show that intermediate embeddings transmitted among participants can leak private prompts. As LLMs evolve into multimodal LLMs (MLLMs), this risk extends beyond text: image prompts contain rich visual and semantic information, making their intermediate embeddings highly privacy-sensitive. Yet, image-prompt leakage in distributed MLLM inference remains largely unexplored. In this paper, we investigate privacy risks to input images caused by intermediate embeddings in distributed MLLM frameworks. We first analyze the information flow from image pixels to intermediate representations. Since image and text embeddings are often intertwined across MLLM layers, we design an image embedding extraction algorithm as a prerequisite for reconstruction attacks, achieving 100% extraction accuracy across almost all MLLM layers in our experiments. Building on this, we develop two passive black-box image reconstruction attacks, MPAA and IEDA, reflecting realistic threats from normal participants with limited knowledge and capability. MPAA performs fine-grained pixel-level reconstruction via patch-wise information extraction and assembly, while IEDA performs coarse-grained semantic reconstruction through embedding-guided diffusion generation. We evaluate our attacks on four representative MLLM families: Gemma 3, Phi 4 Multimodal, Qwen 2.5 VL, and Llama 4 Scout. Results show consistently superior reconstruction performance in various settings. We further analyze the effects of MoE architecture, image preprocessing, model size, and text-image dependency on attack performance. To our knowledge, this is the first study of image reconstruction attacks on MLLMs.
Symplecticity-preserving prediction of parameter-dependent Hamiltonian dynamics by Generalized Kernel Interpolation
arXiv:2606.18937v1 Announce Type: new Abstract: We extend the kernel-based symplectic predictor of [1] to a parameter-augmented setting in which the learned flow-map surrogate depends not only on the state, but also on additional variables such as physical parameters and macro time-step sizes. The method uses a product kernel ansatz on a parameter and macro step augmented domain and constructs the prediction through an implicit symplectic-Euler-type update. Hence, for every fixed admissible parameter and time-step instance, the resulting large-step predictor is symplectic by construction. The training problem is formulated as gradient Hermite--Birkhoff interpolation in a reproducing kernel Hilbert space. Efficient surrogates are obtained by greedy center selection. We show that the convergence analysis from the non-augmented setting carries over to the product-kernel framework and derive corresponding prediction error bounds. Numerical experiments for a pendulum with varying length and time-step size and for a parameter-dependent discretized wave equation illustrate the accuracy and structure-preserving behavior of the proposed approach.
Stitching the Divide: Investigating Mixed Reality as a Bridge Between Paper-Based and Digital Artifacts in UI/UX Design
arXiv:2606.18511v1 Announce Type: new Abstract: UI/UX designers work with both paper-based and digital artifacts but lack tools that seamlessly integrate the two. Mixed Reality (MR) offers under-explored opportunities to combine the strengths of both design environments. To examine these opportunities, we first conducted interviews with 19 professional UI/UX designers to understand their current experiences using paper and digital artifacts. Motivated and informed by the interview insights, we organized nine conceptual-probe user study sessions in which designers engaged with a MR-probe that combined paper and digital prototyping processes and brainstormed MR's potential in UI/UX design. We found that participants valued MR for enabling continuous hybrid design workflows, reducing manual reconstruction, supporting spatially anchored workspaces, and facilitating real-time cross-medium collaboration. They also envisioned future MR tools with AI assistance, richer interactive and dynamic content, and the ability to manage diverse design artifacts within a unified environment. From these findings, we derive four design dimensions for future MR systems that could enable more fluid, creative, and collaborative design practices.
Architectural Bias in Face Presentation Attack Detection: A Comparative Study of Vision Transformers and Convolutional Neural Networks
arXiv:2606.18510v1 Announce Type: new Abstract: Face Presentation Attack Detection (PAD) systems constitute a critical security layer in biometric authentication; however, existing approaches exhibit systematic performance disparities across demographic groups, disproportionately affecting individuals with darker skin tones. This paper presents a comparative empirical investigation of whether Vision Transformer architectures reduce demographic bias in face PAD systems relative to convolutional baselines. Experiments are conducted on the CASIA-SURF Cross-Ethnicity Face Anti-Spoofing (CeFA) dataset. Three architectures are evaluated: a Multimodal ViT-Tiny trained from scratch, a ResNet18 CNN baseline, and a pretrained DeiT-S fine-tuned on CeFA across African, East Asian, and zero-shot Central Asian demographic groups. DeiT-S achieves the highest overall accuracy of 97.27% and the lowest EER of 0.86%, outperforming ResNet18 at 90.15% accuracy. In terms of fairness, DeiT-S reduces the inter-ethnic ACER gap between African and East Asian subjects to 0.13%, compared to 0.75% reported in an LBP-based work [6], representing an 83% reduction. Most notably, while ResNet18 records a BPCER of 10.44% on zero-shot Central Asian subjects, DeiT-S maintains 2.89% on the same unseen group, demonstrating a 3.6x generalization advantage. These results suggest that pretrained Vision Transformers achieve superior PAD accuracy, produce smaller demographic performance gaps, and generalize more equitably across unseen demographic groups, indicating that cross-demographic fairness in PAD may partly be influenced by architectural design.
Confident yet Concerned: Inconsistencies in Computing Students' Attitudes on Cybersecurity
arXiv:2606.18541v1 Announce Type: new Abstract: Today's young adults are most immersed in technology, leading in feelings of powerlessness in managing online privacy across many platforms, and particularly susceptible to phishing attacks. This raises questions about their general, wide-ranging attitudes towards and management of cybersecurity. How do young, tech-savvy adults approach cybersecurity? We seek a better understanding of their cybersecurity knowledge, attitudes and experiences, in particular in addressing deceptive online communications. We surveyed a group of `lead users': computing university students (n = 236). By combining thematic analysis of open-ended responses with quantitative data, we provide insights into their experiences and perceptions. While students demonstrate reasonable cybersecurity awareness, their cybersecurity experiences vary, and inconsistencies exist around their practices, perceptions of responsibility, and support structures. Findings also reveal four key thematic tensions: 1) Computing students are knowledgeable yet have persistent incorrect beliefs, 2) They learn more about keeping safe from sources outside the classroom, 3) They have limited assistance and have fallen victim to cybercrime, and 4) Many are confident, yet others are concerned about their own safety and responsibility. Through cluster analysis of attitudes, we identify two groups, with one feeling less prepared, less confident, yet expressing a desire to learn more. Established measures of intentions and objective knowledge were correlated to preparedness. Self-efficacy correlated to confidence and predicted cluster membership.
Closed-Form and Constant-Time New-Source Selection for Fault-Tolerant Broadcasting in Dense Gaussian Networks
arXiv:2606.18715v1 Announce Type: new Abstract: Fault-tolerant broadcasting in dense Gaussian networks is recovered by re-rooting the broadcast at a new source at maximum graph distance from the faulty nodes. This paper extends the re-rooting framework by replacing its boundary-search source-selection step with a quotient-lattice-aware algebraic construction. The first contribution is a constant-time counting method for valid new sources, formulated as an intersection of two diameter-$k$ boundary sets in the Gaussian quotient. The exact count is obtained by a fixed union of side-pair intervals over nine quotient-lattice copies, giving a closed-form procedure without scanning the network or boundary. The second contribution is a shifted direct selector for two arbitrary faulty nodes. Given faulty nodes $A$ and $B$, the problem is translated to $C=\operatorname{mod}_{G_k}(B-A)$, and the selector finds $P$ satisfying $d(P,0)=d(P,C)=k$. For each of nine quotient-lattice shifts, sixteen signed linear systems are checked. Nonparallel systems are solved via Cramer's rule; parallel systems are handled by interval-endpoint selection. At most $9\times16=144$ shifted sign cases are evaluated, giving $O(1)$ selection under the word-RAM model. Validation reports zero count mismatches over $26{,}623$ tested nodes, $500{,}000$ valid outputs over $500{,}000$ sampled fault pairs, and $40{,}000$ successful re-rooted broadcast trials. The shifted selector achieves a $5.92\times$ speedup over boundary search at $k=200$, remaining stable as $k$ increases. These results make new-source selection algebraic, bounded, and independent of network size.