Forskningsradar

Science Journals

Peer-reviewade publikationer — 54512 artiklar

MuellerPT: Decomposition Driven Pretraining for Dense Learning in Mueller Polarimetry
arXiv:2605.23840v1 Announce Type: new Abstract: Mueller matrix imaging provides rich, physically meaningful contrast for biomedical tissue analysis, but supervised learning is hindered by scarce dense annotations and strong domain shifts across specimens and acquisition settings. We introduce MuellerPT, a physics guided pre-training approach that learns transferable dense representations by predicting Lu-Chipman decomposition maps from per-pixel 4x4 Mueller matrices. To scale pre-training, we collected a new large Multispectral Animal Polarimetric Organ dataset (MAP-Org). The pre-trained encoder is adapted with a segmentation head for grey vs. white matter segmentation in lamb brain. A classification head is used for colorectal cancer vs. non-cancer classification. Both segmentation and classification are evaluated across few-shot learning scenarios. In segmentation, MuellerPT improves label efficiency and cross specimen transfer compared to models without pre-training, achieving an absolute DICE gain of over 20% compared to the baseline trained from scratch when using 5% of the training data. In classification, MuellerPT also enhances label efficiency, improving overall accuracy by 8% compared to the baseline when using 1% of the training data. We demonstrate MuellerPT's robustness to domain shift with a qualitative evaluation of its predicted Lu-Chipman maps on an ex vivo human oesophagus sample. These results suggest that predicting Lu-Chipman decomposition is an effective and practical pretext task for robust biomedical inference from Mueller polarimetry and can pave the way for future work on label efficient Mueller imaging.
Atom-Photon Bound States in Fractal Photonic Lattices: Localization Length and Anomalous Diffusion
arXiv:2605.23625v1 Announce Type: cross Abstract: We study atom-photon bound states seeded by two-level emitters coupled to self-similar photonic lattices. By expressing the photonic Green's function through the heat kernel, we show that the far-field localization length obeys $\xi \sim \Delta^{-1/d_w}$, with the detuning $\Delta$ from the lower spectral edge and the walk dimension $d_w$ of the underlying fractal. This scaling is controlled by anomalous diffusion and does not rely on translational invariance or a band-edge effective-mass approximation. Exact diagonalization on Sierpi\'nski gaskets, pyramids, Vicsek graphs, and Sierpi\'nski carpets confirms the far-field prediction once the bath Hamiltonian is rendered Laplacian-like by compensating the local inhomogeneity in the connectivities with on-site potentials. In the near field, the bound-state amplitude exhibits an additional algebraic variation. For nested finitely ramified fractals, the corresponding exponent agrees with the classical resistance/ first-passage scaling, whereas Sierpi\'nski carpets display clear deviations from this simple law. Our results extend structured-bath waveguide QED to self-similar non-periodic geometries and connect bound-state profiles to transport exponents of the underlying fractal lattice.
The Impact of AI Coding Assistants on Software Engineering: A Longitudinal Study
arXiv:2605.23135v1 Announce Type: new Abstract: AI coding assistants have become prolific in recent years. Through a longitudinal mixed-methods investigation, we examined how professional software engineers perceive the effects of AI coding assistants in regard to task focus, developer experience, and productivity. Two questionnaires were administered six months apart, yielding 158 eligible participants at the first time point, 101 at the second, and a matched longitudinal cohort of 95. Participants reported spending less time on most development tasks, with 82% reporting less on writing code. We find broader shift in focus from creation to verification activities. We propose a new category of work we term supervisory engineering work, encompassing the direction, evaluation, and correction of AI output. We also identified a productivity-experience paradox: productivity perceptions held stable, with 84% reporting improvement at both time points, yet among matched participants, the proportion reporting worsened developer experience in at least one dimension nearly doubled from 14% to 27%, with flow state and cognitive load eroding while feedback loops improved. These findings suggest that AI coding assistants are impacting both the nature of software engineering work and how engineers experience it.
VFM$^{4}$SDG: Unveiling the Power of VFMs for Single-Domain Generalized Object Detection
arXiv:2604.21502v2 Announce Type: replace Abstract: Real-world weather, illumination, and imaging variations often induce severe domain shifts, degrading single-source detectors in unseen environments. Existing single-domain generalized object detection (SDGOD) methods mainly rely on data augmentation or domain-invariant learning, while largely overlooking how domain shift disrupts detector prediction stability. Through analytical experiments, we find that performance degradation is mainly dominated by increasing missed detections. Further analysis shows that this phenomenon stems from reduced cross-domain stability in DETR-style detectors: domain shift disrupts encoder-side object-background and inter-instance relations, and further weakens the semantic-spatial binding between decoder queries and real objects. Motivated by this, we find that vision foundation models (VFMs) still preserve stable relational structures and object responses under severe shifts, making them suitable cross-domain stability priors to compensate for detector degradation. To this end, we propose VFM$^{4}$SDG, a dual-prior learning framework for SDGOD, which introduces a frozen VFM into encoder representation learning and decoder query modeling. Specifically, we propose Cross-domain Stable Relational Prior Distillation to distill stable object-background and inter-instance relations from the VFM into the encoder, compensating for relational degradation. Meanwhile, we propose Semantic-Contextual Prior-based Query Enhancement, which injects category semantic prototypes and global object context into queries before they enter the decoder layer, enhancing semantic-spatial query-object binding stability. Extensive experiments show that VFM$^{4}$SDG significantly outperforms existing advanced methods on standard SDGOD benchmarks and two mainstream DETR-based detection frameworks, demonstrating its effectiveness, robustness, and generality.
Graphene-based Photodetector with Engineered Hot Carrier Cooling Dynamics
arXiv:2605.23646v1 Announce Type: cross Abstract: Graphene has emerged as a promising material for integration into silicon photonics, owing to its ultrafast and broadband photoresponse without the need for an external bias voltage. This photoresponse relies on the photo-thermoelectric effect created by hot carriers. A key factor underlying the performance of graphene photodetectors is the cooling dynamics of these hot carriers. In this work, we engineer these dynamics in a WSe2-graphene-WSe2 waveguide-integrated photodetector. In particular, by introducing proximity screening by a nearby graphite layer to this structure, we prolong the hot-carrier cooling time, leading to an enhanced photoresponse. We characterize the cooling dynamics under continuous-wave laser excitation by employing a photomixing technique, revealing an increase in the cooling time by up to a factor of four. Direct photoresponse measurements show that the internal photoresponsivity improves by approximately 50%. Together, these results demonstrate the potential of proximity screening to enhance the performance of graphene-based photodetectors on an integrated photonics platform.
Minimum Effort Control Using Variational Methods of Analytical Mechanics A New Approach For Optimal Control
arXiv:2605.23813v1 Announce Type: cross Abstract: Modern optimal control theory involves adjoining the already known equations of motion of a dynamic system to the objective function using dynamic costates; this is done in order to constrain the optimal control solutions to satisfy the equations of motion. The use of costates increases the number of variables and hence increases the complexity of the problem. On the other hand, variational methods of analytical mechanics finds the equations of motion by minimizing an action functional of the dynamic system, realizing control forces as external input to the system. In this paper a new disruptive approach for computing the optimal control is presented. This approach adopts the variational methods of analytical mechanics to derive equations for the control, in addition to the equations of motion. This is achieved by recognizing the control actuator as part of the dynamic system. In addition to the kinetic energy and potential energy, the action functional in this new approach includes additional energy terms that represent the control energy of the system. Two different methods are presented to write the modified action functional. The proposed approach is a significant departure from the modern optimal control theory, and it eliminates the need for costates when solving for the control. In this paper, a case study is presented to demonstrate the new approach.
Exact versus tight-binding models in longitudinally modulated $\mathcal{PT}$-symmetric coupled waveguides
arXiv:2605.23853v1 Announce Type: cross Abstract: The tight-binding (TB) model is a widely adopted approximation scheme for describing light propagation in waveguide arrays. Despite its success, its validity in $\mathcal{PT}$-symmetric systems characterized by strong longitudinal modulation has not been rigorously benchmarked against exact analytical solutions. In this work, we address this gap by performing a comparative analysis between exact continuous solutions derived from $z$-dependent supersymmetric (SUSY) transformations and their corresponding discrete TB approximations. To achieve this, we develop a theoretical model for two PT-symmetric coupled waveguides subject to longitudinal modulation. We then evaluate the performance of the TB framework against the exact SUSY benchmark. Our results delineate the specific validity range of the TB approximation, demonstrating its proficiency in reproducing spatial intensity distributions. However, we also identify its limitations in accurately capturing the complex oscillatory phase dynamics inherent to this non-Hermitian evolution.
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
arXiv:2602.04431v2 Announce Type: replace Abstract: LLM-based multi-agent systems have demonstrated impressive capabilities, but they also introduce significant safety risks when individual agents fail or behave adversarially. In this work, we study the automated design of agentic systems that remain safe even when a subset of agents is compromised. Inspired by Stackelberg security games, we formalize this problem as a game between a system designer (the Meta-Agent) and a best-responding Meta-Adversary that selects and compromises a subset of agents to minimize safety. We propose Meta-Adversary-Meta-Agent (MaMa), a novel algorithm inspired by this formalization for automatically designing safe agentic systems. Our approach uses LLM-based adversarial search, where the Meta-Agent iteratively proposes system designs and receives feedback based on the strongest attacks discovered by the Meta-Adversary. Empirical evaluations across diverse environments show that systems designed with MaMa consistently defend against worst-case attacks while maintaining performance comparable to systems optimized solely for task success. Moreover, the resulting systems generalize to stronger adversaries, as well as ones with different attack objectives or underlying LLMs, demonstrating robust safety beyond the training setting.
Extended Resolution Clause Learning via Dual Implication Points
arXiv:2406.14190v5 Announce Type: replace Abstract: We present a new extended resolution clause learning (ERCL) algorithm, implemented as part of a conflict-driven clause-learning (CDCL) SAT solver, wherein new variables are dynamically introduced as definitions for {\it Dual Implication Points} (DIPs) in the implication graph constructed by the solver at runtime. DIPs are generalizations of unique implication points and can be informally viewed as a pair of dominator nodes, from the decision variable at the highest decision level to the conflict node, in an implication graph. We perform extensive experimental evaluation to establish the efficacy of our ERCL method, implemented as part of the MapleLCM SAT solver and dubbed xMapleLCM, against several leading solvers including the baseline MapleLCM, as well as CDCL solvers such as Kissat 3.1.1, CryptoMiniSat 5.11, and SBVA+CaDiCaL, the winner of SAT Competition 2023. We show that xMapleLCM outperforms these solvers on Tseitin and XORified formulas. We further compare xMapleLCM with GlucoseER, a system that implements extended resolution in a different way, and provide a detailed comparative analysis of their performance.
CALAD: Channel-Aware contrastive Learning for multivariate time series Anomaly Detection
arXiv:2605.23139v1 Announce Type: new Abstract: Multivariate time series anomaly detection has become increasingly important in real-world applications, where labeled data are often scarce. Many existing approaches rely on unsupervised learning to model normal patterns, but they often treat all channels equally. This design can dilute anomaly-relevant signals, since not all channels contribute equally to anomaly detection. In this paper, we propose CALAD, a channel-aware contrastive learning framework for multivariate time series anomaly detection. CALAD governs the construction of contrastive samples using estimated channel relevance, allowing the learning process to reflect anomaly semantics rather than generic similarity. Channel relevance is estimated from reconstruction errors of a transformer-based autoencoder and is used to distinguish channels that are more influential to anomalous behaviors. Using this information, we design a channel-wise augmentation strategy in which positive and negative samples are constructed based on whether anomaly-relevant channels are preserved or perturbed. This encourages invariance to changes in irrelevant channels while being sensitive to changes in anomaly-relevant channels. Furthermore, CALAD combines contrastive learning and an auxiliary reconstruction head, allowing the model to learn discriminative representations while retaining normal structures. Experiments on multiple real-world datasets shows that CALAD consistently outperforms existing methods, particularly under distribution shift scenarios. We provide the code for reproducibility at https://github.com/hirundo1218/CALAD
Reflections on the design, applications and implementations of the normative specification language eFLINT
arXiv:2511.12276v2 Announce Type: replace Abstract: Checking the compliance of software against laws, regulations and contracts is increasingly important and costly as the embedding of software into societal practices is becoming more pervasive. Moreover, the digitalised services provided by governmental organisations and companies are governed by an increasing amount of laws and regulations, requiring highly adaptable compliance practices. A potential solution is to automate compliance using software. However, automating compliance is difficult for various reasons. Legal practices involve subjective processes such as interpretation and qualification. New laws and regulations come into effect regularly and laws and regulations, as well as their interpretations, are subjected to constant revision. In addition, computational reasoning with laws requires a cross-disciplinary process involving both legal and software expertise. This paper reflects on the domain-specific language eFLINT developed to experiment with novel solutions to these challenges. Specifically, the language has been developed to experiment with the abstract syntax and semantics of a language supporting different types of reasoning for various applications. The language combines declarative and procedural elements, formalises connections between legal concepts and computational concepts, and is designed to automate compliance checks before, during and after a software system runs. The various design goals and applications areas for the language give rise to (conflicting) requirements. This paper presents and reflects on the current design of the language by recalling applications and requirements. As such, this paper reports on results and insights of an investigation that can benefit language developers within the field of automated compliance.
Heterogeneous Sheaf Neural Networks
arXiv:2409.08036v3 Announce Type: replace Abstract: Heterogeneous graphs, whose nodes and edges can belong to different types and feature spaces, arise in many real-world domains, including biology, recommendation, social networks, and computer systems. Existing heterogeneous graph neural networks typically handle this heterogeneity at the architectural level through relation-specific modules, meta-path machinery or type-aware attention, which often leads to increasingly specialised parameter-heavy designs. In this work, we propose HetSheaf, a framework for learning heterogeneous graphs through cellular sheaves. Instead of encoding heterogeneity solely in the architecture, HetSheaf represents it directly in the underlying data structure by assigning type-aware local feature spaces and learning restriction maps conditioned on node features, node types, and edge types. To support graph-level prediction, we further introduce SheafPool, a universal stalk-space readout that aggregates node representations while being invariant to local changes of basis, thereby making graph classification with sheaf networks well-defined and achieving an F1 Score up to 42 percentage points higher than mean pooling. Across a diverse suite of benchmarks (node classification, link prediction and graph classification). HetSheaf consistently achieves up to 2 percentage points higher performance (up to 94.97% Macro F1 Score on node classification and up to 99.62% on link prediction) on the Heterogeneous Graph Benchmark (HGB) framework against homogeneous (GCN, GAT, GIN, GraphSAGE), heterogeneous (R-GCN, HAT, HGT) and type-agnostic sheaf baselines, while reducing the number of parameters by up to 10$\times$.
Smoothing and spatial verification of global fields
arXiv:2412.00936v5 Announce Type: replace Abstract: Forecast verification plays a crucial role in the development cycle of operational numerical weather prediction models. At the same time, verification remains a challenge as the traditionally used non-spatial forecast quality metrics exhibit certain drawbacks, with new spatial metrics being developed to address these problems. Some of these new metrics are based on smoothing, with one example being the widely used Fraction Skill Score (FSS) and its many derivatives. However, while the FSS has been used by many researchers in limited area domains, there are no examples of it being used in a global domain yet. The issue is due to the increased computational complexity of smoothing in a global domain, with its inherent spherical geometry and non-equidistant and/or irregular grids. At the same time, there clearly exists a need for spatial metrics that could be used in the global domain as the operational global models continue to be developed and improved, along with the new machine-learning-based models. Here, we present two new methodologies for smoothing in a global domain that are potentially fast enough to make the smoothing of high-resolution global fields feasible. Both approaches also consider the variability of grid point area sizes and can handle missing data appropriately. This, in turn, makes the calculation of smoothing-based metrics, such as FSS and its derivatives, in a global domain possible, which we demonstrate by evaluating the performance of operational high-resolution global precipitation forecasts provided by the European Centre for Medium-Range Weather Forecasts.
Edge Assisted Multi-Camera Vehicle Tracking Framework for Real-Time and Scalable Deployment
arXiv:2511.13904v2 Announce Type: replace Abstract: Cameras are a core sensing modality in modern intelligent transportation systems (ITS), providing rich visual information on road-user activities. Multi-Camera Vehicle Tracking (MCVT) uses this data to reconstruct vehicle trajectories across camera networks, supporting applications such as traffic flow prediction and optimisation. However, most existing MCVT studies emphasise tracking accuracy while paying limited attention to real-time performance and scalability, both essential for real-world and city-scale deployment. To address this gap, we propose Edge-Assisted, Scalable and Efficient MCVT (EASE-MCVT), a distributed edge--server framework designed for real-time throughput and scalable operation. On the edge side, each camera stream is processed through object detection, single-camera tracking, geo-mapping and feature extraction, while only lightweight metadata, including vehicle locations and appearance features, is sent to the central server for cross-camera association. To improve both tracking accuracy and system efficiency, EASE-MCVT is optimised from algorithmic and system perspectives. Algorithmically, it introduces a dynamic workload scheme for tracklet-level feature extraction, a server-side re-match module to reconnect fragmented tracklets, and a self-supervised camera link model that learns spatio-temporal constraints to accelerate and stabilise cross-camera association. Systemically, it integrates production-oriented data engineering components to standardise deployment and data exchange for large-scale operation. To the best of our knowledge, EASE-MCVT is the first MCVT framework explicitly designed to address both real-time performance and scalability in a distributed edge--server setting. Experiments on the RoundaboutHD and CityFlow datasets demonstrate real-time throughput with competitive tracking accuracy, paving the way for city-wide real-time traffic management.
Communication Offloading on SmartNIC DPUs: A Quantitative Approach
arXiv:2605.04842v2 Announce Type: replace Abstract: SmartNIC Data Processing Units (DPUs) offer a promising solution for saving high-end CPU resources by offloading tasks to programmable cores near the network interface. In this work, we explore the feasibility of SmartNIC DPUs in supporting an asynchronous communication model called "fire-and-forget", particularly its core message routing service. We design a communication offloading engine called Buddy that decouples communication tasks from the application process. Buddy runs flexibly on SmartNIC DPUs such as the Nvidia BlueField-3 DPU and generic x86 CPUs. Our evaluation results in five applications identify the memory-to-communication ratio as a key predictor of the offloading performance. Host-dominated workloads, such as Quicksilver and Sparse Matrix Transpose, achieved up to 1.55x speedup with communication offloaded to the DPU. We further identify a 625x increase in DRAM traffic due to the absence of Direct Cache Access support on the DPU, highlighting a critical need in future SmartNIC designs.
Mid-infrared single-pixel imaging via two-photon optical encoding
arXiv:2605.23153v1 Announce Type: new Abstract: Mid-infrared (MIR) imaging offers powerful capabilities for label-free chemical analysis, yet its practical deployment remains hindered by the high cost and cryogenic complexity of conventional cameras. Two-photon absorption (TPA) provides a promising route to room-temperature MIR detection, but existing TPA imagers based on raster scanning or array detectors are constrained by slow acquisition speed or limited detection sensitivity. Here we present a scanning-free MIR single-pixel imaging approach based on non-degenerate TPA in a silicon detector. The involved spatial encoding is realized by a near-infrared structured pump with a resolution of 7 $\mu$m, thus allowing high-fidelity MIR optical modulation through the phase-matching-free nonlinear interaction. Consequently, the spatially modulated TPA response is intrinsically integrated in the single-element photodetector, which favors computational reconstruction of the impinging MIR image by correlating measured intensities and predetermined patterns. Notably, the use of advanced algorithms of compressed sensing and deep learning facilitate image recovery under sub-Nyquist sampling with a compression ratio of 10\% and photon-starved illumination with an incident light flux of 0.5 pJ/pulse. Furthermore, a multispectral imaging over 2.5-3.8 $\mu$m is manifested for chemical discrimination of plastic films. The presented architecture would offer a broadband and sensitive alternative for MIR imaging in various fields ranging from biomedical diagnostics to material inspection.
SoK: Critical Evaluation of Quantum Machine Learning for Adversarial Robustness
arXiv:2511.14989v4 Announce Type: replace Abstract: Quantum Machine Learning (QML) integrates quantum computational principles into learning algorithms, offering improved representational capacity and computational efficiency. However, the security and robustness of QML systems remain underexplored, particularly under adversarial conditions. We present the first comprehensive systematization of adversarial robustness in QML, combining conceptual organization with empirical evaluation across black-box, gray-box, and white-box threat models. We implement five representative attacks: a label-flipping poisoning attack under black-box; an encoder-level indiscriminate poisoning attack and a proxy-model clean-label backdoor attack under gray-box; and a circuit-level backdoor attack (QTrojan) and gradient-based evasion attacks (FGSM and PGD) under white-box. We evaluate these attacks using a Quantum Multilayer Perceptron (QMLP) trained on MNIST and AZ-Class across circuit depths of 2, 5, 10, and 50 layers with angle and amplitude encoding schemes. Our evaluations reveal a fundamental accuracy-robustness trade-off. Amplitude encoding achieves the highest clean accuracy (92.6% on MNIST and 67% on AZ-Class) but collapses under adversarial perturbations and depolarizing noise, whereas shallow angle-encoded models remain more stable. QUID is effective under noiseless conditions but weakened by noise, while the proxy-model backdoor persists unless the circuit itself is overwhelmed. Furthermore, the circuit-level backdoor fails in the multi-class setting, indicating a scalability limitation. Finally, QMLP models are more robust than Classical Multi-Layer Perceptron (CMLP) models under label-flipping attacks but substantially more vulnerable to gradient-based evasion. We conclude by proposing a threat-aware and noise-resilient framework for secure QML deployment.
ProOPF: Benchmarking and Improving LLMs for Professional-Grade Power Systems Optimization Modeling
arXiv:2602.03070v5 Announce Type: replace Abstract: Growing renewable penetration introduces substantial uncertainty into power system operations, necessitating frequent adaptation of dispatch objectives and constraints and challenging expertise-intensive, near-real-time modeling workflows. Large Language Models (LLMs) provide a promising avenue for automating this process by translating natural-language (NL) operational requirements into executable optimization models via semantic reasoning and code synthesis. Yet existing LLM datasets and benchmarks for optimization modeling primarily target coarse-grained cross-domain generalization, offering limited, rigorous evaluation in power-system settings, particularly for Optimal Power Flow (OPF). We therefore introduce \textbf{ProOPF-D} and \textbf{ProOPF-B}, a dataset and benchmark for professional-grade OPF modeling: ProOPF-D contains 12K instances pairing NL requests with parameter adjustments and structural extensions to a canonical OPF, together with executable implementations; ProOPF-B provides 121 expert-annotated test cases with ground-truth code, enabling end-to-end evaluation under both concrete and abstract OPF modeling regimes.
Mid-infrared nonlinear pinhole imaging
arXiv:2605.23154v1 Announce Type: new Abstract: Pinhole imaging is the most primitive and simplest lensless imaging paradigm, capable of transcending the physical limitations of conventional lens optics. This modality is particularly attractive for accessing a virtually infinite depth of focus or operating at extreme wavelengths. Here, we devise and implement a mid-infrared (MIR) pinhole imaging system at 3.07 $\mu$m based on nonlinear spatial filtering. Instead of using a physical aperture, the involved pinhole is optically formed by a near-infrared pump at 1.03 $\mu$m within a nonlinear crystal, which allows flexible and precise control over the effective aperture size to optimize imaging performance. Meanwhile, the MIR rays passing through the nonlinear pinhole are spectrally upconverted to facilitate sensitive imaging via a silicon camera. Consequently, the implemented upconversion pinhole imaging enables a large depth of field over 35 cm, beyond the reach of typical lens-based upconversion imagers. Furthermore, depth-resolving imaging across a large depth range is demonstrated in both the reflection and transmission modes based on time-of-flight and trigonometric techniques, respectively. The achieved capabilities -- featuring large operation depth, wide field of view, and flexible adaptability to various illumination conditions -- highlight the potential of the presented MIR imaging architecture for expansive scene detection and motion-aware applications in industrial inspection and night vision.
What Does the Server See? Understanding Privacy Leakage from Large Language Models in Split Inference
arXiv:2605.23158v1 Announce Type: new Abstract: The deployment of large language models (LLMs) on resource-constrained devices remains challenging, spurring interest in split inference, where models are partitioned between client and server to reduce computational burden and enhance privacy by transmitting only intermediate activations. However, the privacy-preserving capabilities of split inference, particularly in the context of LLMs, have not been exhaustively investigated. To fill this gap, we introduce ActInv, which solves an intermediate activation matching problem to reconstruct the client's input. Extensive evaluations demonstrate that ActInv achieves high-fidelity reconstructions, even in the presence of common perturbation-based defenses such as Gaussian noise injection and activation sparsification. To systematically understand this vulnerability, we develop Perturbation Amplification Factor (PAF), a metric for quantifying a layer's inherent resistance to reconstruction. Our analysis reveals that privacy vulnerability is not uniform across layers, with some layers being highly susceptible to leakage while others offer natural resistance. Furthermore, we demonstrate that defense effectiveness can be significantly improved by calibrating perturbation directions to maximize reconstruction error during backpropagation. Building on these insights, we design PriPert and conduct comprehensive evaluations, covering privacy, utility, and computational overhead, to demonstrate its effectiveness.
A European Multi-Center Breast Cancer MRI Dataset
arXiv:2506.00474v3 Announce Type: replace-cross Abstract: Early detection of breast cancer is critical for improving patient outcomes. While mammography remains the primary screening modality, magnetic resonance imaging (MRI) is increasingly recommended as a supplemental tool for women with dense breast tissue and those at elevated risk. However, the acquisition and interpretation of multiparametric breast MRI are time-consuming and require specialized expertise, limiting scalability in clinical practice. Artificial intelligence (AI) methods have shown promise in supporting breast MRI interpretation, but their development is hindered by the limited availability of large, diverse, and publicly accessible datasets. To address this gap, we present a publicly available, multi-centre breast MRI dataset collected across six clinical institutions in five European countries. The dataset comprises 741 examinations from women undergoing screening or diagnostic breast MRI and includes malignant, benign, and non-lesion cases. Data were acquired using heterogeneous scanners, field strengths, and acquisition protocols, reflecting real-world clinical variability. In addition, we report baseline benchmark experiments using a transformer-based model to illustrate potential use cases of the dataset and to provide reference performance for future methodological comparisons.
Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation
arXiv:2412.14642v4 Announce Type: replace Abstract: Recently, Large Language Models (LLMs) have demonstrated great potential in natural language-driven molecule discovery. However, existing datasets and benchmarks for molecule-text alignment are predominantly built on one-to-one mappings, measuring LLMs' ability to retrieve a single, pre-defined answer, rather than their creative potential to generate diverse, yet equally valid, molecular candidates. To address this critical gap, we propose Speak-to-Structure (S^2-Bench), the first benchmark to evaluate LLMs in open-domain natural language-driven molecule generation. S^2-Bench is specifically designed for one-to-many relationships, challenging LLMs to exhibit genuine molecular understanding and open-ended generation capabilities. Our benchmark includes three key tasks: molecule editing (MolEdit), molecule optimization (MolOpt), and customized molecule generation (MolCustom), each probing a different aspect of molecule discovery. We also introduce OpenMolIns, a large-scale instruction tuning dataset that enables Llama3.1-8B to surpass the most powerful LLMs like GPT-4o and Claude-3.5 on S^2-Bench. Our comprehensive evaluation of 31 LLMs shifts the focus from simple pattern recall to realistic molecular design, paving the way for more capable LLMs in natural language-driven molecule discovery. Our codes and datasets are fully accessible through the Github Repository: https://github.com/phenixace/S2-TOMG-Bench and Huggingface Datasets: https://huggingface.co/datasets/phenixace/S2-TOMG-Bench.
SyMerge: From Non-Interference to Synergistic Merging via Single-Layer Adaptation
arXiv:2412.19098v4 Announce Type: replace Abstract: Model merging combines independently trained models into a single multi-task model. However, most existing approaches focus primarily on avoiding task interference. We argue that its greater potential lies in enabling task synergy, where tasks actively improve one another. We identify cross-task performance, defined by compatibility between encoders and predictors across tasks, as a key indicator of merge quality. We demonstrate that adapting only a single task-specific layer is sufficient to induce such synergy. This study proposes SyMerge, a lightweight framework that jointly optimizes merging coefficients and a single task-specific layer. We adopt an expert-guided self-labeling objective, providing stable supervision beyond entropy minimization. Intriguingly, we further show that SyMerge successfully merges models trained from different initializations, a regime where standard methods break down. Our minimalist yet principled method achieves state-of-the-art results across vision, dense prediction, and NLP benchmarks. Our code is available at https://aim-skku.github.io/SyMerge
Scaling-Aware Adapter for Structure-Grounded LLM Reasoning
arXiv:2602.02780v3 Announce Type: replace Abstract: Large language models (LLMs) are enabling reasoning over 2D and 3D structures, yet existing methods remain modality-specific and typically compress structural inputs through sequence-based tokenization or fixed-length query connectors. Such architectures either omit the geometric grounding requisite for mitigating structural hallucinations, or impose inflexible modality fusion bottlenecks that concurrently over-compress and suboptimally allocate structural tokens, thereby impeding the realization of generalized all-atom reasoning. We introduce Cuttlefish, a unified multimodal LLM that grounds language reasoning in geometric cues while scaling modality tokens with structural complexity. First, Scaling-Aware Patching leverages an instruction-conditioned gating mechanism to generate variable-size patches over structural graphs, adaptively scaling the query token budget with structural complexity to mitigate fixed-length connector bottlenecks. Second, Geometry Grounding Adapter refines these adaptive tokens via cross-attention to modality embeddings and injects the resulting modality tokens into the LLM, exposing explicit geometric cues to reduce structural hallucination. Experiments across interdisciplinary all-atom benchmarks demonstrate that Cuttlefish achieves superior performance in heterogeneous structure-grounded reasoning. Code: github.com/zihao-jing/Cuttlefish.
Error Estimation for Adaptive Mesh Refinement in Droplet Simulations
arXiv:2508.15081v2 Announce Type: replace Abstract: We present a one-dimensional shear-force-driven droplet formation model with a flux-based error estimator. The model is derived using asymptotic expansion and a front-tracking method to simulate the droplet interface. The model is then discretized using the Galerkin finite element method in the mixed form. However, the solution gradients exhibit large jumps across element boundaries and can grow rapidly due to the highly convective pinch-off process. This leads to an erroneous droplet interface and incorrect curvature. Therefore, the mesh must be sufficiently refined to capture the interface accurately. The mixed form of the governing equation naturally provides smooth interface gradients that can be used to compute the error estimate. The computed error estimate is then used to drive the adaptive mesh refinement algorithm. The efficacy of the error estimator is illustrated by comparing the droplet profiles obtained with adaptive refinement to those obtained with regular refinement. The adaptive mesh refinement approach reduces the computational cost significantly without compromising accuracy.