Forskningsradar

Science Journals

Peer-reviewade publikationer — 56942 artiklar

Counting Hamiltonian Paths in 3-Regular Planar Graphs
arXiv:2606.07844v1 Announce Type: cross Abstract: We introduce two infinite families of 3-regular planar graphs. Both families are conceptual adversaries to the Pohl-Warnsdorf algorithm for finding Hamiltonians. We provide a closed form calculation of the number of Hamiltonians.
A Polarization-Decomposed Method for Simulating Inhomogeneous Birefringence in Laser-Interferometric Gravitational-Wave Detectors
arXiv:2606.08943v1 Announce Type: new Abstract: Birefringence in test mass substrates is an emerging limitation for current and future laser-interferometric gravitational-wave detectors, particularly as detectors move toward higher circulating power, cryogenic operation, and crystalline optical materials. Spatially varying birefringence alters both the polarization state and spatial mode content of the intracavity field, reducing interference contrast and coupling into length and alignment control signals. Accurate modeling of these effects is complicated by the fact that most frequency-domain simulation tools employ scalar modal propagation and lack native support for polarization and two-dimensional substrate maps. In this work, we present a practical and general method for simulating inhomogeneous birefringence without modifying existing simulation frameworks. The approach represents the two polarization components as independent scalar fields and introduces their coupling through an equivalent triple-Mach-Zehnder construction that reproduces the Jones matrix of a birefringent medium. We demonstrate the method using realistic birefringence maps of the KAGRA sapphire input test masses. The technique is compatible with any frequency-domain interferometer model and enables efficient birefringence studies for next-generation gravitational-wave detectors.
From Hazard Functions to Language Space: Cox-Supervised Distillation of Survival Risk into a Large Language Model
arXiv:2606.08945v1 Announce Type: new Abstract: We investigate whether information about time-to-event risk estimated by a Cox proportional hazards model can be transferred into a generative large language model. We propose a text-based survival modelling pipeline in which structured clinical covariates are converted into text prompts and a Qwen-based large language model is fine-tuned to generate patient-specific survival risk using Cox model predictions as a training target. Across GBSG2, ACTG320, and WHAS500, the model achieves competitive held-out discrimination and calibration despite being trained as a text-generation task rather than with a conventional survival-analysis loss. We further analyse the geometry of the model's hidden states, where t-SNE visualisations reveal smooth risk gradients in latent space, suggesting that the model represents survival risk as a continuous structure rather than isolated risk categories. Together, these findings suggest that large language models can internalise survival-risk structure while supporting calibrated prediction, providing a route towards time-to-event reasoning in language models.
Are Classification Robustness and Explanation Robustness Really Strongly Correlated? An Analysis Through Input Loss Landscape
arXiv:2403.06013v2 Announce Type: replace Abstract: This paper delves into the critical area of deep learning robustness, challenging the conventional belief that classification robustness and explanation robustness in image classification systems are inherently correlated. Through a novel evaluation approach leveraging clustering for efficient assessment of explanation robustness, we demonstrate that enhancing explanation robustness does not necessarily flatten the input loss landscape with respect to explanation loss - contrary to flattened loss landscapes indicating better classification robustness. To deeply investigate this contradiction, a groundbreaking training method designed to adjust the loss landscape with respect to explanation loss is proposed. Through the new training method, we uncover that although such adjustments can impact the robustness of explanations, they do not have an influence on the robustness of classification. These findings not only challenge the prevailing assumption of a strong correlation between the two forms of robustness but also pave new pathways for understanding relationship between loss landscape and explanation loss.
Federated Large Language Models: Current Progress and Future Directions
arXiv:2409.15723v3 Announce Type: replace Abstract: Large Language Models have achieved impressive performance across diverse applications, yet their training typically depends on centralized data collection, raising serious privacy and governance concerns. Federated Learning offers a decentralized alternative by enabling multiple clients to collaboratively train shared models without exposing raw local data. However, integrating FL with LLMs introduces new challenges, including data heterogeneity, convergence instability, communication overhead, and computational constraints. This survey provides a comprehensive and up-to-date overview of Federated Learning for Large Language Models (FedLLM). We systematically review recent advances, with particular emphasis on federated fine-tuning and federated prompt learning, and analyze how existing methods address efficiency, personalization, and security challenges. We further summarize emerging directions such as federated pre-training and federated agents. Our goal is to offer a structured perspective on this rapidly evolving field and to highlight promising avenues for future research.
Coop-WD: Cooperative Perception with Weighting and Denoising for Robust V2V Communication
arXiv:2505.03528v2 Announce Type: replace Abstract: Cooperative perception, leveraging shared information from multiple vehicles via vehicle-to-vehicle (V2V) communication, plays a vital role in autonomous driving to alleviate the limitation of single-vehicle perception. Existing works have explored the effects of V2V communication impairments on perception precision, but they lack generalization to different levels of impairments. In this work, we propose a joint weighting and denoising framework, Coop-WD, to enhance cooperative perception subject to V2V channel impairments. In this framework, the self-supervised contrastive model and the conditional diffusion probabilistic model are adopted hierarchically for vehicle-level and pixel-level feature enhancement. An efficient variant model, Coop-WD-eco, is proposed to selectively deactivate denoising to reduce processing overhead. Rician fading, non-stationarity, and time-varying distortion are considered. Simulation results demonstrate that the proposed Coop-WD outperforms conventional benchmarks in all types of channels. Qualitative analysis with visual examples further proves the superiority of our proposed method. The proposed Coop-WD-eco achieves up to 50% reduction in computational cost under severe distortion while maintaining comparable accuracy as channel conditions improve.
Hardware-aware Low-latency Quantum Compilation with Data-driven Lightweight Error Detection for Early Fault-Tolerant Systems
arXiv:2606.07666v1 Announce Type: cross Abstract: Noisy intermediate-scale quantum (NISQ) processors are entering an early fault-tolerance regime where full quantum error correction carries prohibitive resource costs, yet lightweight error detection can meaningfully improve algorithmic success rates. Existing compilation and error-detection toolchains treat these concerns in isolation, with no principled way to balance detection overhead against success probability under latency constraints. We present an integrated hardware-aware compilation and data-driven quantum error-detection (QED) framework that jointly optimises qubit mapping, SWAP insertion, and syndrome-schedule placement via a noise-weighted cost function and a learned multi-objective scheduler. Simulation experiments on an HPC cluster using GPU-accelerated density-matrix simulation (NVIDIA cuQuantum SDK) across VQE, phase-estimation, and Grover benchmarks, three noise profiles, and circuit sizes of 6-20 qubits (depths 10-160), show that joint co-design raises algorithmic success probability by up to 68 percent (95 percent CI: 60 percent to 76 percent) over SABRE on an 8-qubit VQE instance with post-selection.
Benchmarking Quantum Algorithmic Resilience for CVaR Portfolio Optimization: The Expressibility-Coherence Trade-off
arXiv:2606.07727v1 Announce Type: cross Abstract: Quantum combinatorial optimization offers theoretical advantages for complex financial modeling, but physical implementation on Noisy Intermediate Scale Quantum (NISQ) devices is severely constrained by hardware topology. This study presents a hardware benchmarking analysis between a Hardware Efficient Variational Quantum Neural Network (HE-VQNN) and the Warm Start Quantum Approximate Optimization Algorithm (WS-QAOA) for a hybrid Mean Variance and Conditional Value at Risk (CVaR) portfolio objective. By implementing a novel classical quantum hybrid proxy matrix to bypass the CVaR auxiliary qubit bottleneck, we map up to 16 assets from the NIFTY 50 index onto an IBM heavy hex processor. We systematically quantify algorithmic resilience to the "SWAP tax" incurred during routing. Empirical results reveal a critical operational trade-off: WS-QAOA provides exact theoretical mapping but suffers catastrophic hardware decoherence due to exponential nonlocal gate overhead. Conversely, HE-VQNN preserves hardware coherence but lacks the mathematical expressibility to capture dense tail risk asset correlations. This study exposes the limitations of dense financial optimization on current architectures forces an nonviable choice between algorithmic inexpressibility and hardware decoherence. This is indicative of a deeper limitation as to what can and cannot be done with NISQ computers lacking in all-to-all connectivity.
Feasibility to detect rapid change and disappearance of seagrass: Lessons from nearly 80 years of vegetation change in the Ako, Seto Inland Sea, Japan
arXiv:2606.07949v1 Announce Type: cross Abstract: This study analyses the Ako tidal flat in the Seto Inland Sea, Japan, where nearly all Zostera marina disappeared within a single year in 2025. Using aerial photographs from the 1940s onward, high-resolution satellite imagery, GRUS images (2.5-5 m), and monthly Sentinel-2 composites (10 m), we reconstructed approximately 80 years of seagrass distribution. YOLO-based segmentation using deep learning achieved high accuracy (overall accuracy >= 0.9) across these datasets; although species could not be discriminated, the models captured the major temporal dynamics in vegetation area. The long-term mean seagrass area was 6.8 ha, but values fluctuated widely, from 3.5 ha in 1974 to 41.3 ha in 1989 except 0.2 ha in 2025. Sentinel-2 composites from 2019 to 2026 revealed clear seasonality, with vegetation increasing in early summer and declining from autumn. In 2025, however, the area decreased sharply after summer and remained anomalously low throughout the winter of 2025-2026. Our results, indicating that the 2025 event was not a normal fluctuation but a rapid ecosystem shift involving the loss of the dominant canopy-forming species, most plausibly driven by regionally elevated summer water temperatures. The findings also have implications for seagrass Essential Ocean Variables (EOVs) and the State of Nature (SoN) metrics used in TNFD-aligned nature-related disclosures. Unlike forests, seagrass meadows require finer temporal resolution because both pronounced seasonality and abrupt collapse strongly influence area-based indicators. Therefore, in addition to previously noted issues such as species-level classification accuracy, we recommend that (1) baselines be defined over the longest available record and justified ecologically, (2) seasonal standardization be applied before inter-annual comparisons, and (3) years with extreme area anomalies be flagged rather than used as reference points.
A spectral model of power-law decay in natural and engineered systems
arXiv:2606.08342v1 Announce Type: cross Abstract: We present a first-principles spectral mechanism for the emergence of nonextensive $q$-exponential dilution and power-law relaxation in non-ideal transport systems. By modeling an incompletely mixed reactor as a layered diffusion matrix with an absorbing boundary, we demonstrate that macroscopic power-law tails depend on the geometric interaction between the initial tracer placement and the domain's boundary configuration. For a one-dimensional system, an asymmetric, volumetrically distributed initial concentration profile projects onto the low-wavenumber eigenmodes, generating an emergent Gamma distribution of relaxation rates; at an infinitesimal boundary layer thickness ($\Delta z \to 0$), this profile yields the nonextensive $q$-exponential decay function exactly across the entire temporal domain with $q = 5/3$. Extended to $d$ dimensions under a highly localized, boundary-adjacent singular initial condition, the resulting scaling exponents and corresponding $q$ values depend explicitly on the spatial configuration of the absorbing boundaries. However, in the one-dimensional limit ($d=1$), these distinct initial states and boundary formulations intersect, rendering the $q=5/3$ exponent geometrically invariant. Our approach establishes a clear connection between linear diffusion transport and nonextensive statistical mechanics, showing how heavy-tailed transport can be derived from boundary geometry and spectral dimensionality.
Dust to Dust: Prospects for Passive Technosignatures as Relics of ETI
arXiv:2606.08373v1 Announce Type: cross Abstract: Technological societies are separated in time, not just space -- that is the lesson of the Drake equation. Might the best way to seek them be to find technosignatures that persist long after their creators? I present work I and my collaborators have done on the idea of passive technosignatures, requiring no upkeep from an active society. These range from microscopic to galactic in scale, including specular reflections from shiny artifacts in the Solar System, lens flares from X-ray binaries, and the survivability of Dyson swarms. I discuss prospects for detecting these technosignatures. In the end, what we may be left with are the end products of collisional cascades: dust.
Quantum Kravchuk Transform using $\mathfrak{su}(2)$ fast-forwarding
arXiv:2606.08443v1 Announce Type: cross Abstract: We present a quantum algorithm for the Kravchuk transform that scales logarithmically in both the dimension and the inverse of the error parameter. The quantum Kravchuk transform maps computational basis states to states with amplitudes proportional to Kravchuk functions. We achieve this by combining two key techniques: the structural relationship between the Kravchuk transform and the Lie algebras $\mathfrak{su}(2)$, and a recent fast-forwarding simulation method for $\mathfrak{su}(2)$ operators in the oscillator representation. More precisely, we first establish the map from Kravchuk transform in computational basis to $\mathfrak{su}(2)$ in Fock basis. Then built on this connection, we apply the fast-forwarding to achieve an efficient quantum Kravchuk transform.
Beyond FLOPs: Benchmarking Real Inference Acceleration of LLM Pruning under a GEMM-Centric Taxonomy
arXiv:2606.09080v1 Announce Type: new Abstract: Pruning has emerged as a dominant paradigm for accelerating large language model (LLM) inference, spanning a broad spectrum of methods that remove computation across tokens, layers, heads, dimensions, and attention patterns. Despite sharing the same objective, these pruning approaches induce fundamentally different execution behaviors, causing realized speedups to depend heavily on hardware and kernel implementations. Consequently, the practical acceleration benefits of different pruning families remain poorly understood. In this work, we introduce a GEMM-centric taxonomy that reorganizes existing pruning methods according to the logical \textbf{M}, \textbf{N}, and \textbf{K} dimensions of general matrix multiplication (GEMM). Leveraging this abstraction, we build a unified benchmarking framework that enables implementation-consistent comparison across the pruning design space and systematically characterizes the acceleration--quality Pareto frontier. Our results show that static depth pruning remains the strongest Pareto-optimal baseline and stays closest to its theoretical acceleration upper bound in memory-bounded scenarios. During prefill, the frontier transitions from static depth at low quality loss (0\%--4\%), to dynamic depth at moderate loss (5\%--16\%), and finally to static width pruning at higher loss levels (17\%--26\%). These findings establish the first unified view of the practical limits of pruning-based LLM acceleration and provide guidance for future pruning research.\footnote{Code is available at https://github.com/EIT-NLP/LLM-Pruning/tree/main/PruningInferSim}
A Dual Metastable-State Encoding Architecture for Quantum Processing with $^{171}\mathrm{Yb}$ Atom Arrays
arXiv:2606.08453v1 Announce Type: cross Abstract: Neutral-atom arrays combine scalable qubit registers, long coherence times, flexible optical control, and strong Rydberg-mediated entangling interactions, making them a promising platform for quantum information processing. However, physical error rates remain a challenge, and fault-tolerant quantum error correction (QEC) requires repeated mid-circuit measurement and reset of ancilla qubits without disturbing nearby data qubits. This requirement introduces significant control and architectural overhead, making qubit encoding an important architectural decision. Here, we propose a dual metastable-state qubit encoding for $^{171}\mathrm{Yb}$ atoms that utilizes two independent qubit subspaces in the $(6s6p)\,{}^3\mathrm{P}_0$ and $(6s6p)\,{}^3\mathrm{P}_2$ manifolds. The ${}^3\mathrm{P}_0$ manifold provides a long-coherence nuclear-spin (NS) qubit suitable for storage and arithmetic operations, while the ${}^3\mathrm{P}_2$ manifold provides a hyperfine-spin (HF) qubit, with $\Delta_{\mathrm{HF}} = 2\pi \times 6.7~\mathrm{GHz}$, that enables fast Raman operations and direct state-selective imaging. Coherent shelving between the two metastable manifolds connects the qubit subspaces, allowing operations to be assigned to spectrally distinct processor zones. We simulate single-qubit and two-qubit gate fidelities in ${}^3\mathrm{P}_2$, as well as coherent shelving between the HF and NS qubit subspaces. We incorporate these physical-level estimates into an architectural resource estimation and logical-level simulation. Our approach integrates mid-circuit measurements and fast qubit operations within a single-species platform, providing a versatile framework for future fault-tolerant quantum computing with neutral-atom qubits.
Semantic and Task-Oriented V2X Communications: Pushing the Limits of V2X Networks Scalability
arXiv:2606.09126v1 Announce Type: new Abstract: Scalable Vehicle-to-Everything (V2X) networks are key to support the large-scale deployment of connected and automated mobility. However, the scalability of V2X networks is currently challenged by the limitations of existing V2X communication paradigms, which prioritize the reliable and timely delivery of the transmitted information over a careful message content selection - an approach that can potentially lead to the transmission of unnecessary information and an inefficient usage of communication resources. Semantic and task-oriented V2X communications have recently been proposed to address these scalability challenges by focusing on the content of the transmitted messages, particularly on its relevance to the intended receivers. In this paper, we numerically demonstrate that semantic and task-oriented V2X communications can substantially improve the scalability of V2X networks, increasing by up to a 4.1x factor the number of supported vehicles under high-density conditions. In addition, we show that semantic and task-oriented V2X communications can also decrease the inter-reception time between consecutive messages by up to 67% and lead to a twofold increase in the probability of successfully delivering all required relevant information to the intended receivers.
Multiversion Concurrency Control for Multiversion B-Trees
arXiv:2606.09133v1 Announce Type: new Abstract: Multiversion concurrency control (MVCC) enables scans to read data from a committed snapshot (version), reducing conflicts with write operations compared to traditional concurrency approaches. Currently, versioned records are often managed in a B$^+$-tree using version chains. However, version chains introduce overhead during scans and can still lead to conflicts between scans and writers. The multiversion B-tree (MVBT) was designed for optimal range scan performance on arbitrary versions, but has been considered impractical due to its structural complexity and, until recently, the lack of effective concurrency control. In this paper, we present the concurrent MVBT (cMVBT), a redesign of the MVBT featuring a novel concurrency control protocol that uses optimistic latches for write operations and requires no latches for range scans, while preserving all the optimality guarantees of the original MVBT. Additionally, cMVBT supports continuous garbage collection without activity spikes, seamlessly integrating free-space management. Experiments with mixed workloads derived from a standard benchmark show that the cMVBT achieves low overhead, high write throughput, and excellent range scan performance, outperforming state-of-the-art methods based on version chains.
Functional design of efficient and parallelizable combinatorial generators using convolution
arXiv:2507.03980v4 Announce Type: replace Abstract: The application of program transformation and algebraic methods to the development of efficient combinatorial optimization (CO) algorithms relies on an exhaustive combinatorial generator for the problem specification, followed by the fusion of thinning or filtering processes into this specification. However, the effectiveness of such fusion transformations critically depends on the structural compatibility between the objective function and the generator, which is highly problem dependent. In practice, when the majority of candidate solutions remain unfiltered or are not eliminated-as is the case for most intractable CO problems-the overall efficiency of the resulting fused program is largely determined by the intrinsic efficiency of the combinatorial generator. Consequently, if the specification itself exhibits suboptimal performance, the fused program will inherit a correspondingly inferior level of efficiency. We argue that a genuine designed process should also account for hardware compatibility and parallelizability-particularly the ability to support efficient parallel execution on modern hardware architectures, including multi-level cache hierarchies and GPUs. However, does achieving formal correctness necessarily conflict with designing algebraically elegant algorithms that support fusion? Can we obtain both simultaneously? In this paper, we show that techniques from functional programming, provide powerful formal tools for the systematic construction of such hardware-compatible and parallelizable combinatorial generators. This paper investigates generators for two of the most fundamental combinatorial structures-combinations and permutations-together with their natural extension to nested generators (e.g., combinations/permutations of combinations/permutations).
Region-Wise Correspondence Prediction between Manga Line Art Images
arXiv:2509.09501v4 Announce Type: replace Abstract: Understanding region-wise correspondences between manga line art images is fundamental for high-level manga processing, supporting downstream tasks such as line art colorization and in-between frame generation. Unlike natural images that contain rich visual cues, manga line art consists only of sparse black-and-white strokes, making it challenging to determine which regions correspond across images. In this work, we introduce a new task: predicting region-wise correspondence between raw manga line art images without any annotations. To address this problem, we propose a Transformer-based framework trained on large-scale, automatically generated region correspondences. The model learns to suppress noisy matches and strengthen consistent structural relationships, resulting in robust patch-level feature alignment within and across images. During inference, our method segments each line art and establishes coherent region-level correspondences through edge-aware clustering and region matching. We construct manually annotated benchmarks for evaluation, and experiments across multiple datasets demonstrate both high patch-level accuracy and strong region-level correspondence performance, achieving 78.4-84.4% region-level accuracy. These results highlight the potential of our method for real-world manga and animation applications.
Dynamic Function Configuration and its Management in Serverless Computing: A Taxonomy and Future Directions
arXiv:2510.02404v2 Announce Type: replace Abstract: The serverless cloud computing model offers a framework where the service provider abstracts the underlying infrastructure management from developers. In this serverless model, FaaS provides an event-driven, function-oriented computing service characterised by fine-grained, usage-based pricing that eliminates cost for idle resources. Platforms like AWS Lambda, Azure Functions, and Cloud Run Functions require developers to configure their function(s) with minimum operational resources for its successful execution. This resource allocation influences both the operational expense and the performance quality of these functions. However, a noticeable lack of platform transparency forces developers to rely on expert knowledge or experience-based ad-hoc decisions to request desired function resources. This makes optimal resource configuration a non-trivial task while adhering to performance constraints. Furthermore, while commercial platforms often scale resources like CPU and network bandwidth proportional to memory, open-source frameworks permit independent configuration of function resources, introducing additional complexity for developers aiming to optimise their functions. These complexities have directed researchers to resolve developer challenges and advance towards an efficient server-less execution model. In this article, we identify different aspects of resource configuration techniques in FaaS settings and propose a taxonomy of factors that influence function design, configuration, run-time cost, and performance guarantees. We conduct an analysis of existing literature on resource configuration to present a comprehensive review of current studies on function configuration. We also identify existing research gaps and suggest future research directions to enhance function configuration and strengthen the capabilities of serverless computing environments to drive its broader adoption.
Performance Evaluation of Social Learning
arXiv:2606.09176v1 Announce Type: new Abstract: Social Learning is a decentralized decision-making paradigm in which spatially dispersed agents collect streaming observations regulated by one of a finite number of models (the hypotheses). The agents are interested in assigning probability scores (the beliefs) to the possible hypotheses. To this end, the agents exchange their beliefs according to a certain communication graph. It has been shown that, under reasonable conditions on the identifiability of the decision model and the network connectivity, each agent ultimately places all the belief mass on the true hypothesis governing the data. However, several questions remain unanswered regarding the evaluation of the social learning performance. One recently adopted performance metric is the rejection rate, i.e., the rate at which the beliefs about the erroneous hypotheses vanish. One contribution of this work is to establish that the rejection rate leads to several paradoxes, which make it unsuitable as a valid performance measure. We then focus on studying the error probability measure. For a binary Gaussian problem, we derive an analytical formula characterizing the ratio between the individual agents' probabilities and the optimal Bayesian probability. The formula shows that this ratio is expressed by the product of two terms quantifying the effect of the network connectivity and the role of the prior information. As a result, an irreducible gap emerges between the decentralized and the centralized error probabilities, which is agent-dependent and does not disappear asymptotically.
Resource-aware Computation-Communication Overlap for multi-GPU ML Workloads
arXiv:2606.09200v1 Announce Type: new Abstract: The rapid growth of large-scale machine learning (ML) has made distributed training across multiple GPUs a fundamental component of modern ML systems. As model sizes and computational throughput continue to increase, communication overhead has become a dominant bottleneck in multi-GPU training, particularly when computation and communication are executed sequentially. This work explores concurrent execution of computation and collective communication using two portable runtime controls: shared-memory-driven occupancy shaping for computation kernels and elevated scheduling priority for communication kernels. Our approach regulates computation-kernel residency through per-block shared-memory allocation, leaving sufficient on-chip resources for communication kernels to make progress. In addition, assigning higher priority to communication streams ensures steady communication progress once resources become available. Experiments on NVIDIA A40, A100, H100, and AMD MI250X GPUs demonstrate that the proposed method enables effective computation-communication overlap and reduces total execution time by up to 25.5 percent, without modifying vendor libraries or kernel implementations.
AttnRegDeepLab: A Two-Stage Decoupled Framework for Interpretable Embryo Fragmentation Grading
arXiv:2511.18454v4 Announce Type: replace Abstract: Assessing embryo fragmentation is crucial for predicting IVF success, yet manual grading is prone to subjectivity, and existing AI models struggle with clinical interpretability and segmentation errors. We propose AttnRegDeepLab, a Multi-Task Learning (MTL) framework designed to solve these challenges. The model enhances a DeepLabV3+ decoder with Attention Gates to filter out cytoplasmic noise and retain sharp contour details. It also introduces a Multi-Scale Regression Head with Feature Injection, guiding the segmentation process with global grading priors to eliminate systematic area estimation errors. Based on a two-stage decoupled training strategy and a range-based loss for weakly labeled data, our method resolves MTL gradient conflicts. AttnRegDeepLab yields high grading precision and excellent segmentation quality (Dice coefficient = 0.729), avoiding the trade-off between contour integrity and grading accuracy seen under standard joint optimization. This provides a reliable, clinically interpretable tool balancing visual and quantitative accuracy.
MedVision: Benchmarking Quantitative Medical Image Analysis
arXiv:2511.18676v2 Announce Type: replace Abstract: Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.g., "Is this normal or abnormal?") or qualitative descriptive tasks. However, clinical decision-making often relies on quantitative assessments, such as measuring the size of a tumor or the angle of a joint, from which physicians draw their own diagnostic conclusions. This quantitative reasoning capability remains underexplored and poorly supported in existing VLMs. In this work, we introduce MedVision, a large-scale dataset and benchmark specifically designed to evaluate and improve VLMs on quantitative medical image analysis. MedVision spans 22 public datasets covering diverse anatomies and modalities, with 30.8 million image-annotation pairs. We focus on three representative quantitative tasks: (1) detection of anatomical structures and abnormalities, (2) tumor/lesion (T/L) size estimation, and (3) angle/distance (A/D) measurement. We show that current off-the-shelf VLMs perform poorly on these tasks. However, supervised and reinforcement fine-tuning on MedVision significantly enhances performance across detection, T/L estimation, and A/D measurement. MedVision provides a foundation for developing VLMs with robust quantitative reasoning capabilities in medical imaging.
AutoPot: Automated and massively parallelized construction of Machine-Learning Potentials
arXiv:2601.01185v2 Announce Type: replace Abstract: Machine-learning potentials (MLIPs) have been a breakthrough for computational physics in bringing the accuracy of quantum mechanics to atomistic modeling. To achieve near-quantum accuracy, it is necessary that neighborhoods contained in the training set are rather close to the ones encountered during a simulation. Yet, constructing a single training set that works well for all applications is, and likely will remain, infeasible, so, one strategy is to supplement training protocols for MLIPs with additional learning methods, such as active learning, or fine-tuning. This strategy, however, yields very complex training protocols that are difficult to implement efficiently, and cumbersome to interpret, analyze, and reproduce. To address the above difficulties, we propose AutoPot, a software for automating the construction and archiving of MLIPs. AutoPot is based on BlackDynamite, a software operating parametric tasks, e.g., running simulations, or single-point ab initio calculations, in a highly-parallelized fashion, and Motoko, an event-based workflow manager for orchestrating interactions between the tasks. The initial version of AutoPot supports selection of training configurations from large training candidate sets, and on-the-fly selection from molecular dynamics simulations, using Moment Tensor Potentials as implemented in MLIP-2, and single-point calculations of the selected training configurations using VASP. Another strength of AutoPot is its flexibility: BlackDynamite tasks and orchestrators are Python functions to which own existing code can be easily added and manipulated without writing complex parsers. Therefore, it will be straightforward to add other MLIP and ab initio codes, and manipulate the Motoko orchestrators to implement other training protocols.
Radiation damage to normal mammalian tissue in vivo with laser-driven protons at ultra-high instantaneous dose rate
arXiv:2602.20460v2 Announce Type: replace Abstract: The differential sparing of normal tissues relative to tumor control observed at ultra-high dose rates, referred to as the FLASH effect, has recently gained considerable attention. The therapeutic advantages of FLASH radiotherapy are expected to be further amplified through the use of protons and ions, which enable precise dose deposition at tumor depth while minimizing irradiation of healthy tissues proximal and distal to the target. Nevertheless, the mechanism underlying this sparing effect remains poorly understood. Laser-driven proton accelerators are capable of delivering uniquely high instantaneous dose rates in ultrashort bunches. Here, we report the first in vivo investigation of normal tissue response to laser-driven proton irradiation, with controlled exposures to 8 MeV protons, delivering total doses up to 50 Gy at 2 Gy per laser shot. Our findings reveal a reduction in tissue swelling following laser-driven proton treatment compared with X-ray irradiations at conventional dose rates. RNA sequencing identified differential gene expression associated with immune and epidermal programs following laser-driven proton irradiations at two different dose levels.