Forskningsradar

Science Journals

Peer-reviewade publikationer — 58142 artiklar

Motion Estimation Techniques for Volumetric Video Attribute Compression
arXiv:2607.03576v1 Announce Type: cross Abstract: Point cloud compression relies on techniques to compress both geometry and attributes. Motion-based approaches for dynamic solid point cloud geometry compression within the geometry-based point cloud compression (G-PCC) framework have achieved significant reductions in geometry rate. However, motion-based techniques for attribute compression remain underexplored, making it challenging to achieve significant reductions in the temporal redundancy of attributes. Firstly, this paper proposes a geometry-based inter-coding scheme to compress the attributes of dynamic solid point clouds. Secondly, a graph-based motion-estimation scheme for point-cloud attribute compression is proposed. Thirdly, an interpolation-free fractional-voxel motion estimation method is proposed to refine motion accuracy to fractional-voxel precision. Our experimental results on the MPEG point cloud dataset show that the proposed scheme outperforms G-PCC, GeS-TM, and V-PCC in lossless and lossy geometry conditions. We achieve average bitrate savings of $55.3\%$, $42.3\%$, and $16.5\%$ over G-PCC, GeS-TM, and V-PCC, respectively, under lossy-geometry conditions.
Conflict-Based Lazy Search for Fast Multi-Manipulator Planning
arXiv:2607.04124v1 Announce Type: new Abstract: Employing multiple manipulators can boost efficiency and accomplish tasks that a single manipulator cannot do. However, real-time planning for multiple manipulators in a cluttered workspace still poses significant challenges for planning algorithms. This article proposes a new planning algorithm called Conflict-Based Lazy Search (CBLS) for multimanipulator planning. CBLS is built on Conflict-Based Search (CBS), an efficient multiagent pathfinding (MAPF) algorithm that has shown an order of magnitude speedup over previous approaches [1], [2]. CBS addresses MAPF by solving many single-agent pathfinding (SAPF) problems. Thus, its planning time directly depends on the efficiency of the SAPF algorithm adopted. Our CBLS algorithm enhances CBS with precomputation and lazy search. First, a lazily evaluated graph with controlled sparsity is precomputed for a single manipulator. Second, we propose the Lazy Edged-based A* (LEA*) for efficient SAPF. Since edge evaluation is the computational bottleneck of manipulator planning, LEA* uses lazy search and an edge queue to reduce the number of edge evaluations. We show that LEA* is optimally vertex efficient and has improved edge efficiency compared to A*. We apply the proposed CBLS to multi-manipulator planning problems and show its superior performance by comparing it with CBS and a sampling-based algorithm, namely, RRT-Connect.
MRMS: A Multi-Resolution Memory Substrate for Long-Lived AI Agents
arXiv:2607.04617v1 Announce Type: new Abstract: Long-lived AI agents require continuity across interactions, but continuity cannot be obtained by simply extending the prompt window. An agent must preserve useful prior experience, retrieve it selectively, distinguish personal context from external evidence, and revise memory when the underlying situation changes. We propose an architectural memory substrate organized along two orthogonal axes: a representational axis spanning structured records, vector representations, and graph relations; and a temporal axis spanning short-term traces, medium-term abstractions, and long-term semantic commitments. Its key design constraint is synchronized structured-vector-graph memory: structured records govern eligibility, vector representations support recall, and graph relations adjudicate support, contradiction, and supersession before gated context projection. Its central claim is that reliable personalization is a memory design problem: useful memory is structured, selectively exposed, continuously consolidated, and epistemically labeled rather than stored as undifferentiated conversation history. Beyond the framework, we instantiate MRMS as a lightweight prototype implementing structured records, vector retrieval, temporal policies, and graph-based revision. The prototype exercises the core substrate mechanisms through pre-generation memory selection, revision, boundary enforcement, and evidence attribution under controlled long-lived interaction scenarios with explicit evidence requirements.
Model Confidence-Guided Multi-Image Fusion of Fundus Images for Diabetic Retinopathy Diagnosis
arXiv:2607.03643v1 Announce Type: cross Abstract: Purpose: Early screening for eye diseases is critical in low- and middle-income countries where access to care is limited. We investigate whether a confidence-guided, multi-image diabetic retinopathy diagnosis framework can integrate image filtering with confidence-aware predictions for reliable screening at capture. Methods: We develop a multi-image fusion method that aggregates retinal views to improve confidence and balanced accuracy. Our method uses confidence to identify unreliable predictions, prompting retakes when needed. We compare: (1) a cascaded image-quality and disease diagnosis pipeline using a single image per patient, (2) confidence-based prediction, and (3) our confidence-based multi-image fusion pipeline. All methods are evaluated using a RETFoundGreen backbone on the mBRSET (n = 1,234) and BRSET (n = 7,599) datasets. Results: At 70% coverage, our method achieves 91% balanced accuracy on mBRSET and 97% on BRSET, improvements of ~12% and ~6%, respectively, over cascade filtering. The image-quality cascade reaches sensitivities of 61% on mBRSET and 86% on BRSET, whereas our framework reaches 94% and 96%, respectively, at 50% coverage. Conclusions: Human-annotated quality labels are weakly associated with diagnostic performance, and confidence-based filtering consistently outperforms image quality-based cascaded pipelines. Translational Relevance: Using confidence-based multi-image fusion, patients receive more reliable predictions, reducing incorrect diagnoses during screening. The lightweight backbone and single inference pass per image make the framework compatible with low-latency mobile screening systems in resource-limited settings.
Sequential Correlations Change In-Context Learning: Effective Context Length and Architectural Mismatch
arXiv:2607.03660v1 Announce Type: cross Abstract: Modern sequence models have a striking capacity for in-context learning (ICL); they can perform new tasks based only on examples given in the prompt. Understanding how this ability emerges requires theory that captures important properties of natural data. Linear regression has served as a useful sandbox for ICL theory, but existing work has largely focused on prompts with independent examples. In this work, we extend this setting to sequentially correlated data, a basic feature of real sequences. We present a solvable model based on linear attention and test our predictions on realistic transformer architectures. We identify two distinct effects: First, when the query token is independent of the context, within-context correlations induce an effective context length: correlated prompts behave like shorter i.i.d. prompts. Second, when the query is also correlated with its context, test error is reduced, particularly for softmax attention when compared to linear attention. These results suggest that correlated prompts alter not only the effective sample size of in-context learning, but also which attention architectures are best matched to the task.
Spectral Mixture Modeling with Laboratory Near-Infrared Data II: Effects of Grain Size and Implications for Europa
arXiv:2607.03668v1 Announce Type: cross Abstract: Spectral analysis using linear mixture (LM) and radiative transfer-based (RT) intimate mixture modeling based on Hapke theory at near-infrared wavelengths are applied to estimate the abundance of surface materials on Europa. Previously, Emran (2026) compared these approaches against the laboratory spectra of H$_2$O ice and H$_2$SO$_4$$\cdot$8H$_2$O mixtures with $\sim$100 $\mu$m grains. Here, the effect of particle size on spectral modeling accuracy was assessed using laboratory spectra of H$_2$O ice mixtures with small ($\sim$70 $\mu$m spherical) and coarse ($\sim$1 mm irregular) grains, measured over the $\sim$1.2-2.5 $\mu$m wavelength range at 100 K and 120 K (Stephan et al., 2021). Modeled abundance estimates at both temperatures show consistent trends across all mixing ratios, with only minor temperature-dependent variations. The discrepancy in abundance estimates from both LM and RT models remains within $\pm$10% across all mixtures, with the error reduced to $\pm$5% when fine grains dominate. Across all mixtures, the average difference between RT- and LM-derived abundance estimates remains within $\pm$2% for mixtures containing both small and large grains. In contrast, mixtures composed solely of smaller grains render larger deviations between the models, with RT producing more accurate estimates (Emran, 2026) -- indicating that the presence of coarse H$_2$O ice grains minimizes abundance differences between LM and RT modeling. Thus, I posit that Hapke-based RT modeling is the preferred spectral modeling approach -- regardless of grain size or compositional mixture -- for constraining Europa's surface composition. Nonetheless, LM modeling remains a reliable approach for compositional analysis of terrains containing H$_2$O ice with $\sim$mm-sized grains.
Isotropy and Galilean invariance of Lattice Boltzmann Method: Theoretical and numerical analysis using oblique dipole benchmark *
arXiv:2607.03222v1 Announce Type: cross Abstract: This work focuses on the two-dimensional, nine-velocity (D2Q9) lattice Boltzmann model. First, we show that the D2Q9 scheme cannot achieve secondorder accuracy unless the cubic velocity terms are neglected, and we explain how some of these parasitic terms can be eliminated. Second, we demonstrate that the standard choice of the equilibrium distribution has no effect on the equivalent PDE at second order. Finally, we numerically investigate the effect of these cubic terms and study different choices of equilibrium distributions using a new benchmark called the Oblique Dipole Benchmark, which describes obliquely propagating 2D vortex dipoles with periodic boundary conditions.
Computational Oncology of Chemotaxis-Driven Tumour--Immune Spatial Patterning and Stability
arXiv:2607.03813v1 Announce Type: cross Abstract: Spatial tumour--immune heterogeneity is a key feature of solid-tumour progression, immune infiltration, and immune exclusion. We develop a computational oncology model in which tumour cells, immune effector cells, and a chemokine signal interact through a reaction--diffusion--chemotaxis system on a bounded tissue domain with no-flux boundaries. Chemokine is produced by tumour cells and tumour--immune contact, recruits immune cells, and guides chemotactic migration. After nondimensionalization, we establish positivity, a tumour-density bound, and immune/chemokine mass estimates. We identify the tumour-free equilibrium, derive the immune-control threshold $\sigma_0>\delta$, and reduce coexistence to a scalar equation. Linear stability analysis about coexistence yields a mode-wise dispersion relation in which chemotaxis appears as a wavenumber amplified coupling, producing finite-wavelength instability above a critical sensitivity. A conservative finite-volume scheme with upwind chemotactic flux verifies the thresholds, dominant unstable modes, sensitivity maps, positivity, convergence, and residual consistency.
GLOW-FDG: Generalized cancer LesiOn Whole-body segmentation model for $^{18}$F-FDG-PET/CT
arXiv:2607.03931v1 Announce Type: cross Abstract: Whole-body fluorodeoxyglucose positron emission tomography combined with computed tomography is widely used in cancer care, but manual lesion delineation is slow, subjective, and difficult to scale. We present GLOW-FDG, an open-source artificial intelligence model for whole-body cancer lesion segmentation in fluorodeoxyglucose positron emission tomography and computed tomography. The model was trained on 1,563 scans spanning multiple cancer types and evaluated on 185 external scans from independent institutions. Across breast cancer, nonmetastatic and oligometastatic lung cancer, head and neck cancer, and metastatic melanoma, GLOW-FDG consistently outperformed publicly available benchmark models in lesion detection, while reducing false positives and maintaining strong segmentation accuracy. Quantification of total tumor burden and total lesion glycolysis was robust across cohorts, and performance approached the variability observed between expert radiation oncologists. These results support GLOW-FDG as a generalizable tool for automated cancer segmentation and quantitative imaging biomarker extraction in whole-body imaging.
Orthogonality Edges in Strong-Coupling Quantum Work Statistics
arXiv:2607.03950v1 Announce Type: cross Abstract: Strong coupling to a reservoir can do more than shift, broaden, or dress the work peaks of a driven quantum system. When the reservoir is infrared singular, a sudden change of a local control parameter can alter the boundary condition seen by infinitely many low-energy modes, converting a quasiparticle-like threshold line into a many-body edge. We demonstrate this mechanism for the inclusive work distribution of the biased spin-boson model under a sudden bias inversion. In the independent-boson limit, the problem is exactly solvable and gives a sharp infrared classification: a super-Ohmic bath can retain a finite elastic threshold weight, whereas Ohmic and sub-Ohmic baths extinguish the elastic line through boundary orthogonality. At the Ohmic fixed point, the same exponent controls both the vanishing elastic residue and the low-work continuum. We then ask how this edge is resolved away from the static-boundary limit. Using displaced-basis exact diagonalization of logarithmically discretized baths, we find that finite tunnelling leaves an edge-like continuum over the accessible energy window, while separating two operational diagnostics of the threshold: the cumulative-continuum exponent extracted from $z$-interleaved spectra lies above the elastic-overlap exponent extracted from $z$-averaged overlaps, $\theta_C>\theta_Z$. We interpret this separation as a finite-energy crossover away from the static-boundary fixed point, not as evidence for a new asymptotic fixed point. The separation survives fitting-window variation, oscillator-cutoff checks, spectrum-size checks, and leave-one-$z$-out tests, while time-domain characteristic functions provide a compatible but non-decisive diagnostic. Finally, the same threshold edge controls the sampling cost of Jarzynski-type exponential averages, making rare low-work events increasingly important at low temperature.
Cross-Modal Fusion of OCT and OCT angiography enface for Improved Diagnostics of Diabetic Retinopathy
arXiv:2607.03959v1 Announce Type: cross Abstract: Diabetic retinopathy (DR) is a leading cause of vision impairment worldwide, highlighting the need for accurate and accessible screening tools. Optical Coherence Tomography (OCT) provides high-resolution structural information of the retina, whereas OCT angiography (OCTA) offers complementary vascular information that is highly relevant for DR diagnosis. In this study, we propose a cross-modal fusion of OCT B-scans with single-channel en face OCTA using a bidirectional cross-modal attention network for automated DR classification. Two independent datasets, OCT500 and UIC, comprising 730 subjects in total, were utilized to evaluate performance under within-dataset, combined-dataset, and cross-dataset generalization settings. A ConvNeXt V2 model trained solely on OCT images served as the unimodal baseline. In addition to ground-truth (GT) OCTA, we explored the use of translated (TR) OCTA generated from OCT scans, eliminating the requirement for dedicated OCTA hardware. Experimental results demonstrate that cross-modal fusion consistently outperforms unimodal OCT classification across all evaluation scenarios. Fusion with GT OCTA improved classification accuracy and discriminative performance, while TR OCTA achieved comparable or superior results in most settings. Furthermore, TR OCTA improved sensitivity and cross-dataset generalization, indicating enhanced robustness to domain shifts. These findings demonstrate that attention-based OCT-OCTA en face fusion provides clinically meaningful improvements for DR detection and suggest that computationally generated OCTA can serve as a practical, low-cost alternative to hardware-acquired OCTA, enabling broader deployment of high-performance retinal screening systems in resource-limited clinical environments.
NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization
arXiv:2607.03985v1 Announce Type: cross Abstract: Advanced neural technologies in speech synthesis and voice conversion (VC) have introduced severe risks to personal privacy, necessitating robust Speaker Anonymization Systems (SAS). Existing SAS approaches modify voice characteristics in the hand-crafted feature space or speaker embedding space, often struggling to provide sufficient identity variance across generated voices. In this paper, we propose NouveauVoice, a novel pseudo-speaker generation framework based on a Hierarchical Deep Variational Autoencoder (NVAE). Integrated as a standalone plug-in module on top of state-of-the-art architectures (FACodec and CosyVoice2), our approach leverages tractable sampling and the Evidence Lower Bound (ELBO) objective to synthesize highly expressive pseudo-speaker embeddings with significantly enhanced speaker diversity. Evaluating our framework under a protocol similar to the VoicePrivacy Challenge alongside Maximum Mean Discrepancy (MMD) analysis, we demonstrate that NouveauVoice achieves strong identity concealment, yielding an Equal Error Rate (EER) exceeding 38% against an automatic speaker verification attacker model. Our system shows a reasonable trade-off between strict anonymity, rich pseudo-speaker diversity, and downstream speech utility, such as intelligibility and emotional expressiveness.
Finite-Sample Closed-Loop Stability of Model Predictive Path Integral Control for Linear Time-Invariant Systems
arXiv:2607.04006v1 Announce Type: cross Abstract: We establish finite-sample closed-loop stability guarantees for Model Predictive Path Integral (MPPI) control applied to discrete-time Linear Time-Invariant (LTI) systems with additive Gaussian process disturbances. The key observation is that, for unconstrained LTI/quadratic systems with the DARE terminal cost, the exact finite-horizon MPC law has the same first control action as the infinite-horizon LQR law for every planning horizon. Thus, finite-sample MPPI can be analyzed as a stochastic perturbation of LQR. First, we show that the MPPI control law approximates the LQR feedback with high probability. The approximation error decomposes into a Monte Carlo term that decreases with the sample count and an infinite-sample temperature bias that persists at finite temperature but vanishes as the temperature is reduced. The resulting constants are written in terms of the horizon-dependent stacked cost matrices, making explicit that the finite-sample certificate is parametrized by the selected planning horizon. Second, we use a Lyapunov perturbation argument to prove practical exponential stability in expectation. On sample paths that remain in a compact Lyapunov sublevel set over a finite operating horizon, the expected state norm decays exponentially up to three residual floors: a process-noise floor, an MPPI approximation floor, and a confidence floor from the per-step sampling failure probability. The sufficient sample threshold is explicit and computable from the DARE solution, LQR stability margin, MPPI sampling parameters, temperature, and planning horizon. In the joint limit of infinite samples and vanishing temperature bias, the result recovers the stochastic LQR stability bound.
Coherent quantum control of dark excitons in hybrid metal organic chalchogenolates
arXiv:2607.04080v1 Announce Type: cross Abstract: Artificial atom-like systems are a promising candidate for next generation quantum processing. Among them, dark excitons exhibit one of the longest lifetimes at high temperatures. Here, we demonstrate coherent control of dark excitonic states in metal-organic chalcogenolates (MOChas) by using an ultrafast pulse shaper at room temperature. These dark exciton states are optically accessed via two-photon absorption and directly read out with a four-wave mixing process. The system is described by a non-perturbative, two-photon Hamiltonian based on well-known atomic physics and applied to a three level system comprised of two dark excitons. Empirical and theoretical state specific optical access is shown via a simple optical pulse shape. The developed Hamiltonian-based description is a first step towards a quantum processing platform using three-level systems and two photon transitions, one example being dark excitons in the MOCha silver benzeneselenolate (mithrene). Simple conditions for gate operations are laid out and described.
Deep Learning-Based Characterization of Detonation-Cell Size Distributions in Soot-Foil Records
arXiv:2607.03764v1 Announce Type: cross Abstract: The geometric size and regularity of detonation cells are key physical parameters for characterizing detonation waves. Traditional manual measurement of soot foils is time-consuming and subjective, while existing computer vision techniques often exhibit poor generalization on real experimental images with high noise, blurred boundaries, and severe overlapping. To address this, we propose a novel method for automated recognition and high-order feature extraction of detonation cells based on deep learning instance segmentation (Mask R-CNN). By constructing a custom heterogeneous dataset (numerical simulations and physical experiments) and integrating transfer learning, the model achieves accurate pixel-level mask prediction within highly noisy flow fields. Results indicate high pixel-level agreement in benchmark validations and strong robustness against noise in complex real-world soot foils. Predicted average cell sizes agree well with manual measurements, yielding relative errors under 2% and 3.5% for regular and irregular conditions, respectively. Sensitivity ablation experiments confirm the model's scale adaptability and guided the establishment of a standardized preprocessing paradigm for appropriate image patching. Overcoming the limitation of extracting only global average sizes, this model achieves automated tracking of the transient spatial evolution of cell sizes along the propagation direction. Furthermore, it quantitatively extracts high-order regularity features, such as the irregularity index (RI) and standard deviation of cell deflection angles, demonstrating consistency with theoretical expectations. The proposed method enhances the efficiency and objectivity of statistical analysis, providing a powerful data extraction tool for experimental and numerical soot foils.
A Cross-Platform Analysis of High-Performance Quantum Error Correction Codes
arXiv:2607.04082v1 Announce Type: cross Abstract: The theory of quantum error correction was established decades ago. Yet the limitation of the quantum computing platforms in terms of noise level and available physical qubit count persists, which greatly hinders the development of scalable quantum computing systems. In this paper, we present analytical estimates of logical error rates of advanced QEC codes across leading hardware platforms and distributed quantum computing systems using a simple but unified framework. The analysis captures two dominant contributors to logical error: code structure and two-qubit gate overhead. The framework provides a fast estimate of logical error rates and identification of dominating factors in different hardware platforms, such as circuit volume, routing overhead, inter-QPU operations, or asymmetric noise protection. We show that several qualitative trends observed in larger-scale simulations can be reproduced and interpreted analytically within this framework. We further demonstrate that the framework can be used to find the sweet spot design region of distributed QEC, which is critical for the design of distributed quantum computing systems.
DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech
arXiv:2607.04140v1 Announce Type: cross Abstract: Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness by propagating local errors and hallucinations. This limitation stems from their left-to-right AR commitment: each token must be determined before future speech-token context is available. However, such ordering is not an inherent requirement for TTS, as the full input text is available before synthesis. In this paper, we introduce DELTA-TTS, a lightweight LoRA-based adaptation framework that converts a pretrained AR TTS model into a discrete diffusion language model (dLLM) for confidence-ordered speech-token decoding. To better capture the local structure of speech, DELTA-TTS incorporates a convolution module that injects local acoustic context, together with a $1/t$-weighted training objective and a time-shifted inference schedule that defer low-confidence positions to later steps. Trained on only $585$ hours of LibriTTS, DELTA-TTS achieves a $\textbf{1.75}\%$ WER on Seed-TTS test-en, outperforming its AR backbone while generating tokens $\textbf{3.3}\times$ faster. Further analysis shows that DELTA-TTS produces sharper text--speech alignment, increases overall decoding confidence, and mitigates hallucinations observed in AR generation.
Sequential cable constructions and linear rank-width
arXiv:2607.04141v1 Announce Type: cross Abstract: We introduce split-free cable terms and cable plays, a sequential graph-construction language whose live cables impose uniform GF(2)-row behaviour across the current cut. Every play of width w gives a birth-order layout whose cutrank is at most half of w, rounded down, so the sequential split-free width is at least twice the linear rank-width. At the first nontrivial level we prove an exact characterization: a connected graph with at least two vertices has linear rank-width at most one exactly when it admits a stream, equivalently a singleton-birth play of width at most four. We show that unrestricted term width and sequential width differ unboundedly on trees, calibrate the construction on the net graph, and formulate an affine upper-bound conjecture relating sequential split-free width to linear rank-width. For the rank-two case we prove a two-accumulator scheduling criterion that yields width-six plays under a natural future-uniformity hypothesis.
FedProIn: Mitigating Client Drift for Learnable Prototypes in Federated Medical Imaging
arXiv:2607.04158v1 Announce Type: cross Abstract: Federated learning (FL) is severely hindered by statistical heterogeneity due to variations in scanners, acquisition protocols, and patient populations. Such non-IID data induces client drift during local optimization, leading to unstable convergence and suboptimal global models when parameter-based aggregation is applied. We propose a prototype-based, influence-aware federated learning framework (FedProIn) that uses multiple learnable class prototypes to capture shared semantic structures across heterogeneous clients. We introduce feature divergence loss and prototype contrastive loss to mitigate client drift by decomposing it into feature drift and prototype drift. In addition, we propose a normalized influence aggregation strategy that adaptively weights client prototypes according to their contribution to the global representation, reducing the impact of biased or low-quality updates. Experimental results on two publicly available medical datasets, HAM10000 and Matek-19, demonstrate that FedProIn achieves accuracies of (83.5% IID, 81.1% non-IID) on HAM10000 and (96.2% IID, 95.8% non-IID) on Matek-19, respectively, outperforming existing baselines in both conditions. Our code is available at https://github.com/harsh-kmr/FedProIn.
Three-Phase Evaluation of AI-Assisted Software Development Life Cycle
arXiv:2607.05125v1 Announce Type: new Abstract: This paper presents an exploratory evaluation of how increasing levels of AI autonomy affect software development productivity, requirement adherence, and developer cognitive workload. A team of four developers reimplemented the same full-stack web application across three sequential phases: partial AI-assisted development using GitHub Copilot, an AI-exclusive workflow using GitHub Copilot, and an AI-exclusive workflow using AWS Kiro. Evaluation metrics included development effort (hours), requirement adherence (RITM score), AI-interaction efficiency, and NASA-TLX workload measures. Across phases, higher levels of AI autonomy were associated with reduced development effort, improved requirement adherence, and lower self-reported mental workload, while developer frustration increased modestly. The AWS Kiro phase achieved the strongest overall performance on most measured dimensions, suggesting that tooling architecture may influence outcomes independently of AI autonomy level.
An End-to-End Explainable AI Framework with Automated LLM-Based Natural Language Explanation Generation for Energy Systems
arXiv:2607.04374v1 Announce Type: new Abstract: Explainable AI (XAI) is important for deploying machine learning systems in domains where stakes are very high and where transparency, trust and accountability are critical. Although black box models like deep neural networks often perform with high efficiency, interpreting their decisions remains as a difficult task. This paper proposes a reusable end-to-end XAI framework that is the combination of prediction, explanation generation, evaluation and converting these explanations into natural language text of explanation which can be easily understood by the non-technical stakeholders as well. This framework initially trains deep neural network for both classification and regression tasks. Local and global explanations are generated using XAI algorithms, including Local Interpretable Model-agnostic Explanation (LIME) and SHapley Additive exPlanations (SHAP) respectively. To evaluate these explanations, we use fidelity and stability metrics to know how accurately and consistently explanations reflect the model behavior. The generated explanation includes feature importance scores, prediction specific attributes, and then transformed into a structured input to the Large Language Model (LLM), which generates a natural language explanation through which everyone can understand the explanations generated by XAI algorithms. This framework is tested on power system fault dataset detection dataset and building energy labels dataset. For fault detection, the neural network model achieved 99% accuracy with ROC-AUC score of 1.00. For building energy prediction, model achieves R2 score of 0.67. These findings say that the proposed approach produces a stable and faithful explanations while improving the interpretability of black box model to everyone with the help of LLMs.
SurgAM: Surgical Affordance Map Prediction with Multimodal Feature Fusion for Robot Autonomy
arXiv:2607.04378v1 Announce Type: new Abstract: Surgical automation is being increasingly studied, yet bridging visual scene understanding with autonomous action planning remains a fundamental challenge. While much research effort has been made on scene perception (e.g., tool recognition and scene segmentation), understanding and predicting actionable possibilities for surgical automation is still underexplored. In this paper, we introduce surgical affordance prediction, which identifies actionable regions for fundamental surgical actions from visual data. Specifically, a novel adaptive feature fusion framework is proposed that leverages the complementary strengths of a self-supervised vision transformer encoder for its superior semantic understanding and a large-scale generative model encoder for its spatially-aware capability. Furthermore, we introduce a hierarchical prompt learning mechanism to adapt to varying procedural contexts. Finally, a scene-guided attention decoder is proposed to focus on critical surgical areas while suppressing background distractions. To validate the effectiveness, we established a new dataset, derived from publicly available surgical datasets with affordance annotations for three basic surgical actions: aspiration, clipping, and retraction. Extensive experiments demonstrate that our approach achieves state-of-the-art performance. Moreover, we validate our framework's applicability for downstream automation on a realistic lung and prostate phantom, and results show that the predicted affordance maps successfully enable autonomous surgical actions.
Necklaces and Lyndon words in colexicographic order
arXiv:2607.05324v1 Announce Type: cross Abstract: We present the first constant-amortized-time algorithms for generating all length-$n$ necklaces and Lyndon words over a $k$-letter alphabet in colexicographic order, for arbitrary $k\geq 2$. Our approach introduces a novel class of words called \emph{quasinecklaces}, which serve as an easily generated superset of necklaces through which all necklaces can be efficiently identified. We derive a formula for the number $Q_k(n)$ of length-$n$ quasinecklaces and show that $Q_k(n)$ is proportional to the number of length-$n$ necklaces, which is the key property needed to achieve constant amortized time. We also apply our results to efficiently generate a well-known de Bruijn sequence and efficiently generate necklaces and Lyndon words subject to a weight constraint.
Light Coils: MRI with Fully Optical Data and Power Transmission
arXiv:2607.04211v1 Announce Type: cross Abstract: In MRI, dense receiver coil arrays with a high number of coil elements are used to efficiently detect and encode the signal. Further increasing the number of coils is hampered by electrical cabling and massive electronics that introduce electromagnetic coupling, integration complexity and even safety constraints. Here we introduce the novel Light Coils concept, a fully optical MRI receive architecture in which data transmission, front-end power delivery, and coil detuning are all implemented optically, thereby reducing the massive galvanic cabling to a few optical fibers. For signal encoding, Mach-Zehnder modulators (MZM) are used to convert the MR signal from each coil onto a C-band optical carrier. The preamplifiers are driven via a power-over-fiber (PoF) system that uses a high-efficiency photovoltaic (PV) cell for optical-to-electrical power conversion. A pulse-sequence-triggered optical path controls active detuning. Jointly optimizing modulator bias, optical power and front-end gain under realistic receiver chain conditions, Light Coils can match the signal-to-noise ratio (SNR) of conventional RF coil systems with galvanic cables at MZM input powers of 5-10mW and photonic power converter inputs of 80-100mW. At a clinical 3T MRI system, we show in vivo human brain imaging with a single-channel Light Coil element with an image quality and SNR comparable to a conventional coaxial readout using the identical coil element. Extending the concept to a four-channel array using dense wavelength-division multiplexing over a single fiber, we demonstrate wavelength-selective routing with inter-channel optical isolation exceeding 28dB, reduced noise correlation compared with the galvanic reference, and parallel imaging. These results establish a scalable route towards lightweight, modular, and potentially ultra-dense MRI receive arrays based on integrated photonics and power-over-fiber.