arXiv:2607.08778v1 Announce Type: new
Abstract: Alzheimer's Disease (AD) is a complex neurodegenerative disorder that continues to impact millions of people worldwide. Predicting AD conversion during the prodromal stage remains critical for disease understanding and patient care. As such, survival models are widely used for AD risk prediction, yet they are typically static predictors with limited interpretability and no capacity for natural language reasoning. In this work, we propose iLENS, an interpretable large language model (LLM) guided framework based on mixture-of-experts (MoE) for survival prediction in AD conversion. Our approach uses LLM to synthesize structured neuroimaging measurements and unstructured information to guide expert routing. Our framework demonstrates competitive predictive performance and capability in patient subtyping. Furthermore, our framework provides transparent, biologically grounded rationales for its routing decisions, bridging the gap between high-performance survival analysis and interpretable clinical decision support.
Science Journals
arXiv:2403.14926v3 Announce Type: replace-cross
Abstract: Electronic health record (EHR) systems capture a wealth of multimodal clinical data, encompassing both structured clinical codes and unstructured clinical notes. Yet, many EHR-focused studies have traditionally examined these modalities in isolation or combined them using simplistic methods, overlooking the intrinsic synergy between them. In reality, these modalities are deeply interconnected, each containing clinically relevant and complementary information that, when integrated effectively, can provide a more comprehensive understanding of patient health. Despite the success of multimodal contrastive learning in vision-language applications, its potential remains under-explored in multimodal EHR, particularly in terms of theoretical understanding. To support statistical analysis of multimodal EHR data, we propose a multimodal feature embedding generative model and design a multimodal contrastive loss to learn EHR feature representations. Our theoretical analysis demonstrates the effectiveness of multimodal learning over single-modality learning and connects the solution of the loss function to the singular value decomposition of a pointwise mutual information matrix. This connection leads to a privacy-preserving algorithm tailored for multimodal EHR representation learning. Simulation studies show that the proposed algorithm performs well under a variety of configurations. We further validate its clinical utility using real-world EHR data.
arXiv:2607.09324v1 Announce Type: new
Abstract: Identifying and assigning keywords at scale is a technical, practical, and ethical challenge for crowdsourced collections. This article reports the findings of the "Extracting Keywords from Crowdsourced Collections" project, which used the Their Finest Hour Online Archive, a crowdsourced Second World War digital collection hosted by the University of Oxford, as a case study. The project evaluated three Natural Language Processing approaches to automate keyword extraction: Named Entity Recognition, Keyword Extraction, and Topic Modelling. It tested these approaches across a range of artificial intelligence techniques, from traditional statistical methods to modern GenAI neural networks. Our quantitative and qualitative findings indicate that Natural Language Processing approaches offer real potential for keyword extraction at scale in crowdsourced collections, but that no single method offers a complete solution and that model choice significantly shapes results. We argue that in crowdsourced collections, where metadata is the direct product of engagement with living contributors, automated keyword extraction raises distinct stewardship responsibilities that must be addressed alongside technical performance. Open-weight, extractive models emerge from our evaluation as best placed to support responsible deployment, while generative AI, despite its abstractive potential, introduces accountability risks that anyone managing crowdsourced collections should weigh carefully.
arXiv:2606.14130v2 Announce Type: replace
Abstract: Safe coordination problems surface in multi-agent reinforcement learning when global safety cannot be enforced by any agent unilaterally: the admissibility of one agent's action may depend on the dynamics of other agents. Decentralised shields can enforce safety at runtime, but purely factorised permissions often exclude optimal team behaviour that is safe only through coordination. We study deterministic safety guarantees for agents trained and deployed under decentralised execution, recovering team-optimal safe behaviour without centralised runtime control. Agents have a shared global specification $\phi$ in the safety fragment of Linear Temporal Logic ($\mathsf{LTL}_{\mathsf{safe}}$ ), and select among tuples of local $\mathsf{LTL}_{\mathsf{safe}}$ obligations whose conjunction implies the global specification $\phi$. Each agent may rely on the other agents' local obligations as assumptions because the whole contract tuple is certified simultaneously and allows projection into local action masks. At learning time, a non-stationary multi-armed bandit chooses among a library of local $\mathsf{LTL}_{\mathsf{safe}}$ obligations to select the tuple that optimises team reward, all without forgoing end-to-end safety. We evaluate the approach across 6 environments and 15 algorithmic variants.
arXiv:2607.08791v1 Announce Type: new
Abstract: Designing effective multi-objective Bayesian optimization (MOBO) algorithms requires balancing many interdependent design choices whose optimal configuration is problem-dependent and typically demands deep expertise. We extend the LLaMEA framework to MOBO, using large language models as mutation and crossover operators within evolutionary strategies to generate complete algorithm implementations, with SMAC hyperparameter optimization integrated into the evolutionary loop. Across nine evolutionary runs we generated approximately 900 algorithms and benchmarked them on twelve synthetic problems (ZDT, DTLZ, WFG) and three real-world engineering problems (RE), using a BoFire qParEGO implementation as a state-of-the-art Bayesian-optimization baseline. On the synthetic suite the strongest generated algorithm attains the highest mean normalized hypervolume (0.971, vs. 0.869 for qParEGO) while requiring roughly 60x less wall-clock time; a Friedman test with post-hoc analysis places the two in a single top-performing group, and per-problem tests find the generated algorithm significantly better than qParEGO on 7 of the 12 problems and never worse, matching state-of-the-art accuracy at an order-of-magnitude lower cost. On the three unseen real-world engineering problems a generated algorithm attains the best mean normalized hypervolume (0.985, vs. 0.971 for qParEGO)--significantly better than qParEGO on two of the three problems--at roughly 3.4x lower wall-clock cost, confirming that the gains transfer beyond the synthetic regime. LLM-driven evolutionary search can thus discover algorithm designs that achieve Pareto-efficient trade-offs difficult to reach through manual design.
arXiv:2607.09389v1 Announce Type: new
Abstract: Super-critical collisionless shocks are not static structures but evolve continuously as they reflect incoming ions back upstream. The physical process responsible for this non-stationarity -- whether it is dominated by wave-like corrugation of the shock surface (rippling) or by a cyclic rebuilding of the shock transition (reformation) -- remains debated. We combine Magnetospheric Multiscale (MMS) observations of a nearly perpendicular ($\theta_{Bn}\approx89^\circ$), supercritical ($M_A\approx6$) bow shock with high-resolution two-dimensional hybrid simulations to address this question. MMS reveals repeated ion phase-space holes and intense, localized Hall electric fields. A virtual-spacecraft analysis of the simulation reproduces these signatures and shows that they arise from a self-regulating feedback cycle: strong Hall-field ion reflection builds a reflected-ion foot, which weakens the Hall field and suppresses further reflection until the foot decays and the cycle restarts. This reformation cycle, spatially organized by the two-dimensional shock structure, explains most of the observed non-stationarity.
arXiv:2607.09631v1 Announce Type: cross
Abstract: We study the computational complexity of problems that ask if a given graph admits an edge-coloring that does not contain an edge-colored clique from some fixed finite family. We show that every such problem is poly-time equivalent to a Constraint Satisfaction Problem, yielding a P vs. NP-complete dichotomy. Our main contribution lies in the reduction from the CSP to the coloring problem where we apply methods from Ramsey theory and a novel notion of cut-homotopy.
arXiv:2601.01641v2 Announce Type: replace
Abstract: We present a finite-temperature extension of density matrix embedding theory (FT-DMET) for realistic crystalline systems. We describe a practical framework for constructing extended bath orbitals, solving the embedding problem, and performing DMET self-consistency at finite temperature. To reduce computational cost, we introduce strategies based on mutual-information-guided bath truncation, controlled treatment of the thermal electron number without explicit optimization, and the use of low-temperature impurity solvers and one-shot FT-DMET in the low-temperature regime. We apply this approach to periodic hydrogen chains and square lattices to characterize their finite-temperature phases. We observe the Pomeranchuk-like effect in one dimension and enhanced stability of long-range order in two dimensions.
arXiv:2606.16770v2 Announce Type: replace
Abstract: This qualitative study explores how undergraduate women define what it means to be a physics person alongside how they describe their own physics identity. Our team analyzed 120 survey responses (97%, or 116 of the 120 respondents, identified as women or non-binary) to understand what factors are associated with an individual's physics identity. Those surveyed were attendees of Conferences for Undergraduate Women in Physics, who attended 88 unique institutions in 30 states. Our participants' conceptions of what it means to be a physics person fell into two themes: (1) people who are skilled academically and in research and (2) those who have an interest in physics ranging from an enjoyable interest to a life-devoting passion for the subject. When asked to explain their own physics identities, students responded through a personal and social lens more frequently than through an academic lens. Our analysis revealed three tensions in the development of undergraduate women's physics identity: (1) the tension between identifying entirely as a physics person and having other identities that are important to who they are, (2) the tension between the need for belonging in the physics community and their view of themselves as a physics person, and (3) the tension between their expectations of a physics person and their beliefs about their own abilities or interest.
arXiv:2506.22784v2 Announce Type: replace
Abstract: Point-pixel registration between LiDAR point clouds and camera images is a fundamental yet challenging task in autonomous driving and robotic perception. A key difficulty lies in the modality gap between unstructured point clouds and structured images, especially under sparse single-frame LiDAR settings. Existing methods typically extract features separately from point clouds and images, then rely on hand-crafted or learned matching strategies. This separate encoding fails to bridge the modality gap effectively, and more critically, these methods struggle with the sparsity and noise of single-frame LiDAR, often requiring point cloud accumulation or additional priors to improve reliability. Inspired by recent progress in detector-free matching paradigms, we revisit the projection-based approach and introduce the detector-free framework for direct point-pixel matching between LiDAR and camera views. To further enhance matching reliability, we introduce a repeatability scoring mechanism that acts as a soft visibility prior. This guides the network to suppress unreliable matches in regions with low intensity variation, improving robustness under sparse input. Extensive experiments on KITTI, nuScenes, and MIAS-LCEC-TF70 benchmarks demonstrate that our method achieves state-of-the-art performance, outperforming prior approaches on nuScenes (even those relying on accumulated point clouds), despite using only single-frame LiDAR.
arXiv:2607.09215v1 Announce Type: new
Abstract: AI coding assistants are now widely used in professional development, yet they offer only limited ways for developers to control how they behave. In this paper, we investigate what kinds of configurations experienced developers want in coding assistants, how they prioritize different types of configuration needs, and which interface mechanisms they prefer. We first synthesize product documentation and prior research on trust and personalization to compile a list of 33 configuration options, grouped into four categories: Code suggestions, System & policies, Human-assistant interaction, and Users & their personal context. We then conduct a survey with 56 professional developers and 7 design sessions in which participants arrange configurations into their perfect control board and talk about their needs and experiences in more depth.
Developers report strong interest in configurability: 72.6% of usefulness ratings are positive, while only around a third indicate that the corresponding configuration is known to participants in their tools. Demand is particularly high for task-related controls such as minimum confidence thresholds, visibility of suggestion quality, and response length, whereas many persona-related configurations are seen as unnecessary. In this paper, we discuss the implications for designing more unified and discoverable configuration surfaces for future coding assistants
arXiv:2607.09164v1 Announce Type: new
Abstract: Saliency maps are most useful when they identify the image regions that are sufficient to preserve a model's behaviour. We introduce SEAMS, a sufficiency-based saliency method that directly optimises a soft mask using a preservation objective. Given a frozen differentiable model output, such as a class probability, CLS embedding, or token representation, SEAMS searches for a compact mask that preserves the selected output. The approach relies on a simple optimisation framework based on soft masks, a learnable budget, and a three-way image composite generated entirely from the query image. As a result, it requires no auxiliary distractor dataset, architecture-specific attribution mechanism, or differentiable top-k relaxation. Experiments with frozen ViT-S/16 and ConvNeXt models show that the same optimisation pipeline can generate object-level, class-conditioned, and token-level explanations by changing only the preserved target. The resulting masks are compact, interpretable, stable across random initialisations, and competitive on insertion and deletion benchmarks. Our results also indicate that different architectures often rely on different sufficient evidence while achieving similar preservation fidelity, highlighting the architecture-dependent nature of visual explanations.
arXiv:2604.00130v2 Announce Type: replace
Abstract: Chain-of-Thought (CoT) prompting has significantly improved the reasoning capabilities of large language models (LLMs). However, conventional CoT often relies on unstructured, flat reasoning chains that suffer from redundancy and suboptimal performance. In this work, we introduce Hierarchical Chain-of-Thought (Hi-CoT), a structured reasoning paradigm specifically designed to address the challenges of complex, multi-step reasoning. Hi-CoT decomposes the reasoning process into hierarchical substeps by alternating between instructional planning and step-by-step execution. This decomposition enables LLMs to better manage long reasoning horizons and maintain logical coherence. Extensive evaluations across diverse LLMs and mathematical reasoning benchmarks show that Hi-CoT consistently improves average accuracy by 6.2% (up to 61.4% on certain models and tasks) while reducing reasoning trace length by 13.9% compared to CoT. We further show that accuracy and efficiency are maximized when models strictly adhere to the hierarchical structure. Our code is available at https://github.com/XingshuaiHuang/Hi-CoT.
arXiv:2604.20730v3 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) have shown promising capabilities in generating Scalable Vector Graphics (SVG) via direct code synthesis. However, existing paradigms typically adopt an open-loop "blind drawing" approach, where models generate symbolic code sequences without perceiving intermediate visual outcomes. This methodology severely underutilizes the powerful visual priors embedded in MLLMs vision encoders, treating SVG generation as a disjointed textual sequence modeling task rather than an integrated visuo-spatial one. Consequently, models struggle to reason about partial canvas states and implicit occlusion relationships, which are visually explicit but textually ambiguous. To bridge this gap, we propose Render-in-the-Loop, a novel generation paradigm that reformulates SVG synthesis as a step-wise, visual-context-aware process. By rendering intermediate code states into a cumulative canvas, the model explicitly observes the evolving visual context at each step, leveraging on-the-fly feedback to guide subsequent generation. However, we demonstrate that applying this visual loop naively to off-the-shelf models is suboptimal due to their inability to leverage incremental visual-code mappings. To address this, we first utilize fine-grained path decomposition to construct dense multi-step visual trajectories, and then introduce a Visual Self-Feedback (VSF) training strategy to condition the next primitive generation on intermediate visual states. Furthermore, a Render-and-Verify (RaV) inference mechanism is proposed to effectively filter degenerate and redundant primitives. Our framework, instantiated on a multimodal foundation model, outperforms strong open-weight baselines on the standard MMSVGBench. This result highlights the remarkable data efficiency and generalization capability of our Render-in-the-Loop paradigm for both Text-to-SVG and Image-to-SVG tasks.
arXiv:2606.17970v3 Announce Type: replace
Abstract: We propose ACFK: Auto-correlation Function Keying, a new integrated sensing and communication (ISAC) waveform that carries random communication data while directly controlling the peak sidelobe level (PSL) of the periodic auto-correlation function (P-ACF). In contrast to existing works aiming at controlling the expected sidelobe level (ESL), which fails to characterize realization-specific sidelobe behaviors, we formulate a mutual information maximization problem under PSL and power constraints, and show that a continuous ACF-domain uniform distribution is asymptotically optimal at high signal-to-noise ratio (SNR) over quasi-static frequency-flat channels.
Motivated by this principle, ACFK maps finite-constellation symbols onto auto-correlation function (ACF)-domain sidelobes and uses independent phase symbols to exploit the remaining degrees of freedom. The resulting waveform enables exact control of the nominal P-ACF, which coincides with the actual P-ACF when the power spectral non-negativity condition is satisfied. We further analyze the non-negativity violation probability and bound the corresponding peak sidelobe level ratio (PSLR) degradation. A reference ISAC transceiver and its high-SNR approximate bit error rate (BER) analysis are also provided. Numerical results show that ACFK achieves stronger PSLR control, and improved weak-target detection performance, than a generalized probabilistic amplitude shaping (PAS) baseline at similar data rate and BER.
arXiv:2607.09662v1 Announce Type: cross
Abstract: Current electroencephalography (EEG)-based dream detection relies on power spectral density (PSD) and statistical moment features, achieving a state-of-the-art area under the receiver operating characteristic curve (AUC) of approximately 0.70 on the DREAM database (Wong et al., 2025, Nature Communications). We introduce PHINN-EEG (Persistent Homology Inspired Neural Network for EEG), the first topological time-series framework for dream mentation analysis. Using sliding-window Takens delay embeddings and Vietoris-Rips filtrations on multichannel pre-awakening EEG epochs, we extract Dynamic Betti Curves that characterize the geometric architecture of neural activity, not merely its energy. These topological invariants, combined with topology-conditioned flow matching, are analytically projected to outperform existing PSD and catch22 benchmarks, targeting AUC = 0.82-0.90 on the 1,462-awakening open-access subset of the DREAM database (drawn from a full registry of 3,191 total awakenings from 263 participants across 20 independent laboratories). We further introduce a topology-conditioned rectified flow model for dream-state EEG synthesis-with a spectral-conditioned flow model of comparable feature dimensionality as an additional ablation baseline to isolate the value of topological conditioning specifically-and propose a set of candidate Betti transition archetypes linking topology to phenomenological dream report categories, presented as an exploratory hypothesis space pending empirical validation. If validated, this work represents a paradigm shift from spectral energy to phase-space geometry in neural rare-event detection, with potential future implications for wearable BCI dream monitoring.
arXiv:2607.09507v1 Announce Type: new
Abstract: Global Structure-from-Motion (SfM) is an efficient paradigm for recovering camera poses and sparse 3D structure from unordered images. However, its reliance on scale-ambiguous epipolar geometry makes global positioning sensitive to noisy baseline estimates and weak view-graph constraints, while false edges from visually ambiguous pairs can further degrade reconstruction. We propose DGSfM, a depth-aware global SfM pipeline that uses monocular depth maps as a scalable prior while preserving explicit multi-view optimization. For each image pair, we use a depth-aware relative pose solver to convert scale-ambiguous epipolar constraints into scale-aware relative pose constraints. We further improve robustness through view-graph filtering and depth-consistency-based correspondence pruning, which suppress false edges and matches that remain plausible under epipolar geometry alone. Finally, global scale averaging and depth-guided pose-point initialization align monocular depth maps into a common reconstruction scale and provide stable initialization for global positioning and bundle adjustment. Experiments on ETH3D and IMC2021 show that DGSfM consistently improves over strong global SfM baselines across sparse and dense matching front-ends, achieving substantial gains in pose accuracy. Code is available at https://github.com/sithu31296/DGSfM.
arXiv:2607.09560v1 Announce Type: new
Abstract: Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research tasks. These are powerful capabilities, but they share a structural limitation: the representational frame within which the model operates, including its conceptual vocabulary, the space of admissible solutions it can search, and the criteria by which success is evaluated, is typically fixed and supplied in advance. This paper argues that building stronger intelligent systems capable of open-ended innovation requires additional classes of operations: the creation, stabilization, and reuse of new representational primitives, which alter the space being searched rather than simply searching within it.
We characterize the distance between current AI systems and genuinely open-ended intelligence through two gaps. The first is the vocabulary gap, the difficulty of inventing and stabilizing new representational primitives rather than merely recombining existing ones. The second is the verifier gap, the difficulty of judging the value of a new primitive when its full payoff may be visible only after future reuse. We interpret both gaps through a unified framework of intelligence as cognitive discrepancy reduction. By viewing intelligent behaviors as a sequence of cognitive transformations, we distinguish intra-space transformations which operate within a fixed representational frame, from generative transformations which may modify the frame itself. On this basis, we propose a ladder of innovation autonomy and outline several directions for advancing open-ended AI, including objectives that reward useful representational change, persistent memory architectures for invented primitives, and adaptive verification mechanisms capable of evolving alongside the representations they evaluate.
arXiv:2607.08786v1 Announce Type: new
Abstract: With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge. Pruning techniques that introduce sparsity into weight matrices can accelerate inference. However, maintaining model quality typically limits pruning to moderate unstructured sparsity (around 50\%). At these sparsity levels, none of the existing GPU kernels for sparse matrix multiplication (SpMM) can outperform their dense counterparts. This paper proposes an efficient GPU inference method for LLMs with moderate sparsity. We propose a three-layer matrix storage format comprising: (i) a Sparse-TC layer enabling sparse tensor cores to accelerate SpMM; (ii) a Slot-Filling layer using parallel differential distance for matrix compression while supporting low-cost on-chip decoding; (iii) a lightweight Residual Layer ensuring correct SpMM computation. Building on this format, we design a SpMM kernel that jointly utilizes sparse tensor cores and CUDA cores. This design enables an efficient execution pipeline and overlaps on-chip computation with memory access. Evaluations show that our work is the first to outperform dense matrix multiplication on modern GPUs equipped with high-bandwidth memory (HBM). It achieves up to 1.64x kernel-level speedup over SpInfer (EuroSys'25, Best paper) and up to 1.41x end-to-end speedups over FlashLLM (VLDB'24). Our source code: https://github.com/moui0/cudac.
arXiv:2607.08797v1 Announce Type: new
Abstract: Decentralized federated learning (DFL) dispenses with the central server of classical FL by utilizing peer-to-peer model exchanges among edge devices. This server-free architecture enables ad-hoc, flexible distributed learning in large device-to-device (D2D) networks. However, wireless DFL converges slowly because peer-to-peer model aggregation incurs high delays and errors. Each DFL training round involves many-to-many gradient sharing over wireless channels, resulting in uncoordinated channel access, large communication errors from stragglers, and slow model consensus, especially in large-scale D2D networks with pronounced clustering structures. We address these aggregation bottlenecks by provisioning a few reliable backhaul links at straggling nodes to enhance network connectivity. Building on this idea, our budget-aware, cluster-centric DFL framework first partitions the network into densely connected clusters, and then allocates the limited backhaul budget to selected cluster heads. The resulting two-tier protocol executes fast, parallel model aggregation within clusters and infrequent inter-cluster exchanges among the heads, yielding an O(1/t) convergence rate in t iterations. Numerical experiments on image-classification tasks confirm that our approach accelerates convergence compared to state-of-the-art DFL baselines with only a few strategically placed backhaul links.
arXiv:2607.09102v1 Announce Type: cross
Abstract: Medical imaging models are often deployed without the demographic, acquisition, and quality metadata needed for subgroup auditing. Once those metadata disappear, clinically critical failure modes can be masked by strong aggregate performance, and many robust-learning methods lose the group structure they rely on. We present CAPRA, a calibrated proxy-axis framework for hidden subgroup analysis under missing metadata. CAPRA predicts image-derived semantic axes, calibrates axis posteriors on a small metadata-labeled split via patient-level cross-fitting, and organizes those posteriors into a calibrated subgroup interface that supports both deployment-time failure analysis and downstream robust learning without requiring subgroup labels at deployment. Across fundus, dermoscopy, and chest radiography, CAPRA reveals disparity patterns missed by metadata-only slicing, remains informative under dataset shift, and produces subgroup partitions that align more closely with explicit failure axes than image-only or latent-slice baselines. The same interface can also be reused by downstream robust learners, although those gains are domain-dependent. Overall, CAPRA turns hidden subgroup analysis under missing metadata into a calibrated, interpretable, and reusable subgroup interface for deployment-time analysis and robust transfer.
arXiv:2607.08969v1 Announce Type: cross
Abstract: Numerical modeling of solar plasma dynamics is affected by the resolution of the computational grid. This often requires the estimation of subgrid processes related to the small-scale flow turbulence, as these processes play a critical role in momentum transport and energy dissipation. In this work, we investigate the use of deep learning techniques as surrogate models for subgrid turbulent transport in realistic hydrodynamic simulations of the quiet Sun. We describe the development of a 3D Convolutional Neural Network (CNN) to capture spatial dependencies in 3D velocity fields, leveraging different activation functions, as well as different architectural designs. We specifically focus on the prediction of Reynolds stress tensor components. The resultant model integrates velocity vector components and scalar features, such as plasma density, to enhance prediction accuracy. We compare the 3DCNN model to other types of models, such as a Multilayer Perceptron (MLP) and physics-based Gradient and Smagorinsky models, and show that the final model design reconstructs the Reynolds stress tensor components more accurately. Specifically, a 3DCNN model achieves an average improvement of ~31% on diagonal components and ~8% on the off-diagonal components of the stress tensor. Additionally, we show that applying a logarithmic data transformation of the target stress tensor components, to handle heavily skewed data, improves model performance. Results demonstrate the potential of deep learning, particularly CNNs, to approximate Reynolds stress tensor components for the upper solar convection zone and lower atmosphere, making them a viable candidate for modeling subgrid processes and a promising alternative to traditional turbulence models.
arXiv:2607.09113v1 Announce Type: cross
Abstract: Data scarcity and class imbalance are persistent challenges in machine learning that degrade model generalization and introduce predictive bias. We present a hybrid quantum-classical framework for synthetic data generation using a Quantum Circuit Born Machine (QCBM) to address these limitations. The proposed approach exploits quantum mechanical properties -- superposition and entanglement -- within a parameterized variational quantum circuit to model complex probability distributions that are difficult for classical generative methods to capture. Experiments are conducted on two tabular benchmark datasets: the Iris dataset and the Telco Customer Churn dataset. Preprocessing includes normalization and PCA-based dimensionality reduction to enable efficient basis encoding for quantum circuits. The QCBM is trained by minimizing Kullback-Leibler (KL) divergence between real and generated data distributions using a gradient-based parameter-shift optimization rule. Augmenting training data with QCBM-generated synthetic samples at 40-50% of the minority class improves F1-score by approximately 5-15% and minority-class recall by 10-25%. Cross-domain evaluations (Train on Synthetic, Test on Real; and Train on Real, Test on Synthetic) reveal a performance gap of only 3-10%, indicating strong distributional fidelity. Comparative analysis against classical oversampling methods -- SMOTE, Borderline-SMOTE, KMeansSMOTE, and SVM-SMOTE -- shows that QCBM achieves competitive classification performance and produces lower Maximum Mean Discrepancy (MMD) on the Telco dataset, suggesting superior structural similarity in certain imbalanced settings. These findings establish QCBM as a viable complementary tool for data augmentation, particularly for low-dimensional structured tabular data with class imbalance.
arXiv:2512.17869v2 Announce Type: replace-cross
Abstract: Understanding the fundamental properties that dictate photoexcited polarons in materials is critical to tuning their properties. Theoretical models of polarons have only recently been extended to the excited state. Experimental measurements of polaron formation and transport have been widely undertaken across a range of materials, from photocatalysts and superconductors to soft conducting polymers. Here, we map thermalized excited state experimental measurements of quantities such as polaron strength onto phase diagrams of the Holstein, Hubbard-Holstein, and t-J-Holstein models. This work demonstrates that tuning electron-phonon coupling strength, electron localization, and spin exchange can be leveraged to suppress or control polarons in transition metal oxides. We find that the t-J-Holstein model best describes the measured iron oxides and could be generally applied to a wide range of systems that exhibit polaron formation in the excited state. This work combines experimental data with ground state models to provide a qualitative parameter space for informing photoexcited polaron design, under which excited state polaronic behavior can be classified within ground-state calculable models.
arXiv:2607.09073v1 Announce Type: new
Abstract: Bayesian optimization routinely warm-starts a target experiment with data from related source tasks, and the multi-task Gaussian process is the textbook surrogate for the job. We revisit this default in a controlled setting and find that it misestimates the cross-task correlation even in the simplest non-trivial case, affinely related source and target tasks, where a working transfer learning method should obviously succeed. We trace the failure to two independent structural mechanisms. Per-task standardization, the textbook fix for the affine slice ambiguity, propagates a finite-sample alignment error into the recovered correlation. The marginal likelihood itself identifies the correlation only at a per-sample rate that a Gaussian process at non-overlapping designs further dilutes. We propose three conservative remedies that follow from the analysis: promoting per-task means and scales to model parameters, restricting the task covariance to non-negative correlations, and co-locating part of the source and target designs. Across synthetic multi-task problems and surrogate-based hyperparameter tuning transfer, these remedies recover the target-only baseline on the simple instances, while the broader failure persists on harder instances and across most rank-based and latent-context variants.