Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

Benign Overfitting with Quantum Kernels
arXiv:2503.17020v2 Announce Type: replace-cross Abstract: Kernel methods compare inputs through feature maps. Quantum kernels follow the same principle: input data are encoded into quantum states, which define quantum feature representations in Hilbert spaces. Kernel values are then obtained by estimating inner products between these states using suitable quantum circuit measurements. As a result, quantum kernels may be intractable to compute classically while remaining efficiently computable on quantum hardware, potentially leading to a quantum advantage. However, designing effective quantum kernels remains a major challenge. Many quantum kernels, such as the fidelity kernel, suffer from exponential concentration. This results in near-identity kernel matrices that fail to capture meaningful data correlations and lead to overfitting and poor generalization. In this paper, we propose a novel strategy for constructing quantum kernels that achieve good generalization performance, drawing inspiration from benign overfitting in classical machine learning. We introduce the concept of Local-Global quantum kernels, which combine two components: a local quantum kernel based on measurements of small subsystems, and a global quantum kernel derived from full-system measurements. To support the effectiveness of the proposed construction, we show theoretically and empirically that Local-Global quantum kernels exhibit benign overfitting.
Graph Convolutional Attention: A Spectral Perspective on Graph Denoising and Diffusion
arXiv:2607.06546v1 Announce Type: new Abstract: Denoising graphs is a fundamental problem in graph learning and the core operation of graph diffusion models. Attention-based architectures like graph transformers have recently shown promise in denoising graphs. However, our principled understanding of attention-based graph denoising remains limited, making it unclear whether standard attention is the right mechanism for this task. Here we show that, under a denoising objective, linear attention is suboptimal and can only learn an average spectral denoising filter over the training distribution. This creates a fundamental limitation as graphs often vary spectrally across the distribution. To overcome this limitation, we introduce Spectral Attention, which directly utilizes the input graph spectrum and provably outperforms linear attention by a margin governed by the spectral diversity of the distribution. We then derive Graph Convolutional Attention (GCA), a practical and permutation-equivariant realization of this idea that implements spectral denoising through graph-filtered queries and keys. For stochastic block models, GCA provably matches the idealized Spectral Attention mechanism. We further show that the softmax operation, that follows the attention, provides additional denoising by approximately projecting noisy eigenvectors onto the clean eigenspace. Empirically, replacing linear attention with GCA consistently improves graph denoising and diffusion on synthetic and real datasets, with gains strongly correlated with spectral diversity. In DiGress, GCA matches standard graph-transformer performance without computing expensive structural features, and when combined with the recently proposed PEARL positional encodings, avoids explicit eigendecomposition computations resulting in faster inference without degrading quality. The code can be found here: github.com/shervinkhalafi/graph_conv_att
ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
arXiv:2607.06565v1 Announce Type: new Abstract: Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text-3D interaction remains largely implicit. Existing methods concatenate text and 3D tokens into a flat sequence and rely on self-attention, collapsing coarse structural cues and fine geometric details into one undifferentiated representation. We introduce ELSA3D, a unified 3D model that addresses this with elastic semantic anchoring, structuring language and geometric reasoning jointly along matched abstraction scales. ELSA3D represents geometry with a scale-aware octree tokenizer and introduces Anchor Tokens, sparse cross-modal units that select semantic cues, route them to the most relevant 3D scale, retrieve scale-specific geometric evidence, and write the fused signal back into the unified representation, keeping interaction sparse yet precise. A lightweight per-block router makes both computation and reasoning elastic, choosing which text tokens instantiate anchors at which geometric scale so that cross-modal capacity concentrates where alignment is most needed. ELSA3D achieves state-of-the-art performance across image-to-3D generation, text-to-3D generation, and 3D captioning, outperforming the strongest unified baseline while roughly halving FLOPs and inference latency relative to the non-elastic version of the same model.
Breaking Structural Isolation: Scalable Graph Clustering via Community-Aware Sampling and Structural Entropy
arXiv:2607.05469v1 Announce Type: new Abstract: Unsupervised graph clustering is a fundamental technique for uncovering underlying semantic patterns in large-scale networks. Although Graph Contrastive Learning has demonstrated promising performance, existing methods often suffer from the "structural isolation" issue during mini-batch training, making it challenging to capture cohesive community structures that characterize the global topological distribution. To address these challenges, we propose SCISE, a Scalable unsupervised graph Clustering framework that preserves structural Integrity by synergizing community-aware sampling with constrained Structural Entropy. Specifically, we first introduce the Structural Entropy Community Constraint operator (SECC), which optimizes structural information within a constrained solution space to mitigate community fragmentation and enhance partition cohesion. Second, to prevent global information loss during batch training, we design a Community-Aware Sampling Expansion (CSampE) mechanism that incorporates the community context of target nodes into sampling batches, effectively breaking structural barriers and preserving topological integrity. Finally, we devise a Structural Contrastive Learning (StructCL) module that refines edge weights based on intra-batch structural similarity, guiding the encoder to learn representations in a higher-order structural space. Extensive experiments on six mainstream benchmark datasets demonstrate that SCISE significantly outperforms state-of-the-art algorithms, with ablation studies and robustness analyses further validating its effectiveness and reliability for real-world large-scale graphs.
How Environment and Urbanization Shape Bird Diversity in Sri Lanka
arXiv:2607.00582v2 Announce Type: replace-cross Abstract: This study presents a comprehensive analysis of bird diversity across Sri Lanka by integrating spatial, temporal, and environmental data. Bird observation records were combined with environmental variables, including weather conditions, air pollution, the Normalized Difference Vegetation Index (NDVI), land cover, elevation, and Artificial Light At Night (ALAN), and rigorously preprocessed to ensure data quality. Spatial analyses were conducted on multiple grid scales (2 km, 5 km, 10 km) to evaluate patterns in species richness while minimizing sampling bias through spatial thinning. Temporal trends were assessed using effort-corrected metrics including rarefied richness and occupancy rates to account for variations in observation effort over time. Environmental drivers of bird diversity were examined using multivariate statistical models, including Poisson Generalized Linear Models (GLMs) and correlation analyses, to identify key associations between ecological factors and species richness. Additionally, community structure, dominance patterns, and beta diversity were analyzed to understand variations in species composition across regions and time. The study found that land-cover type is a stronger predictor of bird diversity than individual continuous variables such as NDVI or temperature alone. Urbanization, measured by ALAN, exhibits nuanced scale-dependent effects, supporting high abundances of a few generalist species while reducing overall richness. The findings provide actionable insights into the patterns and drivers of avian diversity in Sri Lanka, offering a scalable and reproducible framework for biodiversity research and conservation planning.
Umm... With Transformers? Insights from Filled Pause Use across Four Slavic Parliaments
arXiv:2607.05964v1 Announce Type: new Abstract: Filled pauses (FPs) are a universal feature of spontaneous speech, yet most studies rely on small, single-language corpora, limiting the generalisability of their findings. We analyse ~4,000 hours of parliamentary speech across four related Slavic languages (Croatian, Czech, Polish, Serbian). FP occurrence is obtained via transformer-based automatic detection, while FP rate is modelled using Generalised Estimating Equations (GEE) with Mundlak correction to distinguish within- from between- speaker effects. We replicate a negative association of age and speech rate with FP rate, but find that gender effects are language-specific and directionally opposite to most prior literature. Novel analyses of sentiment, political orientation, and power status reveal a consistent positive association between sentiment and FP rate, alongside parliament-specific modulation by orientation and power status, with opposition speakers tending toward lower FP rates than governing coalition speakers.
Population-Level Profiling of DSM-5 Depressive Symptoms Among Self-Reported ADHD and ASD Users on Twitter: An Exploratory Study Using Advanced NLP and Statistical Analysis
arXiv:2607.05626v1 Announce Type: new Abstract: Background: Depression frequently co-occurs with ADHD and autism spectrum disorder (ASD), but population-level differences in symptom expression between these groups remain underexplored. Objective: We examined whether social media users with ADHD and ASD differ in how they express DSM-5 depressive symptoms in their tweets, and whether differences persist across varying levels of depressive-content filtering. Methods: We analysed 1,282,437 tweets from 792 users (622 ADHD; 170 ASD) with self-reported diagnoses on Twitter. Tweets were pre-filtered for depressive relevance using zero-shot NLI, then classified into nine DSM-5 symptoms using MentalRoBERTa fine-tuned on ReDSM5. Profiles were mean-centered per user. We applied L1-penalised logistic regression with cross-validation to distinguish ADHD from ASD users, complemented by Pearson correlations for symptom co-occurrence, and tested robustness across five filtering thresholds using bootstrapping. Results: MentalRoBERTa achieved macro-F1 of 0.901 on a held-out set, outperforming the original ReDSM5 benchmark. ADHD vs ASD classification yielded stable but modest performance (cross-validated ROC-AUC 0.645-0.653). Cognitive issues, sleep issues, appetite change, and fatigue leaned toward ADHD, while suicidal ideation and anhedonia leaned toward ASD. A largely shared symptom co-occurrence structure emerged between groups; no pair met our criterion for a robust disorder-specific difference. Conclusions: Population-level differences in depression-related language between ADHD and ASD social media users were consistently observed across thresholds, reflecting reproducibility rather than clinical validity. Findings are exploratory and do not establish differing phenomenology at the individual level.
Axioms for physical reasoning: codifying the Seiberg--Witten solution in Lean
arXiv:2607.06379v1 Announce Type: cross Abstract: Mathematicians have embraced interactive theorem provers with growing enthusiasm -- building large shared libraries and machine-checking a string of landmark results. Theoretical physics is different: most of its results are not theorems but justified by arguments the community trusts without a rigorous proof. For many -- the one we treat here among them -- no rigorous proof is within reach. For 4d Yang--Mills theory, deriving exact rigorous results from first principles would first require constructing the interacting theory nonperturbatively, which is a sizable piece of one of the Clay Millennium prize problems. We argue here that an interactive theorem prover can be used to verify some non-rigorous physics arguments. The method is to postulate a short list of explicit, named physical postulates, which imply the physical results by virtue of a machine-checkable proof. The trust that remains then rests on that short, inspectable list, and the prover can report, for any downstream result, exactly which assumptions it used. We carry this out for the Seiberg--Witten solution of ${N}=2$ $SU(2)$ super-Yang--Mills -- the genus-one case -- formalized in Lean 4; the higher-genus $SU(N)$ generalization is developed in the same repository as an axiomatized skeleton and left to future work. We describe what is proved, what is assumed, how the assumptions are checked -- external review and an independent numerical oracle -- and why this discipline is a sound standard for validating AI-generated results in theoretical physics. What we offer is a discipline, reviewable on its own terms: a reader may take the Seiberg--Witten mathematics on trust and still assess the formalization method.
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability
arXiv:2607.06157v1 Announce Type: new Abstract: Deliberation plays a crucial role in collaboration; when humans work together, they naturally engage in communication to align information and reach an agreement. In this paper, we investigate deliberative large language model (LLM) agents under partially observable joint decision-making tasks. We formalize deliberative collaboration as a cooperative joint decision problem with partial and asymmetric observations, and introduce a scalable benchmark that instantiates this problem across multiple task settings and domains in which agents must exchange information through deliberation to reach a joint decision with a shared reward. We then instantiate a reference scaffold and evaluation protocol for deliberative agents and conduct a systematic evaluation of a range of representative LLMs. The results reveal that complex deliberative collaboration tasks continue to challenge state-of-the-art language models. Even with the aid of external mathematical tools, language models may fail in either the deliberation process for aligning information or the complex reasoning process for making the decision. On the other hand, diagnostic analysis reveals that the deliberation process may also provide opportunities for reflection and error correction, sometimes improving performance over centralized baselines. Altogether, our work establishes a foundation for evaluating and improving LLM agents in deliberative collaboration and provides insights into the strengths, limitations, and properties of current LLM-based multi-agent systems.
Structural Divergence of the Roman--Byzantine Trade Network, 0--1453\,CE: Persistent Homology, Topological Velocity, and Criticality Indicators of Imperial Collapse
arXiv:2607.05695v1 Announce Type: new Abstract: We extend the persistent homology analysis of~\paperone{} to the full Roman--Byzantine trade network (0--1453\,\textsc{ce}), using 2{,}599 nodes and 4{,}503 trimodal edges calibrated against the \textsc{orbis} Geospatial Network Model. Five results are reported. % (i)~The $H_t{=}0$ western sub-network result of~\paperone{} is a data-coverage artifact: with full western representation ($N_{\rm west}=987$, $\beta_1\approx52$ cycles per decade) a baseline East--West entropy gap of $+2.22$ units is present from 0\,\textsc{ce} and grows at $+3.3\times10^{-3}$\,yr$^{-1}$, predating the Theodosian partition by four centuries. % (ii)~A \emph{hub-selection artifact} in degree-heterogeneous networks can reverse the sign of the inferred Phase~III slope, requiring full-coverage or stratified sampling for reliable structural-break detection. % (iii)~Decomposing Byzantine resilience into geographic ($H_{\rm geo}$) and economic ($H_{\rm eco}$) components reveals a peak decoupling ratio $R_d = H_{\rm eco}/H_{\rm geo} = 47.7$ at 620\,\textsc{ce}, falling to 13.9 at 640\,\textsc{ce}, quantifying the McCormick--Ward-Perkins historiographical debate as a contrast between two network layers operating on different timescales. % (iv)~The inter-decade $W_2$ Wasserstein velocity identifies the Late Roman--Early Byzantine transition (495\,\textsc{ce}) as the highest topological-velocity event of the 1,453-year record; the cross-network Wasserstein ratio increases by $150$--$300\times$ after the Chrysobull of 1082\,\textsc{ce}, providing an independent diagram-space analogue of $R_d$. Both the Western collapse (476\,\textsc{ce}) and the Byzantine endpoint (1453\,\textsc{ce}) occur at $H^{\ast}\approx0.524$, interpreted as a candidate topological percolation threshold.
LLM-Driven Neural Network Generation with Same-Family Architecture Guidance: Disentangling Transfer and Adaptation
arXiv:2607.05704v1 Announce Type: new Abstract: Large language models (LLMs) can generate neural-network modifications, but unrestricted generation is often invalid or harmful. This paper studies a narrower setting: improving a weak target model using a stronger same-family source model from a neural-network database. We propose a source-guided candidate-generation protocol with non-source controls, source-conditioned candidates, and a no-LLM hp_copy ablation under equal evaluation budgets. The protocol reports validity separately from accuracy and selects the best valid candidate only when it improves the target. On CIFAR-10, the strongest source-guided candidate reaches 0.5049 accuracy versus 0.2398 for the best non-source candidate, a +0.2651 advantage, while improving a weak target originally at 0.1254; a five-epoch check preserves the gain at 0.7686 versus 0.4839. On SVHN AlexNet with DeepSeek-Coder-6.7B, source-guided transfer reaches 0.7880 versus 0.2254, a +0.5626 advantage; a fresh repeat reaches 0.8069 versus 0.2509, a +0.5560 advantage. Direct source-recipe copy produces 0.1959 on SVHN AlexNet, matching the original target, while hp_transfer reaches 0.7880, showing that the LLM adapts rather than copies the source recipe. Family-level analysis shows the clearest positive signals for AlexNet, with 6/8 wins across SVHN, Imagenette, and CelebA-Gender, and alt_nn1, with 8/10 wins on CIFAR-10.
Uncertainty-Aware Cross-Modal Remote Sensing Image-Text Retrieval via Evidential Learning
arXiv:2607.06032v1 Announce Type: new Abstract: In cross-modal remote sensing image-text retrieval (CMRSITR), test-time remote sensing (RS) images and textual descriptions may deviate from well-curated benchmark conditions due to sensor- and atmosphere-related image degradations and text-side RS-vocabulary heterogeneity. Under such non-ideal conditions, existing CMRSITR methods may produce unreliable retrieval results because they perform retrieval with full certainty for each query and do not distinguish the varying uncertainty across queries. To address this issue, we propose an evidential learning-based CMRSITR (ELC) method for uncertainty-aware retrieval. During the training phase of ELC, evidential learning (EDL) is employed to model the inter-modal correspondences between RS images and textual descriptions as Dirichlet distributions, from which the uncertainty of each query can be obtained. Based on the EDL outputs, uncertainty-correctness alignment learning (UCL) is introduced to align the estimated uncertainty with retrieval correctness, encouraging high uncertainty for incorrect retrieval and low uncertainty for correct retrieval. Furthermore, intra-modal relationship learning (RL) distills the intra-modal similarity structure from pretrained mentor encoders for the trainable encoders, thereby making the Dirichlet distributions modeled by EDL more discriminative. In the test phase of ELC, the estimated uncertainty is compared with a threshold determined by a fixed deferral ratio, where low-uncertainty queries are directly returned and high-uncertainty queries are refined by RS-aware test-time augmentation (RS-TTA). Experimental results demonstrate that ELC achieves competitive retrieval performance compared with state-of-the-art CMRSITR methods and provides stronger robustness under the evaluated RS-specific degradations, including sensor- and atmosphere-related image perturbations and RS-vocabulary heterogeneity.
The Masks We (Think We) Wear: Privacy Threats of Browser-Extension Wallets in the Web3 Ecosystem
arXiv:2607.06141v1 Announce Type: new Abstract: Cryptocurrency wallets are the primary interface for managing pseudonymous blockchain addresses, viewing balances, and interacting with Web3 applications. Although users typically assume that their addresses remain independent of each other unless intentionally revealed, modern wallets routinely communicate with both blockchain infrastructure and decentralized applications (dApps), generating network-side and web-side signals that may undermine this assumption. In this paper, we identify and formalize five privacy threats that arise directly from wallets interacting with the network and the web browser. Using large-scale dynamic measurements of 85 of the most popular Chrome Web Store browser-extension wallets (representing 35.16 million users), we observe that routine remote procedure call (RPC) operations leak structural links between a user's addresses; that the majority of Ethereum wallets implement permission revocation inconsistently and continue to expose previously revoked addresses across sessions; and that many wallets inject their provider interfaces into cross-origin iframes, enabling passive cross-site tracking beyond dApps and potentially real-world identity deanonymization without user interaction. Taken together, our results show that these wallet behaviors leak sensitive information that can be used to link multiple addresses to the same user, track wallet users across sessions and sites, and connect their browsing activity to their on-chain wealth. We discuss practical mitigations and show that many of these threats can be substantially reduced through improved wallet implementation, stronger privacy considerations in ecosystem standards, and stricter controls over provider exposure. Our results highlight the need for standardized, privacy-preserving wallet architectures.
Redundant contacts and force redistribution stabilize limbless vertical climbing
arXiv:2607.06239v1 Announce Type: new Abstract: Animals navigating complex vertical environments must secure stable footholds, a challenge for species without feet. While arboreal climbing has evolved repeatedly in snakes, the physical mechanisms they use to scale broad, nearly flat surfaces remain poorly understood. By measuring three-dimensional body kinematics and per-contact forces on a smooth vertical wall with protruding posts, we show that cornsnakes climb by dynamically balancing forces across a highly redundant network of 5 to 16 simultaneous contacts--far exceeding the three contacts minimally required for physical stability. Using a computational model and a robotic climber, we demonstrate that while simple body undulations and passive friction are mechanically sufficient to climb this terrain, snakes systematically deviate from this passive baseline. While downward climbing relies primarily on friction, ascending snakes actively generate positive mechanical work at their contacts to propel themselves. Furthermore, we found that whenever a snake engages a new contact, it triggers a stereotyped, system-wide redistribution of force that seamlessly integrates the new foothold without disrupting whole-body balance. These results reveal how a continuous, flexible body can transform sparse environmental features into a robust, fault-tolerant network. This mechanism provides a biomechanical framework for understanding the repeated evolution of limbless climbing and offers physical principles for designing agile robots for unstructured terrain.
Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search
arXiv:2607.05970v1 Announce Type: new Abstract: Dataset search depends heavily on metadata, making LLM-generated metadata a consequential form of synthetic content in retrieval systems. We study six metadata-generation settings for RDF datasets, ranging from simple rewriting to profile-grounded and agentic graph-based generation, and evaluate them jointly for retrieval effectiveness and faithfulness. Unconstrained metadata rewriting delivers the strongest retrieval gains over the original metadata, but it is also the least faithful, showing that search improvements can be driven by unsupported semantic expansion. More grounded settings substantially improve faithfulness, and profile-grounded rewriting provides the most balanced trade-off between retrieval effectiveness and grounding. These findings position synthetic metadata as a system-level IR problem in which effectiveness, provenance, and trust must be evaluated together.
Designing Maintainable Hybrid Generative Systems: A Quantum-Inspired Approach to Automated Music Harmony Generation
arXiv:2607.06296v1 Announce Type: new Abstract: This paper presents the design and evaluation of a maintainable hybrid generative architecture for automated music harmony generation from melody. The proposed system combines quantum-inspired candidate exploration over overlapping melodic contexts with explicit rule-based optimization to balance generative flexibility and structural control. The architecture is evaluated using explicit and reproducible metrics covering structural coherence, functional agreement, harmonic similarity, and robustness. The results show that the proposed approach produces harmonizations that preserve tonal structure and cadential behavior while allowing multiple valid harmonic realizations. Furthermore, the optimization layer improves structural coherence, stability, and predictability without requiring a training corpus. The study demonstrates that transparent and controllable hybrid generative systems can be systematically designed and evaluated within the context of Information Systems Development.
No-Regret Gaussian Process Optimization of Time-Varying Functions
arXiv:2512.00517v3 Announce Type: replace-cross Abstract: Sequential optimization of black-box functions from noisy evaluations has been widely studied, with Gaussian Process bandit algorithms such as GP-UCB guaranteeing no-regret in stationary settings. However, for time-varying objectives, no-regret is unattainable under pure bandit feedback unless strong and often unrealistic assumptions are imposed. We propose a novel method for optimizing time-varying rewards in the frequentist setting, where the objective has bounded RKHS norm almost surely. Time variations are captured through uncertainty injection, enabling heteroscedastic Gaussian process regression that adapts past observations to the current time step. As no-regret is unattainable in general in the strict bandit setting, we relax the latter allowing additional queries on previously observed points. Building on sparse inference and the effect of uncertainty injection on regret, we propose W-SparQ-GP-UCB, an online algorithm that achieves no-regret with a vanishing number of additional queries per iteration. To assess the theoretical limits of this approach, we establish a lower bound on the number of additional queries required for no-regret, proving the efficiency of our method. Finally, we provide a comprehensive analysis linking the temporal regime of the function to achievable regret rates, together with upper and lower bounds on the number of additional queries needed in each regime.
Rank-Order N-of-M Codes for Sparse Distributed Memory: Disentangling Representation and Learning Effects in Noise Robustness Against Contemporary Neuromorphic Architectures
arXiv:2607.02967v2 Announce Type: replace Abstract: Large language models remain limited as continual learning systems, motivating renewed interest in Sparse Distributed Memory (SDM) as an explicit online episodic memory. CALM (Nechesov and Ruponen, 2025) identifies its threshold-binary encoder as an open design question. This paper evaluates rank-order N-of-M encoding (Furber et al., 2007) as an alternative. We make three contributions. First, a faithful reimplementation validates the published architecture by confirming exact equivalence between WheelSDM and RankOrderSDM (cosine similarity 1.0000 across 10 seeds) and reproducing the documented divergence of RDLIF neurons under interference. Second, multi-seed capacity experiments show RankOrderSDM outperforming StandardSDM by 13.4 percentage points at saturation in the scaled configuration and by 0.8 percentage points at the published architecture scale. Third, BER robustness experiments disentangle representation and learning effects, showing that the large robustness gain arises primarily from the interaction of rank-order encoding with MAX-Hebbian learning, while the encoder alone provides only a small advantage under matched learning conditions. Experiments on GloVe-100 embeddings confirm this small but consistent encoding benefit on real structured data, whereas sentence embeddings exhibit a ceiling effect at low memory load. A secondary analysis shows that idealized rank-order encoding requires half the component-level encoding energy of SpikingMamba's SI-LIF neurons at four-bit precision, although decoder costs dominate overall system energy. These results identify which components of the original rank-order SDM architecture provide measurable benefits for contemporary memory-augmented AI systems, offering practical guidance for architectures such as CALM.
EOM-CC Excited-State Gradients and Nonadiabatic Couplings on a Consumer GPU from a Contraction-DAG with Laplace-Transform J/K Kernels
arXiv:2607.05622v1 Announce Type: new Abstract: We present a unified, memory-bounded GPU realization of equation-of-motion coupled-cluster (EOM-CC) excited-state gradients and interstate nonadiabatic couplings (NACMEs) on a single 8\,GB consumer GPU. Both are built from one contraction directed acyclic graph: the EOM-CC relaxation is the reverse-mode transpose of the forward density build rather than a per-state re-derivation, and an atomic-orbital-direct Laplace-transform $J/K$ kernel, made non-symmetric ($J^x(A,B)\neq J^x(B,A)$) by the transition densities, resolves every energy denominator with no four-index molecular-orbital tensor; a two-sided Davidson returns both eigenvectors from one device-resident, spin-pure solve. The pipeline is \emph{validated end to end at small scale}: gradients and NACMEs match finite differences across four spin multiplicities and full configuration interaction to $<\!10^{-12}$ for two electrons, and the excited-state gradient matches the independent \textsc{Psi4} code to $\le\!4.6\times10^{-7}~E_h/a_0$ from \ce{H2O} to aromatic benzene. The kernels and the ground-state solve reach chromophores ($\le\!730$ AO) in 8\,GB, and a frozen-natural-virtual compression lets the eigensolver \emph{execute} a complete excited-state gradient and $Q$--$B$ NACME of the chlorophyll-core chromophore \ce{Mg}-porphine (def2-SVP, $439$ AO) on the card. We present that run as a \emph{capability demonstration} -- executed and translationally invariant to machine zero, but anchored only piece-wise and bounded by a direct convergence study at ${\sim}10^{-2}~E_h/a_0$ -- not a converged spectroscopic result. The validated small-scale capability and the memory-bounded implementation are the contribution.
Responsible Personalisation: The Double-Edged Sword of Personalisation in Human-Robot Interaction
arXiv:2607.06344v1 Announce Type: new Abstract: While personalisation is becoming a defining capability in human-robot interaction (HRI), the existing literature on responsible personalisation remains fragmented, offering isolated accounts of ethical risks without a structured understanding of how they emerge across interaction contexts. This gap is particularly critical in HRI, where robots' embodiment and social presence can amplify and reshape such risks or generate new types of risks. We present a lifecycle-based and context-sensitive framework for personalised HRI, grounded in an embodiment-aware perspective. The framework combines stages of the personalisation process with interaction characteristics (short-term vs. long-term, open-domain vs. closed-domain), enabling systematic analysis of how risks arise and evolve. Building on this, we conduct an integrative analysis of key ethical risks, including autonomy erosion, biased user modelling, manipulation, dehumanisation, and privacy violations, and examine how they manifest across contexts. We translate these insights into actionable design recommendations and outline open research challenges. By structuring both the design space and risk landscape of personalised HRI, this work provides a foundation for more systematic, transparent, and ethically grounded approaches to personalised robot behaviour.
EcoVision: AI-Powered Drone Imaging for Salt Marsh Vegetation Monitoring and Dominance Mapping
arXiv:2607.06105v1 Announce Type: new Abstract: High-resolution RGB imagery acquired from low-altitude UAV surveys was processed through a modular pipeline incorporating transformer-based semantic segmentation, connected-component vegetation extraction, fine-grained species classification using a ConvNeXt architecture, and grid-based dominance scoring at 2x2m resolution. The framework targeted two ecologically significant halophytic grasses, Spartina maritima and Puccinellia maritima, and was trained using a curated and manually annotated UAV imagery, along with biodiversity imagery sourced from publicly accessible datasets. In order to identify these plants from the imagery, our segmentation yielded reliable species masks (mean IoU = 0.56; pixel-level accuracy = 0.96), while object-level classification achieved very good discrimination (F1 = 0.99). Dominance estimates closely matched quadrat-based field surveys, with mean absolute differences below 8%, preserving fine-scale spatial structure under realistic survey conditions. The developed system, named EcoVision, establishes a practical foundation for scalable, high-resolution salt marsh monitoring, demonstrating how AI-driven workflows can translate pixel-level predictions into ecologically interpretable metrics.
The Cryogenic System of DMRadio-50L
arXiv:2607.05771v1 Announce Type: new Abstract: The DMRadio-50L experiment is designed to search for axion dark matter in the 5 kHz - 5 MHz frequency range using a lumped-element LC resonator and a toroidal magnet and to serve as a testbed for quantum sensors. This paper describes the custom cryogenic system developed to meet the stringent requirements of the experiment within a standard laboratory environment. The system is designed to cool a 200 kg detector assembly to temperatures as low as 50 mK while providing sufficient cooling power at multiple temperature stages. We present the conceptual design, technical implementation, and measured performance of the hybrid cryogenic system, which combines a horizontal dilution refrigerator with a large vertical payload cryostat.
Controlling Tool Use with Heading-Specific Activation Steering
arXiv:2607.05790v1 Announce Type: new Abstract: Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily. We investigate whether tool-use decisions have any stable internal representation that can be extracted and manipulated, a question that is non-trivial given that tools exist entirely in context at inference time and have no direct encoding in model weights. We show that steering vectors extracted from heading-anchors positions exert bidirectional causal control over tool-invocation behavior across five open-source models and three domains, suppressing unnecessary tool use most effectively in domains where parametric reasoning suffices. However, geometric analysis reveals that this causal effectiveness does not correspond to clean linear structure: tool-invocation steps exhibit diffuse, bimodal alignment with the suppression vector rather than the consistent negative alignment a linear encoding account would predict, and different tool types recruit largely distinct internal signatures with low cross-tool feature overlap. We hypothesize these geometric properties are indicative of the non-parametric nature of tools, and distinguish tool-use steering vectors from those extracted for parametrically grounded concepts. The relationship between this geometric irregularity and the observed causal effectiveness remains an open question.
High-Resolution Artwork Outpainting with Global Blueprint Guidance and Layout Control
arXiv:2607.06162v1 Announce Type: new Abstract: Image outpainting extends an image beyond its original borders, requiring seamless style integration and globally coherent scene completion. Building on the success of diffusion models, recent methods have achieved substantial improvements in visual quality. In practice, however, high-resolution outpainting is commonly performed via progressive expansion around a fixed source image, particularly in artwork scenarios. Despite this progress, existing approaches still suffer from three key limitations: (i) the absence of a reliable global planning mechanism, which leads to structural instability and error accumulation at high resolutions; (ii) limited spatial controllability beyond text prompts, making it difficult to place objects at user-specified locations; and (iii) high inference latency caused by inherently sequential patch generation. To address these issues, we propose a global blueprint-guided two-stage diffusion framework for layout-controllable high-resolution outpainting with efficient parallel synthesis. In Stage 1, we generate a low-resolution global blueprint using a layout adapter that injects bounding-box conditions into a Stable Diffusion inpainting backbone, producing a globally consistent structural plan while extracting global guidance features. In Stage 2, we synthesize high-resolution local patches in parallel by injecting the blueprint-derived global guidance and initializing each patch from the blueprint using the low-frequency preservation property of forward diffusion. This design eliminates sequential dependency while maintaining global coherence. Extensive experiments on large-scale artwork datasets demonstrate improved visual fidelity, stronger semantic consistency, and substantially reduced inference time compared to prior baselines, while uniquely supporting explicit layout control for artwork outpainting.
Life Cycle Assessment of Pre-training the Lucie 7B Open-Source Large Language Model on the Jean Zay Supercomputer
arXiv:2607.05408v1 Announce Type: new Abstract: The environmental impact of training large language models (LLMs) is increasingly scrutinised, yet most published estimates focus on operational energy and disclose little about manufacturing (embodied) emissions, water consumption, or the underlying highperformance computing (HPC) infrastructure. We present a life cycle assessment (LCA) of the pre-training of Lucie 7B, an open-source multilingual Foundation Model developed by the OpenLLM-France consortium and trained on the NVIDIA H100 partition of the Jean Zay supercomputer operated by IDRIS (CNRS). The assessment is framed by the AFNOR SPEC 2314 "Frugal AI" reference and applies the Labos 1point5 methodology for greenhouse gas(GHG) accounting in computing. The study scope extends from data preparation to model validation, and integrates the full life cycle of the hardware infrastructure: manufacturing (including raw-material extraction), use (compute, temporary storage, system administration, cooling), and end-of-life. We report (i) an annual footprint of 417.5 tCO2eq for the Jean Zay H100 partition, split almost equally between manufacturing and operation; (ii) an effective intensity of 36.7 gCO2eq per H100 GPU-hour; (iii) a total training footprint of 21 tCO2eq for Lucie 7B (574 564 H100 GPU-hours), inclusive of amortised hardware manufacturing; (iv) on-site water consumption of approximately 76m3 for the training campaign and an annual Water Usage Effectiveness (WUE) of 0.07 L/kWh for IDRIS; (v) a heat-reuse factor (ERF) of 0.37 thanks to waste-heat recovery into the urban heating network. The study contributes one of the few publicly documented LCAs of an LLM training campaign that explicitly couples operational data with embodied emissions decomposed by subsystem (compute, storage, power chain, cooling), and discusses the implications for the design of frugal-by-construction AI systems in Europe.