arXiv:2607.16008v1 Announce Type: new
Abstract: The concentration of 77,622 spectators during football games at Notre Dame Stadium creates an exceptionally demanding environment for wireless infrastructure. To handle this extreme user density, the stadium deploys concurrent multi-tier networks serving outdoor users: an enterprise 5/6 GHz Wi-Fi network with ~900 outdoor Access Points (APs) alongside high-density multi-carrier 4G/5G networks powered by a neutral-host small-cell Distributed Antenna System (DAS) with up to 129 unique cell identifiers (PCIs) per operator. This study evaluates user-perceived performance and QoE across these networks using commercial smartphones to execute web browsing, WhatsApp messaging, and Instagram media posting workloads. Our empirical results reveal that while cellular networks deliver strong peak downlink performance in an empty stadium, game-day crowd loads heavily strain uplink and latency performance, triggering a severe cellular "uplink gap." Under Non-Standalone (EN-DC) anchor congestion, web browsing handshakes suffer a catastrophic 5,983 ms P90 Time-to-First-Byte (TTFB), and image upload failure rates climb to 46%. Furthermore, while narrow low-band FDD channels (e.g., n5) maintain robust channel quality during uploads, they exhibit a 70% median Block Error Rate (BLER) during active browsing tests, driving a 36.6% page-load failure rate. Conversely, the dense stadium Wi-Fi infrastructure delivers downlink throughput comparable to the best performing 5G Standalone (SA) deployment while providing better uplink and latency resilience, yielding the lowest game-day page-load failure rate (3.9%) and bounding image upload latency degradation to just 2.1x relative to empty-stadium baselines. These insights proves that densification through localized Wi-Fi deployment is essential to absorb severe stadium traffic spikes.
Science Journals
arXiv:2607.16026v1 Announce Type: new
Abstract: We calculate the one-loop self-energy in hydrogenlike atoms using a numerical Green function obtained by solving the radial Dirac equation in an exponential basis set. The self-energy correction in the ground state of hydrogenlike uranium is obtained with about $10^{-5}$ relative uncertainty in the Feynman gauge. Using a convergence acceleration scheme, we extend our calculations to the region of low nuclear charges. Our results allow calculating the self-energy correction for the hydrogen atom with $10^{-4}$ relative uncertainty. Calculations in the Coulomb gauge are also presented, improving the precision to $10^{-5}$. Present limitations and possible improvements of our method are discussed.
arXiv:2607.16050v1 Announce Type: new
Abstract: Pluvial (rainfall-driven) flooding accounts for 45% of National Flood Insurance Program (NFIP) claims in the United States and is harder to predict than its riverine and coastal counterparts, with existing approaches limited to coarse resolution, regional domains, or computationally intensive process-based models unsuitable for daily continental-scale use. We present DELUGE, a multimodal deep learning framework for daily pluvial flood damage prediction at ~1 km resolution and national scale, trained on spatially and temporally corrected NFIP claims (2017-2022) and structured around the hazard, exposure, and vulnerability components of disaster risk. Rather than blanket coverage of the Conterminous United States (CONUS), we model the top 100 highest-claim 75 km cells, distributed nationwide and accounting for ~81% of total pluvial flood claims. Our architectural novelty is a pair of parametric modules in the hydrometeorology branch, a Value Modulator and a Temporal Modulator, conditioned on terrain descriptors and AlphaEarth foundation-model embeddings, that expose directly inspectable hydrological response parameters and provide architecture-level interpretability-by-design. Under a spatial block holdout, DELUGE outperforms tuned Random Forest, XGBoost, and LightGBM baselines by 9% to 30% on a dollar-weighted area under the precision-recall curve (PR-AUC), a metric that emphasizes the rare, high-cost claims of greatest operational interest. Beyond DELUGE, we argue this interpretable conditioning scheme is a transferable pattern for integrating foundation-model embeddings into other geospatial prediction tasks.
arXiv:2607.15447v1 Announce Type: new
Abstract: Recent research in clinical machine learning, focusing on outcome predictions in intensive care unit (ICU), has shifted from bespoke supervised models to foundation models, utilising modern representation learning methods. Here, foundation models are pre-trained on mixtures of complex clinical data modalities, useful for various downstream tasks. Existing works often utilise Electronic Health Records (EHR) to provide rich and diverse patient observations to train clinical foundation models. However, existing methods do not sufficiently explore the shared temporal structures between clinical events and time series (TS) observations recorded in EHRs. This limitation potentially leads to less robust and adaptive clinical foundation models, resulting in reduced performance on downstream tasks. To fully exploit this temporal structure, we propose LLM4EHR, a new clinical foundation model trained on ICU EHR data. Combining domain adapted large language models with a transformer TS encoder, we pre-trained LLM4EHR by temporally aligning the EHR events and TS. For this, we propose a regularised contrastive objective to learn robust EHR TS representations conditioned on EHR event embeddings produced by the domain adapted LLM. Supported by an ablation study, we find that learnt EHR TS embeddings from LLM4EHR improve performance on various downstream clinical tasks with competitive performance. Further, we empirically demonstrate that LLM4EHR learns transferable clinical TS embeddings that can be deployed to new cohorts via k-shot adaptation. These findings provide a step towards building more generalisable and performant clinical foundation models.
arXiv:2502.06818v4 Announce Type: replace
Abstract: Recent works modify CLIP to perform open-vocabulary semantic segmentation in a training-free manner (TF-OVSS). In vanilla CLIP, patch-wise image representations mainly encode homogeneous image-level properties, which hinders the application of CLIP to the dense prediction task. Previous TF-OVSS works sacrifice globality to enhance the locality of CLIP features, by making each patch mainly attend to itself or its neighboring patches within a narrow local window. With their modifications,the ability of CLIP to aggregate global context information is largely weakened. Differently, in this paper, we rethink the global knowledge encoded by CLIP and propose GCLIP to answer how to extract and utilize beneficial global knowledge of CLIP for TF-OVSS. As the representation of each patch is finally determined by the attention weights and the Value embeddings, we propose to reshape the last-block attention and Value embeddings to aggregate useful global context into final features. Firstly, we aim to equip the last-block attention with image-level properties while not introducing homogeneous attention patterns across patches. To realize the goal, we fuse the attention from the global-token emerging blocks with the Query-Query attention. Secondly, we aim to make Value embeddings of the last-block attention module more semantically correlated. To realize this, we design a novel channel suppression strategy.Extensive experiments on five standard benchmarks demonstrate that our method consistently outperforms previous state-of-the-arts.
arXiv:2502.07524v2 Announce Type: replace
Abstract: The study investigates hip-hop music producer Scott Storch's approach to tonality, where the song's key is transposed to fit the Roland TR-808 bass drum instead of tuning the drums to the song's key. This process, involving the adjustment of all tracks except the bass drum, suggests significant production motives. The primary constraint stems from the limited usable pitch range of the TR-808 bass drum if its characteristic sound is to be preserved. The research examines drum tuning practices, the role of the Roland TR-808 in music, and the sub-bass qualities of its bass drum. Analysis of TR-808 samples reveals their characteristics and their integration into modern genres like trap and hip-hop. The study also considers the impact of loudspeaker frequency response and human ear sensitivity on bass drum perception. The findings suggest that Storch's method prioritizes the spectral properties of the bass drum over traditional pitch values to enhance the bass response. The need to maintain the unique sound of the TR-808 bass drum underscores the importance of spectral formants and register in contemporary popular music production.
arXiv:2502.11068v3 Announce Type: replace
Abstract: Anchors is a popular local model-agnostic explanation technique whose applicability is limited by its computational inefficiency. To address this limitation, we propose a memorization-based framework that accelerates Anchors while preserving explanation fidelity and interpretability. Our approach leverages the iterative nature of Anchors' algorithm which gradually refines an explanation until it is precise enough for a given input by storing and reusing intermediate results obtained during prior explanations. Specifically, we maintain a memory of low-precision, high-coverage rules and introduce a rule transformation framework to adapt them to new inputs: the horizontal transformation adapts a pre-trained explanation to the current input by replacing features, and the vertical transformation refines the general explanation until it is precise enough for the input. We evaluate our method across tabular, text, and image datasets, demonstrating that it significantly reduces explanation generation time while maintaining fidelity and interpretability, thereby enabling the practical adoption of Anchors in time-sensitive applications.
arXiv:2502.00400v4 Announce Type: replace
Abstract: Rogue wave formation and enhancement over coastal areas have been documented over the last decade. However, this recent knowledge is in apparent contradiction with the established observation of sub-Gaussian wave statistics near shallow water. Current theories and experiments describe the rogue wave amplification near shallow water regimes, but only for small-amplitude waves, and thus, not accounting for wave-breaking processes. To address this gap, we perform experiments to probe inhomogeneous wave fields nearing the wave-breaking regime. We also show that by increasing the significant wave height towards the breaking limit, the kinetic energy grows faster than the variance of the surface elevation due to nonlinearity, providing a physical explanation why the occurrence of rogue wave first increases in shallower waters and suddenly decreases further shorewards.
arXiv:2503.07769v2 Announce Type: replace
Abstract: We study the problem of computing the diameter and the mean distance of a continuous graph, i.e., a connected graph where all points along the edges, instead of only the vertices, must be taken into account. It is known that for continuous graphs with $m$ edges these values can be computed in roughly $O(m^2)$ time. In this paper, we use geometric techniques to obtain subquadratic time algorithms to compute the diameter and the mean distance of a continuous graph for two well-established classes of sparse graphs. We show that the diameter and the mean distance of a continuous graph of treewidth at most $k$ can be computed in $O(n\log^{O(k)} n)$ time, where $n$ is the number of vertices in the graph. We also show that computing the diameter and mean distance of a continuous planar graph with $n$ vertices and $F$ faces takes $O(n F \log n)$ time.
arXiv:2503.16592v2 Announce Type: replace
Abstract: Robust and precise robotic assembly entails insertion of constituent components. Insertion success is hindered when noise in scene understanding exceeds tolerance limits, especially when fabricated with tight tolerances. In this work, we propose ContactFusion which combines global mapping with local contact information, fusing point clouds with force sensing. Our method entails a Rejection Sampling based contact occupancy sensing procedure which estimates contact locations on the end-effector from Force/Torque sensing at the wrist. We demonstrate how to fuse contact with visual information into a Stochastic Poisson Surface Map (SPSMap) - a map representation that can be updated with the Stochastic Poisson Surface Reconstruction (SPSR) algorithm. We first validate the contact occupancy sensor in simulation and show its ability to detect the contact location on the robot from force sensing information. Then, we evaluate our method in a peg-in-hole task, demonstrating an improvement in the hole pose estimate with the fusion of the contact information with the SPSMap.
arXiv:2607.15876v1 Announce Type: new
Abstract: We present a new ML-like programming language Yarrow with algebraic effects and region-based memory management. Reconciling these programming language features into one language is challenging: the non-local control flow of algebraic effects break the stack discipline of function calls and returns that region-based memory management relies on, and multi-shot effect handlers break the invariant that regions can be exited at most once. We present a program logic, called Yarrow Logic (YL), that supports safe and modular reasoning about regions in the presence of one-shot and multi-shot effect handlers. We prove the logic sound w.r.t. the operational semantics of Yarrow which is inspired by the runtime of OCaml but refined for regions. We use YL to prove correctness of a number of case studies with algebraic effects, including checkpointing, asynchronous computation and a LIFO data structure implementation. Since all memory locations used in these case studies are allocated in regions, these case studies avoid using the less efficient garbage collected heap memory. We have formalized Yarrow's operational semantics, the Yarrow program logic, and all our case studies using the Iris separation logic framework on top of the Rocq Prover.
arXiv:2607.15483v1 Announce Type: new
Abstract: Learning reward functions from human preferences is a widely used approach for aligning robot behavior with user expectations in human-robot interaction. Most existing approaches assume that humans evaluate uncertain outcomes using expected utility (EU), aggregating outcome utilities linearly with their probabilities. However, behavioral evidence shows that humans are systematically risk-sensitive, overweighting rare negative events and exhibiting loss aversion. We study the consequences of this mismatch in social robot navigation, where safety-critical outcomes (e.g., collisions) are rare but highly consequential. We compare EU with Cumulative Prospect Theory (CPT), a nonlinear model of human decision-making, within a Bradley-Terry preference learning framework. Our preliminary experiments show that when preferences are generated by risk-sensitive users, CPT-based learners recover reward functions with substantially lower regret compared to EU-based learners. Our results highlight the importance of modeling human risk sensitivity when learning rewards from preferences over stochastic robot outcomes.
Holistic Fusion: Task- and Setup-Agnostic Robot Localization and State Estimation with Factor Graphs
arXiv:2504.06479v2 Announce Type: replace
Abstract: Seamless operation of mobile robots in challenging environments requires low-latency local motion estimation and accurate global localization. While most sensor-fusion approaches are designed for specific scenarios, this work introduces a flexible open-source solution for task- and setup-agnostic multimodal sensor fusion distinguished by its generality and usability. Holistic Fusion formulates sensor fusion as a combined estimation problem of i) the local and global robot state and ii) a (theoretically unlimited) number of dynamic variables, including automatic alignment of reference frames; this formulation fits countless real-world applications without conceptual modifications, offering a comprehensive solution beyond hard-coded/task-specific approaches. The proposed factor-graph formulation enables direct fusion of an arbitrary number of absolute, local, and landmark measurements expressed with respect to different frames by explicitly including them as states in the optimization and modeling their evolution as random walks. Moreover, local smoothness and consistency receive particular attention to prevent estimation jumps. Holistic Fusion enables low-latency and smooth online state estimation on typical robot hardware while simultaneously providing low-drift global localization at the IMU measurement rate. The efficacy of this released framework [1] is demonstrated in five real-world scenarios on three robotic platforms with distinct task requirements, highlighting the advantages of fusing multiple absolute measurement types [2]. [1] Code: https://github.com/leggedrobotics/holistic_fusion [2] Project: https://leggedrobotics.github.io/holistic_fusion
arXiv:2505.00100v2 Announce Type: replace
Abstract: Background and Context. Generative AI (GenAI) tools are increasingly used in programming courses, but we have limited evidence about how brief instruction can foster responsible, learning-oriented use.
Objectives. We evaluate "AI-Lab", a scaffolded GenAI literacy intervention, asking how students' self-reported GenAI usage and their openness and comfort using GenAI for conceptual, debugging, and homework tasks change after participation.
Methods. Across two semesters in three CS courses and one first-year engineering course at a U.S. university, we deployed the "AI-Lab" (pre-lab orientation, in-class critique of GenAI outputs, and a required homework reflection), collecting paired pre/post surveys (Perception N=831; Usage N=826) and six post-intervention focus groups; primary inferential analyses used the three CS courses (N=778 and 773, respectively). We analyzed survey shifts with paired non-parametric tests and focus groups via thematic analysis.
Findings. Openness increased for conceptual questions and homework help, and comfort increased for conceptual, debugging, and homework scenarios; self-reported frequency of GenAI use for homework and projects remained stable, while self-reported use for debugging increased. Focus group participants described adopting more iterative prompting strategies, becoming more skeptical of correctness, and articulating clearer boundaries around integrity and dependence.
Implications. A short, structured intervention can shift students' reported comfort with and willingness to use GenAI and influence the strategies they describe for engaging with it without increasing overall self-reported use on graded work. These results motivate future work triangulating surveys with behavioral traces and learning measures.
arXiv:2505.01360v3 Announce Type: replace
Abstract: Plate Tectonics requires strain localization over the entire thickness of the plates. However, modelling strain localization in the deep sections of the plates, which deform by ductile processes, remains a challenge, prompting the use of ad hoc schemes to model plate boundaries. We posit that the bottleneck for self-consistent generation of ductile strain localization in geodynamical models is poor representation of the intrinsic mechanical heterogeneity of rocks, in particular at small scales. This prevents its effects from being accounted for at larger scales, notably emergent properties that arise during upscaling, like anisotropy. To model this heterogeneity and its evolution, we adopt a stochastic description of the rheology, which evolves in time and space as a function of the local work-rate. This approach enables to reproduce the full variety of responses observed in nature, from heterogeneous deformation at the local scale, but homogeneous at the system-scale, with or without softening, to spontaneous development of system-scale shear zones. It enables, thereby, the construction of regime diagrams for ductile strain localization using three adimensional parameters. These parameters describe the degree of heterogeneity and potential for evolution of the rheology, function of (1) the constitutive equation, which represents the active deformation processes, (2) the evolution law for the rock mechanical properties, (3) the initial properties, and (4) the energy input to the system. This approach also enables the prediction of the intensity of localization and the resulting system-scale anisotropic softening, paving the way for self-consistent modelling of plate boundaries in geodynamics.
arXiv:2505.16121v3 Announce Type: replace
Abstract: Recommender system is one of the most critical technologies for large internet companies such as Amazon and TikTok. Although millions of users use recommender systems globally everyday, and indeed, much data analysis work has been done to improve the technical accuracy of the system, to our limited knowledge, there has been little attention paid to analysis of users' emotion in recommender systems. In this paper, we create a new theory and metrics that could capture users' emotion when they are interacting with recommender systems. We also provide effective and efficient visualization techniques for visualization of users' emotion and its change in the customers' lifetime cycle. In the end, we design a framework for emotion-based recommendation algorithms, illustrated in a straightforward example with experimental results to demonstrate the effectiveness of our new theory.
arXiv:2607.15982v1 Announce Type: new
Abstract: Learned policies trained end-to-end on large datasets often remain brittle in high-precision tasks and struggle with generalization. We find that these limitations largely stem from a lack of structure and focus in data collection. Our key insight is to leverage dense data collection only for the critical segment of contact-rich tasks and to rely on traditional planning during simple free-space motion. We propose an automated data-collection scheme in combination with offline deep reinforcement learning for the critical segment of the task, eliminating reliance on a teleoperator's skill and on online policy updates. Across four challenging real-world tasks, using only 2 to 2.5 hours of autonomous data collection, we achieve an average success rate of 96%, compared to the strongest baseline at 55%. Notably, performance remains high in out-of-distribution scenarios where end-to-end approaches struggle. Our results pave the way for targeted data collection for contact-rich tasks and for high success rates in precision applications.
arXiv:2607.15018v2 Announce Type: replace-cross
Abstract: High-dimensional categorical data arise in genetics, biomedicine, and the social sciences, yet visualization tools for such data remain far less developed than those for continuous variables. Existing methods either scale poorly, rely heavily on low-dimensional displays detached from the original data matrix, or prioritize predictive accuracy over interpretability. To address this gap, we introduce categorical Generalized Association Plots (cGAP), a visualization framework for nominal, ordinal, and binary data that preserves the original data matrix while augmenting it with interpretable geometric structure. cGAP uses Homogeneity Analysis (HOMALS) to embed subjects and category levels in a three-dimensional Euclidean space and maps the embedding to red-green-blue coordinates so that similar patterns receive similar colors. The framework integrates three coordinated views: a HOMALS-guided heatmap of the raw data matrix, a subject proximity matrix, and a variable proximity matrix. Seriation algorithms are then used to reorder rows and columns to reveal coherent clusters, outliers, and local-to-global structure. We also derive barycentric traceability, projection-distortion, and contrast-preservation properties that clarify how embedding geometry is transferred to the display. We demonstrate the versatility of cGAP through applications to student-animal classification data, mammalian dentition profiles, mushroom records from the UCI Machine Learning Repository, and the Clusters of Orthologous Genes database. These examples show that cGAP supports transparent exploratory analysis by maintaining traceability between derived visual structure and the original categorical observations. cGAP provides a full-matrix, heatmap-based visualization environment for investigating complex categorical datasets across scientific domains.
arXiv:2607.15764v1 Announce Type: new
Abstract: Quantum computing and quantum simulation with ultracold neutral atoms require Rydberg excitation of individual atoms in atomic arrays. Rydberg states are extremely sensitive to external electric fields, therefore precise three-dimensional control of the electric field is essential. We performed a spectroscopic study of three-photon Rydberg excitation of a single Rb atom in an optical dipole trap in the presence of an external DC electric field. The field was generated by eight electrodes deposited on the inner surfaces of an ultrahigh-vacuum glass cell. The used three-photon scheme of laser excitation of Rydberg \textit{nP} states allows the Stark shift and the splitting of the resonances to be observed simultaneously, which simplifies calibration of the electric field. In addition, in the commonly used two-photon Rydberg excitation schemes, the light shifts can complicate accurate determination of the DC Stark shift, particularly when the external electric field is scanned across different spatial directions, and different Stark components are excited. These shifts are absent in the three-photon excitation scheme used in our experiment. We demonstrated the ability to independently tune the electric field along all three spatial directions and to compensate for stray electric fields. The measured three-photon spectra exhibit Stark shifts and splittings of the three-photon resonance that are in good agreement with theoretical calculations. These results are also of interest for Rydberg electrometry.
arXiv:2607.15484v1 Announce Type: new
Abstract: Accurate prediction of toroidal plasma rotation is essential for optimizing confinement and stability in future fusion devices. This work investigates turbulent core momentum transport in the DIII-D tokamak across a transition from ion-temperature-gradient (ITG)- to trapped-electron-mode (TEM)-dominated turbulence. A momentum transport framework previously developed for ASDEX Upgrade is applied to modulated neutral beam injection experiments, separating diffusive, convective, and residual-stress contributions via Fourier analysis of the rotation response. The dataset spans low-rotation conditions, dominant electron heating, and background ExB shearing rates below turbulence growth rates, accessing more reactor-relevant conditions. Gyrokinetic CGYRO and gyrofluid TGLF calculations confirm the scan covers an ITG-to-TEM transition. The analysis yields Prandtl numbers near unity. The pinch number shows no explicit dependence on the transition, instead ordering roughly with the logarithmic density gradient. The normalized residual stress, in contrast, exhibits a non-monotonic, V-shaped dependence across the transition: co-current in deep ITG and deep TEM regimes, near-zero or counter-current in the intermediate mixed-mode regime. This trend collapses onto an approximately linear dependence against electron kinetic profile gradients, suggesting residual stress generation by profile-shearing effects. Weaker background ExB shearing further shifts residual stress toward counter-current values. Linear CGYRO simulations for representative ITG and TEM discharges yield Prandtl and pinch numbers in good agreement with experiment, supporting gyrokinetic momentum-transport predictions in TEM-dominated regimes. These results indicate residual stress plays an important role in core rotation prediction for low-torque plasmas and should be included in predictive models of future reactor scenarios.
arXiv:2607.15485v1 Announce Type: new
Abstract: Score-based generative models exhibit a puzzling behavior: they often appear to cover all modes of a target multimodal distribution and yet may fail to learn the correct relative mode amplitudes, which can be interpreted as mixture weights. We resolve this apparent paradox by relating the diffusion score matching (DSM) loss to the error in estimating mixture weights from generated samples. We show that, even when the target score is insensitive to mixture weights, generated samples can recover the weights accurately if the scores at intermediate noise levels are informative about the weights. Accordingly, we define the diffusion score sensitivity index (DSSI) as the variation in the DSM loss relative to changes in a parameter. We then show that the DSSI governs the accuracy with which the parameter of the target distribution can be estimated from generated samples. For Gaussian mixtures in arbitrary dimensions, we prove that the mixture weight estimation errors are on the same order as the DSM loss under mild conditions. Empirically, we show the emergence of sensitivity during the noising process of benchmark data distributions under typical noise schedules, and that these sensitivity values predict how well a well-trained model recovers mixture weights. Furthermore, we show that the choice of noise schedule can reduce diffusion sensitivity, leading to mode amplification. Although we focus on mixture weights, the proposed sensitivity framework governs the recovery of any qualitative parameter of the target distribution.
arXiv:2607.15489v1 Announce Type: new
Abstract: Important statistical properties of velocity circulation in homogeneous isotropic turbulence (HIT) have been unveiled in recent years, raising the question of whether they persist, or are modified, in other classes of turbulent flows. Motivated by the dominant role of small-scale structures in the circulation fluctuations of HIT, we investigate their relevance in direct numerical simulations of Rayleigh-B\'enard convection at a Rayleigh number of $10^9$ and Prandtl number of $O(1)$. Within the thermal boundary layer (TBL), the distribution of elementary vortices is found to be strongly correlated with the temperature field, while the statistics of their core aspect ratios is significantly altered. Additionally, the probability distribution functions of circulation, computed for planar contours which are parallel to the walls, display salient features closely akin to those observed in HIT, with the {\it Area Rule} (a connection between circulation statistics and minimal surfaces) remaining particularly well satisfied -- except for a possible transitional region at a distance of a few TBL thicknesses. Away from the TBL, as expected, the overall statistical behavior of these structures likewise resembles that of HIT, where the intermittent spatial distribution of small vortex tubes is instead determined by the energy dissipation field.
arXiv:2607.15492v1 Announce Type: new
Abstract: The heart's contractions are triggered by action potential waves, which propagate through the cardiac muscle and exhibit diverse spatio-temporal dynamics during different heart rhythms. The dynamics are modeled with partial differential equations (PDEs) in cardiac electrophysiology simulations. However, fitting such models to measurement data to develop digital twins or patient-specific computer models is challenging. Here, we introduce differentiable cardiac electrophysiology simulations that can be fitted automatically to spatio-temporal measurement data of action potential waves in cardiac tissue. By comparing the simulated dynamics with the observation data, we define a loss function that is minimized via gradient-based optimization. Backpropagating the loss gradient through the differentiable PDE solver enables us to learn the parameters and recover the full dynamics, even with sparse, noisy, or partial observations. Implemented using both the finite-difference and smoothed particle hydrodynamics methods, our simulation framework can be applied to pixel-, voxel-, or point-based data, such as 2D or 3D slabs, or arbitrary shapes, such as the heart's ventricles. Using this methodology, we locate early activation sites inside a 3D bi-ventricular simulation geometry and fit a phenomenological model to imaging data of a voltage spiral wave in a cardiac monolayer cell culture. With experimental data, we employed a perceptual loss based on the Video Joint-Embedding Predictive Architecture, which enables fitting to noisy imaging data, and a generative diffusion model to estimate initial conditions and constrain solutions. Differentiable cardiac electrophysiology simulations could improve the diagnosis of rhythm abnormalities in patients and facilitate the development of personalized models or digital twins of the heart.
arXiv:2607.15495v1 Announce Type: new
Abstract: Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous functional distinction has emerged in large language models. Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing. These representations, which we collectively call the J-space, exhibit the functional properties characteristic of a global workspace: their contents can be reported, deliberately summoned and held, used to carry the intermediate steps of silent reasoning, and passed as arguments to arbitrary downstream computations, while automatic processing such as text parsing and routine inference proceeds without them. The J-space also has structural signatures that global workspace theory associates with conscious access: it carries coherent content only in an intermediate band of layers, holds on the order of tens of concepts at a time, and is broadcast by the model's weights more widely than other representations. These properties make it a practical window into a model's unspoken thinking. In alignment audits, it reveals strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in the model's outputs. We find that post-training installs the Assistant's point of view in the workspace, and we introduce counterfactual reflection training, which improves behavior by training only what a model would say if interrupted and asked to reflect. These results indicate that language models maintain a small, privileged set of representations bearing some of the functional hallmarks of conscious access, and that decoding these representations sheds light on ongoing cognitive processes.
arXiv:2607.15835v1 Announce Type: new
Abstract: Big data clustering remains challenging: the Minimum Sum-of-Squares Clustering (MSSC) problem underlying K-means is NP-hard, and existing methods either reach poor local minima or require prohibitive metaheuristic hybrids. We target arbitrarily tall data: a fixed feature space may contain arbitrarily many, possibly infinitely many, observations, while the algorithm accesses only finite random samples. We propose Big-means++, an algorithm achieving scalability and global-search quality by curating inputs to MSSC optimization on big data. It orchestrates local K-means refinements into a data-native global search for big data clustering. Rather than optimizing the full-data MSSC objective, Big-means++ traverses sample-induced surrogate landscapes. Each sample defines a distinct empirical MSSC approximation with a perturbed local-optimum structure, turning sample-to-sample variation into a global-search mechanism. Unlike Big-means, a flowing-incumbent strategy propagates centroid state across empirical landscapes through K-means refinements on fresh samples without rollback to a best-so-far solution. This increases mobility and favors stable, high-quality configurations across approximations of the full-data structure. A new shaking mechanism varies sample size geometrically, broadening the surrogate landscapes explored across resolution scales, accounting for cluster imbalance, and improving solution quality. A competitive multi-agent system asynchronously explores independent sampled landscapes, transforming diverse stochastic trajectories into collective search intelligence. Automatic convergence detection stops each agent after attaining a high-quality solution but before further search risks degrading it, while providing a universal speed-quality control. Experiments on 22 datasets against 11 competing algorithms demonstrate the effectiveness, efficiency, and robustness of Big-means++.