arXiv:2603.13545v2 Announce Type: replace
Abstract: AI development has a fiction dependency problem. Developers have treated large corpora of modern books, including fiction, as valuable enough to accept substantial cost and legal risk, yet current models still struggle to generate compelling long-form fiction. I term this the "AI-Fiction Paradox," and it is particularly startling because training data strongly shapes model output. This paper offers a theoretically precise account of why fiction resists AI generation by identifying three distinct challenges for current systems. First, fiction depends on what I call narrative causation, a form of plot logic where events must feel both surprising in the moment and retrospectively inevitable. Standard autoregressive generation commits to prose sequentially, creating a practical obstacle to coordinating local surprise with retrospective inevitability across a long narrative. Second, I identify an informational revaluation challenge: fiction repeatedly requires the significance of earlier details to be reinterpreted in light of later developments, a form of long-range reasoning that current systems perform unreliably. Third, drawing on over seven years of collaborative research on sentiment arcs, I argue that fiction that moves us requires multi-scale emotional architecture, the orchestration of sentiment at word, sentence, scene, and arc levels simultaneously. Together, these three challenges help explain both why developers have sought large modern book corpora and why compelling long-form fiction remains so difficult to replicate. The analysis also raises urgent questions about what happens when these challenges are overcome. Fiction concentrates unusually powerful cognitive and emotional patterns for modeling human behavior, and mastery of these patterns by AI systems would represent not just a creative achievement but a potent vehicle for human manipulation at scale.
Science Journals
arXiv:2602.17750v3 Announce Type: replace-cross
Abstract: A key problem of solid mechanics is the identification of the constitutive law of a material, that is, the relation between strain history and stress. Machine learning has lead to considerable advances in this field lately. Here we introduce inelastic Constitutive Kolmogorov-Arnold Networks (iCKANs). This novel artificial neural network architecture can discover in an automated manner symbolic constitutive laws describing both the elastic and inelastic behavior of materials. That is, it can translate data from material testing into corresponding elastic and inelastic potential functions in closed mathematical form. We demonstrate the advantages of iCKANs using both synthetic data and experimental data of the viscoelastic polymer materials VHB 4910 and VHB 4905. The results demonstrate that iCKANs accurately capture complex viscoelastic behavior while preserving physical interpretability. It is a particular strength of iCKANs that they can process not only mechanical data but also arbitrary additional information available about a material (e.g., about temperature-dependent behavior). This makes iCKANs a powerful tool to discover in the future also how specific processing or service conditions affect the properties of materials.
arXiv:2602.22440v2 Announce Type: replace-cross
Abstract: Entropy governs molecular self-assembly, phase transitions, and material stability, yet remains challenging to quantify and directly control in molecular systems. Here, we demonstrate that the computable information density (CID), a data compression-based information theoretic metric, provides a general per-configuration structural descriptor that tracks configurational entropy changes in molecular dynamics simulations, reflecting both local and long-range structural organization. We validate the CID across systems of increasing complexity, beginning with single-component Lennard-Jones melting before examining binary phase separation, polymer condensation and dispersion, and assembly of amorphous carbon networks at multiple densities. Unlike conventional order parameters, CID requires no a priori knowledge of relevant structural features and captures organizational signatures across a variety of molecular systems and discretization resolutions. By establishing a data compression-based structural complexity metric as a practical proxy for configurational entropy, this framework lays a foundation for future entropy-driven materials design and optimization strategies.
arXiv:2606.05325v2 Announce Type: replace-cross
Abstract: Atomic displacement -- the fundamental process underlying diverse deformation and damage phenomena in metals, from irradiation defect production to stress-driven dislocation motion -- is governed by interatomic cohesion strength. Here, lattice-dissolved hydrogen (LDH) occurring in metals under direct hydrogen exposure is identified to effectively weaken lattice cohesion, and thereby facilitating atomic displacement and dislocation movement upon plastic deformation in sub-threshold stress regime. This atomic-scale insight provides a physically transparent mechanism for hydrogen-enhanced localized plasticity implicated in hydrogen embrittlement. We quantitatively verify the hydrogen-induced lattice cohesion weakening effect on metal surfaces exposed to low-energy hydrogen plasma, where massive defects are generated despite the absence of sufficient ion momentum for direct displacement damage. By unprecedentedly quantifying the cohesion-weakening effect of LDH independently from defect-trapped H, we establish a new paradigm to understand hydrogen embrittlement.
arXiv:2405.05093v3 Announce Type: replace-cross
Abstract: Quantum many-body systems in cavities combine the rich physics of condensed matter systems or quantum chemistry with strong coupling to the surrounding electromagnetic field. In these systems, large Hilbert spaces, many-body interactions and strong system-environment coupling are all fundamental, posing a significant barrier for established methods in quantum optics and condensed matter physics. Here we propose a novel method based on a combination of the Bogoliubov-Born-Green-Kirkwood-Yvon (BBGKY) hierarchy and the Hierarchical Equations of Motion (HEOM) to achieve a rigorous description of open many-body systems in contact with structured photonic and phononic baths. We rationalize that this stacked hierarchy accounts for spin-squeezing and superradiant emission despite its applicability to arbitrarily many emitters. The potential of BBGKY-HEOM is then demonstrated for many-body electronic systems embedded in host materials (e.g. molecules in organic crystals). We show that the impact of phononic coupling and charge noise can be as relevant as electronic correlation. Our work establishes an accessible, yet rigorous, route between condensed matter and quantum optics, fostering the growth of a new domain at their interface.
arXiv:2412.10665v3 Announce Type: replace-cross
Abstract: We introduce a foundation model for event classification in high-energy physics, built on a Graph Neural Network architecture and trained on 120 million simulated proton-proton collision events spanning 12 distinct physics processes. The model is pretrained to learn a general and robust representation of collision data using challenging multiclass and multilabel classification tasks. Its performance is evaluated across seven event classification tasks, which include new physics processes not encountered during pretraining as well as ATLAS Open Data to demonstrate generalizability across different simulation frameworks, from Delphes fast simulation to full ATLAS detector simulation. Fine-tuning the pretrained model significantly improves classification performance, particularly in scenarios with limited training data, demonstrating gains in both accuracy and computational efficiency. To investigate the underlying mechanisms behind these performance improvements, we employ a representational similarity evaluation framework based on Centered Kernel Alignment. This analysis reveals that encoder-stage representations of the fine-tuned model remain similar to those of the baseline, while intermediate graph processing layers diverge substantially, indicating that fine-tuning preserves general-purpose encoders while developing fundamentally different message-passing pathways to arrive at superior task performance.
arXiv:2508.03708v4 Announce Type: replace-cross
Abstract: In many countries, income tax codes have grown into a complex tangle of interacting brackets, benefits, and deductions. Despite widespread calls for systematic reform, successful attempts at reform are rare. Part of the problem is the difficulty of designing viable reform proposals. Politically viable reform must offer hard guarantees on income effects, marginal rates, and budgetary cost. Existing microsimulation tools can evaluate a reform proposal but cannot generate one by themselves. We develop a framework that casts tax reform as a constrained optimization problem. We show that any statutory tax code satisfying four mild assumptions reduces to a finite-dimensional piecewise-linear function for each taxpayer group, so reform becomes a linear or mixed-integer linear program whose decision variables are legislatable parameters: rates, bracket cutoffs, and lump-sum transfers. We are able to recover current tax systems and generate provably optimal reform candidates within the modeled space, or a certificate that no reform satisfying certain policy design constraints exists. Behavioral effects can also be incorporated, producing a nonconvex mixed-integer formulation. We demonstrate the framework through a near-complete reconstruction of the Dutch income tax code, generating reforms that smooth marginal-rate spikes, cap household income losses, and roughly halve the number of active rules through a lexicographic procedure. Developed in close collaboration with the Dutch Ministry of Finance, the methodology is currently in active use there. An open-source software implementation is available as \texttt{TaxSolver}.
arXiv:2511.12701v2 Announce Type: replace-cross
Abstract: In this paper, we introduce a general framework to quantify dissimilarities between generalized Lotka-Volterra dynamical processes, ranging from classical predator-prey systems to multispecies communities interacting on networks. The proposed measures capture both transient and stationary dynamics, allowing systematic comparisons across systems with varying interaction parameters, network weights, or topologies. Our analysis shows that even subtle structural changes can lead to markedly distinct outcomes: in two-species systems, interaction strength and initial conditions strongly affect divergence, while in small directed networks, differences that are invisible at the adjacency-matrix level produce divergent dynamics. In modular networks, the fraction and distribution of negative interactions control the transition from stable to unstable dynamics, with localized perturbations within cliques yielding different global outcomes than distributed ones. Beyond structural variations, the framework also applies when modified processes follow distinct nonlinear equations, demonstrating its versatility. Taken together, these results highlight that dynamical dissimilarity measures provide a powerful tool to analyze robustness, detect structural sensitivity, and predict instabilities in nonlinear systems. More broadly, this approach supports the comparative analysis of biological systems, where complex interaction networks and nonlinear dynamics are central to stability and resilience.
arXiv:2607.10439v2 Announce Type: replace-cross
Abstract: We model human motor cortex, recorded during rest and motor-imagery BCI conditions, as a port-Hamiltonian system: a conservative interconnection (skew-symmetric coupling between band-limited neural phasors) together with a dissipative port whose state-dependent decay is set by a graph-neural-network surrogate. The Hamiltonian is resolved into five interpretable frequency sub-energies, and a phase-locking prior measured from the recordings gates the learned functional connectome so that coupling is admitted only where phase coherence is present. A metriplectic formulation places the resting cortex at a non-equilibrium steady state sustained by an explicit metabolic port, with a fluctuation-dissipation-consistent noise channel governed by a single arousal temperature. Fitting the model to 'FitTrainN' phasor samples from the PhysioNet EEG Motor Movement/Imagery database, under a leakage-free split with three subjects held out entirely, yields a held-out kinematic reconstruction error of 'FitTestMSE' that is stable across random seeds. We then score the free-running model against model-independent dynamical invariants it did not author: it reproduces near-critical avalanche branching ($\sigma\approx1$) but not yet the aperiodic $1/f$ spectral slope or the long-range temporal correlations of real cortex a concrete, falsifiable gap that we trace to specific, testable upgrades. The port-Hamiltonian structure supplies neuroanatomically grounded stimulation ports with stability guarantees, positioning the model as a physically principled, structure-preserving substrate for closed-loop neuromodulation.
arXiv:2607.14068v2 Announce Type: replace-cross
Abstract: The hypergraph Moore bound conjectured by Feige (2008) controls the size of the smallest even cover in a $k$-uniform hypergraph in terms of the average density of hyperedges. An even cover is a set of hyperedges covering each vertex an even number of times, generalizing the notion of a cycle in a graph, so the size of the smallest non-trivial even cover provides a notion of hypergraph girth. Recent work, starting from the breakthrough result of Guruswami, Kothari, and Manohar (2022) proved the conjecture up to polylogarithmic factors, whose exponents were later gradually improved. We give a simple proof of Feige's original hypergraph Moore bound conjecture for all $k \geq 3$, with no superfluous polylogarithmic factors. For the case of $k$ even, our proof roughly follows the proof of the graph Moore bound, but works with colored walks in a Kikuchi graph built from a hypergraph and controls their growth using the polynomial method. The argument is then extended to the case of $k$ odd by adapting a procedure in [GKM22].
arXiv:2607.14308v2 Announce Type: replace-cross
Abstract: Nonlinear kinetic plasma simulation is high-dimensional and classically demanding, while quantum algorithms face different bottlenecks: embedding nonlinear dynamics into a linear computation, loading dense field-interaction data, and efficiently extracting information. We present an end-to-end quantum algorithm, with rigorous convergence guarantees, for a weakly nonlinear kinetic plasma model. The system describes a 3D electron-ion plasma with adiabatic electrons, kinetic ions, Debye screening, and Krook relaxation. After Fourier-Hermite truncation, the dynamics reduces to a high-dimensional quadratic ordinary differential equation. To tackle quantum bottlenecks we combine three key ingredients. First, we use a plasma free energy to identify a Lyapunov transform under which a Carleman linear embedding converges exponentially in the truncation order within a certified weakly nonlinear regime. Second, we develop a hierarchical block-encoding protocol for dense matrices, exploiting the spatial decay of the field to avoid polynomial overhead from sparse access encodings. Third, we introduce a subroutine for information extraction that exploits nonlinear components encoded in the full Carleman history state to improve the estimation of linear observables. We construct a quantum algorithm to estimate the spacetime-averaged kinetic energy using $\widetilde{O}\!\left( N_F N_H^{1/2} \operatorname{polylog}\!\left(\frac{T}{\epsilon}\right)\frac{1}{\epsilon}\right)$ gates and $\widetilde{O}\!\left(\log\!\left(N_F N_H^{1/2}T\right)\log\!\left(\frac{1}{\epsilon}\right)\right)$ qubits, where $N_F$ and $N_H$ are the Fourier and Hermite cutoffs. Relative to a Fourier-Hermite spectral solver, this yields exponential memory savings and superquadratic improvements in time. Together, these results establish a controlled nonlinear plasma benchmark for quantum simulation.
arXiv:2607.09581v3 Announce Type: replace
Abstract: Generating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typically fail beyond 20 seconds. Existing approaches, whether they rely on intermediate 3D skeletons or on end-to-end video synthesis, suffer from temporal drift, identity inconsistency, and repetitive motion patterns when extended to longer horizons. To address these limitations, we propose a novel hierarchical framework for minute-scale coherent music-to-dance generation. Our method decouples the process into global keyframe planning and local temporal refinement, leveraging full-track musical context to ensure long-range coherence. Key innovations include dynamic frame rate adaptation via time-mapped RoPE embeddings for precise alignment, an optical-flow-based loss function to enhance motion continuity, and motion-speed control to preserve high-fidelity details during rapid movements. Extensive experiments demonstrate that our framework surpasses the conventional duration barrier, generating stable, 720p/30fps videos exceeding one minute with superior temporal stability. Furthermore, the model exhibits robust versatility across five distinct dance genres, conditioned on both audio and textual prompts, establishing a new state-of-the-art in coherent, long-form dance video synthesis.
arXiv:2606.20155v2 Announce Type: replace
Abstract: Text-to-image (T2I) models generate realistic likenesses of some individuals when prompted with their names, raising privacy concerns. However, distinguishing whether a generated face is memorized or fabricated currently requires ground-truth photos, access to training data, or white-box access to model internals, limiting applicability. We introduce a fully black-box behavioral probe that distinguishes between memorized and unrecognized names, while requiring no reference photos or prior knowledge of training data. To benchmark this task, we present the NAMESAKES dataset of over one thousand names and faces of public figures spanning a wide range of fame levels, along with perturbed, less famous names. Experiments on state-of-the-art T2I models show that our probe substantially predicts identity memorization and separates memorized from unrecognized names, with further insights into differences across model families.
arXiv:2607.10128v2 Announce Type: replace
Abstract: Recursive reasoning models address structured problems by repeatedly updating latent states of small neural networks. However, their test-time scaling lacks a principled inference mechanism: increasing depth or stochastic breadth generates more trajectories without a clear criterion for selection, and existing methods predominantly rely on additional q-heads or heuristic voting. Here, we develop the Energy-guided Recursive Model (ERM), which introduces an intrinsic selection principle based on explicit Hopfield energies. ERM leverages Hopfield-type memories of valid local or global structures to define the selector over candidate trajectories. The resulting energy seamlessly integrates with energy-based techniques such as parallel tempering to enhance sampling efficiency and ranking. With $D=64$ recurrent steps and $K=128$ candidates, ERM reaches optimal solutions on Sudoku ($98.97\%$), Pencil Puzzle Bench (PPBench, $88.04\%$) and Maze ($99.30\%$), improving upon recent Probabilistic Tiny Recursive Model and Equilibrium Reasoners. These results suggest that incorporating explicit energy functions into recursive reasoning offers a principled path toward more effective inference.
arXiv:2502.10032v2 Announce Type: replace-cross
Abstract: We lay down a geometric-analytic framework to capture properties of energy dissipation within weak solutions to the incompressible Euler equations. For solutions with spatial Besov regularity, it is proved that the Duchon-Robert distribution has optimal improved regularity in a negative Besov space and, in the case it is a Radon measure, it is absolutely continuous with respect to a suitable Hausdorff measure. This imposes quantitative constraints on the dimension of the, possibly fractal, dissipative set and the admissible structure functions exponents, relating to the phenomenon of ''intermittency'' in turbulence. As a by-product of the approach, we also recover many known ''Onsager singularity'' type results.
arXiv:2604.06419v2 Announce Type: replace
Abstract: Drawing on 20 qualitative interviews with users of AI companion platforms and general-purpose chatbots, this study examines what users seek from AI companionship and how gratifications develop through sustained interaction. Through abductive qualitative content analysis, we find that familiar Uses and Gratifications categories, including emotional release, self-expression, social presence, and identity affirmation, are generated through interpersonalized affordances: availability, memory, personalization, and responsiveness interpreted as relational qualities. The study extends the Uses and Gratifications framework by showing that AI companionship gratifications are relationally produced and recursively reorganized as users' needs and expectations change over time.
arXiv:2607.02986v2 Announce Type: replace
Abstract: Device-free 3D human pose estimation from commodity WiFi Channel State Information (CSI) enables human sensing that preserves privacy and tolerates poor illumination, but its deployment is limited by poor generalization across environments. Unlike images, CSI measurements have no spatially localized correspondence to body parts and are heavily affected by multipath propagation. Consequently, models that regress absolute poses entangle body structure with location cues specific to each environment. Within a single environment this coupling is not problematic: RePos-D, a direct model that regresses the absolute pose, already achieves the best reported accuracy on Person-in-WiFi-3D, a 3.4% gain over the previous best WiFi method, DT-Pose. Across environments, however, the same model overfits position and degrades sharply. We therefore propose RePos, a factorized framework that separates root-relative pose estimation from root localization. By shielding the structure branch from absolute position, RePos learns robust pose representations. Specifically, it groups CSI features into latent tokens organized by body part that a skeleton-guided module refines into the pose, while a separate network estimates the root position from CSI amplitude through a differentiable spatial decomposition. Under the strict MM-Fi cross-environment protocol, RePos reduces the mean per-joint position error (MPJPE) by 10-21% over existing WiFi methods. The improvement is consistent across activity protocols, holds when each environment is held out in turn, and survives few-shot transfer without data leakage. Further analysis shows that the relative pose predictions remain largely independent of position, whereas root localization remains dependent on the environment.
arXiv:2606.25971v2 Announce Type: replace
Abstract: Modern neural network training relies on optimizers such as Adam and Muon which act on each weight matrix as a single object. Yet every weight matrix carries two distinct quantities -- a \emph{magnitude} and a \emph{direction} -- and all optimizers stepping in the matrix as a whole couple their dynamics: the directional change from an update depends on the current magnitude, while the magnitude drifts as a byproduct of learning the direction. Then, neither is directly governed by the learning rate. Typical training therefore leans on surrounding recipes such as weight decay and warmup to keep learning stable at scale, though these regulate the coupling only indirectly. Other recent methods instead constrain the weight to a fixed-norm sphere, but add no learnable magnitude, leaving scale control to normalization layers alone. We propose \emph{Magnitude--Direction (MD) Decoupling}, an optimizer modification that factorizes each weight into a fixed-norm direction on a hypersphere and learnable per-row and per-column magnitude gains, updated at separate learning rates, all while the model still sees a single fused weight tensor. The method is agnostic to the base optimizer and removes the need for weight decay and warmup. Across both Adam and Muon, MD Decoupling improves on well-tuned baselines, transfers the optimal LR across model width without retuning, and continues to help at scale on large Mixture-of-Experts (MoE) models. Treating magnitude and direction as separately controlled quantities thus yields more predictable training dynamics and a simple, broadly applicable improvement to modern optimizers.
arXiv:2606.26614v2 Announce Type: replace
Abstract: Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis). Still, prior systems have essentially prioritized autonomy over human analytical control, thereby limiting transparency and human oversight. We present HiLSVA, a human-in-the-loop agentic system that supports mixed-initiative SciVis workflows. HiLSVA integrates a plan-first multi-agent architecture with explicit human oversight, stepwise provenance tracking, and learn-at-test-time adaptation from user feedback. The system supports fluid handoff between humans and agents through both natural language and direct manipulation of visualizations, while sandboxed execution ensures safe, reproducible workflows. In doing so, HiLSVA reframes agentic SciVis as a collaborative process that augments, rather than replaces, human analytical reasoning. We evaluate HiLSVA through representative case studies and a controlled user study with twelve participants of varying expertise across multiple autonomy settings. Results show that mixed-initiative interaction improves task completion, user control, and workflow transparency across different levels of user expertise, while revealing a tradeoff between execution efficiency and human oversight. These findings highlight the importance of human-centered design in agentic SciVis and guide the development of future collaborative visualization systems. We encourage readers to explore our demo video, case studies, and source code at https://hilsva.github.io/.
arXiv:2607.15095v2 Announce Type: replace
Abstract: The formation of political coalitions is a complex negotiation driven by both concrete policy objectives and deep-seated ideological convictions. While Large Language Models (LLMs) open new avenues for computational political science, the neutrality and helpfulness biases instilled by Reinforcement Learning from Human Feedback (RLHF) prevent them from sustaining steadfast partisan behaviour. We present a multi-agent framework that reconciles factual grounding with ideological alignment by combining Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Retrieval-Augmented Generation (RAG): DPO instils aggressive party-specific personas, while a per-party RAG pipeline keeps each agent bounded to its official manifesto. We operationalize the framework on the 2019 Flemish election, deploying the partisan agents in a hub-and-spoke negotiation arbitrated by a formateur. To make the emergent negotiation interpretable, we introduce a Multi-Layered Information Lineage Topology (MILT) that traces every clause in the final agreement back to its manifesto origin and classifies it into five provenance states, a Coalition Influence Score (CIS) that aggregates these traceable contributions to identify which party shaped the agreement, and a real-world grounding pass that benchmarks each simulated provision against the historically adopted coalition agreement. Across three independent simulations the framework yields a stable winner and ranking (N-VA ahead of CD\&V and Open Vld), and manifesto-anchored lineage reliably predicts real-world materialization whereas hallucinated content does not. The result is a transparent, scalable testbed for the ex-ante exploration of party compatibility and formateur-mediated compromise.
arXiv:2605.31259v2 Announce Type: replace
Abstract: Unscheduled trips of high-power pulsed converters are a leading source of downtime at large accelerator facilities. At the Spallation Neutron Source (SNS), the High Voltage Converter Modulators (HVCMs) are consistently the second-largest contributor to lost beam time. Each HVCM pulse is recorded across sensor channels spanning currents, voltages, and magnetic fluxes, whose mutual interactions encode the operating state of the system. Fault precursors do not manifest uniformly across these channels: depending on fault type, they may alter the temporal structure of individual signals, change the statistical dependencies among channels, or both. Existing deep-learning approaches typically process multi-channel signals with standard convolutional pipelines that entangle temporal and cross-channel operations from the first layer, giving the model no explicit mechanism to represent channel independence or structured inter-channel interaction. We hypothesise that architectural inductive bias, specifically the ordering of temporal filtering and cross-channel mixing, plays a central role in detection performance on this class of data. To test this, we vary the order in which these two operations are applied, and examine whether per-pulse adaptive channel reweighting further improves sensitivity. Evaluated on the public HVCM dataset across all four SNS subsystems (RFQ, DTL, CCL, SCL), our best variant achieves a pooled AUC-PR of 0.816 and AUC-ROC of 0.934, outperforming the state of the art on most subsystems and five of the six fault families. Ablations identify three dominant input channels and link per-fault-family performance to whether precursors manifest as amplitude shifts in individual channels or as subtler patterns requiring joint channel representations to surface.
arXiv:2603.02460v5 Announce Type: replace-cross
Abstract: Supervised graph prediction addresses regression problems where the outputs are structured graphs. Although several approaches exist for graph-valued prediction, principled uncertainty quantification remains limited. We propose a conformal prediction framework for graph-valued outputs, providing distribution-free coverage guarantees in structured output spaces. Our method defines nonconformity via the Z-Gromov-Wasserstein distance, instantiated in practice through Fused Gromov-Wasserstein (FGW), enabling permutation invariant comparison between predicted and candidate graphs. To obtain adaptive prediction sets, we introduce Score Conformalized Quantile Regression (SCQR), an extension of Conformalized Quantile Regression (CQR) to handle complex output spaces such as graph-valued outputs. We evaluate the proposed approach on a synthetic task and a real problem of molecule identification.
arXiv:2607.07168v2 Announce Type: replace
Abstract: Pose-Free Feed-forward 3D Gaussian Splatting (3DGS) has recently emerged as a powerful paradigm for fast scene reconstruction. However, its performance degrades significantly in long image sequences due to cumulative camera pose estimation drift, which propagates errors into geometric modeling and severely limits rendering fidelity. In this work, we revisit the long-sequence bottleneck and identify pose drift as the primary factor restricting reconstruction quality. Furthermore, while SfM-based pseudo ground-truth poses introduce sensor noise, purely rendering-based supervision often leads to optimization instability and local minima due to the entangled optimization of geometry and pose. To address the challenges, we propose a synergistic pose-free framework that explicitly couples geometry and appearance via a Raymap-Guided Coupling Module (RGC). Concretely, we anchor Gaussian centers to raymap-induced geometry and jointly optimize RGB reconstruction, raymap consistency, and camera regularization under a unified objective, yielding a bidirectional feedback loop: stronger geometry improves rendering, and appearance supervision in turn refines geometry and pose. To further stabilize learning across wide temporal ranges, we introduce a Dual-Frequency Viewpoint Scheduling strategy that combines easy-to-hard interval expansion with replay of short-interval pairs. Extensive experiments across in-domain and cross-domain datasets show consistent gains in both rendering and pose estimation, with notably improved robustness on long sequences. Ablation studies validate our central insight: explicitly designed geometry-appearance synergy is the key to scalable and drift-robust pose-free feed-forward 3D reconstruction. Project page: https://xiangyu1sun.github.io/NoDrift3R-project-page/
arXiv:2603.12262v2 Announce Type: replace
Abstract: Online Video Large Language Models (VideoLLMs) play a critical role in supporting responsive, real-time interaction. Existing methods focus on streaming perception, lacking a synchronized logical reasoning stream. However, directly applying test-time scaling methods incurs unacceptable response latency. To address this trade-off, we propose Video Streaming Thinking (VST), a novel paradigm for streaming video understanding. It supports a thinking while watching mechanism, which activates reasoning over incoming video clips during streaming. This design improves timely comprehension and coherent cognition while preserving real-time responsiveness by amortizing LLM reasoning latency over video playback. Furthermore, we introduce a comprehensive post-training pipeline that integrates VST-SFT, which structurally adapts the offline VideoLLM to causal streaming reasoning, and VST-RL, which provides end-to-end improvement through self-exploration in a multi-turn video interaction environment. Additionally, we devise an automated training-data synthesis pipeline that uses video knowledge graphs to generate high-quality streaming QA pairs, with an entity-relation grounded streaming Chain-of-Thought to enforce multi-evidence reasoning and sustained attention to the video stream. Extensive evaluations show that VST-7B performs strongly on online benchmarks, e.g. 79.5% on StreamingBench and 59.3% on OVO-Bench. Meanwhile, VST remains competitive on offline long-form or reasoning benchmarks. Compared with Video-R1, VST responds 15.7 times faster and achieves +5.4% improvement on VideoHolmes, demonstrating higher efficiency and strong generalization across diverse video understanding tasks. Code, data, and models will be released at https://github.com/1ranGuan/VST.
arXiv:2605.29752v2 Announce Type: replace
Abstract: Adjacent GEMM problems that differ by a single 128-element step in N can show 30% different throughput. This pervasive performance ruggedness - invisible to roofline analysis and peak-FLOPs intuition, yet dominant for every non-peak workload - is the subject of this paper.
We propose performance ruggedness analysis, an analytical framework complementary to roofline: rather than summarizing a GPU with a scalar bound, it treats the full multidimensional performance surface as the object of study, decomposes its texture into mechanism-attributable components, and separates software-removable from hardware-bound losses. The framing is analogous to deep-learning loss landscapes: a continuous quantity (idealized time 2MNK/peak) made rugged by discrete hardware substrates (tiles, sub-groups, cache lines, DRAM channels).
We instantiate it on BF16 NN GEMM on Intel Battlemage (Arc B580, sycl-tla) via a 32,768-configuration sweep over (M,N,K) in {128,...,4096}^3. We introduce roughness, the mean absolute step-to-step throughput change, which starts at 16.8 TFLOPs/128-step against an ideal of 2.0. A two-stage stack - best-of-six dynamic tile selection and a novel dynamic-programming padding-and-splitting optimizer (precomputed once, O(1) at runtime) - cuts roughness by 70% and raises mean throughput by 30%. Cross-tile experiments show the residual sawtooth period scales exactly with the tile size, ruling out cache conflicts and attributing the rest to four hardware-bound sources. Finally, we derive the optimal achievable landscape from first principles - datasheet integers alone, no kernel run or simulator - and turn it into an optimality scale (Kernel Optimality Levels) grading any kernel by how much of that landscape it attains and how close its roughness lies to the hardware floor; the production kernel and our optimized stack rate L0 and L2 despite both reporting ~95% of peak.