arXiv:2601.19128v2 Announce Type: replace Abstract: In terrestrial laser scanning (TLS)-based mechanical, electrical, and plumbing (MEP) point cloud segmentation, safety-critical components such as reducers and valves are persistently misclassifed, blocking reliable engineering knowledge extraction. This stems from a dual crisis--extreme class imbalance (215:1) compounded by geometric ambiguity, since most tail classes share cylindrical primitives with dominant head classes--that existing frequencybased re-weighting methods cannot resolve. We propose spatial context constraints that exploit neighborhood prediction consistency to disambiguate locally similar structures. Our approach extends Class-Balanced (CB) Loss with two architecture-agnostic mechanisms: Boundary-CB, an entropy-based constraint that emphasizes ambiguous boundaries and encodes an MEP assemblytopology prior, and Density-CB, a density-based constraint that compensates for scan-dependent variations and encodes TLS sensor-physics knowledge. Both operate at the loss level and integrate into existing pipelines without backbone modifcations. On the Industrial3D dataset (612.7M labelled points from water treatment facilities), our method achieves 55.74% mIoU, exceeding the strongest of three representative fully supervised backbone baselines (39.83-52.48% mIoU), with a 21.7% relative improvement on tail-class performance (29.59% vs. 24.32%) while preserving head-class accuracy (88.14%). Components with primitive-sharing ambiguity show strong gains: reducer improves from 0% to 21.12% IoU, and valve improves by 24.3% relative. These results show that spatial context constraints reduce primitive-sharing errors in the target industrial MEP setting and support more reliable identifcation of safety-critical components for Digital Twin and Scan-to-BIM applications. Code: https://github.com/PointCloudYC/LongTail3D.git.
Science Journals
arXiv:2603.08990v2 Announce Type: replace-cross Abstract: We design and evaluate an edge-side measurement procedure for auditing service tiering and quota-based throttling in Starlink. Using a 232.8-hour plan-hopping campaign on a UK residential terminal, we align 1 Hz terminal telemetry with host-side probes to obtain portal-labeled traces spanning priority, post-quota throttling, stay-active operation, and residential service. These regimes manifest as distinct signatures in goodput, PoP RTT, and an internal-to-user ratio \(R=C_{\mathrm{int}}/T_{\mathrm{user}}\). We further show that high-speed \(R\) is stable over 30-minute sub-windows, that low-rate clusters have no aligned persistent obstruction or PoP-loss signature, and that clean high-speed dips do not move \(R\) into the low-rate band. A lightweight rule on windowed medians separates high-speed from low-rate operation on this trace without operator visibility.
arXiv:2603.14830v3 Announce Type: replace Abstract: Dataset distillation, a training-aware data compression technique, has recently attracted increasing attention as an effective tool for mitigating costs of optimization and data storage. However, progress remains largely empirical. Mechanisms underlying the extraction of task-relevant information from the training process and the efficient encoding of such information into synthetic data points remain elusive. In this paper, we theoretically analyze practical algorithms of dataset distillation applied to the gradient-based training of two-layer neural networks with width $L$. By focusing on a non-linear task structure called multi-index model, we prove that the low-dimensional structure of the problem is efficiently encoded into the resulting distilled data. This dataset reproduces a model with high generalization ability for a required memory complexity of $\tilde{\Theta}$$(r^2d+L)$, where $d$ and $r$ are the input and intrinsic dimensions of the task. To the best of our knowledge, this is one of the first theoretical works that include a specific task structure, leverage its intrinsic dimensionality to quantify the compression rate and study dataset distillation implemented solely via gradient-based algorithms.
arXiv:2603.14844v2 Announce Type: replace Abstract: The ocean remains one of the least instrumented parts of Earth, and many geophysical, biological, and anthropogenic signals go undetected for lack of instrumentation. Distributed acoustic sensing (DAS) can transform submarine fiber-optic cables into dense seafloor sensor arrays, but extracting diverse signals from massive DAS recordings remains challenging. Here we present DASNet, a deep learning framework that detects, classifies, and picks arrival times of diverse marine signals in continuous DAS data. Applied to nearly four years of Seafloor Fiber-Optic Array in Monterey Bay recordings, DASNet identifies more than 620,000 events. These detections reveal local earthquakes; distant earthquake- and volcanic-eruption-generated T-waves from the southwestern Pacific and mid-ocean ridge systems; more than 510,000 blue and fin whale calls with seasonal and interannual variability consistent with hydrophone records; and vessel traffic near the cable. Together, these results show that submarine fiber-optic cables combined with deep learning enable scalable, high-resolution ocean monitoring.
arXiv:2603.03695v2 Announce Type: replace Abstract: Reliable localization is essential for sustainable forest management, as it allows robots to revisit and monitor the status of individual trees over long periods. In modern forestry, this management is structured around Digital Forest Inventories (DFIs), which encode stems using compact geometric attributes rather than raw data. Despite their central role, DFIs have been overlooked in localization research, and most methods still rely on dense gigabyte-sized point clouds that are costly to store and maintain. To improve upon this, we propose TreeLoc++, a global localization framework that operates directly on DFIs as a discriminative representation, eliminating the need to use the raw point clouds. TreeLoc++ reduces false matches in structurally ambiguous forests and improves the reliability of full 6-DoF pose estimation. It augments coarse retrieval with a pairwise distance histogram that encodes local tree-layout context, subsequently refining candidates via DBH-based filtering and yaw-consistent inlier selection to further reduce mismatches. Furthermore, a constrained optimization leveraging tree geometry jointly estimates roll, pitch, and height, enhancing pose stability and enabling accurate localization without reliance on dense 3D point cloud data. Evaluations on diverse forests across four countries show that TreeLoc++ achieves precise localization with centimeter-level accuracy. We further demonstrate robustness to long-term change by localizing data recorded in 2025 against inventories built from 2023 data, spanning a two-year interval. The system represents 15 sessions spanning 7.98 km of trajectories using only 250KB of map data and outperforms both hand-crafted and learning-based baselines that rely on point cloud maps. This demonstrates the scalability of TreeLoc++ for long-term deployment. TreeLoc++ is open-sourced at \bl{https://github.com/minwoo0611/TreeLoc-plusplus}.
arXiv:2603.27044v3 Announce Type: replace Abstract: Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality and substantial functional redundancy inherent to the policy parameter space. A recent framework, which we refer to as Action-based Policy Compression (APC), mitigates this issue by compressing the parameter space $\Theta$ into a low-dimensional latent manifold $\mathcal Z$ using a learned generative mapping $g:\mathcal Z \to \Theta$. However, its performance is severely constrained by relying on immediate action-matching as a reconstruction loss, a myopic proxy for behavioral similarity that suffers from compounding errors across sequential decisions. To overcome this bottleneck, we introduce Occupancy-based Policy Compression (OPC), which enhances APC by shifting behavior representation from immediate action-matching to long-horizon state-space coverage. Specifically, we propose two principal improvements: (1) we curate the dataset generation with an information-theoretic uniqueness metric that delivers a diverse population of policies; and (2) we propose a fully differentiable compression objective that directly minimizes the divergence between the true and reconstructed mixture occupancy distributions. These modifications force the generative model to organize the latent space around true functional similarity, promoting a latent representation that generalizes over a broad spectrum of behaviors while retaining most of the original parameter space's expressivity. Finally, we empirically validate the advantages of our contributions across multiple continuous control benchmarks.
arXiv:2505.15720v2 Announce Type: replace-cross Abstract: In this paper, we introduce a new family of codes relevent for rank and sum-rank metrics. These codes are based on an effective Chinese remainders theorem for linearized polynomials over finite fields. We propose a decoding algorithm for some instances of these codes.
arXiv:2506.15020v2 Announce Type: replace-cross Abstract: We propose persistent discrete homology as a tool for topological data analysis and discuss its advantages over the existing methods. In particular, we provide empirical evidence that persistent discrete homology is more noise-resistant than persistent homology of the Vietoris-Rips complex for data coming from non-metric settings.
arXiv:2606.21185v2 Announce Type: replace-cross Abstract: There is a precise sense in which drawing causal inferences from observational data is hard, even when identifiability is assumed. In particular, Robins and Ritov (1997) and Robins et al. (2003) showed that causal effects can be discontinuous as a function of the data distribution: two arbitrarily close data distributions might correspond to different causal effects. This is a fact independent of the choice of estimator; however, not all estimators are equally unstable. Our contribution is to surface a second layer of instability that depends on the choice of estimator. We show that many standard point estimates can be read as point summaries of multimodal distributions over the space of structural causal models. As such, estimators can jump discontinuously in the data distribution. This defines a taxonomy of estimators that admits a decision-theoretic reading: stability depends on whether the implicit loss function an estimator optimizes is aligned with the causal effect itself. Specifically, inverse propensity weighted estimators and regression estimators are examples of discontinuous summaries, while explicit posterior means and medians are shown to be continuous.
arXiv:2602.17176v4 Announce Type: replace-cross Abstract: Crystal structure prediction (CSP), which aims to predict the 3D atomic arrangement of a crystal from its composition, is central to materials discovery and mechanistic understanding. Crystal symmetry plays a crucial role in CSP, but given the composition in a unit cell, existing methods either struggle with the NP-hard combinatorial challenge of enforcing symmetry rigorously or rely on retrieving known templates, inherently limiting both physical fidelity and the discovery of genuinely new materials. To address this challenge, we introduce NextCrystal, a symmetry-driven generative framework that employs large language models to encode chemical semantics and directly generate fine-grained Wyckoff site patterns from atomic stoichiometry, eliminating reliance on database lookups. To overcome the combinatorial complexity of site assignments, we incorporate domain knowledge via an efficient, linear-complexity heuristic beam search, rigorously enforcing algebraic consistency between site multiplicities and atomic stoichiometry. By integrating this symmetry-consistent template into a diffusion backbone, the framework constrains the stochastic generative trajectory to a physically plausible geometric manifold. NextCrystal achieves state-of-the-art performance on stability, uniqueness, and novelty (SUN) benchmarks, as well as superior structural matching, establishing a rigorous paradigm for exploring previously unexplored crystallographic space without relying on prior structural templates. As a representative application, first-principles screening of HfO2 candidates generated by NextCrystal identifies a previously unreported dynamically stable Pnma phase, 0.056~eV/atom lower in energy than the conventional high-pressure Pnma phase.
arXiv:2602.18180v2 Announce Type: replace-cross Abstract: Achieving near-unity fidelity in conventional continuous-variable quantum teleportation schemes based on two-mode squeezed vacuum states is fundamentally unattainable. To overcome this limitation, alternative approaches utilizing ensembles of two-dimensional entangled qubits have been proposed. In this work, we investigate continuous-variable quantum teleportation employing entangled qutrit resources under realistic noise effects. The results demonstrate that the proposed scheme performs well in both ideal and noisy conditions, enabling high-fidelity teleportation with a reasonable success probability.
arXiv:2606.23327v2 Announce Type: replace Abstract: Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and domain-specific tasks. They face two critical limitations: i) inability to handle diverse video comprehension and editing operations, and ii) lack of long-video understanding for coherent narrative creation. We propose VideoAgent, an all-in-one agentic framework addressing these challenges through two key innovations. First, we develop automated video shot creation with shot planning agents for coherent narratives and cross-modal retrieval for aligned visual content. Second, we design a multi-agent orchestration framework integrating over thirty specialized editing agents. Intent parsing filters relevant tools while textual-gradient graph optimization assembles complex editing pipelines. Extensive experiments on our newly-proposed VideoEdit benchmark and public datasets demonstrate VideoAgent's superiority over existing multimodal LLMs and agentic systems. VideoAgent achieves 87-95% orchestration success rates while reducing API costs by 60%. Human evaluation across six video categories shows VideoAgent produces professional-quality content approaching human-level performance, with ratings only 4% below human-created videos. We release our code at https://github.com/HKUDS/VideoAgent.
arXiv:2606.30543v2 Announce Type: replace Abstract: With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly important. Emotional entrainment is shaped by social relationships and conversational context, influencing affective coordination over time. We introduce DyadEE, a dataset for emotional entrainment detection in dyadic speech interactions, containing both emotionally entrained conversations and synthetic interactions where entrainment is disrupted through partner swapping and emotion resynthesis. We further propose TRACE, a window-level framework that models dyadic interaction as ordered sequences of acoustic embeddings derived from emotion fine-tuned Whisper representations, treating each sample as an interaction trace rather than pooled utterances. Experimental results on DyadEE show that incorporating conversational context and relationship information improves emotional entrainment detection, with TRACE achieving the best accuracy of 97.01%.
arXiv:2606.25264v2 Announce Type: replace-cross Abstract: Information measures acquire operational meaning through coding theorems. The quantum conditional mutual information (QCMI) is nonnegative due to strong subadditivity, yet a direct connection with channel coding has remained elusive. In this work, we propose a quantum communication task-conditional quantum communication-that fills this gap. We show that the optimal rate for establishing quantum correlation between two parties, assisted by a third system, is given by half the QCMI. This result naturally extends the classical key generation capacity of Csisz\'ar and Ahlswede to the quantum domain. We place our model within the family tree of quantum protocols and compute the conditional capacity for several example channels. Our results provide new insights for code design in reliable quantum information processing.
arXiv:2606.31928v2 Announce Type: replace Abstract: Accurate aerodynamic modeling of satellites in very low Earth orbit (VLEO) requires gas-surface interaction (GSI) models that capture the full velocity spectrum from thermal to orbital speeds. Atmospheric particles initially strike spacecraft surfaces at hypersonic velocities of 6 000 - 10 000 m/s. Due to surface roughness and complex geometries, especially within air-breathing electric propulsion (ABEP) intake systems, multiple collisions occur, progressively reducing the particle velocities. A recent machine learning framework for deriving scattering kernels from molecular dynamics (MD) simulations has shown promise, but remains limited to high-velocity single impacts and possibly violates fundamental equilibrium principles such as detailed balance. This work extends this machine learning based scattering kernel to cover the complete velocity range using conditional normalizing flows trained with physics-informed constraints, enabling accurate modeling of multi-bounce scenarios in realistic VLEO applications. We train a conditional Real-valued Non-Volume Preserving (cRealNVP) model on expanded molecular dynamics simulations covering velocities from thermal to hypersonic speeds, incorporating a detailed balance loss term. The resulting model demonstrates improved accuracy compared to previous approaches even in the original high-velocity regime, while successfully capturing thermal-velocity scattering. Quantitative assessment shows that thermalization is approximated within acceptable tolerances. This framework provides essential capabilities for accurate ABEP intake optimization and VLEO mission planning while offering a general methodology applicable to broader rarefied gas dynamics problems requiring thermodynamic consistency.
arXiv:2607.01469v2 Announce Type: replace Abstract: Agentic Video Question Answering (VideoQA) systems invoke tools during inference, but their tool libraries are fixed, so recurring procedures are rebuilt from primitives on every question. Synthesizing composite tools could remove this overhead, but whether such expansion helps is hard to assess: final-answer accuracy, the standard metric, ignores inference effort, so it cannot reveal how a system shifts cost. We propose a cost-aware, paired protocol for auditing tool-augmented video agents. The protocol pairs two complete systems on the same input for each question and reports their net difference across accuracy and cost jointly. For each question, it sorts the paired outcome into one of six groups defined by joint correctness and by the change in visible tool calls, separating accuracy-preserving efficiency gains from harmful regressions. Significance is reported with McNemar's test and paired bootstrap confidence intervals. We instantiate the protocol on Dynamic-SAGE, an agentic VideoQA framework that synthesizes, validates, and persistently registers executable composite tools for reuse on unseen questions, and evaluate it against the SAGE baseline on SAGE-Bench. The audit reveals a multi-axis profile that a scalar accuracy comparison would miss: Dynamic-SAGE improves accuracy by 7.5 points (p < 0.001) and reduces reasoning turns and visible tool calls by roughly 28%, while shifting rather than reducing inference cost, as token usage rises 34% and cost 26%. Gains are largest on visual and open-ended questions and neutral on verbal and multimodal ones, and residual failures concentrate on hard, open-ended questions where the pipeline does the most work. By measuring accuracy and cost jointly, the protocol shows where the pipeline-level difference is reliable and where it is not. The code is available at https://github.com/KurbanIntelligenceLab/Dynamic-SAGE.
arXiv:1811.05336v2 Announce Type: replace-cross Abstract: Inference for factor models is often hampered by the lack of tractable and accurate variance estimates, which can materially distort downstream analyses. In practice, uncertainty in the residual covariance matrix is frequently either ignored or addressed through computationally intensive resampling methods that tend to be unstable. This paper develops a unified analytical framework for inference in exploratory factor analysis under several widely used extraction rules, including least-squares, principal-factor, iterative principal-component, alpha, and image factoring. By treating these estimators as implicitly defined functions of the sample covariance matrix, we derive closed-form Jacobians that translate perturbations in the covariance matrix into changes in the resulting factor solutions. Combined with the delta method and consistent estimators of the sample covariance matrix, the proposed approach yields standard errors that are straightforward to compute and remain valid under non-Gaussianity, heteroskedasticity, and serial or cross-sectional dependence. Simulation evidence confirms that the analytical standard errors accurately capture finite-sample variability while avoiding both the instability of bootstrap procedures and the restrictive assumptions underlying Fisher information-based inference. An application to a factor-augmented structural vector autoregressive (SVAR) model further demonstrates how accounting for this source of uncertainty can substantially affect impulse-response inference. Taken together, the results provide a practical and general tool for propagating estimation uncertainty in settings where factor extraction serves as an intermediate step.
arXiv:2501.15380v2 Announce Type: replace-cross Abstract: Understanding the nature of the accretion disk, its interplay with the X-ray corona, and assessing black hole spin demographics remain open challenges in astrophysics. In this paper, we examine the predictions of the standard $\alpha$-disk model, origin of the puzzling soft X-ray excess, and measure the black hole spin parameter by applying an updated high-density disk reflection model to the XMM-Newton/NuSTAR broadband (0.3$-$78 keV) X-ray spectra of a sample of 11 Type-1 AGN. Our Bayesian analysis confirms that a variable-density relativistic disk reflection model with a broken power-law emissivity profile can simultaneously fit the soft X-ray excess, broad iron K line emission, and Compton hump in 3 out of 11 AGN. For the remaining sources, a distinct warm Comptonization component is still required, which supports a hybrid origin for the soft X-ray excess. The measured temperature and optical depth of the warm corona span nearly the entire theoretically allowed range, with median values of $0.43_{-0.18}^{+0.40}$ keV and $12.5_{-3.9}^{+3.1}$, respectively. Our first systematic calculation of the disk-to-corona power transfer fraction reveals that the fraction of power released from the accretion disk into the hot corona spans a wide range, with a sample median of $0.68_{-0.25}^{+0.25}$. The sample median values for the hot coronal plasma temperature and optical depth are $54_{-12}^{+11}$ keV and $0.98_{-0.28}^{+0.22}$, respectively. Finally, through both hard X-ray (3$-$78 keV) and broadband (0.3$-$78 keV) relativistic reflection spectroscopy, we systematically constrain the black hole spin parameter across the mass scales of $\log(M_{\rm BH}/M_{\odot}) \sim 5.5-9.0$, thereby increasing or refining the available spin measurements in the AGN population by $\sim$20%.
arXiv:2604.25777v2 Announce Type: replace-cross Abstract: Federated inference enhances LLM performance in edge computing through weighted averaging of distributed model predictions. However, autoregressive LLM inference requires frequent full-model forward passes across workers, severely limiting decoding throughput. Distributed deployment further aggravates this due to a communication bottleneck: each worker must transmit full token probability distributions per draft token, dominating end-to-end latency. To address these challenges, we introduce speculative decoding to enable parallel LLM processing and propose a top-K compressed transmission scheme with two server-side reconstruction strategies. We theoretically analyze the robustness of our method in terms of local reconstruction error, aggregation bias, and acceptance-rate bias, and derive corresponding bounds. Experiments demonstrate that our scheme achieves high generation fidelity while significantly reducing communication overhead.
arXiv:2607.02124v1 Announce Type: new Abstract: Conventional traction control architectures intervene only after the adhesion limit of a tire has already been breached. This paper investigates whether Rolling Split Conformal Prediction , monitoring the volatility of non-conformity residuals from a per-driver Random Forest model of expected slip behavior , can serve as a statistically grounded pre-incident warning signal, ahead of gross traction loss. Unlike an earlier internal draft of this work, the evaluation reported here corrects a confound in the slip proxy (vehicle speed is included as an explicit model feature, not left implicit in the target's denominator), uses every racing lap for each driver rather than only the fastest lap, and is scored against real, timestamped incident labels extracted from FIA Race Control Messages and track-limits lap deletions rather than narrated post-hoc. The result is negative: across 19 drivers and 55,563 test-phase telemetry samples, the rolling-volatility detector achieves a mean precision of essentially 0.0 and mean recall of 0.0 against 14 ground-truth incidents, while flagging on average 15.3% of all samples as anomalous , too high a false-alarm rate for any early-warning use. A static 95th-percentile threshold baseline performs no better in any way that would justify the added complexity of the conformal-volatility formulation. Residual autocorrelation diagnostics show the split-conformal exchangeability assumption is violated for every driver (Ljung-Box p < 0.001, n = 19/19), which is one plausible driver of the high false-alarm rate. We report this as a methodologically rigorous negative finding, diagnose its likely causes, and outline what a genuinely predictive version of this approach would require.
arXiv:2605.06634v2 Announce Type: replace Abstract: We describe libwignernj, a freely available, BSD-licensed library that evaluates Wigner 3j, 6j, and 9j symbols, Clebsch--Gordan, Racah $W$, and Fano $X$ coefficients, and Gaunt coefficients over both complex and real spherical harmonics in standards-compliant C99. libwignernj represents factorials by the vector of their signed prime-exponent decomposition - a prime-factorization technique introduced for the angular-momentum coefficients by Dodds and Wiechers (Comput. Phys. Commun. 4, 268 (1972)) and refined in a long line of subsequent work - and combines that representation with the multiword-integer Racah sum of Johansson and Forss\'en (SIAM J. Sci. Comput. 38, A376 (2016)), under which every intermediate quantity is an exact rational and all rounding is confined to the final floating-point conversion. Single-, double-, and long-double-precision results are correct to the last representable bit, and IEEE 754 binary128 evaluation through libquadmath and arbitrary-precision evaluation through the GNU Multiple-Precision Floating-Point Reliable (MPFR) library are optionally exposed. libwignernj has no mandatory runtime dependencies and no caller-side initialization step, making it easy to embed across the atomic, molecular, nuclear, and electromagnetic-scattering applications in which these coefficients arise. C++, CPython, and Fortran 90 bindings ship alongside the C library. Half-integer angular momenta are encoded exactly via integer $2j$ arguments throughout the application programming interface (API). CMake-package and pkg-config files ship for drop-in integration into downstream projects, and a continuous-integration (CI) pipeline runs the full test suite on Linux (shared and static), macOS, and Windows on every push.
arXiv:2602.05999v4 Announce Type: replace Abstract: How does the amount of compute available to a reinforcement learning (RL) policy affect its learning? Can policies using a fixed amount of parameters, still benefit from additional compute? The standard RL framework does not provide a language to answer these questions formally. Empirically, deep RL policies are often parameterized as neural networks with static architectures, conflating the amount of compute and the number of parameters. In this paper, we formalize compute bounded policies and prove that policies which use more compute can solve problems and generalize to longer-horizon tasks that are outside the scope of policies with less compute. Building on prior work in algorithmic learning and model-free planning, we propose a minimal architecture that can use a variable amount of compute. Our experiments complement our theory. On a set 31 different tasks spanning online and offline RL, we show that $(1)$ this architecture achieves stronger performance simply by using more compute, and $(2)$ stronger generalization on longer-horizon test tasks compared to standard feedforward networks or deep residual network using up to 5 times more parameters.
arXiv:2605.15062v3 Announce Type: replace Abstract: Gastrointestinal cancers cause approximately 3.4 million deaths annually, and early small-bowel lesions are easily missed at wireless capsule endoscopy (WCE). RGB-trained WCE classifiers conflate hemoglobin contrast with bile staining and illumination falloff, limiting sensitivity to small-vessel vascular findings such as Lymphangiectasia. We introduce a physics-informed framework that injects an analytic, Monte-Carlo-inspired hemoglobin prior into a standard classifier purely at training time -- to our knowledge the first use of an explicit optical light-transport prior in WCE classification. On Kvasir-Capsule (47,238 frames, 43 patients, 11 evaluable classes; patient-disjoint split) we evaluate, across six seeds against an RGB-only EfficientNet-B0 baseline, a five-channel input-fusion variant feeding the prior alongside RGB, a distillation variant that runs on plain three-channel RGB at inference, and a three-stream extension adding a temporal Transformer and an autoencoder-residual stream; we replicate across ResNet-18 and ConvNeXt-Tiny and assess cross-vendor zero-shot transfer on the public Galar cohort. Input fusion lifts cross-seed macro-AUC from 0.760 to 0.783 (5/6 seeds positive); distillation reaches 0.773; the three-stream model reaches 0.804 (+0.044 over baseline, paired DeLong p < 0.0001). Lymphangiectasia AUC rises from 0.238 to 0.337, sign-consistent across all six seeds. A four-variant ablation reveals a parameterization-mechanism boundary: only the spatial-channel form lifts. Cross-vendor zero-shot on Galar retains about 60% of the lift. The distillation variant deploys on plain RGB with a free interpretability heatmap, and we release GalKva-2026, a paired cross-vendor benchmark.
arXiv:2006.08444v4 Announce Type: replace Abstract: Many modern asymmetric encryption methods rely on prime numbers, as they have distinctive properties. For instance, the security of RSA cryptosystem relies on the computational difficulty of factoring a large composite number in its prime factors, a problem that remains challenging for classical computers but potentially solvable using quantum algorithms. On the other hand, generating large prime numbers is also challenging due to their irregular distribution among integers, necessitating the use of primality testing algorithms to verify candidate primes. In this paper, we intensively review and classify various classical and quantum algorithms for factorization and primality testing, highlighting their advantages, limitations, speed/accuracy tradeoffs, time complexities, along with a brief summary. Furthermore, we apply and compare these algorithms to gain practical insights and conduct a comprehensive performance comparison. The insights from this paper show that while quantum factoring algorithms, particularly Shor's algorithm and its refinements, have introduced significant advancements over their classical counterparts, quantum primality testing algorithms have not demonstrated comparable advantages.
arXiv:2607.01312v1 Announce Type: new Abstract: Visual narratives are central to storyboards, comics, children's media, and film previsualization, where viewers understand stories from images alone. Recent generators such as StoryDiffusion produce coherent sequences, but visual coherence does not guarantee that source-story transition meaning remains recoverable. Existing benchmarks assess visual quality, content faithfulness, and scene coherence, but miss a critical failure mode: storyboards where scenes appear visually coherent while the semantic link between scenes disappears. We introduce KathaTrace, a generator-agnostic protocol for diagnosing semantic trajectory collapse, defined as the loss of transition meaning needed to understand how one scene follows another. KathaTrace evaluates transitions under three evidence conditions: text-only, image-only, and text-plus-image, and filters ambiguous items. We contribute KathaBench-25K, with 5,000 narratives from classical collections including Aesop, Panchatantra, and Kathasaritasagara, 20,000 transitions, and 28,712 recoverability questions. We define Semantic Trajectory Gap, or STG, as text-only minus image-only recoverability, measuring transition meaning lost during visualization. Human validation yields Fleiss' kappa = 0.845. Experiments across state-of-the-art generators show substantial STG of 23.5 +/- 1.3. Semantic Compass, an actionability probe, uses KathaTrace signals for post-generation repair and improves storyboard selection.