Forskningsradar

Science Journals

Peer-reviewade publikationer — 53899 artiklar

GeoRouteNet: A Geometry-Aware Non-Autoregressive Neural Solver for the Euclidean Traveling Salesman Problem
arXiv:2606.22776v2 Announce Type: replace Abstract: Non-autoregressive neural solvers amortize computation across traveling salesman problem (TSP) instances, but models trained on random Euclidean instances can degrade when the number or spatial distribution of nodes changes. We study whether explicit geometric features and a richer within-instance training signal improve transfer across graph sizes and spatial distributions. We introduce GeoRouteNet, which augments a non-autoregressive TSP solver with centered node offsets and radii, learnable radial distance bases, distance-aware graph attention, explicit edge messages, and cross-layer representation mixing. We also introduce multi-candidate self-comparison reinforcement learning (MCS-RL), which trains on several sampled tours per instance using a leave-one-out adaptive baseline, winner-candidate guidance, and annealed entropy regularization. In a single-seed study, all neural variants are trained only on random TSP-50 instances and evaluated with the same greedy and beam-search decoders. Under Beam-1000 decoding, GeoRouteNet-MCS-RL obtains gaps of 0.32% on the TSP-50 validation set used for checkpoint selection, 1.26% on a correlated TSP-100 size diagnostic, and 3.60% across 27 TSPLIB EUC_2D instances. The NAR4TSP-PG gaps on the same evaluations are 0.42%, 2.73%, and 17.12%. A 2x2 comparison crosses encoder and training choices. Under PG, the geometry-aware encoder has lower gaps than the reproduced encoder on the TSP-100 diagnostic and TSPLIB. With the geometry-aware encoder, MCS-RL is associated with a further reduction; with the reproduced encoder, it has a higher TSPLIB gap.
Apparatus for quantum-mixture research in microgravity
arXiv:2508.20820v2 Announce Type: replace-cross Abstract: Experiments with ultracold quantum gases are a rapidly advancing research field with many applications in fundamental physics and quantum technology. Here, we report on a high-flux generation of Bose-Einstein condensate mixtures of $^{41}$K and $^{87}$Rb, using a fully integrated sounding rocket setup. We compare the release and the free expansion of the quantum mixtures obtained with the apparatus placed on ground or in free fall in an Einstein-Elevator. The release dynamics are governed by the intra- and interspecies interactions as well as the decaying magnetic field during the release. The latter can be minimized by a dedicated switch-off protocol of the trap generating currents where an exact model enabled us to characterize the interaction effects. Our results establish a new benchmark for generating ultracold mixtures on mobile platforms, with direct relevance for future experiments on interacting quantum gases and tests of the equivalence principle in space.
Quantization of Polaritons Confined in Dielectric Structures
arXiv:2510.14566v2 Announce Type: replace-cross Abstract: Light-matter interaction in the regime of strong quantum coupling is usually treated within the framework of the Hopfield model. However, the picture of coupling well-defined modes of light and matter is correct only as long as the shapes of these eigenmodes are not substantially modified by the interaction. Moreover, parameters of theoretical models are usually obtained by fitting to experimental data. To date, there is no straightforward method to determine a quantum master equation corresponding to a system with specific dielectric structure, which may lead to incompatibility of theoretical descriptions and physical realizations. In this work, a recipe for obtaining a quantum model in the polariton eigenmode basis is presented, based on Bogoliubov transformation in the conservative case and third quantization technique in the dissipative case. It is shown how this method can be used for boosting interaction strength and engineering nonlocal many-body interactions in carefully designed nanostructures, resulting in strongly nonclassical correlations of emitted light.
Quantum Search With Generalized Wildcards
arXiv:2511.04669v2 Announce Type: replace-cross Abstract: In the search with wildcards problem [Ambainis, Montanaro, Quantum Inf.~Comput.'14], one's goal is to learn an unknown bit-string $x \in \{-1,1\}^n$. An algorithm may, at unit cost, test equality of any subset of the hidden string with a string of its choice. Ambainis and Montanaro showed a quantum algorithm of cost $O(\sqrt{n} \log n)$ and a near-matching lower bound of $\Omega(\sqrt{n})$. Belovs [Comput.~Comp.'15] subsequently showed a tight $O(\sqrt{n})$ upper bound. We consider a natural generalization of this problem, parametrized by a subset $\cal{Q} \subseteq 2^{[n]}$, where an algorithm may test whether $x_S = b$ for an arbitrary $S \in \cal{Q}$ and $b \in \{-1,1\}^S$ of its choice, at unit cost. We show near-tight bounds when $\cal{Q}$ is any of the following collections: bounded-size sets, contiguous blocks, prefixes, and only the full set. All of these results are derived using a framework that we develop. Using symmetries of the task at hand we show that the quantum query complexity of learning $x$ is characterized, up to a constant factor, by an optimization program, which is succinctly described as follows: `maximize over all odd functions $f : \{-1,1\}^n \to \mathbb{R}$ the ratio of the maximum value of $f$ to the maximum (over $T \in \cal{Q}$) standard deviation of $f$ on a subcube whose free variables are exactly $T$.' To the best of our knowledge, ours is the first work to use the primal version of the negative-weight adversary bound (which is a maximization program typically used to show lower bounds) to show new quantum query upper bounds without explicitly resorting to SDP duality.
Improving Backward Conformal Prediction via Non-Conformity Score Transformation
arXiv:2602.01733v3 Announce Type: replace-cross Abstract: Conformal Prediction (CP) provides a statistical framework for uncertainty quantification that constructs prediction sets with coverage guarantees. While CP yields uncontrolled prediction set sizes, Backward Conformal Prediction (BCP) inverts this paradigm by enforcing a predefined upper bound on set size and estimating the resulting coverage guarantee. However, the looseness induced by Markov's inequality within the BCP framework causes a significant gap between the estimated coverage bound and the empirical coverage. In this work, we introduce ST-BCP, a novel method that introduces a data-dependent transformation of nonconformity scores to narrow the coverage gap. In particular, we develop a computable transformation and prove that it outperforms the baseline identity transformation. Extensive experiments demonstrate the effectiveness of our method, reducing the average coverage gap from 4.20\% to 1.12\% on common benchmarks.
Quantum Matrix-Element Estimators for Spin-Coupled Generalized Valence Bond Wavefunctions (SCGVB)
arXiv:2603.12045v3 Announce Type: replace-cross Abstract: Valence-bond wavefunctions such as spin-coupled generalized valence bond (SCGVB) provide compact and chemically interpretable descriptions of strong correlation, but their nonorthogonal determinant structure makes quantum evaluation of overlaps and Hamiltonian matrix elements difficult. Here we introduce an ancilla-free, measurement-driven framework that reformulates these quantities as vacuum expectation values of Pauli-string operators, enabling evaluation with shallow circuits, local Clifford rotations, and computational-basis measurements, without controlled operations. We validate the method for H$_4$ along a dissociation pathway and for a 70-determinant C$_2$ active-valence benchmark. The reconstructed matrices agree with the classical L{\"o}wdin-based reference, while Chirgwin-Coulson weights remain chemically consistent. These results establish the proposed estimators as a low-depth, circuit-compatible formulation for evaluating SCGVB matrix elements, while making no claim of quantum computational advantage.
Gravitational-wave astronomy requires population-informed parameter estimation
arXiv:2604.15885v2 Announce Type: replace-cross Abstract: Gravitational-wave events are interpreted in terms of Bayesian posteriors for their source properties inferred under unphysical reference priors. Though these parameter estimates are important intermediate data products for downstream analyses, we demonstrate that they are generically biased and therefore should not be used for astrophysical interpretation directly, as is common. Hierarchical parameter estimation is the solution, as joint analysis of the entire catalog of observations reduces statistical uncertainties and actually informs the correct prior, with population-informed event parameters now appropriate for astrophysical interpretation. As an example, we show how the most extreme measurements from a catalog can be derived and used to identify exceptional events from previous and ongoing observing runs, pointing out they are more informative about the population than any individual event. Using LIGO-Virgo-KAGRA data, we thus demonstrate that population inference is not optional to interpret gravitational-wave observations.
Evidence-Aware MapReduce for Forkable Compute
arXiv:2607.09689v3 Announce Type: replace Abstract: Snapshot-backed sandboxes make branching cheap while leaving evidence dependence unchanged. Branches can reuse a model, prompt, repository, tests, observations, or execution ancestor, so counting outputs can amplify one repeated error into high-confidence consensus. We introduce an \emph{evidence-aware reduction contract}: each worker reports an estimate, estimated information, evidence identifiers, fork lineage, and execution metadata. For independent workers estimating one common parameter, we use standard inverse-information pooling in its Gaussian/Wald form. The fixed-dimensional numeric summary can merge in any tree order; evidence IDs and lineage follow separate rules. The residual $\Delta$ measures disagreement, becomes Cochran's $Q$ in the scalar inverse-variance case, and appears in the product integral. A reference implementation validates serialized records, rejects repeated nonempty evidence identifiers, carries evidence and lineage through tree reduction, and uses Cholesky-based numerical linear algebra. Unit tests and seeded synthetic checks exercise the algebra, unequal information, and forged precision; one four-worker named-snapshot trace exercises the end-to-end path. Platform logs document the exercised execution paths. A central open systems challenge is to turn evidence identity and fork lineage into a dependence model for correlated and adaptively selected AI branches.
Length Penalties Make Chain-of-Thought Less Monitorable
arXiv:2607.09786v2 Announce Type: replace Abstract: Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In our experiments, training with length penalties does not stop misleading hints from steering models, even though the models' chains of thought mention the hint much less often. A token-accuracy evaluation would count these runs as successful because they use fewer reasoning tokens with little accuracy loss; it would miss whether the remaining trace still shows what drove the answer. We train Qwen3-4B and Qwen3-14B variants with different target chain lengths, then evaluate them with biasing-hint interventions on held-out MMLU-Pro-R and four transfer benchmarks. Compression sharply cuts reasoning tokens, preserves most multiple-choice accuracy, and leaves hint influence near baseline. At the strongest target, lower-bound faithfulness falls to 63.1% of baseline for Qwen3-14B and 69.4% for Qwen3-4B; the raw rate at which a monitor catches hint use falls from 69% to 49% and from 60% to 48%. To separate length from content, we randomly delete sentences from uncompressed baseline chains until the remaining text matches the compressed length. Even after this length matching, compressed chains disclose the hint 7-35 percentage points less often than baseline chains that we shorten at random, for both Qwen3 sizes and all five evaluation distributions. Compression therefore does more than shorten reasoning, preferentially removing the cues a monitor needs to see what influenced the answer. Together, these results reveal a compression-monitorability frontier in which cheaper reasoning can preserve answers while making the influences behind them harder to detect.
The Post-Silicon Semiconductor Era: A Review of Physics, Synthesis, and Architectural Integration of Carbon Nanotube Field-Effect Transistors
arXiv:2509.00947v3 Announce Type: replace-cross Abstract: Silicon CMOS scaling is approaching a set of hard physical limits. Direct source-to-drain quantum tunneling, an unscalable subthreshold swing, and the thermal ceiling known as Dark Silicon motivate the search for a new channel material that can carry logic scaling forward. This review builds the case for single-walled carbon nanotubes (SWCNTs) as that material. We follow a continuous narrative from electronic-structure theory through synthesis to integration. The SWCNT bandgap and its near-ballistic transport limits are derived from the graphene zone-folding framework and the Landauer-Buttiker formalism. We benchmark these theoretical limits against ideal coaxial electrostatic bounds to evaluate how well the geometry suppresses short-channel effects before quantum tunneling takes over. Comparing this analytical framework against published 5 nm experimental data illustrates the aggressive subthreshold degradation driven by source-to-drain tunneling. Furthermore, we derive the exact areal-density equivalence between 1D and 2D quantum capacitance. This demonstrates that even close-packed arrays cannot fully close the dimensional gap to 2D materials, underscoring why superior carrier velocity and electrostatics must carry the CNTFET advantage. Next, we examine CoMoCAT growth and aqueous two-phase extraction against the semiconducting-purity demands of logic fabrication, alongside contact engineering and reversible chemical doping. A closing techno-economic analysis weighs this physics and process picture against IEEE IRDS roadmap projections and environmental health constraints. Taken together, the evidence points to materials purification, contact reliability, and bias temperature instability as the remaining practical barriers to commercial CNTFET adoption.
Manifold Dimension Estimation via Local Graph Structure
arXiv:2510.15141v5 Announce Type: replace-cross Abstract: Most existing manifold dimension estimators rely on the assumption that the underlying manifold is locally flat within the neighborhoods under consideration. More recently, curvature-adjusted principal component analysis (CA-PCA) has emerged as a powerful alternative by explicitly accounting for the manifold's curvature. Motivated by these ideas, we propose a manifold dimension estimation framework that captures the local graph structure of the manifold through regression on local PCA coordinates. Within this framework, we introduce two representative estimators: quadratic embedding (QE) and total least squares (TLS). Experiments on both synthetic and real-world datasets demonstrate that these methods perform competitively with, and often outperform, state-of-the-art approaches.
Worlds in One Demo: A Synthetic Data Engine for Learning Open-World Mobile Manipulation
arXiv:2607.13154v2 Announce Type: replace Abstract: Learning open-world mobile manipulation policies requires vast data to achieve spatial generalization, long-horizon robustness, and scene generalization. Current prevailing data collection paradigms, teleoperation and UMI, demand prohibitive human effort and cost at scale. To scale beyond the limits of manual data collection, we seek to maximize the value of each human demonstration by scalable data generation. To this end, we introduce WANDA: learning open-World mobile mANipulation from one demonstration via a synthetic DAta engine. WANDA first reconstructs background Gaussian splats and robot-object interaction trajectories from source RGBD observations, as a world substrate for later planning and rendering. It then rearranges contact-rich robot-object interaction segments into extensive spatial configurations, utilizing whole-body motion planning to chain them into new trajectories. To enhance long-horizon robustness, it applies Corrective State Expansion to increase the robot and object state diversity at different stages of mobile manipulation. To unlock cross-environment generalization, trajectories are synthesized on diverse generated 3D worlds from everyday photos. Furthermore, we synthesize photo-realistic observations by compositing rendered robot and object meshes with Gaussian splatting backgrounds. We evaluate our approach on extensive simulation and real-world tasks in various scenes. Experiments show that policies trained with WANDA achieve long-horizon robustness, broad spatial generalization and cross-environment generalization from one real demonstration. Moreover, WANDA naturally supports cross-embodiment data generation, validated by zero-shot deployment on another mobile manipulator with a distinct morphology.
What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
arXiv:2607.13162v3 Announce Type: replace Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers only a handful of traits. We present the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and labeling every trait in two open-weight models as natural (expressed at baseline), steerable latent but amplifiable, or intractable (resistant to standard extraction). Both models default to helpful, task-oriented behavior: all nine agentic traits are natural, and their default clinician behavior matches a board-certified psychologist's independent desirability judgments on 16 of 17 traits. Steering produces its largest gains on traits these defaults exclude: hyperbole, hallucination, and sycophancy. The same asymmetry holds across all 171 generic-trait pairs: two steerable traits can collapse the composition, but pairs involving a default never do. Where standard extraction fails on a trait like "evil," a vector transferred from a fine-tuned variant still recovers it, with the residual refusals appearing inside the model's chain-of-thought. Persona vectors are most informative not as a set of controls but as a probe of behavioral organization.
Faithful Autoformalization of Natural Language Assertions
arXiv:2607.13303v2 Announce Type: replace Abstract: Formal contracts are essential for software testing and verification, yet writing them remains labor-intensive and error-prone. LLMs offer a promising path toward autoformalization: synthesizing executable assertions from natural-language specifications and thereby bridging the gap between informal developer intent and formal executable specifications. We present Monty: an autoformalization framework for assertions that tackles the challenges of expectations of validity of assertions and ambiguity in natural-language. Our techniques are based on filtering formalizations using a novel conformance score metric and validity scores obtained from testing the code against formalized assertions. We evaluate our approach on 541 assertion-generation tasks derived from 22 collection-like Java classes, and show that our technique produces the ground truth more reliably (improving upto 20 points in precision on average) than when using LLMs naively to translate assertions.
Constraint-Driven Model Optimization: An Industry Framework for Selecting Compression and Acceleration Techniques in Modern Machine Learning Systems
arXiv:2607.13735v2 Announce Type: replace Abstract: The rapid deployment of machine learning systems across cloud, edge, and enterprise environments has brought model optimization to the forefront of systems-engineering. Despite a rich literature spanning quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference-time optimization, practitioners are often left navigating these techniques through heuristics rather than principled methodology. We argue that optimization should be formulated as a constraint-driven, multi-objective engineering decision and introduce a unified framework that characterizes any production deployment along five interacting constraint dimensions: data availability, latency budget, memory budget, accuracy tolerance, and retraining budget. Building on this taxonomy, we synthesize empirical gains reported across the research literature and map them to operational constraints rather than algorithmic categories. To ensure practical relevance, we selected these techniques by reviewing recent literature for methods that report measurable improvements against critical deployment bottlenecks. We propose a prescriptive decision framework and provide optimization pipelines for four representative industrial scenarios to illustrate it in practice. To the best of our knowledge, this work provides one of the first structured attempts to formalize model optimization as a constraint-aware, multi-objective engineering process, synthesizing quantitative evidence from the research literature.
Stop Means Stop: Measuring and Repairing the Enforcement Gap in Agent-Framework Control Primitives
arXiv:2607.14166v2 Announce Type: replace Abstract: Production LLM-agent frameworks ship control primitives -- human-in-the-loop approval gates, run cancellation, and execution timeouts -- whose names and documentation imply barrier semantics: while a run is paused, cancelled, or timed out, no gated side effect executes. This contract holds on none of six widely used open-source frameworks. Model-free differential probes isolate a recurring sibling leak -- an approval gate suspends its own branch while a sibling's effect executes during the pause, defeating rejection -- in every framework shipping a pre-execution gate (five of six, four execution models, two language runtimes), and confirm replay double-execution, cancellation orphans, and timeout zombies. The hazard is reachable: frontier models emit the leak-triggering plan shape at rates up to 14%, and live models driving unmodified frameworks leak 215 of 1,200 runs (P(leak | emitted)=1.00); on naturalistic tau-bench episodes models serialize writes -- the everyday gap is latent -- while injection induces it deterministically and a 13-incident public corpus corroborates the replay and cancellation failures. We repair the gaps with SOUNDGATE, an environment-external Rust gate through which every side effect must be admitted, enforcing hold-until-decided, reject-cancels, dedup-on-replay, and fence-on-cancel under a stated complete-mediation contract, discharged for network egress by two kernel-enforced routes. The admission core is mechanically verified (Verus; TLA+/TLC to 7.5e7 states; TLAPS; Loom on the deployed Rust) and bridged to code by differential conformance over 1.2e7 operations with zero divergences. Under that contract SOUNDGATE blocks every measured violation on all six frameworks while releasing legitimate effects: gated tau-bench episodes complete with zero refusals at ~1 ms per write, and durable admission sustains ~12k admissions per second.
Acoustic Imaging for UAV Detection: Dense Beamformed Energy Maps and U-Net SELD
arXiv:2508.00307v4 Announce Type: replace-cross Abstract: We introduce a U-net model for 360{\deg} acoustic source localization formulated as a spherical semantic segmentation task. Rather than regressing discrete direction-of-arrival (DoA) angles, our model segments beamformed audio maps (azimuth & elevation) into regions of active sound presence. Using delay-and-sum (DAS) beamforming on a custom 24-microphone array, we generate signals aligned with drone GPS telemetry to create binary supervision masks. A modified U-Net, trained on frequency-domain representations of these maps, learns to identify spatially distributed source regions while addressing class imbalance via the Tversky loss. Because the network operates on beamformed energy maps, the approach is inherently array-independent and can adapt to different microphone configurations and can be transferred to different microphone configurations with minimal adaptation. The segmentation outputs are post-processed by computing centroids over activated regions, enabling robust DoA estimates. Our dataset includes real-world open-field recordings of a DJI Air 3 drone, synchronized with 360{\deg} video and flight logs across multiple dates and locations. Experimental results show that U-net generalizes across environments, providing improved angular precision, offering a new paradigm for dense spatial audio understanding beyond traditional Sound Source Localization (SSL). We additionally validate the same beamforming-plus-segmentation formulation on the DCASE 2019 TAU Spatial Sound Events benchmark, showing that the approach generalizes beyond drone acoustics to multiclass Sound Event Localization and Detection (SELD) scenarios.
Faster than the Team, Faster than the Customer: Tool Integration, Collaboration, and Organisational Lag in AI-assisted RE
arXiv:2606.01772v2 Announce Type: replace Abstract: The impact of applying generative AI tools to requirements engineering (RE) in industrial practice remains poorly understood. This paper examines how AI-assisted RE tools are used in industrial practice at XITASO, a medium-sized enterprise for high-tech software engineering, and how they reshape workflows, tool integration, and PO--developer relationships. We combine a 2024 company-wide use-case survey with two rounds of semi-structured interviews with eight product owners (POs) in late 2025 and spring 2026, covering an in-house chatbot and seven commercial AI tools. We identify 15 distinct use cases across four categories: product backlog management, tender management, requirements and domain understanding, and document and artifact creation. Three findings emerge. First, the effect of AI on PO--developer interaction is mixed: the prevailing single-user interaction model can substitute for collaborative dialogue, and developers do not always welcome AI-generated artefacts. Second, tool integration -- not tool capability -- is the binding constraint: where integration is in place, time savings are dramatic; where it is missing, POs fall back on manual workarounds. Third, AI advances faster than the surrounding organisational systems, so its benefits accrue to individual POs while team processes and customer readiness remain the bottleneck. The empirical GenAI-RE literature remains dominated by early-stage, lab-oriented evaluations of isolated tasks while practice has moved into territory it has not yet studied: practitioners are already assembling cross-tool integrations, navigating customer governance, and renegotiating role boundaries. From these patterns we derive a set of questions practitioners considering AI-assisted RE may ask of their own situation.
CANN-EUCLID: unsupervised constitutive artificial neural network model discovery from full-field data
arXiv:2606.14565v2 Announce Type: replace Abstract: Constitutive artificial neural networks (CANNs) provide interpretable material model discovery, but have so far been used in stress-supervised settings based on apparent stress-strain data from homogeneous tests. Because each test samples only a narrow loading path and provides homogenized rather than local stress information, robust discovery typically requires multiple loading modes to constrain the multidimensional response. This is challenging for soft biological tissues, where repeated testing, damage, and sample variability limit reliable information from a single specimen. Here, we combine CANNs with the stress-unsupervised full-field discovery framework EUCLID to identify sparse hyperelastic laws directly from displacement fields and reaction forces in one heterogeneity-inducing loading case. CANN-EUCLID minimizes equilibrium imbalance with sparsity-promoting regularization selecting compact active terms, without local stress measurements or a prescribed law. We evaluate the approach on isotropic and anisotropic benchmarks with prescribed ground-truth laws. When the ground truth is representable by the chosen CANN basis, our method recovers the correct terms with near-exact accuracy, including exponential terms with embedded parameters. When it is not contained in the basis, the method retains shared terms and approximates missing contributions using available basis functions. Generalization depends strongly on sampled deformation states: exponential strain-stiffening terms can be recovered accurately when sufficiently probed, but can produce large extrapolation errors when the stiffening regime lies outside the sampled domain. Forward FE validation simulations show that the discovered behavior accurately replicates the ground truth. These results establish stress-unsupervised CANN discovery as a promising framework for interpretable full-field constitutive model identification.
Don't Predict, Prioritize: Rethinking GPU Reliability Assessment
arXiv:2607.15115v1 Announce Type: new Abstract: The reliability of Graphics Processing Units (GPUs) is a criticalbottleneck for modern large-scale AI infrastructure, where a sin-gle node failure can disrupt synchronous training jobs and causesignificant financial losses. While predictive maintenance is widelyused in other hardware domains, we demonstrate that accuratelypredicting the exact timing of GPU failures is inherently difficult.Through an in-depth analysis of telemetry data from a productioncluster, we find that major GPU failures, including Double Bit Er-rors (DBEs) and GPU Lost events, exhibit strong stochasticity andlow signal-to-noise ratios in time-series telemetry, which makesconventional time-based prediction ineffective. This insight motivates a paradigm shift: instead of attempting topredict the absolute timing of a failure, we propose a more robustapproach focused on ranking nodes by their relative failure risk. Wepropose HeaRank (Health Rank), a Learning-to-Rank (LTR) frame-work that leverages stable historical failure patterns to computea global risk ranking of GPU nodes. Evaluated on a production-scale cluster with thousands of GPUs, HeaRank achieves an AUCof 0.83, significantly outperforming both heuristic baselines andstate-of-the-art ranking algorithms. In online deployment, HeaRanksuccessfully captures 64% of future failures within the top 5% ofranked nodes, compared to only 21% by the incumbent productionsystem. These results suggest that relative risk ranking can serveas a robust alternative in environments where absolute failure pre-diction is inherently limited. Our work highlights the importanceof risk-aware scheduling and proactive resource management inmodern GPU clusters.
Auditing Inference-Time Defense Evaluation for Multimodal Large Language Models
arXiv:2606.10904v3 Announce Type: replace Abstract: Comparisons of inference-time defenses for multimodal large language models (MLLMs) depend on more than defense code: the image payload, model text, proxy labels, and judge protocol must refer to the same evaluated event. We audit an experimental archive covering two fixed prompt wrappers and a Gaussian image-perturbation adapter, recorded under aliases derived from RapGuard, AdaShield, and SmoothVLM, across eight InternVL and Qwen-VL models. The planned suite contained seven safety benchmarks and 9,000 inputs. Three benchmark branches fail provenance checks: the MM-SafetyBench renderer used the wrong payload field, the JailBreakV loader admitted text-only fallback, and the adversarial-patch branch did not generate the stated attack. We restrict comparative safety results to FigStep and three text-only corpora, totalling 4,820 inputs per configuration. These are descriptive outputs of a legacy keyword protocol, not validated harmlessness rates; the protocol counted empty strings as safe, and raw responses are unavailable to recover the effect. We also audit 274 adversarial responses; after excluding 28 garbage outputs, 246 comparable responses contain 13 keyword false negatives. This search is nonrandom: it detects evaluator failures but cannot estimate their prevalence or family-level rates, and cannot validate per-cell rankings. The benign audit covers 38,500 stored outputs and 11,637 unique judged responses. Pooled refusal estimate is 0.52%, largest cell 3.24%, sampling-only intervals reach 10.92%; two cells remain not estimable. The archive does not support the earlier mass-refusal interpretation. Measured cost instead appears in batch processing time and defensive-preamble contamination. This work contributes a traceable comparative audit and provenance requirements for future MLLM defense evaluation; it does not claim one defense or an adaptive router is universally superior.
A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi
arXiv:2607.14568v1 Announce Type: cross Abstract: A companion study ran a 35B mixture-of-experts model on a 2011 NVIDIA Tesla C2075 (Fermi, sm_20, 6GB) as a GPU-prefill/CPU-decode hybrid, because the 4-bit model did not fit in device memory (arXiv:2606.24031). This report keeps the hardware and asks what a model that fits can do: we deploy MiniCPM-V-4.6, a modern multimodal assistant pairing a SigLIP2 vision encoder and window-attention merger (16x visual token compression) with a compact hybrid gated-delta-net backbone, entirely on the GPU. Three results. (i) An all-GPU engine built on measured foundations: projections that dequantize 8-bit weights once and call the vendor SGEMM still in the last Fermi toolchain (64% of FP32 peak; our best hand-written GEMM hit 37%, wrongly called the ceiling); a chunked delta-rule rewrite of the recurrent layers, 2.8x faster than the sequential scan once attribution exposed one bad kernel; and a measured negative: 4-bit weights make decode slower than 8-bit here, since Fermi issues nibble-unpacking shifts at half rate. (ii) The vision side is a port with a proof obligation: we translate tower, merger, and projector to sm_20 CUDA, validating every stage against a locally generated reference forward (full tower 1.4e-5). One failure, position-embedding bucketization differing on exact rational ties, generalizes to a rule: float tie-breaking in index arithmetic is implementation-defined; call the reference operator, do not reimplement it. (iii) Long context exposes an O(N^2) wall short benchmarks hide: prefill falls from 114 tok/s at 2k tokens to 21 at 10k in a naive attention kernel; per-head vendor-GEMM calls writing into the existing score buffer (zero extra memory) restore a flat profile (408 at 2k, 361 at 10k; 17x), verified by exact needle retrieval from 60% depth. The same rewrite cuts image encoding 6x, to 0.93s. The system answers an image question end-to-end in 1.7s.
Modular molecular toolkit for photochemical energy conversion in a self-assembling nanocontainer
arXiv:2606.27238v2 Announce Type: replace Abstract: Production of useful chemicals using photoelectrochemical biohybrid devices offers an environmentally friendly alternative to existing energetically demanding processes. These devices exploit light-driven charge separation, e.g. by a photosystem, and require efficient electron transfer to a tailored redox enzyme cascade. Here we demonstrate that electron transfer efficiency can be increased by confining the photosystem with the redox protein inside a self-assembling, virus-based nanocontainer. The photosynthetic system from the phototrophic bacterium Cereibacter sphaeroides and cytochrome c were conjugated to a bacteriophage P22 scaffolding protein and co-incorporated into the 50 nm diameter virus shell in vitro. The porous shell confined the macromolecular components for efficient electron transfer while allowing free exchange of small electron mediators. Sustainable and accelerated light-driven electron transfer between the encapsulated components was confirmed by optical spectroscopy. This self-assembly system presents a versatile platform for developing nanoreactors that combine photosystems with complex redox pathways.
Closed-Loop Knowledge Dynamics: An Operational Framework for Saturation and Escape
arXiv:2607.14185v1 Announce Type: new Abstract: Feedback-driven loops support iterative improvement in large language models, reinforcement learning, and autonomous discovery, yet their gains often diminish under repeated internal feedback. We study why closed-loop knowledge systems saturate and what external information can move them beyond their current attractors. We introduce a three-level operational framework in which knowledge states $x_t$ evolve through transition kernels $K_{\theta}$ indexed by a structural parameter $\theta$. The governing structure is defined as the observational equivalence class of $\theta$ induced by these kernels, while attractors and basins are properties of the fixed-$\theta$ dynamics. A structural intervention changes $\theta$ and produces a detectable kernel discrepancy on pre-specified probe states, making structural change falsifiable. Using a Lyapunov drift condition, we show that stable internal dynamics approach bounded stability regions with exponentially attenuated transients and a noise-controlled residual floor. We characterize escape through a metric condition on intervention-induced attractor displacement and a baseline-relative KL lower bound for increasing escape probability. This analysis also explains why conditional mutual information alone cannot certify escape: it measures variation among intervention-conditioned updates rather than departure from the no-intervention law. Case studies in LLM code repair, sparse-reward reinforcement learning, and Bayesian optimization use matched continuation controls to illustrate how feedback strength and alignment affect quality-improving escape. Our contribution is an operational connection among stability tools, measurable intervention effects, and cross-domain diagnostics.
The Planar Case of Thomas Positive Circuits Conjecture
arXiv:2607.14143v1 Announce Type: cross Abstract: The notion of circuit refers to a cyclic oriented influence between the elements of a dynamical system. There are two classes of circuit: positive and negative. R. Thomas conjectured that a necessary condition of multi stationarity is the existence of positive circuits. In this paper we use dynamical system tools and planar analysis to find conditions for which the conjecture holds for planar systems.