arXiv:2605.13553v2 Announce Type: replace Abstract: Description Logics (DLs) are a family of formal languages used for representing and reasoning about structured knowledge in terms of concepts and their relationships. The expressive power of a DL depends on the constructors available for building complex concepts. In this work, we investigate subsumption in the restricted description logic $\mathcal{FL}_{\bot\mathit{reg}}$ and the related fragments $\mathcal{FL}_{\mathit{reg}}$, $\mathcal{FL}_\bot$, and $\mathcal{FL}_0$. These formalisms support value restrictions over role names, where the subscript $\mathit{reg}$ indicates the use of regular expressions over roles. Subsumption between two concept descriptions in $\mathcal{FL}_{\bot\mathit{reg}}$ and $\mathcal{FL}_{\mathit{reg}}$ is PSpace-complete. When subsumption is considered with respect to a TBox (i.e., a set of axioms), the complexity increases to ExpTime-complete. These results can be derived either from complexity bounds established for more expressive logics or from algorithms designed for harder reasoning problems. We reprove the PSpace-completeness result and provide a new proof of ExpTime-completeness for $\mathcal{FL}_{\mathit{reg}}$ and $\mathcal{FL}_{\bot\mathit{reg}}$ with TBoxes via a novel reduction to parity pushdown games. Our algorithm relies only on the constructs available in these logics and may therefore be implemented more easily.
Science Journals
arXiv:2606.24765v1 Announce Type: new Abstract: Two-electron processes can generate high harmonics beyond the conventional single-active-electron cutoff. Motivated by recent experimental evidence of an extended secondary plateau in the helium high-harmonic spectrum [S. Wang et al, Optica, (2023); S. Wang et al, In Print in Nature Photon., (2026)], we present a two-electron generalisation of the strong-field approximation. We analyse the resulting expressions using the saddle-point method and determine the extended cutoff. We find good agreement with classical predictions of cutoff scalings of $4.7$ and $5.5$ times the ponderomotive energy, which significantly exceed the established single-electron scaling of 3.17. We calculate high-harmonic spectra generated via a two-electron process in helium atoms driven by an intense few-cycle infrared laser pulse. Our results demonstrate that the harmonic spectrum extends far beyond the water window, reaching photon energies up to $\approx 1.2\,\mathrm{keV}$ in the soft x-ray region. The large spectral bandwidth can support the generation of sub-attosecond soft x-ray pulses, which are of particular interest for probing ultrafast dynamics across matter, including applications in core-level spectroscopy and biological imaging.
arXiv:2606.24178v1 Announce Type: new Abstract: Pretrained vision models often misclassify inputs that are rotated, scaled, or sheared, even though these affine transformations leave the object class unchanged. Robustness is usually restored either by building equivariance into the architecture or by retraining with augmentation, both of which require changing or retraining the model. Test-time canonicalization instead leaves the classifier untouched. It undoes the transformation of each input, mapping it to a canonical form near the training distribution before classification. Existing canonicalizers, however, rely on a narrow set of logit-based energy scores and bespoke search procedures, leaving the design space of scoring functions and optimizers unexplored. We reframe canonicalization as out-of-distribution (OOD) detection, which lets any OOD score serve as the energy minimized over transformations. Across benchmarks ranging from handwritten characters and sketches to natural images and 3D point clouds, we systematically evaluate around twenty OOD scores and nine search algorithms, finding that distance-based scores paired with random search and local refinement perform best overall. Because canonicalizing an already-aligned input can hurt accuracy, we add a gated mechanism that transforms an input only when its OOD score indicates this is needed, preserving most in-distribution accuracy while retaining the robustness gains on transformed inputs. Code is available at github.com/johschm/its.
arXiv:2606.24849v1 Announce Type: new Abstract: Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still struggle with structure-aware prompt following, where object counts, spatial relations, attribute bindings, and coarse layouts must be preserved. We attribute this limitation in part to the entanglement of structural planning and appearance rendering within a single conditioning stream. To address this issue, we propose Implicit Visual Chain-of-Thought (IV-CoT), a latent visual reasoning framework for query-conditioned image generation. IV-CoT decomposes the visual conditioning queries into a structural-to-semantic cascade, where structural queries first form a latent visual plan and semantic queries then render appearance conditioned on this plan. To guide the structural queries, we introduce training-only sketch supervision, which encourages them to capture structure from sketches without requiring sketch extraction or intermediate decoding at inference time. IV-CoT performs implicit CoT reasoning in a single forward pass and achieves superior results on GenEval and T2I-CompBench. Visualizations and analyses demonstrate that the learned structural and semantic queries play complementary roles in structure-aware generation.
arXiv:2606.03867v2 Announce Type: replace Abstract: Multi-Document Summarization (MDS) plays a critical role in distilling essential information from collections of textual data. Existing approaches often struggle to capture complex inter-document relationships, rely heavily on large amounts of labeled data for supervised training, or exhibit limited generalization across domains and languages. To address these limitations, we present a training-free mixture-of-agents framework for MDS that leverages the complementary strengths of large language models (LLMs) and knowledge graphs. Our approach decomposes summarization into specialized agent tasks: extractive selection, knowledge-aware abstraction, and iterative refinement, each operating without task-specific fine-tuning. We unify their outputs using a multi-perspective consistency mechanism guided by LLMs. Experiments across four datasets in English and Vietnamese demonstrate state-of-the-art or competitive performance, validating the effectiveness and adaptability of our modular design.
arXiv:2606.24252v1 Announce Type: new Abstract: Atomically sharp edges are essential for future high-index nanophotonic structures, yet conventional lithography and dry etching methods inevitably introduce edge roughness that limits optical confinement and reproducibility. Recently, anisotropic wet etching of multilayer van der Waals crystals, such as transition metal dichalcogenides (TMDs), has enabled crystallographically defined, atomically sharp zigzag edges, eliminating the edge-roughness problem. However, the process is intrinsically limited to confined geometries such as isolated triangular or hexagonal features dictated by crystal stacking symmetry. Here, we demonstrate a lithography-guided anisotropic etching framework that drives TMDs etching beyond isolated confined geometries by enforcing controlled interaction of neighboring etched nanoholes regions. In multilayer 2H-WS2, merging of anisotropic etch fronts enables sustained long-range propagation of zigzag facets, introducing a previously inaccessible 180-degree edge alignment and a crystallographically defined design space combining 120-degree and 180-degree junctions. Using this approach, we fabricate extended nanophotonic structures with ultrasharp sidewalls, including sub-100-nm-gap one-dimensional gratings, waveguides, defect-engineered photonic cavities, angle programmed photonic lattices, and diffractive zone plates. Back-focal-plane reflection spectroscopy of atomically sharp 1D periodic 2H-WS2 gratings demonstrates their photonic functionality, revealing symmetry-protected bound states in the continuum (SP-BICs) and strong exciton-photon coupling in multilayer WS2. Finally, we fabricate ultrathin, ultranarrow, and ultralong nanoribbons with record-high aspect ratios. Together, these results demonstrate edge merging as a generic route to fabricate edge-defined, atomically sharp nanophotonic and nanoelectronic architectures in layered van der Waals platforms.
arXiv:2606.24279v1 Announce Type: new Abstract: In Description Logics (DLs), reasoning under Rational Closure (RC) is a well-known and widely accepted non-monotonic formalism to handle defeasible knowledge. In this paper, we study the application of RC to the core and horn variants of the DL-Lite family of lightweight description logics. We analyze both entitlement (instance checking) and Conjunctive Query (CQ) answering under RC. Our main contribution is providing a plug-in architecture that builds upon existing standard classical reasoners, establishing that reasoning and CQ answering under RC for DL-Lite can be done efficiently with minimal computational overhead.
arXiv:2602.04588v2 Announce Type: replace-cross Abstract: Coordination in distributed systems is often hampered by communication latency, which degrades performance. Quantum entanglement offers fundamentally stronger correlations than classically achievable without communication. Crucially, these correlations manifest instantaneously upon measurement, irrespective of the physical distance separating the systems. We investigate the application of shared entanglement to a dual-work optimization problem in a distributed system comprising two servers. The system must process both a continuously available, preemptible baseline task and incoming customer requests arriving in pairs. System performance is characterized by the trade-off between baseline task throughput and customer waiting time. We present a rigorous analytical model demonstrating that when the baseline task throughput function is strictly convex, rewarding longer uninterrupted processing periods, entanglement-assisted routing strategies achieve Pareto-superior performance compared to optimal communication-free classical strategies. We prove this advantage through queueing-theoretic analysis, non-local game formulation, and computational certification of classical bounds. Our results identify distributed scheduling and coordination as a novel application domain for near-term entanglement-based quantum networks.
arXiv:2606.14936v2 Announce Type: replace Abstract: Spectral encoding enables single-shot measurements of ultrafast transients by mapping temporal information onto the spectrum of a chirped probe. This encoding allows dynamics to be recorded that are beyond the response limits of conventional electronic detectors. However, because the measurements record only spectral intensity, the phase of encoded signals is lost, and dispersion in the detection process introduces waveform distortions that complicate reconstruction and quantitative interpretation of spectra. In single-shot terahertz time-domain spectroscopy (THz-TDS), these distortions manifest as a tradeoff between temporal resolution and the measurement window of signals and can produce spectral null frequencies that limit the recoverable THz bandwidth. To address this challenge, a Bayesian inversion framework is developed to recover the underlying waveform from the squared spectral observable by inferring the THz field, the modulation coefficient, and a low-dimensional empirical parameterization of the probe spectrum jointly, while a Gaussian process prior regularizes the waveform. The framework is validated using single-shot THz-TDS experiments spanning two probe spectral profiles and three chirp conditions with $\alpha$ ranging from 14.5 to 40 ps$^{-2}$. Across all cases, the inversion reconstructs both the time-domain waveform and spectral null frequency structure within the credible interval of a delay-line reference measurement. These results establish a pathway to eliminate penalties that are associated with the detection process in spectral encoding methods without adding additional optics or alignment complexity.
arXiv:2606.15043v3 Announce Type: replace Abstract: Weakest pre-expectations are the probabilistic program analogue to weakest preconditions in classical programs. Deductive verification approaches aim to establish bounds on these quantitative expectations. Their automation has been successful in analysing a variety of discrete probabilistic programs. Key routines in that automation require reasoning about (partially unrolled) loops, however, the logical representation of weakest pre-expectations on such unrollings often explodes. In this paper, we develop typed extended decision diagrams (TEDDs), inspired by various extensions to binary decision diagrams. We demonstrate computing WPs represented as TEDDs, SMT-based pruning to further shrink their representation, and we lift some proof rules to operate on TEDDs. Finally, we demonstrate that TEDDs boost the scalability of deductive probabilistic program verification by orders of magnitudes over the state of the art.
arXiv:2606.24331v1 Announce Type: new Abstract: Transformer-based language models have become the default substrate for natural language processing and the pace of new releases has made it hard for practitioners to separate durable ideas from the noise of incremental announcements. This review works at two levels. At the level of mechanism, we organise the main transformer families into a working taxonomy, covering encoder-only, decoder-only, encoder-decoder, long-context, permutation-based, and generator-discriminator variants. We then extend the discussion to post-2023 developments that changed the picture in practice: instruction tuning, reinforcement learning from human feedback, direct preference optimisation, mixture-of-experts scaling, retrieval augmentation and the current flagship model families from OpenAI, Anthropic, Google, Meta, Mistral and DeepSeek. At the level of use, we survey deployments across healthcare, finance, legal, education, customer service, creative writing and scientific work. Based on this we link each to the specific capabilities that make a transformer the appropriate tool. The contribution of this paper is a critical assessment that is based on the survey. We compare architectures on four axes that matter to deployment decisions, we quantify the trade-off between parameter count and energy cost. We also discuss how alignment methods, data provenance and benchmark saturation change what it means to call a model "state of the art". The final section lists the research questions that we think deserve more attention.
arXiv:2605.19248v2 Announce Type: replace-cross Abstract: We consider $(n,k)$ MDS-coded distributed storage over $\mathbb{F}_q$ with per-node storage $\alpha$ symbols. For the oblivious update problem, where a single message symbol changes and neither helpers nor the stale node know which, the classical lower bound is $\alpha k \log_2 q$ bits. We prove that when the $k$ contacted helpers share prior quantum entanglement, the update bandwidth is $\lceil \alpha/2 \rceil \cdot k \log_2 q$ bits-equivalent, a factor approaching 2 reduction. For $\alpha = 2$, a $[[k, k-2]]_q$ CSS code achieves bandwidth $k \log_2 q$ with one qudit per helper. For general $\alpha$, a $[[\lceil \alpha/2 \rceil k, \lceil \alpha/2 \rceil k - \alpha]]_q$ CSS code achieves the bound with $\lceil \alpha/2 \rceil$ qudits per helper. The matching converse uses the superdense coding bound: the stale node holds all transmitted qudits and hence the entangled partners, so each helper's channel supports at most $D^2$ distinguishable signals for dimension $D$. The result holds for all $(n,k)$ pairs with sufficiently large prime $q$.
arXiv:2606.19192v3 Announce Type: replace-cross Abstract: High-field conditioning is the process by which radio-frequency structures in particle accelerators and other high-gradient devices reach their operating fields, yet the underlying physical mechanism remains an open question. Models and indirect measurements point to subsurface dislocation dynamics, but large-area structural measurements have been missing. We present electron backscatter diffraction measurements spanning millimeter-scale regions on a copper cathode conditioned at pulsed direct-current fields up to $\sim$80~MV/m in a sloped-anode geometry, which imposes a known gradient of field exposure across a single electrode. Across nine regions of interest spanning this exposure range, the mean intragrain misorientation of field-exposed regions exceeds that of unexposed references by $\sim$75\%; the difference is reproduced by three independent misorientation metrics and confirmed by Kolmogorov--Smirnov tests. To our knowledge, this is the first large-area observation of structural differences between conditioned and unconditioned regions of a high-field electrode. The misorientation separates into three tiers (high-field center and edge, low-field periphery, and unexposed reference) that match the spatial profile of the conditioning-state variable $E_S$ predicted by Monte Carlo simulations. These observations point to the evolving subsurface dislocation population as a candidate physical basis of conditioning.
arXiv:2606.21545v2 Announce Type: replace Abstract: Altermagnets host momentum-dependent spin splitting without net magnetization, a symmetry-enforced band phenomenon whose photonic analogues have so far been realized only in square lattices governed by fourfold rotation. Here we introduce a photonic altermagnet on a hexagonal lattice whose helicity splitting is governed by mirror rather than rotational symmetry. Elliptical chiral elements of alternating handedness, placed at the vertices of a regular hexagon, leave the two opposite-chirality sublattices connected only by chirality reversal combined with a mirror reflection. Full-wave simulations reveal mirror-related splitting of the two opposite-helicity branches in the band structure and isofrequency contours, with the channels exchanged when the ellipse orientation is reversed. Using a finite photonic crystal slab, we show that such splitting separates a linearly polarized beam into handedness-resolved channels, thus enabling beam splitting and direction-selective helicity filtering with target-helicity output fractions above 0.85 and output paths continuously tunable through the ellipse rotation angle. These results extend photonic altermagnetism to a previously unexplored lattice-symmetry class and establish mirror-symmetric chiral textures as building blocks for altermagnetism-inspired on-chip chiral photonics.
arXiv:2606.22314v2 Announce Type: replace Abstract: Path-based attribution methods such as Integrated Gradients (IG) are widely adopted for their strong axiomatic properties and effectiveness in attributing model predictions to input features by integrating gradients along a path from a baseline to the input. However, the choice of the attribution path largely affects the quality of explanations, and existing approaches rely on fixed or hand-crafted paths that often produce noisy or distorted attributions. To address this limitation, we propose Diffusion Integrated Gradients (DiffIG), a novel method that reformulates path generation as a conditional generative modeling problem. DiffIG first trains a diffusion model to learn a distribution over paths generated from a Stick-Breaking Process, then employs guided sampling to embed user guidance during the sampling procedure. We demonstrate that DiffIG quantitatively matches or outperforms existing path-based methods, achieving perceptually aligned explanations. This work introduces a new generative perspective for flexible, inference-time controllable Explainable Artificial Intelligence (XAI) methods.
arXiv:2606.22366v2 Announce Type: replace Abstract: Advances in the semiconductor industry are driven by the development of increasingly compact devices featuring intricate etched geometries, the characterization of which essentially requires ultraprecise, label-free, and real-time metrology. However, non-destructive and alignment-free optical metrology of sub-wavelength structures with nanometric resolution remains a major challenge. Here, we demonstrate a novel single-shot, label-free, and alignment-free optical metrology approach for determining the 1D position of sub-wavelength nanostructures, achieving lambda/110 (7.2 nm) precision. The high precision benefits from utilizing structured illuminations of Laguerre-Gaussian (LG) or Hermite-Gaussian (HG) beams, and the AI analyzing method can retrieve the information when such structured light interacts with sub-wavelength objects. Instead of relying on phase singularities in superoscillatory microscopy, our approach leverages spatially distributed phase jumps in HG and LG beams interacting with the nanostructures, providing an alignment-robust solution to the challenges in optical metrology. Such an alignment-free, non-destructive, and high-precision metrology technique enables real-time machine vision, semiconductor inspection, and advanced manufacturing.
arXiv:2606.22406v2 Announce Type: replace Abstract: Attention mechanisms have demonstrated remarkable empirical success in identifying relevant information from large collections of tokens, yet the theoretical principles underlying this behavior remain poorly understood. We study a stylized softmax-attention model in which a query vector is learned by stochastic gradient ascent from a collection of informative and nuisance tokens. Exploiting the symmetry of the model, we derive a population objective and characterize the limiting ordinary differential equation governing the learning dynamics. Using tools from stochastic approximation and dynamical systems theory, we establish a rigorous connection between the stochastic learning algorithm and its deterministic limit. Our main result shows that, under suitable high-dimensional scaling assumptions and standard step-size conditions, the learned query converges almost surely to the one-dimensional signal subspace spanned by the latent informative direction. Equivalently, the query asymptotically recovers the latent signal up to the intrinsic sign ambiguity. These results provide a rigorous theoretical foundation for understanding attention mechanisms as signal extraction procedures in high-dimensional noisy environments and offer a dynamical-systems perspective on how attention discovers relevant information in the presence of substantial noise.
arXiv:2307.13127v3 Announce Type: replace-cross Abstract: Data used to train predictive models via empirical risk minimization (ERM) often contain sensitive personal information. While differential privacy (DP) provides mathematically provable bounds to protect such data, previous work has focused almost exclusively on unweighted ERM. We consider weighted ERM (wERM) -- an important generalization where individual contributions to the objective function vary. We propose the first DP algorithm for general wERM with formal privacy guarantees and derive both its empirical and population excess risk bounds. Crucially, this general wERM framework provides a pathway for deriving privacy-preserving learning methods for individualized treatment rules, including the popular outcome-weighted learning (OWL) approach. We evaluate DP-wERM applied to OWL in simulated and real data experiments. Our empirical results demonstrate that training OWL models via wERM provides strong DP guarantees while maintaining robust performance, proving the method is practical for sensitive, real-world data.
arXiv:2311.16707v2 Announce Type: replace-cross Abstract: Dense prediction is a fundamental requirement for many medical vision tasks such as medical image restoration, registration, and segmentation. The most popular vision model, Convolutional Neural Networks (CNNs), has reached bottlenecks due to the intrinsic locality of convolution operations. Recently, transformers have been widely adopted for dense prediction for their capability to capture long-range visual dependence. However, due to the high computational complexity and large memory consumption of self-attention operations, transformers are usually used at downsampled feature resolutions. Such usage cannot effectively leverage the tissue-level textural information available only at the full image resolution. This textural information is crucial for medical dense prediction as it can differentiate the subtle human anatomy in medical images. In this study, we hypothesize that Multi-layer Perceptrons (MLPs) are superior alternatives to transformers in medical dense prediction where tissue-level details dominate the performance, as MLPs enable long-range dependence at the full image resolution. To validate our hypothesis, we develop a full-resolution hierarchical MLP framework that uses MLPs beginning from the full image resolution. We evaluate this framework with various MLP blocks on a wide range of medical dense prediction tasks including restoration, registration, and segmentation. Extensive experiments on six public well-benchmarked datasets show that, by simply using MLPs at full resolution, our framework outperforms its CNN and transformer counterparts and achieves state-of-the-art performance on various medical dense prediction tasks.
arXiv:2407.20353v2 Announce Type: replace-cross Abstract: Many structures in mathematical physics and dynamics exhibit intricate fractal geometry. Such behavior appears prominently in quantum mechanics and materials science through spectra of aperiodic and quasicrystalline operators, where questions of ``size'' (Lebesgue measure, fractal dimension, spectral gaps, etc.) are central. Yet the lack of rigorous computational tools for analyzing these quantities limits both theory and application. Na\"ive truncation often fails, and there is no overarching framework to explain what can, and cannot, be computed. We develop a unified program for the rigorous computation of spectral size for bounded self-adjoint operators, based on local spectral exclusions and adaptive covers. This constructive framework yields algorithmically optimal methods (under natural computational assumptions) that bridge spectral theory with computation to address problems previously deemed intractable. Their complexity is classified within the Solvability Complexity Index (SCI) hierarchy, extending Smale's program on the limits of computation. Sharp computational lower bounds are established through impossibility results for limit-periodic Schr\"odinger operators constructed from adversarial potentials. The methods enable state-of-the-art rigorous computations for one- and two-dimensional aperiodic systems, and pinpoint problems where numerics can feed directly into computer-assisted proofs. Beyond spectral analysis, they apply broadly to computing measures of size for general closed sets, opening new directions in the computational study of complex geometric structures.
CrossFusion: A Multi-Scale Cross-Attention Convolutional Fusion Model for Cancer Survival Prediction
arXiv:2503.02064v2 Announce Type: replace-cross Abstract: Cancer survival prediction from whole slide images (WSIs) is a challenging task in computational pathology due to the large size, irregular shape, and high granularity of the WSIs. These characteristics make it difficult to capture the full spectrum of patterns, from subtle cellular abnormalities to complex tissue interactions, which are crucial for accurate prognosis. To address this, we propose CrossFusion, a novel multi-scale feature integration framework that extracts and fuses information from patches across different magnification levels. By effectively modeling both scale-specific patterns and their interactions, CrossFusion generates a rich feature set that enhances survival prediction accuracy. We validate our approach across six cancer types from public datasets, demonstrating significant improvements over existing state-of-the-art methods. Moreover, when coupled with domain-specific feature extraction backbones, our method shows further gains in prognostic performance compared to general-purpose backbones. The source code is available at: https://github.com/RustinS/CrossFusion
arXiv:2506.13724v2 Announce Type: replace-cross Abstract: Implementing large-scale quantum algorithms with practical advantage will require fault-tolerance achieved through quantum error correction, but the associated overhead is prohibitive. This overhead can be reduced by engineering physical qubits with fewer errors, and by shaping the residual errors to be more easily correctable. In this work, we demonstrate quantum error correcting codes and logical qubit circuits in a metastable ytterbium-171 nuclear spin qubit with a noise bias towards erasure errors. These errors can be located separately from any syndrome information diagnosing the error, and we demonstrate adaptive circuit execution based on erasure information. We show that dephasing errors on the qubit during coherent transport can be strongly suppressed, and implement entangling gates that maintain a high fidelity in the presence of gate beam inhomogeneity or pointing errors. Furthermore, we demonstrate logical qubit encoding in the [[4, 2, 2]] code, with error correction during decoding based on mid-circuit erasure measurements despite the fact that the code is too small to correct any Pauli errors. Finally, we demonstrate logical qubit teleportation between multiple code blocks with conditionally selected ancillas based on mid-circuit erasure checks, a key part of leakage-robust error correction schemes using neutral atoms.
arXiv:2507.11768v3 Announce Type: replace-cross Abstract: Bayesian accounts of in-context learning face a direct objection: exact posterior predictives for exchangeable data are invariant to task-preserving order, yet transformers change next-token probabilities when the same examples are serialized differently. We show this objection targets a structural invariant rather than the quantity scoring online prediction. For any Bayesian reference, excess prequential code length is exactly cumulative predictive KL. For unordered support sets that must be serialized, the expected regret of a single admissible ordering decomposes into that of the order-averaged predictor plus an order-averaging gain. Exchangeability violations are therefore not binary refutations; they are priced by log loss. We instantiate the theory with KT/Dirichlet finite-alphabet prediction and coarsened Bayesian linear-regression (BLR) predictive distributions. On Qwen2.5-7B/14B, floored candidate distributions at support $256$ have one-step excess code lengths of $0.020/0.011$ bits for Bernoulli and $0.039/0.022$ bits for four-way categorical prediction, with candidate mass above $0.999$; coarsened BLR continuations increasingly match the posterior-predictive digit distribution as support grows. A frequentist plug-in baseline sharpens the reading: the predictive distributions sit closer to the Bayesian posterior predictive than to the maximum-likelihood plug-in, by a margin largest at small support, where the plug-in is degenerate, and vanishing as the references converge. Position interventions and a from-scratch ablation localize order sensitivity to the positional encoding, activation patching tests causal use of decoded sufficient statistics, and permutation mixtures quantify the downstream log-loss cost of arbitrary orderings. Transformers need not realize exchangeable posterior predictives for every serialization to be Bayes-competitive prequential predictors.
arXiv:2512.24777v2 Announce Type: replace-cross Abstract: This paper introduces a novel framework for analysing equilibrium in structured production systems that incorporate a static social division of labour, distinguishing between consumption goods traded in competitive markets and intermediate goods exchanged through bilateral relationships. We develop the concept of viability -- the requirement that all producers earn positive incomes, as a foundational equilibrium prerequisite. Our main theorem establishes that a structured production system is viable if and only if it is coherent, admitting no circular conversion processes that yield no net output, which in turn is equivalent to the non-singularity of its matrix representation. We further investigate completely viable systems, in which viable prices exist for all consumption good price vectors. We show that complete viability stands exactly between two input restrictions: it is guaranteed in coherent systems where no consumption good is used as an input in production, and it requires that no consumption good is used as an input in the production of another consumption good. The analysis thus reveals fundamental relationships between the architectural design of production systems and their economic sustainability. Our framework also contributes to the literature on the existence of positive output price systems and the Hawkins-Simon condition in input-output analysis.
arXiv:2602.10330v3 Announce Type: replace-cross Abstract: Context: The characterization of exoplanetary atmospheres has been transformed by the James Webb Space Telescope (JWST), whose infrared sensitivity enables transmission spectroscopy at unprecedented precision. However, stellar heterogeneities (e.g., spots and faculae) remain a dominant source of contamination that can bias atmospheric retrievals if not properly corrected. Aims: We present a methodology for reducing stellar contamination and instrument-specific noise from exoplanet transmission spectra using neural networks, in particular the so-called Denoising AutoEncoders (DAE). Our goals are to enable fast, accurate corrections that improve the reliability of atmospheric parameter retrievals and to promote the use of unsupervised algorithms for efficient data processing. Methods: We designed and trained DAE architectures using large synthetic datasets of terrestrial (TRAPPIST-1e analogues) and sub-Neptune (K2-18b analogues) planets. Atmospheric retrieval experiments were then performed on contaminated spectra in order to compare our deep-learning approach against standard correction methods in terms of accuracy and computational cost. Results: Our autoencoders successfully reconstruct uncontaminated spectra, preserving essential molecular features even in low-S/N regimes. In retrieval tests, the denoising autoencoder pre-processing reduces bias in retrieved abundance parameters compared to uncorrected observations. Notably, our method matches the accuracy of simultaneous stellar-contamination fitting while maintaining a much lower computational cost, typically one order of magnitude smaller. Conclusions: These results demonstrate that DAEs outperform conventional correction methods in computational efficiency while maintaining high accuracy, paving the way for their integration into future atmospheric characterization pipelines for both rocky and giant exoplanets.