arXiv:2607.08065v1 Announce Type: new
Abstract: LLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles (Verga et al., 2024) or "mixture-of-experts" (Shazeer et al., 2017) panels of judges. These systems share a key assumption: that consistency -- agreement among judges, or among a model's own samples -- indicates correctness. We show this assumption is unreliable. Agreement is not accuracy: a model can agree with itself, and different models can agree with each other, out of shared bias, a memorized heuristic, or an option-position prior rather than truth. We ask when agreement is nonetheless a usable proxy, in a large-scale cross-runner study: 53 runners drew K=50 samples for assigned overlapping cases across comparisons of model tier, prompting, and scale on GPQA Diamond and AIME -- 265,000 samples. Using majority-correctness as the deployment label and a hierarchical runner-clustered bootstrap, agreement is a positive but weak predictor (rho 0.20-0.59, all positive under item-clustered resampling) whose usefulness is regime-dependent: best for unsaturated mid-tier models and for allocating compute, and worst -- over-confident yet no more accurate -- for the most consistent frontier model (agreement >=0.8 on 77% of GPQA case-result entries, 48% of those wrong). An exploratory cross-family check on three Claude tiers shows the same frontier over-confidence, with confident errors recurring across providers above a marginal-preserving null. Self-consistency is thus a conditional proxy for correctness, not a standalone confidence score. We publicly release the de-identified per-run rows and answer distributions.
Science Journals
arXiv:2607.07925v1 Announce Type: cross
Abstract: Turbulence in the solar interior and atmosphere plays a crucial role in energy transport, yet modeling its subgrid-scale effects remains a major challenge. This study leverages machine learning (ML) models to predict components of the Reynolds stress tensor using high-resolution StellarBox simulations of the quiet Sun. Previously, we have compared a Multi-Layer Perceptron (MLP) and a 3D Convolutional Neural Network (CNN) against physics-based baselines to achieve a lower Mean Squared Error (MSE) and better generalization across various heights and depths in the solar atmosphere. To enhance learning, in this work, we investigate cluster-weighted training using K-Means and Hierarchical Agglomerative Clustering (HAC). By weighing the loss function based on cluster-specific prediction errors, we direct the model's attention to high-error regions. It significantly improves CNN performance, achieving 34% lower MSE and a significantly higher R2 score indicating that integrating deterministic clustering with ML is a promising technique for modeling subgrid turbulence, in particular, and regression in diverse environments, in general.
arXiv:2607.08071v1 Announce Type: new
Abstract: Online ads are essential to all businesses and ad headlines are one of their core creative component. Existing methods can generate headlines automatically and also optimize their click-through-rate (CTR) and quality. However, evolving ad formats and changing creative requirements make it difficult to generate optimized & customized headlines. We propose a novel method that uses prefix control tokens along with BART fine-tuning. It yields the highest CTR and also allows users to control the length of generated headlines for use across different ad formats. The method is also flexible and can easily be adapted to other architectures, creative requirements and optimization criteria. Our experiments demonstrate a 25.82% increment in Rouge-L and a 5.82% increment in estimated CTR over previously published strong ad headline generation baseline.
arXiv:2607.07932v1 Announce Type: cross
Abstract: We show the effect of gate-induced strain on the valence band of a silicon (Si) metal oxide semiconductor (MOS) confined two-dimensional hole gas (2DHG). Increasing aluminum gate thickness, and thereby the strain in the channel, results in the onset of a second subband contributing to Shubnikov-de Haas oscillations. Temperature-dependent magnetotransport measurements reveal distinct cyclotron masses of $m_c^*=(0.36\pm0.04)m_0$ and $m_c^*=(0.49\pm0.02)m_0$. The measured cyclotron masses differ from those expected for an idealized heavy-hole (HH)/light-hole (LH) picture, reflecting the combined influence of quantum confinement, strain, and HH-LH mixing on the valence band.
Multimodal 3D LUT Generation via StatLUT with Statistical Features for Photorealistic Style Transfer
arXiv:2607.08227v1 Announce Type: new
Abstract: Photorealistic Style Transfer (PST) aims to transfer the color and tonal style of a reference to a content image while strictly preserving its structural integrity. However, existing deep learning-based methods inherently suffer from semantic entanglement caused by pre-trained image encoders, leading to unnatural spatial distortions. Moreover, current pixel-level mapping paradigms often ignore color gamut topology, resulting in color banding, while also lacking the multimodal capability for intuitive text-driven control. To address these bottlenecks, we propose StatLUT, an innovative multimodal framework for 3D LUT generation. First, we bypass traditional encoders and introduce a Lab-Extractor to derive spatially-agnostic statistical features, fundamentally decoupling color distributions from structural semantics to ensure artifact-free rendering. Second, we formulate LUT generation as a Transformer-based Seq2Seq translation task, utilizing a Multi-dimensional Residual Mapper (MR-Mapper) to predict topologically smooth 3D LUTs. Finally, to break the single-modal barrier, we propose the H-Diffuser, a lightweight Diffusion Transformer that directly synthesizes statistical features from natural language prompts, enabling flexible text-driven color grading. Extensive experiments on standard benchmarks demonstrate that StatLUT significantly outperforms state-of-the-art methods in both visual quality and quantitative metrics, pioneering a highly robust and flexible paradigm for multimodal photorealistic style transfer.
arXiv:2607.08744v1 Announce Type: new
Abstract: Forecast aggregation aims to combine information from multiple Bayesian experts' forecasts into an aggregate forecast. In much of this literature, however, the aggregate forecast is optimized for a particular loss or robustness criterion and need not itself be calibrated with respect to the outcome. We introduce and study expert aggregation, where the goal is instead to aggregate Bayesian experts into a new expert that continues to provide calibrated forecasts. In particular, we consider a setting where each input expert reports calibrated predictions, and the aggregator observes the prior distribution over states, and the input experts, but not the underlying Bayes probabilities of the states. We ask whether one can (i) construct a calibrated output expert that Blackwell refines a target expert and cannot be further Blackwell improved using the available information; and (ii) when a proper loss is specified, compute a nearly loss-optimal expert among all such refinements.
We formulate calibrated experts as reduced-form information structures and measure refinement by Blackwell dominance of the induced prediction distributions. We characterize the constructible output experts through observable linear information: the input experts generate a linear system whose row space determines which calibrated output predictions are identifiable, and a new expert is constructible exactly when its predictions lie in the associated observable nonnegative cone. We establish a sharp algorithmic picture. When randomized output experts are allowed, both questions above admit efficient algorithms. In contrast, deterministic output experts are computationally intractable: deciding whether a deterministic calibrated refinement exists is $\mathsf{NP}$-hard, and deterministic proper-loss optimization admits no multiplicative PTAS unless $\mathsf{P}=\mathsf{NP}$.
arXiv:2607.07950v1 Announce Type: cross
Abstract: Pairwise comparisons are fundamental in the analytic hierarchy process. Various consistency indices have been proposed to assess inconsistencies in these comparisons. Since Saaty first proposed his consistency index, the assessment of the degree of consistency in pairwise comparison matrices has remained an open and hot topic in the study of the analytic hierarchy process. The consistency indices CI and CR proposed by Saaty are defined using the principal eigenvalue of the pairwise comparison matrix. In our previous study, we introduced an alternative index derived from the relationship between the coefficient of the characteristic polynomial and the consistency of comparisons.
Saaty proposed a fixed threshold of 0.1 for CI or CR as a guideline for an acceptable level of consistency, regardless of the matrix size. However, whether this threshold represents an equivalent level of consistency across different matrix sizes, that is, across different numbers of evaluation items, remains unclear. This study analysed the relationship between consistency and matrix size by examining pairwise comparison matrices constructed from subsets of evaluation items. Based on this analysis, we propose the fundamental property to be satisfied by a size-independent consistency index.
Furthermore, we refine our previously proposed index to ensure that it satisfies this property, demonstrating that it coincides with the existing consistency index. Finally, we visualise the relationship between the matrix size and consistency index values using randomly generated pairwise comparison matrices, thereby providing insights into the impact of matrix size on consistency evaluation.
arXiv:2607.08574v1 Announce Type: new
Abstract: Intracellular luminescence thermometry has long promised to reveal how heat is generated, dissipated, and regulated inside living cells. Yet, despite substantial progress, the field remains shaped by disagreement over the magnitude and physical plausibility of reported intracellular temperature gradients. In this manuscript, we discuss luminescence thermometry as a powerful approach for probing temperature at subcellular length scales, while emphasizing the experimental care required to make such measurements meaningful. After outlining the field's development, we outline the relevant heat-transfer concepts, before introducing luminescence thermometry and the performance metrics used to describe precision and accuracy. We then examine how thermometer design, intracellular localization, calibration, microscopy configuration, and data treatment influence the final thermal readout. Particular attention is given to two recurrent sources of error: bias, arising from measurement conditions and optical distortions, and cross-sensitivity, arising when the probe responds to parameters other than temperature, such as pH, viscosity, ionic strength, or biomolecular interactions. Finally, we outline practical directions for improving reproducibility, including multi-feature readouts, machine-learning-assisted analysis, and FAIR data practices, while suggesting future research directions.
arXiv:2607.08581v1 Announce Type: new
Abstract: Extreme Learning Machine (ELM) computes output weights analytically using the Moore-Penrose pseudoinverse. Although this leads to fast training, its numerical stability depends strongly on the conditioning of the hidden layer matrix. This paper studies pseudoinverse-based ELM from a spectral perspective. We show that the smallest singular value governs perturbation amplification in the output weights, while the condition number provides a quantitative measure of hidden-layer instability. We compare SVD-based pseudoinverse computation with iterative hyperpower methods and discuss width-dependent conditioning through a random feature interpretation. Experiments on synthetic matrices and ELM benchmarks show that SVD-based methods remain the most reliable under ill conditioning, while iterative methods are more sensitive to spectral properties. The results suggest that ELM stability is fundamentally governed by the singular value structure of the hidden layer matrix.
arXiv:2607.08084v1 Announce Type: cross
Abstract: Radiomic features derived from medical images and segmentation masks are used to support decision making in clinical imaging pipelines. In practice, these features are often computed from predicted masks, but segmentation models can be overconfident or poorly calibrated, making derived measurements appear more reliable than they are. Conformal prediction (CP) provides distribution-free prediction intervals with finite-sample marginal coverage guarantees, but black-box intervals for segmentation-derived radiomics can be inefficient because they ignore test-time information about image appearance, mask geometry, and segmentation uncertainty. We propose ConRad, a conformal framework for scalar radiomic targets that uses covariates derived from the predicted mask, input image, predicted radiomics, and boundary uncertainty to construct adaptive intervals while maintaining coverage. Across five 2D medical imaging datasets and 171 retained radiomic targets, we show that ConRad improves feature-level efficiency compared to baselines while maintaining near-nominal empirical coverage. Ablation results further indicate that segmentation boundary uncertainty features are the largest contributors to interval efficiency.
arXiv:2607.08361v1 Announce Type: cross
Abstract: The development of organic ferroelectric materials through scalable and simplified fabrication routes remains a major challenge for next-generation energy-harvesting technologies. Here, polycrystalline croconic acid (CA) thin films are fabricated by vacuum sublimation onto Ar plasma-treated flexible substrates and stabilized by in situ encapsulation with an adamantane-based remote plasma polymer. This solvent-free strategy effectively suppresses surface degradation under ambient conditions, providing long-term stability. Piezoresponse force microscopy confirms robust ferroelectricity with an oblique polarization orientation, well-defined domains, and low nanoscale coercive fields. The films were integrated into multilayer piezoelectric and pyroelectric devices. The piezoelectric performance strongly depends on film thickness, while embedding the CA layer between dielectric polymeric films significantly improves the macroscopic response, reaching power densities of up to 37 microW m-2 for ca. 2 micrometer CA films. Despite the common assumption that high crystallinity is required to sustain ferro-, piezo-, and pyroelectricity, these polycrystalline CA films exhibit remarkable RT pyroelectricity, a property not previously demonstrated in CA-based devices. A pyroelectric coefficient of ca. 10 microC m-2 K-1 highlights a functional response comparable to that of well-established organic and inorganic pyroelectric materials, demonstrating the potential of CA thin films for thermal energy harvesting. Beyond their functional performance, the proposed low-T fabrication route combines deposition and encapsulation in a single in situ process, simplifying device fabrication. Its compatibility with scalable vacuum technologies, flexible substrates, and further process optimization makes this approach highly promising for developing low-cost, lead-free, multisource energy-harvesting systems.
arXiv:2607.08295v1 Announce Type: cross
Abstract: Reflection of waves at interfaces is conventionally governed by Snell's law, which follows from conservation of momentum parallel to the interface. Here we show experimentally that caustic spin-wave beams in anisotropic media obey a fundamentally different reflection mechanism. Applying time-resolved Kerr microscopy to a yttrium iron garnet waveguide, we observe that reflected beams are selected by transitions between caustic points on the anisotropic iso-frequency contour rather than by momentum conservation. As a consequence, the reflected carrier wave vector and wavefront orientation exhibit trends opposite to those predicted by Snell's law. By tuning the magnitude and orientation of an external magnetic field, we continuously control the resulting reflection process and beam routing. Our results establish caustic-point transitions as a distinct reflection law for anisotropic wave beams and provide a route towards reconfigurable magnonic beam steering.
arXiv:2607.08386v1 Announce Type: cross
Abstract: A novel parallel approach is proposed for QEC decoding based on Belief Propagation with Ordered Statistics Decoding. The main idea is to pre-process the error vectors obtained from Belief Propagation by applying Singular Value Decomposition locally to sub-regions of the lattice. The proposed approach is applied to distributed quantum computers and evaluated in terms of complexity, accuracy, and scalability.
arXiv:2607.08410v1 Announce Type: cross
Abstract: The solar corona exhibits a pronounced temperature inversion, with plasma temperatures increasing by nearly two orders of magnitude from the chromosphere to the corona. We investigate how spatially sparse and temporally intermittent stochastic heating at the base of the transition region shapes the temperature and density structure of coronal loops within a kinetic framework. Stochastic thermal boundary conditions and surface coarse graining are introduced. Analytical solutions are derived in the collisionless limit for heating-event time scales shorter or longer than the particle crossing time, and Coulomb collisions are incorporated through a reduced kinetic model describing the thermalization of suprathermal particles. In the short-time-scale regime, spatial filling factor and temporal intermittency combine into a single effective parameter controlling the suprathermal population, producing a transition region and a hot corona both within individual loops and after coarse graining. Collisions preserve this thermal structure while reducing the coronal density through progressive thermalization. In the long-time-scale regime, individual loops are nearly isothermal and the temperature inversion emerges only after coarse graining, depending solely on the spatial filling factor. Here, Coulomb collisions and optically thin radiative losses have only minor effects, while density and temperature profiles remain broadly consistent with coronal observations. These results show that sparse, intermittent heating naturally generates suprathermal particle distributions and reproduces the observed thermal structure of the solar corona within a kinetic framework, highlighting the different sensitivity of the two regimes to collisional effects.
arXiv:2607.07758v1 Announce Type: new
Abstract: Foundation models (FMs) have transformed machine learning from isolated task-specific model development toward general-purpose models pretrained on broad data and adapted to multiple downstream tasks. Earth observation (EO) is an important domain for this paradigm because satellite and airborne archives are large, high-revisit, and increasingly multimodal, while reliable field labels are often sparse. Remote sensing foundation models (RSFMs) cannot be transferred reliably/optimally without domain-specific adaptation. This is because EO data are governed by measurement physics and operational decision constraints. This chapter reviews the design principles arising from these domain-specific constraints. It first defines the FMs paradigm in remote sensing (RS), then synthesizes the current model landscape, pretraining objectives, architecture designs, downstream adaptation and trustworthiness requirements. The chapter also incorporates recent benchmark evidence showing that no single geospatial foundation model is universally best and that inconsistent evaluation remains a major issue to fair comparison and reliable deployment. In addition, two brief environmental monitoring case studies; physics-informed spectral targeted masking for harmful algal bloom prediction and reinforcement learning for adaptive environmental monitoring station selection to illustrate the FMs domain-guided principles in practice. This chapter posits that next-generation RSFMs should be evaluated not only by benchmark accuracy, but also by modality-aware transfer and physically plausible representations for trustworthy EO decisions.
arXiv:2602.04718v2 Announce Type: replace
Abstract: A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not substantially alter the effects of others. In practice, however, feature entanglement leads to interference such that localized interventions can have unintended downstream effects. Motivated by the \textit{Independent Causal Mechanisms} principle, we propose to constrain internal features to be almost orthogonal. We argue that this promotes modular representations amenable to causal intervention. We formalize this problem by characterizing the gap between an idealized isolated intervention and its realized effect on model outputs in terms of feature interference. We upper-bound the propagation of feature interference in terms of the self-coherence of the feature dictionary, and relate this discrepancy to an explicit orthogonality regularization on the dictionary itself. Empirically, we show that this regularization enables more isolated interventions on mathematical reasoning concepts while preserving model performance. Our code is available under \texttt{https://github.com/mrtzmllr/sae-icm}.
arXiv:2606.15920v2 Announce Type: replace
Abstract: Reinforcement learning for multimodal large language models (MLLMs) is often hindered by severe reward sparsity in complex reasoning tasks. This challenge is particularly pronounced in human-centered scenarios involving states, emotions, intentions, and behaviors, where heterogeneous multimodal signals and subjective human factors make high-quality chain-of-thought (CoT) annotations expensive and difficult to obtain. Although many multimodal datasets provide expert-annotated ground-truth labels, directly using these labels for supervised fine-tuning may encourage shortcut learning in multimodal perception and provides limited transparency for safety-critical human--AI interaction. To address these limitations, we propose OmniOPSD, a Rationale-Privileged On-Policy Self-Distillation framework that uses frontier-generated rationales as teacher-side privileged evidence rather than student imitation targets. OmniOPSD uses frontier-generated evidence-aware rationales only as training-time privileged evidence context for a local teacher. The student samples its own rollout from the original multimodal input, while the rationale-privileged teacher scores the same tokens and provides dense token-level supervision. Thus, the student learns on its own trajectory distribution without directly imitating frontier-model completions, and inference requires no labels, rationales, CoT annotations, or closed-source model access. Experiments on MER-UniBench show that OmniOPSD achieves state-of-the-art performance with an average score of $84.19$, and ablations further support the value of rationale-privileged teacher guidance.
arXiv:2607.08143v1 Announce Type: new
Abstract: We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents. OCR post-correction remains a long-standing challenge in digital heritage: large-scale collections of digitized documents are affected by legacy OCR errors, while re-digitization at scale remains impractical. Large language models (LLMs) offers a major opportunity to revisit this challenge, yet their effectiveness across languages, document types, and noise conditions - and their tendency to hallucinate - remains insufficiently understood. HIPE-OCRepair-2026 pursues two objectives: (i) to evaluate the capabilities of modern OCR post-correction systems, and (ii) to provide a reproducible evaluation framework anchored in the HIPE-OCRepair-2026 dataset, a harmonized multilingual resource consolidating existing and newly curated historical datasets. Participants were tasked with correcting noisy OCR transcripts from historical newspapers and printed works in English, French, and German (17th-20th century), working at the level of coherent transcription units (paragraphs or articles) without access to source images. The evaluation adopts a retrieval-oriented rather than diplomatic scoring approach, reflecting the practical use case of search and access over digitized collections. Four teams submitted systems ranging from zero-shot prompting to continued pre-training and fine-tuning, offering insights into the merits of different adaptation strategies. Results show that modern LLM-assisted systems can significantly improve OCR quality, but performance varies across datasets, languages, and noise levels. Over-correction on low-noise inputs emerges as a recurring challenge, highlighting the importance of evaluation beyond character error reduction. The dataset, scorer, and evaluation pipeline are publicly released to support future research.
arXiv:2607.08127v1 Announce Type: new
Abstract: Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on robotic action data. We refer to this discrepancy as the video-action generalization gap. In this paper, we systematically investigate this gap by evaluating a comprehensive design space of VAMs, demonstrating that standard design choices yield no emergent explanation pattern. To explain this behavior, we introduce the Temporal Ratio (TR), an attention-based measure of how strongly the action head relies on future latent rollouts relative to the anchored current frame. TR has two key properties: first, a model's structural reliance on future-predictive latents, measured via TR, acts as a predictor of its compositional generalization capacity; second, it natively fluctuates based on task phase, shifting attention to future frames during planning and reverting to the present frame for precise manipulation. Finally, based on these findings, we propose an inference-time adaptive guidance method, which exploits this intrinsic feature attention pattern to dynamically amplify compositional video conditioning signals precisely when the policy relies on future rollouts. Evaluated on the LIBERO benchmark and real-world tasks, our approach mitigates the OOD-ID compositional generalization gap. More details: https://umishra.me/temporal-ratio/
arXiv:2607.07962v1 Announce Type: new
Abstract: Inferring latent physical properties from sensory observations is a fundamental challenge in machine perception. Among available sensing modalities, thermal imaging is particularly promising because temperature evolution is directly governed by heat-transfer physics and therefore encodes information about underlying thermophysical properties of a scene. Recovering spatially resolved thermophysical properties from thermal observations could transform applications ranging from digital twins and infrastructure monitoring to robotics and scientific imaging. However, existing thermal scene reconstruction methods can recover temperature fields in complex 3D environments without identifying the thermophyiscal properties that govern thermal evolution, whereas inverse methods provide physically interpretable parameter estimation but typically rely on simplified geometries and controlled experimental conditions.
Here we introduce ThermoField, a framework that unifies thermal scene reconstruction and thermophysical parameter estimation through differentiable heat-transfer simulation. The proposed framework represents these quantities as spatially varying neural fields and constrains them through scene geometry, governing heat-transfer physics, and temporal thermal observations. We demonstrate that ThermoField jointly reconstructs geometry, estimates spatially varying thermal diffusivity, and predicts thermal evolution under previously unseen environmental conditions. By integrating neural scene representations with differentiable heat-transfer solver, the framework enables physically interpretable parameter inference in complex 3D scenes. Our results establish a bridge between thermal scene reconstruction and inverse heat-transfer analysis, providing a unified approach for geometry reconstruction, thermophysical property estimation, and predictive thermal simulation from thermal observations.
arXiv:2607.08149v1 Announce Type: new
Abstract: Capturing the ultrafast structural dynamics that occur at the solid-liquid interface is key to understanding adsorption, desorption, diffusion, and aggregation processes in catalysis and interfacial chemical reactions. Hard-X-ray scattering in grazing-incidence geometry can, in principle, access interfacial structural changes with angstrom-scale structural sensitivity and ultrafast temporal resolution. However, the long optical paths of the optical pump and hard-X-ray pulses inside the liquid sample pose significant challenges to the temporal resolution, signal-to-noise ratio, and overall stability of such an experimental scheme. Here, we report a method for creating and characterizing ultrathin surface-attached free-flowing liquid sheets, whose submicrometer thickness enables ultrafast temporal resolution and reduces the bulk-liquid scattering contribution. The impinging-jet geometry produces stable micrometer-scale sheets whose morphology depends systematically on incidence angle, jet velocity, and capillary diameter. Gas-assisted shaping using a second capillary further narrows and thins the sheet, producing an extended ultrathin region and reducing the measured minimum thickness below 500~nm for acetonitrile. The resulting platform provides a reproducible, continuously flowing, surface-attached liquid geometry for grazing-incidence scattering experiments.
arXiv:2601.21688v2 Announce Type: replace
Abstract: Disentangled representation learning aims to map independent factors of variation to independent representation components. On one hand, purely unsupervised approaches have proven successful on fully disentangled synthetic data, but fail to recover semantic factors from real data without strong inductive biases. On the other hand, supervised approaches are unstable and hard to scale to large attribute sets because they rely on adversarial objectives or auxiliary classifiers.
We introduce \textsc{XFactors}, a weakly-supervised VAE framework that disentangles and provides explicit control over a chosen set of factors. Building on the Disentangled Information Bottleneck perspective, we decompose the representation into a residual subspace $\mathcal{S}$ and factor-specific subspaces $\mathcal{T}_1,\ldots,\mathcal{T}_K$ and a residual subspace $\mathcal{S}$. Each target factor is encoded in its assigned $\mathcal{T}_i$ through contrastive supervision: an InfoNCE loss pulls together latents sharing the same factor value and pushes apart mismatched pairs. In parallel, KL regularization imposes a Gaussian structure on both $\mathcal{S}$ and the aggregated factor subspaces, organizing the geometry without additional supervision for non-targeted factors and avoiding adversarial training and classifiers.
Across multiple datasets, with constant hyperparameters, \textsc{XFactors} achieves state-of-the-art disentanglement scores and yields consistent qualitative factor alignment in the corresponding subspaces, enabling controlled factor swapping via latent replacement. We further demonstrate that our method scales correctly with increasing latent capacity and evaluate it on the real-world dataset CelebA. Our code is available at \href{https://github.com/ICML26-anon/XFactors}{github.com/ICML26-anon/XFactors}.
arXiv:2604.11305v3 Announce Type: replace
Abstract: Conformal selection (CS) uses calibration data to identify test inputs whose unobserved outcomes are likely to satisfy a pre-specified minimal quality requirement, while controlling the false discovery rate (FDR). Existing methods fix the target FDR level before observing data, which prevents the user from adapting the balance between number of selected test inputs and FDR to downstream needs and constraints based on the available data. For example, in genomics or neuroimaging, researchers often inspect the distribution of test statistics, and decide how aggressively to pursue candidates based on observed evidence strength and available follow-up resources. To address this limitation, we introduce post-hoc CS (PH-CS), which generates a path of candidate selection sets, each paired with a data-driven false discovery proportion (FDP) estimate. PH-CS lets the user select any operating point on this path by maximizing a user-specified utility, arbitrarily balancing selection size and FDR. Building on conformal e-variables and the e-Benjamini-Hochberg (e-BH) procedure, PH-CS is proved to provide a finite-sample post-hoc reliability guarantee whereby the ratio between estimated FDP level and true FDP is, on average, upper bounded by 1, so that the average estimated FDP is, to first order, a valid upper bound on the true FDR. PH-CS is extended to control quality defined in terms of a general risk. Experiments on synthetic and real-world datasets demonstrate that, unlike CS, PH-CS can consistently satisfy user-imposed utility constraints while producing reliable FDP estimates and maintaining competitive FDR control.
arXiv:2607.08154v1 Announce Type: new
Abstract: Intrinsically-typed presentations of type theory often use equality in the meta-language to represent object-language judgmental equality. In such equational syntax, proof-relevant logical relations define computability predicates on judgmental equivalence classes of types and terms. This approach, however, does not directly account for reduction, which is directed and plays a central role in many logical-relations arguments. This paper develops a directed version of proof-relevant logical relations in simplicial homotopy type theory, where reductions are internalized as \emph{inequality types}. We construct object syntax as a directed quotient inductive type. The central observation is that contravariant families in simplicial type theory provide exactly the proof-relevant form of closure under expansion for logical relations: computability evidence can be transported backward along reductions, with the required functoriality and universal property built in. Using this observation, we construct a unary logical relations model with contravariant computability predicates and prove directed Boolean canonicity: every closed Boolean term reduces to either true or false. We then extend the construction to dependent types and universes, where a comonadic flat modality provides the discreteness needed for type conversion and universe predicates. Finally, we adapt the method to binary logical relations, separating vertical reduction from horizontal parametricity and obtaining a proof-relevant account of representation independence.
arXiv:2607.08758v1 Announce Type: new
Abstract: Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, much like biological genomes. Current benchmarks still say little about whether AI systems can follow this inheritance structure. We present IdeaGene-Bench (IG-Bench), a benchmark for scientific lineage reasoning and lineage-grounded idea generation. IG-Bench is organized around the IdeaGene framework: each paper or proposal is represented as a set of minimal, typed, evidence-grounded Idea Genome objects, and a GenomeDiff aligns these objects to record inheritance, mutation, loss, external import, and novel insertion under six operational evolutionary dynamics. The benchmark contains 1,961 golden lineage traces, 1,085 curated Idea Genome objects, and 920 pairwise GenomeDiff records across 10 scientific domains. It supports two evaluations. IG-Exam (42 task types, 1,029 instances) tests closed-form lineage reasoning across Idea Genome abstraction, inheritance tracing, evolutionary reasoning, and lineage verification. IG-Arena evaluates generation with a lineage-conditioned Population-Evolution Score(PES), asking whether a proposal can be inserted as a coherent descendant of a given lineage population: it should inherit the right Idea Genome objects, vary meaningfully from nearby work, and offer selection value for future research. Experiments on 14 LLM-based scientists expose a compositional bottleneck. The strongest system reaches only 27.3% exact accuracy on lineage reasoning, and structured lineage context reshuffles system rankings rather than helping every participant uniformly.