Forskningsradar

Science Journals

Peer-reviewade publikationer — 57198 artiklar

Semitopological Barycentric Algebras
arXiv:2512.12865v5 Announce Type: replace-cross Abstract: Barycentric algebras are an abstraction of the notion of convex sets, defined by a set of equations. We study semitopological and topological barycentric algebras, in the spirit of a previous study by Klaus Keimel on semitopological and topological cones (2008), which are special cases of semitopological and topological barycentric algebras. For example, the space of all continuous valuations (a very close cousin of measures) over a topological space is a topological cone, while probability valuations form a topological barycentric algebra, and subprobability valuations form a pointed topological barycentric algebra. Among other results, we show the existence of free semitopological cones over semitopological barycentric algebras and over pointed semitopological algebras, we investigate which semitopological barycentric algebras embed into semitopological cones and which pointed semitopological barycentric algebras embed strictly into semitopological cones. We study notions of local convexity, which split into weak local convexity, local convexity, local affineness and local linearity. We show that the weakly locally convex topological barycentric algebras are exactly the affine retracts of locally affine topological barycentric algebras. On locally convex barycentric algebras, we show sandwich theorems, extending theorems by Roth and Keimel on cones. A running theme of this paper is the notion of barycenters, which we progressively generalize until we reach a general notion of barycenters of continuous (resp., subprobability, probability) valuations, inspired by a definition of Choquet. We conclude with a general barycenter existence theorem, whose proof relies on the study of the Smyth poweralgebra, namely the topological barycentric algebra of all non-empty convex compact saturated subsets of a topological barycentric algebra.
Spatially Grounded Concept Bottleneck Models via Part-Factorized Attention
arXiv:2606.04364v2 Announce Type: replace Abstract: Concept bottleneck models (CBMs) predict a layer of human-named attributes before predicting a class, which makes their decisions auditable. On fine-grained recognition tasks the concept heads are usually free to attend anywhere in the image, so a head named for one body region can be satisfied by evidence on another. This work studies a part-factorized CBM that removes that freedom by construction. The method has three components built on a frozen DINOv3 vision transformer. A learned foreground gate, trained on DINOv3 patch features, suppresses background patches inside the part attention. A set of part queries cross-attends to patch features and each of the 312 CUB attributes is routed, through a fixed concept-to-part map, to read only from the part token its name implies. A learnable two-dimensional Gaussian prior, injected additively in log space into the attention logits, breaks the permutation symmetry among part queries; its means are initialized from the dataset-average keypoint location of each part, which requires no per-image keypoint supervision at training or test time. On CUB-200-2011 the spatial-prior model matches a fully supervised baseline (88.85% versus 88.95% top-1) while raising pointing accuracy by 16 points (52.6% versus 36.4%). Replacing bounding-box supervision with a PCA foreground target and combining it with the Gaussian prior removes all per-image supervision and reaches 88.6% top-1 at about 70% pointing accuracy. A keypoint-fraction sweep shows that 0.5% of the training set (about 27 images) suffices to initialize the prior with no measurable loss. Removing part identity entirely is the harder case: without any spatial prior, pointing accuracy collapses to $2.9\%$.
CRAFTIIF: Cross-Resolution Analytic Four-Type Interpretable Isolation Forest for Multivariate Time Series Anomaly Detection
arXiv:2606.13486v1 Announce Type: new Abstract: Anomaly detection in multivariate time series is challenged by four structurally distinct anomaly types -- point (isolated spikes), distributional (level shifts), temporal (rhythm changes), and collective (inter-sensor correlation breakdowns) -- each requiring different feature representations. Most unsupervised methods target only one or two types and provide limited interpretability. We present CRAFTIIF (Cross-Resolution Analytic Four-Type Interpretable Isolation Forest), a fully unsupervised framework targeting all four types without dataset-specific tuning. CRAFTIIF generates K=500 random analytic wavelet feature draws across four families (Morlet, DOG, Haar, Coiflet), each targeting a specific anomaly type, feeding five structured Isolation Forests -- one per type plus a meta-IF for compound anomalies. An adaptive Otsu/MAD threshold calibrates detection automatically across anomaly rates from 0.1% to 69.2%. Because each IF is trained exclusively on type-specific features, branch firing provides direct anomaly-type attribution by construction, without post-hoc explanation. Evaluated on all 19 datasets of the mTSBench benchmark (Zhou et al., TMLR 2026), CRAFTIIF achieves mean F1=0.228 (all 19 datasets) and F1=0.322 (13 detectable datasets), ranking first among all 25 evaluated methods on VUS-PR (0.463 vs. previous best 0.329, +40.7%). A diagnostic framework -- oracle F1, detectability limits, and branch separation ratios -- identifies 6 of 19 datasets as fundamentally undetectable by any unsupervised method. Ablation over 11 conditions confirms adaptive thresholding (+38% F1), four-branch structure (+20%), and meta-IF (+23%) are each essential. Code: https://github.com/smitswil/craftiif
Stable, bidirectional electro-optic transduction in thin film lithium tantalate
arXiv:2606.12726v1 Announce Type: cross Abstract: Efficient and stable microwave-optical transduction is a key enabling technology for distributed superconducting quantum computing and heterogeneous quantum networks. Electro-optic transducers based on thin-film lithium niobate (TFLN) have shown strong promise, but demonstrations to date have been limited by various factors such as low frequency bias drift, low efficiency, fabrication complexity, and scalability. Here we demonstrate the first integrated electro-optic microwave-optical transducers realized in thin-film lithium tantalate (TFLT), a material platform offering Pockels nonlinearity comparable to TFLN together with improved bias stability and high-power handling. We fabricate superconducting microwave resonators coupled to tunable photonic-molecule optical resonators using wafer-scale deep ultraviolet lithography, offering high-throughput production of hundreds of devices per wafer. Across six devices we observe coherent bidirectional conversion between C-band optical photons and 4.9-5.5 GHz microwave photons, with measured on-chip efficiencies and inferred single-photon coupling rates g_0/2{\pi} ~ 1 kHz consistent with theory. Continuous operation over multiple days is achieved using a static bias field with minimal feedback, demonstrating a major operational advantage. We further characterize optical loss statistics, microwave resonator performance, and optically induced added noise under pulsed pumping, finding less than one added photon for 100 microsecond pulses at the highest measured efficiencies. These results establish TFLT as a scalable and robust electro-optic platform for future quantum interconnects and modular quantum processors.
A conformally-Euclidean Line Element for evaluating color differences
arXiv:2606.12588v1 Announce Type: new Abstract: Starting from our previously proposed line element and considering more ``surface color'' datasets, we derive a simplified version which matches experimental datasets equally well and resulted into a conformally-Euclidean line element, which is conceptually much simpler than any existing color difference metrics. The color difference is written as an Euclidean difference multiplied with a simple factor which depends on the luminance only. In a subspace with constant luminance, as considered by MacAdam, this factor becomes constant and the subspace is flat. The same holds for sufficiently large luminances. Based on this LE we derive perceptual coordinates $\left(A,l_{c},s_{c}\right)$ very similar to the CIELab $\left(L^{*},a^{*},b^{*}\right)$.
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
arXiv:2606.12736v1 Announce Type: new Abstract: AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchmarks for AI agents rarely capture the complexity, heterogeneity, and extended reasoning required by scientific work, whereas benchmarks for scientific tasks often reduce research to static, direct problems and provide limited support for interactive evaluation. Here, we introduce SciAgentArena, a systematic benchmark for evaluating AI agents in real-world scientific research scenarios drawn from emerging needs across multiple domains. SciAgentArena comprises approximately 200 tasks with stepwise verification and an interactive, agent-agnostic environment for assessing diverse AI agents. Using this benchmark, we find that current agents can contribute effectively to well-specified data-analysis workflows, particularly when the task structure and evaluation criteria are clear. However, their performance remains uneven across scientific contexts: agents struggle to generate genuinely novel insights, sustain self-directed exploration, and formulate robust solutions for open-ended research questions. We further characterize common failure modes across agents and identify opportunities for improving their reliability, autonomy, and scientific reasoning. Together, SciAgentArena provides a practical framework for measuring progress in AI agents for science and for guiding the design of future agents capable of addressing complex scientific challenges. Full codes, tasks, and datasets can be accessed via this link: https://sciagentarena.github.io/.
Learning-Augmented Approximation for Unrelated-Machines Makespan Scheduling
arXiv:2606.13133v1 Announce Type: new Abstract: Recently, Antoniadis et al. (ICLR 2025) proposed a framework for incorporating predictions to approximate NP-hard selection problems. Despite its simplicity, this approach tightly matches theoretical lower bounds, making its generalization highly compelling. We address an open question raised in the work of Antoniadis et al., concerning the extension of this approach to other important problems outside the class of selection problems, such as scheduling. We develop a learning-augmented algorithm for the makespan minimization problem on unrelated machines, denoted by $R\|C_{\max}$. By using predictions of heavy job assignments, we achieve a polynomial-time $(1+\varepsilon)$-approximation for accurate predictions that smoothly degrades to a worst-case 2-approximation as the error increases. We conclude our work with an empirical analysis of our method.
Hyperstatistical thermodynamics of the one-dimensional Klein-Gordon and Dirac oscillators: a closed-form q-generalized Boltzmann factor and a quantitative comparison with Beck's superstatistics
arXiv:2606.12454v1 Announce Type: new Abstract: We revisit the thermodynamics of the one-dimensional Klein-Gordon (KGO) and Dirac (DO) oscillators within two frameworks of generalized statistics: Beck's asymptotic superstatistics and the recently introduced hyperstatistics. In hyperstatistics, a $\gamma$-distribution of domain Boltzmann factors yields, after Laplace transformation and averaging over a normalisable density $f(\beta)$, the closed-form q-generalized Boltzmann factor $B_q(\varepsilon) = \exp_q(-\langle\beta\rangle\varepsilon)$, independent of $f(\beta)$. We compute the partition function, entropy $S$, and specific heat $C_v$ for both 1D oscillators using excitation energies $\varepsilon_n = E_n - E_0$ to remove the rest-energy shift and enforce third-law behaviour $C_v \to 0$ as $T = 1/\langle\beta\rangle \to 0$. Appropriate degeneracies ($g_n = 1$ for KGO; $g_0 = 1$, $g_n = 2$ for $n \geq 1$ for DO) are applied. Hyperstatistics successfully (i) reproduces the high-temperature Boltzmann limit $C_v \to 2k_B$, (ii) is structurally independent of $f(\beta)$, (iii) avoids the unphysical negative regions of the Beck polynomial bracket, and (iv) systematically distinguishes KGO from DO by capturing the enhanced entropy and sharper specific-heat structure caused by spin-induced degeneracy. The frameworks agree quantitatively for $q - 1 \ll 1$ and $\langle\beta\rangle E \lesssim 2$, but diverge at high temperatures where Beck's polynomial expansion loses validity and the exact hyperstatistical q-exponential remains positive, monotonic, and analytic. Ultimately, hyperstatistics provides a numerically stable and analytically tractable alternative to asymptotic superstatistics for relativistic oscillators, naturally extensible to higher dimensions and external magnetic fields.
AI SciBrief as a Gateway to Research: A Framework for Onboarding Students into New Research Areas
arXiv:2606.12413v1 Announce Type: new Abstract: Students at all levels of higher education face a significant barrier in the form of information overload, which often paralyzes the initial stages of the research process and suppresses motivation. In response, this article introduces a pedagogical framework that leverages AI SciBrief, a platform powered by a Large Language Model (LLM) designed to automatically generate digests of scientific trends. We describe how this multidisciplinary tool - with initial coverage in finance, medicine, and education - can be integrated into the curriculum to overcome this "entry barrier." The framework provides concrete methodologies for utilizing these digests to facilitate topic selection for term papers, accelerate literature reviews for dissertations, and enable postgraduate students to continuously monitor emerging trends. We conclude that AI SciBrief functions as a "gateway to research" effectively reducing students' cognitive load and empowering them to transition more rapidly from information searching to knowledge creation.
The Khipu Problem: Institutional Legibility Under Distributed Cognition
arXiv:2606.12414v1 Announce Type: new Abstract: AI governance still tends to assume that the relevant object is a bounded model or a bounded agent. That assumption is getting weaker. Real systems increasingly distribute cognition across models, tools, humans, context stores, retrieval layers, runtime policies, authorization boundaries, and delegated institutional roles. In such systems, the central governance problem is no longer only what the system did, but whether later institutions can still read what the system was. This paper introduces the khipu problem for distributed AI: the record can survive while the reading practice needed to interpret it decays. Logs, traces, model versions, tool calls, outputs, and approval artifacts may remain available while the institutional capacity to read them as parts of one coherent cognitive episode disappears. We argue that this failure is better understood as loss of interpretive continuity than as ordinary lack of observability. The result is a distinct governance failure. Institutions must classify, trust, audit, and constrain systems whose relevant identity is distributed across components and whose legibility depends on surrounding interpretive scaffolding. The problem is not merely missing data. It is a structural mismatch between what can be represented and what must still be decided under consequential conditions. We therefore argue that governance for distributed AI requires preservation of interpretive continuity, not only trace retention. The paper distinguishes missing evidence, ambiguous evidence, and structurally unreadable evidence; argues that many consequential outcomes are better understood as distributed cognitive episodes than as bounded model outputs; and proposes governance workspaces together with receipt-bearing governance surfaces as interpretive infrastructure for preserving action identity, authority, boundary truth, evidential scope, and consequential outcomes.
Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science
arXiv:2606.12426v1 Announce Type: new Abstract: LLM annotators are increasingly used in computational social science (CSS), but it is unclear whether their alignment-shaped errors preserve the empirical conclusions a researcher would report. We audit three open-source 7B instruction-tuned models (Zephyr, Mistral-Instruct, Qwen2.5-Instruct) across six TweetEval tasks under four prompt conditions (72 cells) and find that social-desirability failures do not run in a single direction. Zephyr exhibits leniency bias, systematically under-applying harmful labels (offensive language: false benign rate 0.729, false alarm rate 0.031). Mistral and Qwen exhibit overcorrection, over-applying the same labels (Mistral hate-speech FAR = 0.604). All three models exhibit neutrality bias on abortion stance, underestimating opposition prevalence by 24 to 40 percentage points and inflating the neutral label. None of the four prompting interventions we test (neutral, safety framing, depersonalized, chain-of-thought) corrects these failures across models; safety framing can worsen stance distortion. Strikingly, Zephyr's hate-speech prevalence estimate matches the gold rate exactly while its class-conditional errors are large in both directions, an accidental cancellation that misleads aggregate validation. We translate these patterns into a three-part taxonomy with diagnostic FBR/FAR signatures and a lightweight gold-sample validation protocol. The headline for trustworthy CSS: a model that looks calibrated on aggregate metrics can still flip the substantive empirical conclusion a researcher would report.
Muse Spark Safety & Preparedness Report
arXiv:2606.12429v1 Announce Type: new Abstract: Muse Spark is the latest large language model developed by Meta. In this report, we first present evaluations for catastrophic risk domains under Meta's Advanced AI Scaling Framework, along with the evidence that informed our launch decision. We then discuss additional considerations, such as Muse Spark's broader content safety and behavioral profile, that are relevant to overall safety but fall outside the catastrophic risk domains governed by the Framework. Our preparedness results covering Chemical and Biological, Cybersecurity, and Loss of Control risks assess Muse Spark's deployment within Meta AI as presenting acceptable levels of residual risks under our Advanced AI Scaling Framework. We conducted a broad set of evaluations targeting dual-use and high-risk capabilities across these catastrophic risk domains. Those evaluations identified elevated risks prior to mitigations, with Chemical and Biological capabilities assessed as likely reaching the "high risk" category under the Advanced AI Scaling Framework before safeguards were applied. We have implemented a multi-layered set of mitigations that address the identified risks, and Muse Spark demonstrates state-of-the-art refusal across a range of benchmarks related to hazardous workflows in chemistry and biology. We therefore release Muse Spark as the underlying model of Meta AI.
AI Debris: Residual Risk and the Afterlife of Failed AI Systems
arXiv:2606.12432v1 Announce Type: new Abstract: AI governance frameworks primarily focus on risks during the development and deployment phases, implicitly treating system withdrawal as a technical shutdown. This paper argues that decommissioned AI systems generate residual risk, termed AI debris, that persists after model removal and continues to shape institutional behaviour, accountability, and trust. AI debris is defined as the post-withdrawal socio-technical residue of AI systems, including workflow dependency, data contamination, capability displacement (deskilling), legitimacy erosion, and accountability breakdown. The paper develops a typology of debris domains and identifies mechanisms through which debris persists, including institutional memory, path dependency, blame avoidance, and feedback effects in organisational data. To operationalise the concept, the paper proposes an evaluator-ready AI Debris Decommissioning Protocol (AIDP), a stepwise checklist specifying auditable evidence for freezing decision footprints, incident review, remediation, contestability, and post-withdrawal accountability assignment. A brief vignette of Amazon's discontinued hiring tool illustrates how algorithmic decision categories and screening heuristics can persist after system rollback. The paper contributes a practical governance instrument for regulators, auditors, and organisations seeking to prevent paper compliance, strengthen AI lifecycle governance, and improve institutional resilience in high-stakes decision environments.
Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots
arXiv:2606.12439v1 Announce Type: new Abstract: Large language model (LLM) answer engines are increasingly used for information seeking, shifting visibility from ranked lists to synthesized answers. This enables Generative Engine Optimization (GEO), which targets LLM answer engines' evidence pool and generation. We analyze the search engine optimization (SEO) to GEO transition to identify two risks: (i) concentrated influence from low contestability and system sensitivity, and (ii) undisclosed commercial influence embedded in evidence and reasoning. We then formalize a general GEO pipeline to locate where optimization acts and compare academic and industry practices, revealing a third risk: (iii) academic-industry blind spots driven by visibility and evaluation asymmetries between offline setups and deployed systems. This position argues the need for answer-level governance and measurement: stronger contestability, high-precision disclosure, black-box auditing of material influence, and deployment-aligned metrics for exposure persistence.
Strategic Decision Support for AI Agents
arXiv:2606.12587v1 Announce Type: new Abstract: Traditionally, decision support studies how humans use machine learning models to make better decisions. In modern agentic systems, this division of roles is increasingly reversed: AI agents act on behalf of users, while humans and tools becomes support mechanisms around them. This role reversal brings reliability concerns to the forefront, since agentic errors can be consequential and agent behavior must remain aligned with human goals and constraints. Departing from the classical view of decision support, we revisit its two basic principles, the cost--value tradeoff of seeking support and the role of uncertainty quantification, in a setting where AI agents are the central actors. We propose a framework for strategic decision support for AI agents through an optimization problem that minimizes support usage subject to controlling a counterfactual missed-support error: the probability that the agent acts alone on instances where support would have materially improved its output. At the population level, we show that the optimal policy is a threshold rule on the value of support. Building on this structure, we develop an online algorithm that adaptively thresholds such a score and uses randomized exploration to control missed-support error without distributional assumptions. We further introduce a calibration-on-the-fly method that reduces unnecessary support calls online. We instantiate this framework across diverse scenarios, including information gathering, human--AI collaboration, and tool use, showing how each can be modeled through the same strategic decision-support lens. Experiments across these settings show that our method reliably controls the target error while substantially reducing support usage in practice.
Network-Based Multi-Layer Model Using Machine Learning for Optimal Vaccine Prioritization in Heterogeneous Populations
arXiv:2606.12456v1 Announce Type: new Abstract: This work advances epidemic control beyond traditional mass vaccination models by integrating population heterogeneity, network structure, and machine-learning-based decision policies. Using the Email-Eu-core contact network, we compare classical centrality-driven vaccination strategies with graph neural network (GNN) and reinforcement learning (RL) approaches. Across 30 stochastic simulations, classical heuristics, including degree, betweenness, and layer-based vaccination, exhibit similar performance, reflecting the network's dense connectivity and modest community structure. In contrast, the GNN-based strategy substantially reduces peak infection, final epidemic size, and time to peak, demonstrating its ability to identify structurally critical nodes that classical metrics overlook. These results show that learning-based vaccination policies can significantly outperform traditional heuristics by exploiting higher-order relational patterns in real-world networks, offering a powerful framework for targeted epidemic intervention.
Kinematic Probes of Type-II MMG: Pad\'e Cosmographic Analysis of VCDM
arXiv:2606.12464v1 Announce Type: new Abstract: We study the late-time expansion history of the Universe within the VCDM model, a Type-II MMG realization that preserves the successes of General Relativity while extending beyond constant vacuum energy through a minimal Hamiltonian modification, generating a time-dependent vacuum sector without introducing additional degrees of freedom. We investigate this framework within a cosmographic approach by employing a Pad\'e $P_{(2,1)}$ approximation for the Hubble parameter and luminosity distance, allowing the cosmographic parameters to be expressed directly in terms of the underlying VCDM model parameters and enabling a data-driven reconstruction of the expansion history. The model is constrained within a Bayesian framework using the MCMC technique, implemented via the affine-invariant ensemble sampler, with a joint analysis of cosmic chronometers, DESI BAO, and Type Ia supernova datasets (Union3, Pantheon+, and DESY5). We find that the model parameters are tightly constrained and consistent across different dataset combinations, with the jerk parameter remaining very close to its $\Lambda$CDM value, $j_0 \simeq 1$, indicating no significant deviation at the level of higher-order cosmography. Furthermore, the transition feature previously reported in VCDM is not observed within the Pad\'e $P_{(2,1)}$ cosmographic reconstruction, suggesting that it is not a robust requirement of current observational data but is sensitive to the choice of parametrization. Overall, our results indicate that the VCDM model effectively mimics $\Lambda$CDM at the background level when constrained through a cosmographic approach, underscoring the importance of model-independent reconstructions in assessing alternative cosmological scenarios.
A systematic review of COVID-19 epidemic models with endogenous human behaviour. What's next?
arXiv:2606.12465v1 Announce Type: new Abstract: Human behaviour and epidemic dynamics are intertwined, yet accounting for this feedback remains one of the key challenges of epidemiological modelling. The COVID-19 pandemic was an opportunity to overcome the traditional limitations of the field, raising expectations that data-informed endogenous approaches to behaviour modelling would advance substantially. To quantify the progresses made, we conducted a systematic review of SARS-CoV-2 transmission models endogenously including human behaviour in response to epidemic dynamics. The COVID-19 pandemic saw great strides in terms of the expanded use of empirical data in epi-behavioural modelling. However, it also showed shortcomings with respect to limited use of behavioural empirical data, lack of innovation in model structure, and limited engagement with other disciplines and decision-makers. Overall, our results suggest that identifying priorities in model design and behavioural data, building an adequate data collection infrastructure, leveraging on AI advancements, and fostering interdisciplinarity are strategies of utmost importance for pandemic preparedness.
Experimental updates on development of accelerator-driven ion source at TRIUMF to benchmark Ba-tagging techniques for future neutrinoless double beta decay searches
arXiv:2606.12468v1 Announce Type: new Abstract: Neutrinoless double beta decay ($0\nu\beta\beta$) could provide a way to probe physics beyond the Standard Model of particle physics. The proposed nEXO experiment aims to search for $0\nu\beta\beta$ in $^{136}$Xe using a tonne-scale liquid xenon (LXe) time projection chamber. The projected half-life sensitivity for nEXO for 10 years of livetime is $>$10$^{28}$ years. Efforts are ongoing to further suppress backgrounds and increase the experiment's sensitivity. One approach pursued is Ba-tagging, which entails extracting and identifying the daughter nuclide from the $\beta\beta$-decay of $^{136}$Xe, $^{136}$Ba. Once successful, this technique has the potential to separate background events from true $\beta\beta$ events. While different extraction and identification methods are being investigated by different groups, a Ba-ion source is required for testing, quantifying and optimizing them. An accelerator-driven ion source is currently being developed at TRIUMF, where radioactive ions will be be injected into and stopped in an LXe volume, collected electrostatically and detected using $\gamma$ spectroscopy. In this contribution, an experimental status update on the commissioning of this Ba-ion source at TRIUMF is provided.
A Type Theory of Sense: Witnessed Choice in Stratified Semantic Spaces
arXiv:2606.12504v1 Announce Type: new Abstract: We introduce TTS, a dependent type theory in which semantic composition is represented by horn filling and distinctions between possible completions are witnessed relative to explicit measurement regimes. TTS replaces globally canonical composition with regime-indexed indiscernibility and constructive apartness, allowing filler spaces to be classified as canonical when all completions are observationally connected and forked when two warranted completions are positively separated. Separation witnesses enter the calculus only through measurement contexts recording actual instrument outputs, yielding conservativity, provenance, and a no-fork-from-the-empty-record result. We prove that forks persist under refinement while canonicity may fail, and characterize exactly when an identification made by one regime can consistently coexist with a separation made by another. This framework supports a geometric account of Fregean sense as a choice of filler, reference as the boundary constraining that choice, and hyperintensional difference as measured apartness, while providing a falsifiable bridge to stratified representation spaces and branching behaviour in language-model generation.
Rubric-Guided Self-Distillation: Post-Training Without Rubric Verifiers
arXiv:2606.12507v1 Announce Type: new Abstract: Rubrics have emerged as an alternative to RLVR in open-ended domains where a single ground-truth final answer is not available. Existing rubric-based training methods rely on an LLM verifier that scores each rollout against rubrics. This introduces substantial training-time overhead, exposes optimization to verifier-specific biases, and reduces rubric feedback to a sparse end-of-trajectory signal. We propose Rubric-Guided Self-Distillation (RGSD), a verifier-free training method in which the base policy, conditioned on the rubric, serves as the teacher for the unconditioned student. RGSD distills the rubric-conditioned teacher distribution into the student token-by-token, replacing sparse trajectory-level rewards with dense per-token learning signals and removing the LLM judge from the training loop entirely. Across Qwen-2.5 (3B, 7B) and Qwen3-Thinking (4B, 8B) models on medical and science domains, RGSD achieves rubric satisfaction comparable to judge-based GRPO while using one on-policy rollout per prompt and no training-time verifier calls. Ablations show that raw rubrics provide a stronger teacher enrichment signal than self-generated reference responses, while a stronger GRPO judge can outperform RGSD in some settings, positioning RGSD as a complementary verifier-free alternative when verifier cost or reliability is the bottleneck.
Crossing the Validation Crisis: Cross-Validation Reduces Benchmarking Variance Surprisingly Well
arXiv:2606.12552v1 Announce Type: new Abstract: Modern machine learning progresses through empirical work, benchmarking new methods to evaluate relative performance. However, the statistical variability inherent to evaluation - exacerbated by the stochastic nature of many algorithms - often makes performance estimation unreliable due to the limited test samples available, leading to a validation crisis in which genuine advances are difficult to discern. In this work, we show that cross-validation improves markedly confidence when evaluating and comparing learning algorithm performances. We introduce the concept of sample gain, which quantifies the virtual data augmentation achieved by using multiple cross-validation splits to reduce benchmarking variance. Experiments on both synthetic and real-world datasets (histopathologic scans and NLP fine-tuning) demonstrate that multiple splits can substantially improve the reliability and stability of performance estimates, with diminishing returns often setting in later than expected. We also introduce a procedure to dynamically early-stop cross-validation by estimating from the first few folds if subsequent folds will bring large sample gains. Our findings highlight the value of pushing cross-validation on available samples to achieve robust and reliable benchmarking.
Feature-preserving Latent-EnKF for Data Assimilation of Flows with Shocks
arXiv:2606.12559v1 Announce Type: new Abstract: The ensemble Kalman filter (EnKF) is widely adopted for sequential data assimilation, but fails for solutions with discontinuities, such as shocks in compressible flows. Uncertainty in shock location induces multimodal ensemble statistics that violate the Gaussian assumptions underlying the EnKF, producing large-scale spurious oscillations in the analysis state. We introduce a feature-preserving latent-EnKF that performs the ensemble update in a learned low-dimensional latent space, where shock and flow features admit a smooth manifold representation, thereby preserving sharp features during EnKF analysis. The updated latent state is mapped back to physical state through a shared decoder for all ensemble members. The algorithm eliminates the member-specific ordered training and positivity flooring used in prior approaches. Numerical experiments on a Sod shock tube and Mach 2 shock interaction with a 2D cylinder, using sparse and noisy observations, show accurate feature recovery of shocks and contact discontinuities without spurious oscillations.
SCAR dynamics of adolescent substance use: peer influence, dropout, and bifurcation structure in a school-based model
arXiv:2606.12564v1 Announce Type: new Abstract: We develop a four-compartment susceptible--casual--addicted--resistant (SCAR) model for adolescent substance use in a high-school setting. The model divides students into susceptible non-users, casual or experimental users, students with sustained or substance-use-disorder (SUD)-level involvement, and resistant students in protective anti-use environments. It includes peer-driven initiation, escalation from casual to problematic use, protective peer influence, school disengagement, and partial re-entry after rehabilitation. Qualitative analysis and bifurcation diagrams show three main results. First, the return parameter \(\phi\) separates two regimes: when \(\phi=1\), the total population is conserved and interior equilibria may exist; when \(\phi<1\), problematic use causes net school-population loss, so positive scaled equilibria may not represent true endemic equilibria. Second, initiation and escalation are governed by distinct thresholds, meaning first use and progression to problematic use are dynamically different. Third, the model can exhibit multistability, including bistability between a substance-free state and a stable high-use state, so long-term outcomes may depend on initial conditions. These findings suggest that effective school policy should combine universal prevention, early intervention for casual users, targeted support for students at risk of problematic use, recovery-supportive environments, and strong school re-engagement pathways.
Hierarchical Framework of Runaway Electrons using Deep Learning
arXiv:2606.12567v1 Announce Type: new Abstract: We present an adjoint deep learning framework describing the evolution of fluid moments and the energy distribution of the runaway electron (RE) population. We demonstrate that a careful formulation of the adjoint problem allows for the temporal evolution of these quantities for arbitrary initial electron distributions, and in combination with a physics-informed neural network (PINN), we show that the resulting surrogates can resolve a broad range of plasma parameters. This combination of the adjoint formulation and rapid inference of neural networks enables orders of magnitude faster predictions of RE kinetics than traditional methods. Here, we detail the mathematical formulation and the design of three PINNs which recover the temporal evolution of the RE current, average energy and energy distribution. Predictions are validated against a traditional RE solver, with good agreement across a broad range of scenarios.