arXiv:2606.18020v1 Announce Type: cross Abstract: We study equilibrium clustering in a finite one-dimensional lattice gas of $L$ sites with periodic boundary conditions, as a minimal model for adsorption and binding on small ring-like substrates. Using a grand-canonical formulation with nearest-neighbor coupling, we derive exact finite-size expressions for the mean occupancy, the mean number of domain walls, and the mean number of clusters. Building on exact $k$-site correlation functions, we further derive expressions for the mean number of clusters of size $k$ and for two complementary size statistics: the cluster-size distribution, and the site-weighted cluster-size distribution. These observables characterize how spatial organization changes across attractive (cooperative) and repulsive (anticooperative) interactions, and highlight finite-size and parity-dependent effects of the underlying lattice, the latter being particularly pronounced near half filling in small systems. To access larger lattices without enumerating all $2^L$ microstates, we also develop a cluster-based combinatorial formulation in which configurations are classified by cluster counts and sizes, reducing the effective state space to a set whose size scales with integer partitions, $\approx e^{\sqrt{L}}$, rather than with $\approx e^{L}$. Taken together, our results provide exact benchmarks for finite periodic systems and suggest experimentally relevant cluster observables that complement occupancy-based measures of cooperativity, with particular relevance for binding on ring-like substrates for biological assemblies.
Science Journals
arXiv:2606.18025v1 Announce Type: cross Abstract: Improving the resolution of telescope systems will provide the opportunity to study new physical phenomena in previously unobserved environments. Spatial mode de-multiplexing (SPADE) based imaging is a promising and rapidly evolving technique for pushing the resolution of optical telescopes beyond the diffraction limit. A key application of this technique is for near-optimal hypothesis testing for the presence of secondary and extended sources in the sub-diffraction regime. We present the first demonstration of a binary-SPADE based hypothesis testing instrument deployed on-sky. In our proof-of-principle experiment, based on mode demultiplexing with a double clad fiber coupler, we demonstrate detection of a binary star system separated below the diffraction limit. We perform measurements in the photon-starved regime where no image can be formed by traditional direct imaging. We find the scaling of the system's type II error rate (the ``binary source miss" chance) was heavily limited by unbalanced loss in our double-clad fiber coupler when compared to the idealized quantum limits. Despite this the evaluated type II error is always lower than a perfect direct imaging measurement. We expect that if this instrument is scaled to larger aperture telescope systems the effects of atmospheric turbulence will further degrade this system's performance.
arXiv:2606.17557v1 Announce Type: new Abstract: Image restoration seeks to recover high-quality images from degraded inputs but becomes highly ill-posed under complex, mixed degradations. While unified all-in-one models are common, their performance declines as degradation complexity increases. Recent works adopt Chain-of-Thought (CoT) reasoning for multi-round restoration using specialized modules. However, this approach faces two key limitations: (i) increased computational cost due to multi-step processing, and (ii) weak modeling of interactions between degradations during stepwise inference. We introduce CoTIR, a universal image restoration framework that internalizes CoT reasoning within a single model. Concretely, we view image restoration as a specialized subtask of image editing, which implies that a large-scale pre-trained editing model provides a more favorable optimization starting point. Building on this, we fine-tune the model for restoration and further encode structured CoT-style reasoning into the learning objective via a differentiable formulation inspired by Lagrangian optimization, enabling holistic restoration without chaining specialized restorers. To facilitate training and evaluation, we further present CoTIR-Bench, a large-scale benchmark comprising 5.2 million samples with CoT-style reasoning traces. Extensive experiments on CoTIR-Bench and broad real composite degradation scenes show that CoTIR achieves stronger perceptual quality and more competitive fidelity than both all-in-one models and multi-round restoration methods. The source code is available at https://github.com/gy65896/CoTIR.
arXiv:2606.18010v1 Announce Type: new Abstract: Tabletop role-playing games provide a unique environment for interaction with artificial intelligence (AI) due to their complex and collaborative nature. We analyze Adventure AI, a podcast featuring human-AI interactions in Dungeons & Dragons play, to examine how AI is and can be used in tabletop role-playing gaming and how players perceive this use. We complete a qualitative analysis of three seasons of this podcast, from 2023 to 2025, reporting on the overarching themes of roles of AI, roles of humans, the evaluations and failures of AI, and its treatment as a person and character at the table. There are many aspects of the game where artificial intelligence succeeds, while there are others where it is less appropriate. This analysis gives a basis for future work on where artificial intelligence should and should not be used in gaming spaces.
arXiv:2606.18018v1 Announce Type: new Abstract: The paper deals with new formulations of a Lagrange interpolant polynomial based on the nodes of the well-known anti-Gauss rule. A first representation is given in terms of the classical Christoffel-Darboux kernel appropriately modified. The second one closely follows the barycentric form of the classical Lagrange polynomial, while the third formulation represents the interpolant as a combination of an orthonormal family of polynomials with respect to the discrete anti-Gauss inner product. A numerical test shows the performance of the explored forms.
arXiv:2606.17523v1 Announce Type: cross Abstract: Information-Geometric Optimization (IGO) provides a unified framework for black-box optimization by interpreting the adaptation of a search distribution as a natural gradient update. Despite its conceptual importance, the convergence theory of IGO remains limited: most existing results concern continuous-time idealizations such as the IGO flow, rather than discrete-time updates with non-infinitesimal learning rates. In this paper, we study discrete-time IGO in continuous spaces, formulated as natural gradient updates in the expectation-parameter coordinates of an exponential family. In particular, we analyze IGO over the multivariate Gaussian family on strongly convex quadratic objective functions. Our analysis covers a setting that simultaneously incorporates full covariance adaptation, a fixed positive learning rate, and quantile-based weights. In this setting, we prove that the covariance matrix converges to the zero matrix. We further show that the mean vector converges to the global optimum, provided that the condition number of the appropriately scaled covariance matrix is bounded at sufficiently frequent iterations. These results advance the convergence theory of IGO and help bridge the gap between the mathematical theory of IGO and practical covariance-adaptive search methods such as CMA-ES.
arXiv:2606.17562v1 Announce Type: new Abstract: LiDAR sensors are widely deployed in autonomous systems for 3D perception and safety-critical decision-making. We identify a previously unexplored attack surface in which dormant malware embedded in the LiDAR sensing pipeline remains inactive during normal operation and can be externally triggered after deployment, without requiring access to sensor hardware or networking at attack time. To operationalize this threat, we design malware capable of low-level point-cloud manipulation and embed it into LiDAR firmware. This malware was developed in a closed research test environment with vendor technical support, rather than by exploiting an inherent production supply-chain vulnerability. To selectively trigger attack activation, we design and implement an optical trigger that remotely activates the malware by delivering a modulated signal into the sensing environment. Once triggered, the malware performs real-time point cloud manipulation, and we demonstrate false object injection and real object suppression on static and mobile victim platforms. Our evaluation first establishes attack feasibility, including static operation at 300~ft and recorded drive-by runs reaching 35~mph. We then illustrate quantitatively that injected person-like artifacts can remain semantically detectable by a state-of-the-art 3D object detector. Finally, we demonstrate multiple modes of safety-critical impact on a deployed tactical autonomous vehicle. Together, these results highlight the need for stronger integrity guarantees throughout the LiDAR sensor development and deployment pipeline.
arXiv:2606.17564v1 Announce Type: new Abstract: Standardized evaluation protocols are indispensable for robust benchmarking in remote sensing, particularly as foundation features are increasingly transferred across diverse sensors and complex imaging geometries. In satellite multi-view reconstruction, conventional evaluations relying on unconstrained 2D global matching are often misleading. The Rational Function Model (RFM) and its Rational Polynomial Coefficients (RPC) dictate a curved, height-dependent epipolar geometry that render flat 2D search spaces physically inconsistent. We propose a geometry-faithful and reproducible protocol tailored for the RPC framework. Our approach integrates an RPC-projected 3D consistency metric with a geometry-constrained dense matching proxy, specifically evaluating whether similarity responses remain localized and unique under physically plausible search manifolds. A pivotal finding of our joint reporting strategy is the decoupling of semantic agreement and geometric localization: high cross-view similarity at a projected 3D point does not guarantee reliable matchability in practical inference. Our benchmark demonstrates that incorporating geometric constraints is fundamental to the problem definition in satellite imagery. Furthermore, we show that state-of-the-art 2D backbones remain remarkably competitive against specialized 3D-aware models when subjected to this RPC-consistent evaluation.
arXiv:2606.17300v1 Announce Type: new Abstract: In this article, we investigate the construction of linear codes over a finite ring $\mathcal{S}$, where $\mathcal{S}$ is taken to be an extension of a commutative non-unital ring $I$ of order $p^2$. Our approach is based on the defining set method. The defining sets considered in this work are derived from general simplicial complexes that may contain multiple maximal elements. We determine the parameters of these codes over $\mathcal{S}$ and study their Gray images. We also study the corresponding subfield-like codes. We show that these Gray image codes and subfield-like codes produce several families of divisible codes. Furthermore, we establish sufficient conditions under which these codes are minimal, optimal, and self-orthogonal. As applications of our results, we obtain several families of projective few-weight codes, and locally recoverable codes with small locality. We also study the minimal access structures of secret-sharing schemes associated with the duals of these minimal codes. Moreover, we construct several families of strongly regular graphs from projective two-weight codes and determine their parameters explicitly.
arXiv:2606.17375v1 Announce Type: new Abstract: Boron pebble aggregate was tested for the first time as a high-heat-flux granular plasma-facing material in a tokamak divertor. Exposures of up to $q_{\parallel} = \SI{80}{\MW\per\m\squared}$ incident heat flux were conducted in the DIII-D tokamak. Single protruding rods of pebble aggregate composed of sintered amorphous boron pebbles bound with carbon binder were mounted in the Divertor Material Evaluation System (DiMES) sample holders and exposed to L-mode lower single null (LSN) plasmas. Under these heat loads, significant boron dust emission from the boron spheres was observed, and this dust dominates the divertor boron ionization source. Only about half of the released boron was recovered locally as mm-sized particles; with the rest presumably lost mainly as dust into the plasma and vacuum chamber. Preliminary estimates suggest that the rate of surface recession of $\sim$1 cm/s in the pebble conglomerate within the plasma divertor is consistent with the recession rates observed in laser bench tests subjected to normal-incidence heat loads. Although core performance was not adversely affected by the high boron dust emission, future work will need to improve the boron pebble aggregate design to reduce boron dust emission at high heat loads.
arXiv:2606.17419v1 Announce Type: new Abstract: We develop approximation and generalization error estimates for multi-input neural operators, with the output error measured in Sobolev norms. In contrast to standard operator-learning settings with a single input function, our framework allows multiple input functions defined on possibly different domains, with different dimensions and Sobolev regularities. The derived rates explicitly quantify the contribution of each input space to the final error bound. In particular, in the balanced regime, the approximation and generalization rates are governed by the interaction between the input dimensions, regularities, and Sobolev orders, while the dependence on the model complexity retains a \(\log\log/\log\)-type structure. Our analysis provides a general theoretical framework for multi-input operator learning, including Sobolev training, and is applicable to operator learning problems arising from partial differential equations and scientific computing.
arXiv:2606.17971v1 Announce Type: new Abstract: Parametric PDE-constrained optimal control with pointwise state constraints requires repeated solution of restricted Schur-complement systems on parameter-dependent inactive sets. In a primal active-set method, each inactive-set system is symmetric positive definite, but the active set can change nonsmoothly with the parameter. The resulting operator may vary in dimension, sparsity pattern, and spectrum, limiting reuse of sparse factorizations, multigrid hierarchies, and Krylov information. We propose a reusable spectral-deflation strategy anchored to one full-domain reference Schur complement. Low reference eigenmodes are computed once, restricted online to each inactive set, and used as an A-DEF2 deflation basis for Jacobi-preconditioned CG. The framework also supports POD enrichment, Rayleigh-Ritz reselection, coarse-grid or analytical reference modes, and conditioning safeguards. Given the active set, the method preserves the high-fidelity inactive-set system and solves it to the prescribed CG tolerance; it accelerates the linear algebra rather than replacing the optimal-control solve with a surrogate. We explain the method through a spectral-coherence view, motivated by interlacing and perturbation arguments and assessed with principal-angle diagnostics. Across diffusion, convection-diffusion, nonlinear thermal, and conjugate-heat-transfer benchmarks, deflation reduces CG iterations by about 55 to 98 percent. GPU deployments also show wall-time gains over CPU sparse-direct and algebraic-multigrid baselines, because the reference basis is built once whereas competing solver structures are rebuilt per instance. Coarse-grid or analytical modes amortize the offline cost within a single parameter sweep; fine-grid eigensolves remain more precompute-limited. Timings isolate the inactive-set linear-solve kernel; reducing the active-set outer loop is outside the present scope.
arXiv:2606.17565v1 Announce Type: new Abstract: Massive controlled DC resources (CDCRs), such as battery energy storage systems, are connected to AC power systems through bidirectional inverters for power balance requirements. This study investigates converter-driven stability (CDS) issues in the sub-synchronous frequency range caused by large-scale bidirectional inverter-based stations (IBSs). The impacts of the AC and DC connections of IBSs on subsynchronous oscillations (SSOs) are compared by examining three factors: the number of CDCRs, power flow direction, and control parameters of the inverters. For AC connections, IBSs may induce instability as the number of CDCRs increases, regardless of the power flow direction. To maintain stability, the maximum power amplitude of the IBS is calculated. It is found that switching to DC connections can reduce these instability risks if the DC line resistance is much less than the AC line reactance. Moreover, the method of tuning control parameters is demonstrated to be more effective in improving power-related critical stability under DC connections. Therefore, The DC-IBS is preferred for high-voltage transmission. Finally, the conclusions are validated in power systems connected with both AC- and DC-IBSs under various network topologies and system scales.
arXiv:2603.18492v3 Announce Type: replace Abstract: Mixture-of-Experts (MoE) language models increase parameter capacity without proportional per-token computation, yet deployment still requires storing the full expert pool, making expert pruning important for reducing memory and serving overhead. Existing task-agnostic expert-pruning methods are typically calibration-dependent: they estimate expert importance from routing or activation statistics on a calibration set, making pruning decisions sensitive to calibration-data variation while introducing substantial preprocessing cost. We propose AIMER (\textbf{A}bsolute mean over root mean square \textbf{IM}portance for \textbf{E}xpert \textbf{R}anking), a simple calibration-free criterion that identifies more distinct experts by capturing the concentration pattern of expert weights, making it well suited for task-agnostic expert pruning. Across 7B to 47B MoE language models with distinct architectures and 16 diverse benchmarks, AIMER consistently delivers stronger capability balance across diverse tasks than existing calibration-free methods. Surprisingly, AIMER also achieves better balance than strong calibration-based expert-pruning baselines calibrated on the widely used task-agnostic C4 corpus, while requiring only 0.22--2.06 seconds to score all experts.
arXiv:2606.00024v3 Announce Type: replace Abstract: Long-context decoding in Large Language Models (LLMs) is constrained by the cost of accessing and processing the Key-Value (KV) cache. Despite evidence that attention outputs depend jointly on keys and values, most existing KV management methods rely on key-only pruning, since incorporating values incurs prohibitive overhead. In this paper, we propose Attention Run-time Termination (ART), a lightweight run-time mechanism that tracks accumulated attention outputs during kernel execution and terminates subsequent KV block accesses once further contributions become negligible. Rather than replacing KV selection, ART dynamically terminates redundant KV traversal on top of existing dense or sparse attention policies. We introduce a stability-based criterion that monitors both magnitude and directional changes of intermediate attention outputs and provideds a theoretical characterization of the resulting truncation error. Experiments on the LongBench and RULER Needle-in-a-Haystack tasks show that ART increases the generation throughput of existing KV-cache methods by up to 20%, without compromising the result quality.
arXiv:2606.18108v1 Announce Type: cross Abstract: We develop a text-to-SQL (structured query language) system based on large language models (LLMs) using in-context learning and apply it to the Automatic Learning for the Rapid Classification of Events (ALeRCE) astronomical database. ALeRCE is a community broker for the Zwicky Transient Facility and the Vera C. Rubin Observatory. The system enables users to query the database in natural language (NL) and generates executable SQL queries. To develop and evaluate the system, we constructed a dataset of 110 NL/SQL pairs. We propose a step-by-step generation framework comprising four modules: schema linking, query classification, prompt decomposition, and self-correction. The performance of thirteen LLMs is evaluated using in-context learning and prompt engineering techniques. Text-to-SQL performance is assessed using the perfect-match (PM) rate for row identifiers (e.g., object identifiers) and column identifiers (i.e., column names). The proposed step-by-step framework consistently outperforms a direct-inference baseline, while the self-correction module consistently reduces execution errors. For Claude Opus 4.6, PM performance on row (column) identifiers is high for simple queries, reaching 0.97 (0.94), and decreases with query complexity to 0.44 (0.72) for medium queries and 0.59 (0.49) for hard queries. Among the thirteen evaluated models, the best-performing LLMs for the text-to-SQL task are Claude Opus 4.6, Gemini 2.5 Pro, Gemini 3 Flash, and GPT-5.2-Codex.
arXiv:2606.18150v1 Announce Type: cross Abstract: Channel state information (CSI)-based neural positioning learns a mapping from CSI measurements to user equipment (UE) positions using neural networks. However, most existing performance evaluations utilize randomly partitioned train/test CSI-dataset splits, which fail to reflect the generalization requirements of practical deployments and present optimistic results. In this paper, we study the spatial and temporal generalization of neural positioning with standard-compliant Wi-Fi and 5G NR systems for three real-world CSI datasets acquired in indoor and outdoor environments. We assess generalization with two different architectures, a conventional multilayer perceptron (MLP) and a novel transformer architecture, to unseen spatial regions, unseen UE trajectories, and CSI measurement campaigns separated by one week. Our experiments show that both architectures generalize well in space and time, and the proposed transformer consistently outperforms the MLP in positioning accuracy while requiring fewer model parameters.
arXiv:2606.17399v1 Announce Type: new Abstract: When small transformers grok modular multiplication, prior work reports that the learned embedding has a "dense" Fourier spectrum requiring all frequencies. This contrasts with modular addition, where only a sparse set of key frequencies suffices. We show this density is an artifact of analyzing in the wrong basis. The natural Fourier transform for multiplication is not the standard additive DFT but the multiplicative character transform, which decomposes functions on the multiplicative group $(\mathbb{Z}/p\mathbb{Z})^*$ into its irreducible representations. Applying this transform to a grokked transformer trained on $a \cdot b \bmod 113$, we find the embedding spectrum becomes highly sparse (Gini coefficient 0.58 vs. 0.07 in the additive basis) with only 4 key frequencies carrying significant energy. Furthermore, 96.9% of MLP neurons are cleanly tuned to a single multiplicative frequency, and neuron activation heatmaps reveal 2D-periodic structure when reordered by the discrete logarithm. These results demonstrate the transformer reduces multiplication to addition in discrete-log space, implementing a "Discrete-Log Clock" algorithm analogous to Nanda et al.'s Clock algorithm for addition. The methodology generalizes: matching the analysis basis to the algebraic structure of the task reveals interpretable structure where standard tools see noise.
arXiv:2606.17071v1 Announce Type: new Abstract: Einstein field equations allow cosmological dynamics to depart from the Friedmann-Lemaitre-Robertson-Walker (FLRW) idealisation in several physically different ways. Matter may become spatially inhomogeneous, the local expansion scalar may vary across a hypersurface, the expansion may acquire anisotropic components through shear, and the free gravitational field may be encoded in nonzero Weyl curvature. The key question is not only how far a model is from FLRW, but which geometric mechanism is responsible. A single departure from FLRW number cannot distinguish these mechanisms. This paper introduces a compact geometric diagnostic framework that keeps them separate while using standard quantities in general relativity. The framework is observer-explicit and domain-explicit, intended as a practical tool for comparing analytic and numerical solution families rather than as a new invariant classification of spacetime. Buchert's kinematical backreaction is retained as a derived explanatory quantity rather than a separate axis, since it is already fixed by the expansion-variance and shear contributions. A single curvature normalisation is used for all Weyl diagnostics. The method is tested on six benchmarks, namely FLRW, Bianchi-I, Kasner, Lemaitre-Tolman-Bondi dust, scalar-perturbed FLRW, and tensor-perturbed FLRW. These benchmarks occupy distinct regions of the diagnostic space, and the magnetic Weyl contribution appears only in the tensor case. The classification remains stable under changes of perturbation amplitude, spatial resolution, averaging domain, constraint reliability, and a leading-order observer tilt. The curvature expressions for the exact benchmarks are verified symbolically against metric-derived Weyl invariants, and the supporting computer code, numerical results, tables, and figures are publicly available.
arXiv:2606.17405v1 Announce Type: new Abstract: Clinical decision support AI systems (CDSASs) must adapt to evolving patient conditions in real-time while adhering to strict safety constraints. We present an online adaptive framework that integrates Treatment Effect (TE) estimation to quantify clinical benefits, a patient Digital Twin (DT) to simulate treatment trajectories, and Reinforcement Learning (RL) for sequential decision-making. The AI system is initially trained on historical medical records and operates in a continuous learning loop. To ensure safety, a rule-based module monitors vital signs and blocks contraindicated treatments. Cases with strong internal model disagreement are flagged for clinician review, simulated in our experiments via a pre-trained outcome model. We validate our framework using both a synthetic clinical simulator and a real-world ovarian cancer dataset from The Cancer Genome Atlas (TCGA). In both simulated and clinical settings, our method demonstrated superior effectiveness and stability in recommending treatments compared to standard computational baselines. Furthermore, the AI system maintains low latency and requires expert consultation for only a minority of cases in our experimental validation, demonstrating its potential as a safe, clinician-supervised tool for personalized medicine that continuously improves through practical use.
arXiv:2606.17376v1 Announce Type: new Abstract: Respiratory-rate (RR) monitoring is a critical component of remote triage and victim assessment in emergency response, disaster recovery, and infectious-disease scenarios, where minimizing physical contact can reduce responder risk and improve operational safety. However, field deployment of contactless RR monitoring remains challenging due to variable illumination, posture changes, platform heterogeneity, and the impracticality of wearable sensors in hazardous environments. In this paper, we present a modality-adaptive contactless RR monitoring framework for heterogeneous mobile robots with onboard edge computing. The proposed system combines brightness-adaptive sensor selection across RGB, thermal, near-infrared (NIR), and low-light cameras, keypoint-guided chest ROI extraction for posture-robust monitoring, and a signal-quality-index (SQI)-based filtering mechanism for reliable respiratory estimation. We implement and evaluate the framework on three robotic platforms spanning quadruped and wheeled locomotion and multiple edge-computing architectures. Experiments conducted across diverse lighting conditions, subject poses, and robot-to-subject distances demonstrate that the framework generalizes across platforms without per-platform algorithmic retuning, while revealing modality-specific operational boundaries. RGB provides the broadest coverage up to 8m, NIR remains effective up to 6m, thermal is reliable only at short range, and low-light sensing supports monitoring in complete darkness up to 8m. Overall, the results demonstrate the feasibility of multimodal contactless RR monitoring on mobile robots and support its use as a foundation for autonomous triage and victim assessment in hazardous search-and-rescue settings.
arXiv:2606.17387v1 Announce Type: new Abstract: In recent decades, privacy-enhancing technologies (PETs) have been recognized as a means of meeting regulatory and user privacy requirements in software systems that process personal data. Despite substantial research efforts, support from regulators, contributions by large technology companies such as Google and Microsoft, and growing interest among software practitioners, the practical adoption of PETs remains limited. Existing research consistently identifies recurring challenges to PETs adoption in SE, such as technical complexity and insufficient training. Despite ongoing research efforts, these challenges largely remain unresolved in practice. In this industrial challenge paper, we apply a practical, requirements engineering (RE)-driven perspective to examine challenges to PET adoption across multiple stakeholder groups (PET developers, integrators, and adopters) as well as across different disciplinary perspectives (engineering, law, and business). We argue that RE can facilitate the adoption of PETs by systematically addressing each of the complementary engineering, business, and legal viewpoints on privacy. Neglecting challenges in any of these viewpoints (e.g., the impact of PETs on software architecture, their business implications, and their contribution to regulatory compliance) can increase the impediments or even lead to implementation failure. In practice, explicit specification of these viewpoints within RE can enable meaningful coordination among stakeholders to more effectively realize the benefits of PETs in software engineering.
arXiv:2606.17588v1 Announce Type: new Abstract: Several studies have examined the use of large language models (LLMs) for title-abstract screening in systematic reviews (SRs), reporting mixed accuracy. However, questions of reliability remain largely unaddressed. In this study, we go beyond quantitative LLM-human agreement metrics and qualitatively investigate how and why LLMs fail. We also propose actionable recommendations. We analyzed disagreements between LLMs and researchers across six software engineering SRs and over 1,000 primary study papers. For each SR, papers were screened independently by human experts and LLMs in zero-shot mode, resulting in Kappa values ranging from 0.52 to 0.77. Qualitative analysis suggests that human-LLM disagreement results from recurring, identifiable causes, such as boundary ambiguity in key terms, keyword overemphasization, and incorrect topic inference. Based on these findings, we propose recommendations such as validating semantic understanding before deployment, running multiple LLMs, and focusing validation efforts on borderline cases. Future studies are needed to validate the impact of our recommendations, and community efforts are needed to develop normative guidelines on LLM usage in SRs.
arXiv:2606.17078v1 Announce Type: new Abstract: Binary collision approximation (BCA) codes are potentially powerful tools to simulate ion irradiation ejecta properties, such as the composition and the angular and energy distributions of the sputter yield. However, recent advances in the sputtering of minerals have highlighted the low predictive fidelity of BCA codes such as SDTrimSP when compared to experimental measurements. We demonstrate how a sputtering model that underestimates the forward sputtering on a flat surface at large ion incidence angles from surface normal will lead to an erroneous result for rough and porous surfaces, where most ejected particles are directed along the surface normal. We demonstrate how this is the case for an existing model, which reliably predicts sputtering mass yields from a flat enstatite surface but fails to accurately reproduce the angular distribution of sputtered particles. We then compare this to a BCA model incorporating higher surface-binding energies$-$based on a molecular dynamics description of plagioclase$-$which underestimates mass yields but significantly reduces back-sputtering and better reproduces laboratory sputter angle distributions measured at large ion incidence angles. We conclude that the BCA model cannot simultaneously reproduce both the sputter yield and the sputter angle distribution arising from He irradiation of mineral targets, either due to the inherent geometric simplicity of the BCA or because the model neglects yield-enhancing processes such as molecule and cluster sputtering. This demonstrates a structural limitation of current BCA-based models when realistic surface morphologies are considered, rather than a problem that can be resolved by parameter tuning alone.
arXiv:2606.16070v2 Announce Type: replace Abstract: World-model synthesis aims to turn interaction experience into an internal model of environment dynamics. Existing symbolic approaches often fit observed transitions or mixtures of local rules, but they do not produce a complete executable program that can run independently of the real environment. We present Mind-Studio, a framework that synthesizes executable pygame-style world models from state-action-next-state trajectories using large language models. Mind-Studio combines entropy-selected traces with a lightweight game skill file containing object, action, and static scene information extracted from screenshots. We evaluate synthesis quality with a K-step lookahead fidelity protocol that compares generated world-model rollouts against Real-ALE rollouts from the same state. On Montezuma's Revenge, Mind-Studio improves chosen-action next-state prediction from 0.3% for PoE-World to 48.7% while verifying 5 of 8 subgoals; across Alien, Assault, and Skiing, it achieves stronger branch-level fidelity than prior learned lookahead sources.