arXiv:2606.31411v1 Announce Type: new Abstract: Rapid advancements in generative speech technology have compromised the reliability of voice biometrics. While current spoofing detectors excel when assessed under in-domain conditions, generalisation to out-of-domain settings is often poor. We show that this can be due to linguistic bias. A reliance on linguistic cues observed in training data can then compromise robustness to cross-data. We propose a linguistic-invariant spoofing detection framework utilizing teacher-student adversarial learning. The linguistic-aware teacher model, pre-trained on linguistic content of an external dataset, guides the student detector via gradient reversal to minimize the linguistic information. To prevent the inadvertent removal of non-linguistic cues, we incorporate a Variational Information Bottleneck to enable suppression of principal cues. Across nine DF Arena datasets, our method achieves up to a 36.2% relative reduction in the EER compare to the baseline.
Science Journals
arXiv:2606.31247v1 Announce Type: new Abstract: Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame rates (e.g., 25 or 12.5 Hz), ignoring the time-varying information density of speech and offering no flexibility to trade off quality for speed at inference time. Recent audio tokenizer research has proposed dynamic frame rate speech coding, which exploits this non-uniformity and enables two new capabilities: very low average frame rates and frame rate controllability. However, this technique has not yet been applied to SLMs. We introduce Flexible Spoken Language Model (FlexiSLM), the first SLM that supports dynamic and controllable frame rates on both speech input and output. Using dynamic frame rate representations, FlexiSLM outperforms fixed-frame-rate 7B models including Qwen2.5-Omni and Kimi-Audio at its high-quality operating points. We further verify that FlexiSLM can be accurately steered down to 4.0 Hz; at 6.25 Hz, it roughly halves inference time relative to 12.5 Hz while retaining strong speech-to-speech quality. Audio samples are available at https://flexislm.github.io .
arXiv:2606.31379v1 Announce Type: new Abstract: We introduce P3MaZe, a real-space particle-mesh electrostatic method that combines the standard short-range/long-range decomposition of Particle-Particle Particle-Mesh (P3M) electrostatics with the Mass-Zero constrained dynamics (MaZe) framework. In this formulation, the smooth long-range electrostatic potential is represented on a mesh as a zero-inertia auxiliary field, while the discretized Poisson equation is enforced as a holonomic constraint during molecular dynamics. By retaining the standard P3M decomposition, P3MaZe preserves the systematic accuracy controls associated with the real-space cutoff, the Ewald splitting, the mesh spacing, and the charge-assignment procedure, while replacing the conventional multigrid Poisson solver by a constrained correction problem. The method is validated for molten NaCl and simple point-charge flexible water (SPC/Fw). Structural, translational, collective, and rotational dynamical observables are in quantitative agreement with those obtained with established electrostatic methods, including real-space P3M, and Ewald summation. The constrained formulation consistently requires fewer multigrid iterations than the corresponding real-space P3M solver while retaining the expected linear scaling with system size. These results establish P3MaZe as a promising new direction for scalable real-space electrostatics in large-scale molecular simulations.
arXiv:2606.31680v1 Announce Type: new Abstract: Despite advances in indoor scene generation, synthesizing coherent building exteriors consistent with generated interiors remains largely unexplored. Existing methods can generate floor plans and wall layouts but typically stop at a structural shell, lacking stylistically consistent facades and roofs. Completing these exteriors is challenging because the footprint, wall geometry, and opening semantics must remain fixed-constraints that unconstrained generative models often violate. We introduce ShellMaker, a language-guided exterior completion framework that operates under these structural constraints. Given a building scaffold and a text style prompt, ShellMaker generates a complete exterior mesh with PBR materials by combining parametric roof generation, LLM-based part-aware prompt refinement, joint wall-roof material retrieval, and geometry-aware assembly. Operating on a format agnostic scaffold representation, ShellMaker generalizes to indoor generators, CityGML, and CAD inputs, while maintaining structural consistency and improving architectural coherence over retrieval and unconstrained generative baselines. The project page is available at https://ruiqixu37.github.io/ShellMaker_web/
arXiv:2606.31200v1 Announce Type: new Abstract: Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing VLM-based methods rely on visual similarity for object matching, neglecting physical affordances such as handle graspability and material fragility, and operate open-loop without spatial reasoning or failure recovery, limiting their effectiveness when objects are densely packed or physically diverse. We present Agentic RAG-VLM, a unified framework that bridges VLM-based semantic understanding and physically grounded grasp execution by integrating retrieval-augmented generation (RAG) with vision-language models (VLMs) and agentic self-reflective planning. Agentic RAG-VLM introduces three tightly coupled components: (1) a Hierarchical Affordance-Aware RAG (HAA-RAG) that encodes four-dimensional affordance descriptors, including type, material, fragility, and graspable region, and retrieves strategies by functional affordance compatibility rather than visual appearance; (2) a Scene Graph Constraint Reasoner that constructs spatial relationship graphs from VLM perception and translates proximity, occlusion, and support constraints into concrete grasp parameter adjustments; and (3) an Agentic Self-Reflective Pipeline with a 14-type failure taxonomy and three-level adaptive retry for closed-loop grasp refinement. Evaluated on a 12-task benchmark spanning single-grasp, interactive, and long-horizon scenarios with 360 trials per configuration, Agentic RAG-VLM achieves 78.3 percent overall success, a 53.3 percentage-point absolute gain over VLM-only baselines, demonstrating that affordance-aware retrieval, scene graph reasoning, and agentic recovery are jointly essential for robust manipulation.
arXiv:2606.31479v1 Announce Type: new Abstract: Pulse-resolved spectral phase measurement of mid-infrared (MIR) pulses is essential for many applications, from precise waveform control to ultrafast quantum optics. However, conventional MIR pulse characterization techniques are typically limited to sub-kHz-rate operation, leaving a substantial speed mismatch with MIR sources operating at kHz or MHz rates. Here, we introduce time-stretch upconversion-based mid-infrared pulse evaluation (TSUBAME), a technique that enables pulse-to-pulse spectral phase characterization of ultrashort MIR pulses at the laser repetition rate. TSUBAME combines MIR-to-NIR (near-infrared) upconversion, time-stretch, and spectral interferometry to achieve scan-free high-speed spectral phase measurements. We validated the technique by measuring MIR pulses spanning 4.98-5.30 um while introducing well-defined dispersion, obtaining excellent agreement with theoretical predictions. Operating at a measurement rate of 1 MHz, TSUBAME achieves the fastest single-pulse-resolved spectral phase characterization of MIR pulses reported to date. As a further demonstration, we captured dynamic spectral phase variations on a microsecond timescale. TSUBAME provides a powerful tool for real-time monitoring and optimization of high-repetition-rate MIR pulses, with potential applications in strong-field physics, high-harmonic generation, and coherent molecular control.
arXiv:2606.31700v1 Announce Type: new Abstract: Biological neural circuits obey Dale's principle: each neuron's synapses are uniformly excitatory or inhibitory. Artificial networks that respect this constraint must coordinate separate excitatory and inhibitory populations, fundamentally changing how credit is assigned during learning. Several biologically plausible learning rules avoid backpropagation's weight transport requirement, but it has been difficult to achieve strong performance under Dale's principle beyond MNIST. Error Diffusion (ED) was originally proposed in a dual-stream excitatory/inhibitory architecture, where learning is driven by routing global error signals to all layers without transporting transposed forward weights or relying on random feedback matrices. Whether such a rule can scale under Dale's principle across both supervised classification and reinforcement learning remains unknown. Here, we introduce modulo error routing to extend Error Diffusion beyond binary classification, and show that a dual-stream excitatory/inhibitory architecture trained with this method achieves 96.7% on MNIST and establishes a 61.7% baseline on CIFAR-10, demonstrating that representation learning is possible even when strictly enforcing Dale's principle. For the classification setting, we introduce three domain-specific innovations: layer-specific sigmoid widths, batch-centered class error signals, and asymmetric initialization, and ablation analysis reveals that their relative importance reverses between MNIST and CIFAR-10, exposing task-dependent credit-assignment bottlenecks invisible to single-benchmark evaluation. In reinforcement learning, we integrate ED with Proximal Policy Optimization (PPO) and evaluate it on continuous-control tasks in Google Brax and on Craftax, an open-ended exploration task. We show that ED-PPO achieves competitive performance relative to Direct Feedback Alignment, a backpropagation-free baseline.
arXiv:2606.31407v1 Announce Type: new Abstract: Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions. Common entropy-based methods, such as Semantic Entropy (SE), rely on output diversity. Yet our analysis shows that overconfident visual embeddings suppress output diversity under stochastic decoding, causing SE to underestimate uncertainty in such cases. Recent methods instead probe output diversity through input perturbations, including textual paraphrasing or joint text-image perturbations, and show improved performance. We study these approaches and reveals that the resulting variability is often dominated by textual changes rather than visual evidence, causing uncertainty estimates to reflect prompt sensitivity rather than visual ambiguity. We therefore propose Visual Semantic Entropy (VSE), which perturbs only the image to probe nearby visual variations while keeping the text query fixed. VSE measures uncertainty by clustering generated answers into semantic prototypes and computing the mass-weighted dispersion among them. Extensive evaluation across five modern vision-language models and five diverse VQA benchmarks demonstrates that VSE effectively captures visual ambiguity, establishing a new state-of-the-art for VLM uncertainty estimation.
arXiv:2606.31478v1 Announce Type: new Abstract: Autonomous research agents can now draft hypotheses, write code, run experiments, and produce papers, but they remain brittle when experiments fail. Under the prevailing paradigm, failure recovery is usually delegated to a single free-form reflection: a rich trajectory of metrics, logs, and design choices is compressed into one verbal critique, which often leads either to localized trial-and-error or to hard pivots that discard useful context. We propose SAGE, a Self-correcting, Autonomous, Grounded Experimenter, to tackle this failure-recovery bottleneck. Its core mechanism, Multi-Hypothesis Failure Attribution (MHFA), treats recovery as a structured causal diagnosis. By analyzing dynamic trajectory features, MHFA systematically generates multiple evidence-grounded explanations for a failure, independently evaluates their severity, and deterministically routes the verified root cause to the correct intervention level (hypothesis, experimental design, or implementation). To guarantee scientific honesty, SAGE further employs a grounded reporting mechanism that explicitly constrains drafted results to actual measured values, redacting hallucinated numbers. On a 12-topic, 5-domain benchmark, SAGE increases metrics-bearing outputs from 42% to 92% over a reflection baseline, improves artifact quality from 5.00 to 6.75/10, and blindly outscores AI-Scientist-v2 (52.0 vs. 48.2), with gains concentrated in code development and execution. While fully autonomous scientific writing and generating conference-ready papers remain notoriously difficult open problems for the entire field, SAGE successfully produces significantly more reliable and higher-quality scientific artifacts. Ultimately, by coupling structured recovery with explicit grounding constraints, SAGE significantly outperforms monolithic reflection paradigms, establishing a highly trustworthy foundation for future autonomous research.
arXiv:2606.31408v1 Announce Type: new Abstract: Large Language Models (LLMs) have rapidly proliferated, driving widespread adoption of AI applications. Most deployments rely on centralized infrastructures such as Microsoft Azure, Google Cloud, or AWS, requiring users to share sensitive data and training or fine-tuning code. This dependence raises significant security and privacy concerns, as cloud providers must be trusted to ensure confidentiality and integrity. Trusted Execution Environments (TEEs) e.g., Intel SGX/TDX, AMD SEV-SNP, and ARM CCA have been introduced to mitigate these risks. More recently, NVIDIA has developed GPU TEEs (e.g., H100/H200), yet comprehensive evaluations of end-to-end workflows that integrate CPU and GPU TEEs remain limited. Critical aspects, including performance overhead, remote attestation, and security guarantees for AI/LLM applications, have not been sufficiently studied. This paper addresses this gap by presenting an end-to-end workflow that combines CPU and GPU TEEs. We propose mechanisms to ensure confidentiality and integrity at both the VM level (via Intel TDX and AMD SEV-SNP) and the application level, highlighting vulnerabilities such as Kubernetes administrators' ability to access confidential VM contents. Finally, we evaluate the performance overhead of our system using industry benchmarks, focusing on configurations that integrate Intel TDX with NVIDIA H200 GPUs.
arXiv:2606.31433v1 Announce Type: new Abstract: The modulation of galactic cosmic rays, driven by the evolution of the heliospheric magnetic field, strongly influences the intensity of cosmic rays reaching near-Earth space. Characterizing this process is crucial both for advancing our understanding of cosmic-ray transport and for assessing radiation exposure and related hazards in space environments. Here we present a newly developed forecasting framework built on a numerical description of charged particle transport in the heliosphere and its dependence on solar activity, designed for the long-term forecasting of galactic cosmic-ray fluxes. It solves a one-dimensional, spherically symmetric form of the Parker transport equation, including diffusion, solar-wind advection, and adiabatic energy losses. The model has been validated using multi-species flux measurements from space-based experiments: PAMELA, AMS-02, and ACE. Its strategy is based on Hilbert-Huang transform filtering and cross-correlation between delayed solar proxies and effective model parameters. Our charge-sign- and rigidity-dependent parametric description of the diffusion-advection processes yields good overall agreement with the data, as shown by the reconstruction uncertainty. The robustness of this approach is validated across a broad set of multichannel datasets covering different particle species, energy ranges, and phases of solar activity, supporting its applicability to space radiation monitoring and forecasting. Furthermore, when coupled with solar-proxy forecasting models, it enables decadal-scale predictions of galactic cosmic-ray fluxes, thereby supporting long-term planning and radiation-risk assessment for future space missions.
arXiv:2606.31338v1 Announce Type: new Abstract: Recent music audio-language models achieve high accuracy on instrument question-answering benchmarks, but it remains unclear whether this reflects robust audio grounding or benchmark-specific shortcuts. In this paper, we introduce an OpenMIC-derived diagnostic benchmark sequence for instrument grounding in music audio-language models, extending binary instrument-presence QA to genre-prior-reduced examples, confusable instrument discrimination, longer audio context, and temporal localization. Across these settings, high binary QA accuracy often fails to predict model behavior: models can exhibit option-position bias, confusable-instrument errors, and temporal response bias. These results suggest that instrument grounding should be evaluated with multi-axis diagnostic benchmarks rather than a single aggregate accuracy.
arXiv:2606.31414v1 Announce Type: new Abstract: Packet routing on scale-free networks faces a fundamental trade-off: shortest-path routing is efficient at low demand but funnels traffic through hubs and jams early, whereas congestion-aware routing postpones jamming at the price of a sharper collapse. Since neither paradigm dominates across the full range of traffic load, here we ask whether the appropriate balance can emerge endogenously rather than being imposed by design. To answer this, we recast adaptive packet routing on networks as an evolutionary game letting a heterogeneous population of strategies compete for prevalence under selection pressure generated by their own performance. We study this competition under two formalisms (strategy anchored to the packet or to the generating node), global and local update rules, and two payoff metrics. Across every implementation the evolutionary dynamics yield the same outcome: the jamming transition is delayed relative to shortest-path routing while the violent collapse of fixed congestion-aware routing is avoided. This improvement emerges spontaneously, without centralized coordination or global information. Crucially, under local update rules, the node-level volatility of strategy choices peaks sharply at the transition, furnishing a purely local early-warning signal of imminent jamming that requires no global monitoring.
arXiv:2606.31722v1 Announce Type: new Abstract: Automatic speech recognition (ASR) systems often perform poorly in dysarthric speech, limiting their usefulness to affected speakers in everyday communication. This paper presents a personalized ASR system for a dysarthric speaker, built by adapting a foundation ASR model to speaker-specific data. Using the TEQST tool, we collected 92 hours of read speech and later added 8.8 hours of user corrections gathered through a deployed mobile application. Starting from Whisper, fine-tuning reduced word error rate to 15.8% with only 1.4 hours of adaptation data, reached 10.7% with 22.5 hours, and achieved the best result of 9.7% when using all available data including the corrections. Using LoRA adaptation and/or Qwen3-ASR as foundation model performed worse in this setting. The results show that personalized fine-tuning can make foundation ASR models substantially more effective for dysarthric speech and suitable for practical deployment.
arXiv:2606.31415v1 Announce Type: new Abstract: Embedded systems that combine hardware interrupts, buffering, and distributed communication are often perceived as inherently asynchronous and difficult to analyze. However, such systems can exhibit a deterministic timing structure when modeled using explicit logical-time semantics. This paper presents a Global Navigation Satellite System (GNSS) correction-data pipeline implemented as a federated Lingua Franca (LF) application. The federated LF program decomposes the end-to-end pipeline into reactors with explicit time semantics, including a time-triggered GNSS receiver, a UART interrupt stream derived from baud rate and First-In First-Out (FIFO) buffer characteristics, a periodic forwarding task, and downstream processing with jitter monitoring. Federated execution and runtime logs validate the analytically derived deterministic timing structure-including interrupt cadence, ring-buffer evolution, packetization behavior, and physical--logical jitter-yielding a reproducible and predictable timing profile.
arXiv:2606.31740v1 Announce Type: new Abstract: Evolutionary game theory provides a framework by which to study the emergence of cooperation in a population of self-interested actors. In such a framework, players' decisions on whether or not to cooperate evolve according to decision rules called population dynamics. However, often games are studied under the assumption that all individuals play under the same conditions, and many common choices of update rule are not well suited for a heterogeneous population. In this paper, we categorise and compare four different population dynamics in such a population as ``extrinsic'', where players learn by looking outward at the payoffs of other players, and ``intrinsic'', where players look inwardly at their own attributes or potential payoffs. We show that extrinsic population dynamics admit a ceiling on the rate of cooperation which can be exceeded by intrinsic population dynamics, and demonstrate this using the public goods game with heterogeneous contributions.
arXiv:2606.31404v1 Announce Type: new Abstract: Human swarm intelligence demonstrates remarkable collective accuracy but faces scalability constraints in cost, coordination, and time. We investigate whether large language models (LLMs) can approximate swarm intelligence effects through artificial swarms, addressing a critical gap in understanding AI-based aggregation mechanisms. We conducted a controlled experiment with 960 manually executed prompts across three proprietary models (GPT-5, Gemini 2.5 Pro, Claude Sonnet 4.5), testing intra-model sampling and inter-model aggregation on eight estimation tasks. Results reveal consistent error reduction through intra- and inter-model aggregation, with significant error reductions up to 37 percentage points in MAPE across different aggregation strategies. We observed small to large effect sizes for positive correlations (Spearman's $\rho=0.242-0.568$, all $p<0.001$) between relative confidence interval widths and relative estimation errors, suggesting LLMs possess metacognitive awareness when assessing uncertainty. We discuss implications for research and practice, providing actionable insights for deploying LLM swarms in organizational decision-making.
arXiv:2606.31769v1 Announce Type: new Abstract: We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees together with data-dependent regret bounds. Recent work (Dann et al., 2023; Li et al., 2026) has shown that policy optimization can adapt to both adversarial and stochastic losses with first-order, second-order, and path-length bounds, but only under known transitions, leaving open whether such data-dependent guarantees are achievable by policy optimization when the transition kernel is unknown. We resolve this by developing a new algorithm based on optimistic follow-the-regularized-leader that attains these guarantees under unknown transitions. The key ingredient is a new design of optimistic $Q$-function estimators together with a data-dependent transition bonus that controls estimator bias through the loss-prediction error. Our analysis further identifies an unavoidable transition-dependent complexity term that captures the intrinsic cost of estimating the transition kernel. As a result, we obtain first-order, second-order, and path-length bounds with the transition-dependent complexity term while simultaneously achieving gap-dependent $\mathrm{polylog}(T)$ regret in the stochastic regime.
arXiv:2606.31781v1 Announce Type: new Abstract: Log parsing is a fundamental step in automated log analysis, transforming raw system logs into structured event templates for downstream tasks such as anomaly detection and system monitoring. Existing log parsing methods range from rule-based and clustering-based approaches to neural models that learn semantic representations from log messages. However, neural approaches typically rely on dense matrix multiplications, which can result in high computational cost and energy consumption. This paper presents SpikeLogBERT, a spiking neural network framework for energy-efficient log parsing. The proposed model integrates a spiking transformer architecture with knowledge distillation from a BERT teacher model, enabling spike-driven computation while preserving semantic representation capability. By leveraging sparse spike activations and event-driven processing, the number of active operations during inference can be significantly reduced. As an initial benchmark study, experiments on the HDFS dataset demonstrate that SpikeLogBERT outperforms ANN-based neural log parsing models with a parsing accuracy of 0.99997, while reducing estimated theoretical energy consumption by up to 62.6% under standard 45nm CMOS assumptions.
arXiv:2606.31881v1 Announce Type: new Abstract: This paper presents two logic systems, SCP and its extension SCP1, to distinguish between different types of impossibility. The semantics use a stratified structure that partitions worlds into logic-normal worlds (N) and anti-logic worlds (I). By defining a metaphysical accessibility relation within N, the systems separate metaphysical impossibility from logical contradiction. This allows for logical reasoning to be maintained even when dealing with metaphysical impossibilities. To address the challenge of vacuism, SCP1 implements a non-empty constraint on the selection function, ensuring that counterpossibles with impossible antecedents are not trivially true but depend on the connection between the antecedent and the consequent. We provide proofs for the soundness, completeness, and decidability of both systems. Finally, we indicate the possibility of applying this stratified approach to other modal domains, such as deontic or epistemic logic.
arXiv:2606.31891v1 Announce Type: new Abstract: In their prior work, Wang and Fan proposed conditional knowing-value logic and provided a complete axiomatization. However, in natural language scenarios and logic puzzles, knowing-value reasoning often appears together with arithmetic operations, which motivates us to enrich knowing-value logic with arithmetic function symbols. In this paper, we extend the language of conditional knowing-value logic with equality and the successor function. Due to the failure of compactness over the class of standard models, we additionally introduce non-standard models to facilitate the technical analysis. Our main results establish the finite model property and provide an axiomatization that is strongly complete with respect to the class of non-standard models and weakly complete with respect to the class of standard models. Furthermore, we extend our logic with public announcement operators and use the resulting system to formalize and solve the "Consecutive Numbers" puzzle. This work provides a novel framework for integrating epistemic logic with arithmetic.
arXiv:2606.31066v1 Announce Type: new Abstract: Federated Learning (FL) is highly susceptible to stealthy backdoor attacks, which aim to force a model into predicting an attacker-chosen target class for inputs containing a specific trigger. However, existing statistical defenses primarily focus on the early stages of model convergence. In this paper, we identify a fundamental vulnerability termed ``Late-stage Failure.'' We demonstrate that as the global model converges, decaying gradient norms render malicious and benign updates morphologically indistinguishable. This vanishing statistical variance effectively blinds traditional defenses, enabling adaptive adversaries to remain dormant and subsequently hijack the training process. To overcome these constraints, we propose Secure-CHG, a hybrid framework that pivots the defense paradigm from superficial morphological detection toward intrinsic semantic contribution verification. Secure-CHG employs an adaptive defense pipeline: a cascaded statistical filter stabilizes optimization during the early oscillatory phase, while a novel CHG-Shapley mechanism takes over during late-stage convergence. By leveraging sample hardness (i.e., local training loss) to project updates into a composite Hardness-Gradient space, it effectively amplifies adversarial semantic traces, enabling the isolation of stealthy attackers even as gradient norms vanish. Furthermore, we derive a closed-form solution for CHG-Shapley, facilitating low-complexity, retraining-free node valuation and trust-modulated aggregation. Extensive evaluations on CIFAR-10, MedMNIST, and NEU-SDDB demonstrate that Secure-CHG effectively mitigates Late-stage Failure. Specifically, it significantly suppresses advanced backdoor attacks, reducing their attack success rate by 2.3$\times$ and 2.0$\times$ relative to the mainstream Krum and Trimmed Mean baselines, respectively.
arXiv:2606.31892v1 Announce Type: new Abstract: "Any fool can know; the point is to understand." A well-known remark often attributed to Einstein captures a widely shared intuition: understanding is more than merely knowing. Yet epistemic logic has paid relatively little attention to understanding, despite its central role in contemporary epistemology, philosophy of science, and recent debates about AI. A recurring theme in the philosophical literature is that, unlike knowledge, understanding comes in degrees: one may understand something more or less well, and one's understanding may be better than another's. We introduce a comparative epistemic logic of understanding with level-indexed understanding modalities and a comparative connective for saying that one agent understands why a proposition better than another agent does. Semantically, we enrich multi-agent epistemic models with agent-indexed graded explanation structures and a justification-style term algebra. This yields a unified framework for representing minimal, ordinary, more demanding, and ideal understanding, together with comparisons between agents with respect to the same formula at issue. We distinguish a finitary bounded-level calculus from an infinitary full-language companion system. We establish soundness and strong completeness, and show that each fixed finite-level fragment is decidable.
arXiv:2606.31423v1 Announce Type: new Abstract: Real-world data analysis is a multi-step process over heterogeneous inputs rather than merely producing a final answer. A practical system should autonomously organize multi-step workflows, execute generated code in a sandboxed and controllable environment, and remain inspectable through visible action traces and intermediate artifacts. Existing LLM-based analysis tools, however, often emphasize isolated subtasks, leaving limited support for complete execution-grounded workflows. We present DA-Studio (Data Analysis Studio), an interactive web-based demo system for end-to-end data analysis that is autonomous, sandboxed, and inspectable. DA-Studio integrates an action-structured analysis backend, a sandboxed execution workspace, and a browser interface for task setup, streamed action traces, artifact preview, code editing and rerunning, and report export. Through iterative action generation, code execution, and feedback incorporation, it incrementally constructs executable analysis steps from raw files and natural-language requests while exposing intermediate results and artifacts throughout the process.
arXiv:2606.31893v1 Announce Type: new Abstract: We introduce a generalized form of Medvedev logics obtained by removing the greatest element from finite products of rooted Kripke frames with a top. We show that, before removing the top, the intermediate logic characterized by such finite products is exactly KC. Classical Medvedev logic is characterized by topless products of 2-chains, and a theorem of Maksimova, Skvortsov and Shehtman establishes that it is not finitely axiomatizable. Motivated by this result, Nick Bezhanishvili conjectured that non-finite axiomatizability extends to topless products of arbitrary finite chains and, more generally, to topless products of finite rooted frames with a top. We prove that every such generalized Medvedev logic is not finitely axiomatizable, thereby settling both conjectures in the affirmative. In 2003, van Benthem, Guram Bezhanishvili, and Gehrke introduced Cheq, the logic of chequered sets, and we show that whenever Cheq is a sublogic of a generalized Medvedev logic, the latter is not finitely axiomatizable over Cheq. Finally, we investigate the order structure of generalized Medvedev logics. We prove that there are at least countably many distinct generalized Medvedev logics and that no least such logic exists. These results extend the classical theory of Medvedev logic and clarify the behaviour of intermediate logics generated by topless product constructions.