arXiv:2606.18298v1 Announce Type: cross Abstract: We study the space of splines $\mathcal{S}^{\mathbf{r}}(\Sigma^\mathscr{A})$ where ${\mathbf{r}}$ denotes a smoothness distribution and $\Sigma^\mathscr{A}$ is the fan of a central hyperplane arrangement $\mathscr{A}$ in $\mathbb{R}^3$. This is the first step in the analysis of splines on three-dimensional cross-cut partitions, which naturally generalize planar cross-cut partitions. We show that the Hilbert function of $\mathcal{S}^{\mathbf{r}}(\Sigma^\mathscr{A})$ is bounded by an expression that involves the dimensions of specific Koszul homology modules constructed from the defining equations of the hyperplane arrangement $\mathscr{A}$ and the smoothness distribution function. By exploiting this connection with Koszul homology, we are able to: 1) compute the dimension of the spline space in high degrees, 2) compute all values of the dimension of the spline space if $\mathscr{A}$ is generic with five or fewer hyperplanes, and 3) compute the Hilbert function of the spline space if $\mathscr{A}$ is a generic arrangement with sufficiently many hyperplanes and ${\mathbf{r}}$ is a constant distribution. As an application of our methods, we compute $\dim \mathcal{S}^0_d(\Sigma^\mathscr{A})$ and $\dim \mathcal{S}^1_d(\Sigma^\mathscr{A})$ for all values of $d$ when $\mathscr{A}$ is a generic arrangement.
Science Journals
arXiv:2606.18718v1 Announce Type: cross Abstract: Perfect sphere packing in the Boolean space is a fundamental and complex problem with significant implications for coding theory, cryptography, and discrete mathematics. The classical solution to the perfect sphere packing problem was provided by Hamming via his well-known perfect codes. However, a major limitation of the traditional Hamming metric is its strict applicability, as it allows perfect partitioning only for spaces with specific, highly constrained dimensions. To address this structural limitation, this article introduces a novel distance metric specifically designed for Boolean hypercubes. The proposed metric modifies the topological properties of the space, making it mathematically viable to partition a Boolean space of any arbitrary dimension into disjoint, perfect spheres. We rigorously define the algebraic properties of this new distance function and demonstrate its consistency across various dimensions. Furthermore, we explore the structural characteristics of the resulting packings. This approach bypasses the classical dimensional constraints of Hamming codes, potentially opening new avenues for designing error-correcting codes and cryptographic primitives in non-traditional dimensions.
Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
arXiv:2606.17846v2 Announce Type: replace Abstract: Foundation models in language and multimodality achieve strong generalization by aligning heterogeneous data under a unified formulation and training at scale. In this report, we investigate whether this scaling recipe can be applied to robotic manipulation to achieve genuine generalization. This is challenging because, unlike text, manipulation data is heterogeneous by nature, expensive to collect, and narrow in diversity, making alignment and scale simultaneously difficult. We present Qwen-RobotManip, a generalizable Vision-Language-Action foundation model built on Qwen-VL. Qwen-RobotManip introduces a unified alignment framework across the representation, motion, and behavioral dimensions of manipulation, making large-scale multi-source training coherent rather than conflicting. This alignment capability in turn enables Qwen-RobotManip to absorb manipulation data at a scale that prior training regimes could not sustain. A human-to-robot synthesis pipeline converts egocentric hand demonstrations into robot trajectories across 15 platforms, and a rigorous curation pipeline harmonizes heterogeneous datasets. Using only open-source datasets and human videos without proprietary data collection, Qwen-RobotManip constructs a ~38,100-hour pretraining corpus and exhibits emergent generalization capabilities, including zero-shot instruction following, robustness to perturbations, reactive error recovery, and cross-embodiment transfer. We find that standard benchmarks fail to capture pretraining quality and instead adopt OOD settings including RoboCasa365, LIBERO-Plus, EBench, RoboTwin-Clean2Rand, RoboTwin-IF, and RoboTwin-XE. Qwen-RobotManip substantially outperforms prior state-of-the-art models, including $\pi$0.5, across all OOD settings, ranks 1st in RoboChallenge with a 20% relative improvement, and is validated on real-robot platforms including AgileX ALOHA, Franka, UR, and ARX.
arXiv:2606.17862v2 Announce Type: replace Abstract: In this article, we develop a multi-dimensional two-layer thin film model extending the thin film model proposed in \cite{barthwal2025hyperbolic}. The model considered in \cite{barthwal2025hyperbolic} considered a very specific Marangoni scale by choosing Marangoni numbers in both layers to be $1$. We relax this condition here and prove that the obtained system possesses a full set of Riemann invariants. Based on these findings, we develop a Riemann Invariant-based Local Characteristic Decomposition WENO (RI-WENO) method for the two-layer thin film model in one and two dimensions. The method is built upon a specially designed variable transformation constructed from the derived Riemann invariants of the system. This transformation partially diagonalizes the governing equations and yields a sparse structure in the transformed eigenvector matrices. As a result, the proposed RI-WENO framework significantly reduces the computational cost of the standard Local Characteristic Decomposition WENO approach while retaining its strong capability to suppress spurious oscillations. Numerical experiments, including new benchmark test cases, demonstrate that the RI-WENO method achieves an effective balance between accuracy and computational efficiency, making it a promising and practical choice for solving the two-layer thin film model.
arXiv:2606.18871v1 Announce Type: cross Abstract: Transitioning quantum magnetometry from laboratory environments to real-world applications has been limited by a persistent trade-off between sensor miniaturization and magnetic sensitivity. While bulky systems can achieve high sensitivity, endoscopic probes commonly suffer from inefficient fluorescence collection and reduced performance. Here we resolve this trade-off and present a miniaturized diamond quantum magnetometer with a 6 mm diameter endoscopic sensor head, achieving a magnetic-field sensitivity of 91 pT/sqrt(Hz) with a 2 kHz measurement bandwidth in a magnetically unshielded environment. The fluorescence collection bottleneck is overcome by separating excitation and collection into different cores of a fused multi-core fiber bundle, coupled to the diamond through a custom high-numerical-aperture micro-objective. A compact FPGA-based backend performs microwave control, lock-in detection and real-time resonance tracking, enabling robust operation during magnetic-field imaging. To demonstrate the practical utility of the miniaturized sensor, we image the magnetic field of a commercial lithium-ion pouch cell during charge and discharge and reconstruct depth-integrated current-density maps of the current flow. These results show that endoscopic diamond magnetometers can combine high sensitivity with a probe geometry suitable for confined, unshielded measurements, opening new avenues in battery technology and beyond.
arXiv:2606.18921v1 Announce Type: new Abstract: We introduce epistemic pairwise maximin share (EPMMS), a new fairness notion for fair division of indivisible goods. Two fundamental notions in this setting are envy-freeness up to any item (EFX) and pairwise maximin share (PMMS), with PMMS being stronger than EFX. While EFX has been extensively studied, far less is known about PMMS. Recent work shows that relaxing EFX via an epistemic perspective leads to substantial progress on the EFX problem, raising the question of whether a similar approach can advance our understanding of PMMS. Motivated by this, we initiate the study of EPMMS, the epistemic relaxation of PMMS. EPMMS is more challenging than EEFX: the key approaches underlying recent progress on epistemic EFX inherently fail to extend to EPMMS. We establish the following results. (1) For additive valuations, $4/5$-EPMMS allocations exist and can be efficiently computed. (2) For bivalued valuations, EPMMS allocations exist and can be efficiently computed; in fact, we obtain the stronger guarantee of epistemic groupwise maximin share (EGMMS), which also strengthens the existence of MMS allocations for this setting. (3) We prove that EPMMS allocations exist in two settings where MMS allocations need not exist: instances with three additive agents or two types of additive agents.
arXiv:2606.18555v1 Announce Type: new Abstract: In the realm of computer vision, indoor image recognition presents challenges due to the intricate interplay of lighting conditions, occlusions, and diverse object arrangements within confined spaces. To address the lacks of training indoor images, we introduce a novel approach leveraging Stable Diffusion (SD) for the generation of synthetic images, which serve as a powerful data augmentation tool. The utilization of SD offers a principled framework for synthesizing diverse and realistic indoor scenes, thereby enriching the training data pool for robust indoor image recognition models. Experimental findings on the MIT Indoor Scene dataset reveal the potential of our proposed approach in enhancing the training of deep models when authentic data is limited. Furthermore, to prevent the misuse of SD synthetic images, we introduce a counter measure based on DIffusion Reconstruction Error (DIRE). The powerful DIRE presentation enables training robust classifiers only using lightweight deep models. Experiments show that our approach can perfectly recognize SD generated images with the accuracy of 100% using MobilenetV3.
arXiv:2508.16177v2 Announce Type: replace Abstract: In rank aggregation, the task is to aggregate multiple weighted input rankings into a single output ranking. While numerous methods, so-called social welfare functions (SWFs), have been suggested for this problem, all of the classical SWFs tend to be majoritarian and are thus not acceptable when a proportional ranking is required. Motivated by this observation, we design SWFs that guarantee that every input ranking is proportionally represented by the output ranking. Specifically, our central fairness condition requires that the number of pairwise comparisons between candidates on which an input ranking and the output ranking agree is at least proportional to the weight of the input ranking. As our main contribution, we present a simple SWF called the Proportional Sequential Borda rule which satisfies this condition. Moreover, we introduce a more involved variant of this rule, the Flow-adjusting Borda rule, which satisfies a stronger fairness condition that applies to arbitrary groups of rankings. Many of our axioms and techniques are inspired by results in approval-based committee voting and participatory budgeting, where the concept of proportional representation has been studied in depth.
arXiv:2601.07052v2 Announce Type: replace Abstract: Simulation is crucial in real-world robotics, offering safe, scalable, and efficient environments for developing a variety of robotic applications. While the Robot Operating System (ROS) has been widely adopted as the backbone of these robotic applications in both academia and industry, its asynchronous, multi-process design complicates reproducibility, especially across varying hardware platforms. Deterministic callback execution cannot be guaranteed when computation times and communication delays vary. This lack of reproducibility complicates scientific benchmarking and continuous integration, where consistent results are essential. To address this, we present a methodology to create deterministic simulations using ROS 2 nodes. Our ROS Simulation Library for C++ (RSLCPP) implements this approach, enabling existing nodes to be combined into a simulation routine that yields reproducible results, usually without requiring any source code changes. We demonstrate that our approach produces identical results across various CPUs and architectures when testing both a synthetic benchmark and a real-world robotics system. RSLCPP is open-sourced at https://github.com/TUMFTM/rslcpp.
arXiv:2606.18606v1 Announce Type: new Abstract: It is essential for large language model (LLM) technology to serve many different cultural sub-communities in a manner that is acceptable to each community. However, research on LLM alignment has so far predominantly focused on predicting a unified response preference of annotators from certain regions. This paper aims to advance the development of alignment models with a more global outlook, that are able to accurately represent the preferences of subcommunities and do not exhibit excessive bias towards any of them. We focus on the development of reward models for this purpose and present a novel reward model training algorithm (SCPO) that can incorporate diverse cultural preferences in a balanced manner. Our method results in performance increases of the minority reward model of up to 7 points over the baseline model across two datasets, PRISM and GlobalOpinionQA, and across 7 countries. SCPO is up to 280% more training data-efficient than full-data finetuning of reward models. In addition, we perform analysis of bias by separately evaluating on the preference of subcommunities and show that excessive bias is mitigated via our weighting method. Our code is available at https://github.com/minsik-ai/Steerable-Cultural-Preference
arXiv:2606.18636v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have empowered home assistants with natural language interaction capabilities. However, current assistants overlook the progressive omission that occurs in human dialogue as shared context accumulates, leading to more elliptical expressions for efficient communication. Thus, current assistants still struggle to interpret such elliptical expressions accurately, which limits their effectiveness in real-world applications. In practical smart home scenarios, assistants face two major challenges caused by elliptical commands: (1) referential ambiguity caused by different environmental expectations among multiple users; and (2) intention ambiguity resulting from user preferences that evolve over time or change with the environment. To address these challenges, we introduce PEC-Home, the first simulated home dataset specifically designed for interpreting progressively elliptical commands in smart homes. Extensive experiments on various LLMs, including GPT-4o, show that existing home assistants struggle to execute user-intended operations based solely on elliptical commands. Even when equipped with tools for storing and retrieving user dialogue history, execution accuracy remains below that achieved with complete commands.}.
arXiv:2601.17462v4 Announce Type: replace Abstract: Atmospheric Methane Removal (AMR) is a third class of climate intervention, along with Carbon Dioxide Removal (CDR) and Solar Radiation Management (SRM). We show that, unlike CDR, the avoided warming by AMR is not durable due to methane's short atmospheric lifetime, although its temperature rebound upon termination is less abrupt than that of SRM. AMR's impact on tropospheric ozone can be further modulated by background pollutant levels.
arXiv:2602.06774v2 Announce Type: replace Abstract: State Space Models (SSMs) have emerged as an efficient alternative to the Transformer architecture. Prior work shows that, when trained under comparable conditions, SSMs can match or surpass Transformers on code understanding tasks. However, their internal mechanisms remain a black box. We present the first systematic analysis of what SSM-based code models learn along with the direct comparison between SSM and Transformer models in this domain. Our analysis shows that SSMs capture syntactic and semantic structure more effectively than Transformers during pretraining but forgets certain relations during fine-tuning on some tasks. To investigate this behavior, we introduce SSM-Interpret, a frequency-domain framework that exposes a spectral shift toward short-range dependencies during fine-tuning. Guided by these findings, we propose architectural modifications that significantly improve the performance of SSM-based code model by upto +6 MRR on NLCodeSearch. This demonstrates that our analysis not only explains model behavior but also leads directly to better designs.
arXiv:2603.27714v2 Announce Type: replace Abstract: We present a discrete Helmholtz--Hodge decomposition for H(div)-conforming Brezzi--Douglas--Marini (BDM) finite elements on triangulated surfaces of arbitrary topology. The divergence-free BDM subspace is split L2-orthogonally into rotated gradients of a continuous streamfunction space and a finite-dimensional space of discrete harmonic fields whose dimension equals the first Betti number of the surface. Consequently, any incompressible flow discretized on this subspace can be reformulated with a scalar streamfunction and finitely many harmonic coefficients as the only unknowns. This eliminates the pressure and the saddle-point structure while ensuring exact tangentiality, pointwise divergence-freeness, and pressure-robustness. We present a randomized algorithm for constructing the harmonic basis and discuss implementation aspects including hybridization, efficient treatment of the harmonic unknowns, and pressure reconstruction. Numerical experiments for unsteady surface Navier--Stokes equations on a trefoil knot and a multiply-connected sculpture surface demonstrate the method and illustrate the physical role of the harmonic velocity component.
arXiv:2606.18101v2 Announce Type: replace Abstract: Graphical user interface (GUI) grounding requires vision-language models (VLMs) to identify small target elements in high-resolution screenshots and predict precise screen coordinates. On-policy self-distillation (OPSD) is a promising post-training approach for this coordinate-sensitive task, since it provides dense token-level teacher signals beyond hard coordinate labels. However, naive OPSD is not well suited to GUI grounding: OPSD evaluates the teacher on student-generated prefixes, the quality of coordinate-token teacher signals can degrade when the prefix has already deviated from the target coordinate, leading to unreliable teacher signal. To mitigate this, We propose quality-aware self-distillation for VLM-based GUI grounding, which improves coordinate-token teacher-signal quality through soft correctness-aware gating and teacher-probability scaling. The soft correctness-aware gate checks whether the teacher's current coordinate-token prediction can still be completed into the ground-truth box under the student-generated prefix. If not, the corresponding teacher signal is down-weighted. Teacher-probability scaling then uses the teacher's confidence as a lightweight factor to further calibrate the strength of the gated supervision. A key empirical finding is that neither component alone improves overall performance, whereas combining them consistently improves performance. This suggests that the two mechanisms play complementary roles: correctness-aware gating suppresses unreliable coordinate-token supervision, while teacher-probability scaling calibrates the strength of the remaining signals. Experiments across six GUI grounding benchmarks show that our method consistently improves the base model and outperforms strong baselines.
arXiv:2606.18782v1 Announce Type: new Abstract: Large Language Models are increasingly applied to sensitive domains that require redaction of personally identifiable information (PII). While redacting PII is a data cleaning prerequisite, existing benchmarks conflate extraction mechanics with privacy semantics. A public phone number is not equivalent to a phone number in a medical record. Whether information constitutes a violation depends heavily on who holds it, why, and in what context, fundamentally differentiating redaction from simple entity recognition. Grounded in contextual integrity, we introduce RedactionBench, a manually annotated benchmark comprising 200 diverse documents across 11 domains, mostly seeded from real-world sources. We also introduce R-Score, a novel character-level metric that treats semantically similar redactions equally and nullifies shallow formatting choices, such as varying masking styles for phone numbers. Evaluations across Named Entity Recognition models, entity extraction Small Language Models, and frontier models equipped with agentic tools demonstrate that contextual redaction remains an unsolved problem. A human evaluation with over 80 users on RedactionBench reveals a stark dichotomy in privacy perceptions. Annotators show consensus with target labels for mandatory redactions (89.4 percent) and safe text preservations (94.1 percent), but fail to agree on contextual redactions (47.7 percent). This variance demonstrates the subjective nature of contextual privacy and motivates R-Score, which decouples contextual ambiguity from strict precision. We compare 35 models across families and report their performance in redacting PII. Finally, we release RedactionBench to establish a baseline for future privacy-preserving systems, hoping to inspire efficient model design and standardized evaluations.
arXiv:2606.18810v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in training LLMs for reasoning tasks, but representative methods such as GRPO assign uniform credit across all tokens, wasting gradient on routine tokens while under-crediting pivotal reasoning steps. Existing token-level credit assignment methods require resources beyond the model's own rollouts. GRPO variants rely on process reward models or ground-truth answers. Knowledge distillation assigns credit through per-token divergence but requires external teachers (On-Policy Distillation) or privileged information (On-Policy Self Distillation). However, these dependencies limit applicability in the pure RLVR setting. We observe that conditioning the model on its own verified trajectories induces a measurable per-token KL divergence between the original and conditioned distributions, and prove that distilling from a self-teacher constructed by verified trajectories leads to infeasible weighted-average solutions when multiple verified trajectories exist. We propose SC-GRPO (Self-Conditioned GRPO), which uses KL divergence mentioned before as a multiplicative weight on GRPO gradients. Across five benchmarks spanning math, code, and agentic tasks, SC-GRPO consistently outperforms 8.1% over GRPO and 5.9% over DAPO with stronger OOD performance. Moreover, SC-GRPO achieves higher performance than OPD.
arXiv:2606.18105v2 Announce Type: replace Abstract: Network planning optimization is a fundamental problem across diverse domains, including transportation systems, communication networks, and power grids. It requires simultaneous optimization of multiple competing objectives under complex constraints. Existing network planning optimization frameworks rely on mixed integer programming (MIP) solvers, heuristics, and deep reinforcement learning (DRL) models to compute planning decisions. However, they lack effective adaptability to diverse and dynamic user intents, thus leading to the trade-off between execution time and optimality. In this paper, we propose OmniPlan, an adaptive framework that achieves both timeliness and near-optimality in network planning optimization. To achieve the adaptability lacking in existing solutions, OmniPlan employs a large language model (LLM)-based interpreter to convert heterogeneous natural-language intents into a unified and quantifiable user-preference vector. Then it employs a mixture-of-experts architecture that integrates MIP solvers, heuristics, and DRL models as specialized experts, where OmniPlan adapts to diverse intents by dynamically selecting timely and near-optimal experts. Finally, it incorporates a DRL-based expert configuration module that fine-tunes optimization objective weights to align planning decisions with user-specific preferences. We evaluate OmniPlan with a representative real-world workload, i.e., distributed machine learning (ML), where we leverage OmniPlan to offload a wide spectrum of ML inference tasks, e.g., decision trees, SVM, naive Bayes, XGBoost, and random forests, onto a network of hardware devices. Our experiments on a real-world testbed indicate that OmniPlan achieves near-optimal and low-execution-time offloading for real-world ML inference tasks, reducing latency by up to 97.8\% and network device resource consumption by up to 11.5\%.
arXiv:2604.06367v2 Announce Type: replace Abstract: Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries. While current benchmarks evaluate general-purpose performance~(e.g., WebArena) or safety against malicious actions~(e.g., SafeArena), no existing framework assesses an agent's ability to successfully execute user-facing website security and privacy tasks, such as managing cookie preferences, configuring privacy-sensitive account settings, or revoking inactive sessions. To address this gap, we introduce WebSP-Eval, an evaluation framework for measuring web agent performance on website security and privacy tasks. WebSP-Eval comprises 1) a manually crafted task dataset of 200 task instances across 28 websites; 2) a robust agentic system supporting account and initial state management across runs using a custom Google Chrome extension; and 3) an automated evaluator. We evaluate a total of 8 web agent instantiations using state-of-the-art multimodal large language models, conducting a fine-grained analysis across websites, task categories, and UI elements. Our evaluation reveals that current models suffer from limited autonomous exploration capabilities to reliably solve website security and privacy tasks, and struggle with specific task categories and websites. Crucially, we identify stateful UI elements are a primary reason for agent failure, with toggles causing more than 45% task failure across many models.
arXiv:2505.15215v3 Announce Type: replace-cross Abstract: Data fusion, the process of combining observational and experimental data, can enable the identification of causal effects that would otherwise remain non-identifiable. Although identification algorithms have been developed for specific scenarios, do-calculus remains the only general-purpose tool for causal data fusion, particularly when variables are present in some data sources but not others. However, approaches based on do-calculus may encounter computational challenges as the number of variables increases and the causal graph grows in complexity. Consequently, there exists a need to reduce the size of such models while preserving the essential features. For this purpose, we propose pruning (removing unnecessary variables) and clustering (combining variables) as preprocessing operations for causal data fusion. We generalize earlier results on a single data source and derive conditions for applying pruning and clustering in the case of multiple data sources. We give sufficient conditions for inferring the identifiability or non-identifiability of a causal effect in a larger graph based on a smaller graph and show how to obtain the corresponding identifying functional for identifiable causal effects. Examples from epidemiology and social science demonstrate the use of the results.
arXiv:2507.14839v5 Announce Type: replace-cross Abstract: With rapid advancements in quantum computing, it is widely anticipated that scalable quantum hardware may threaten classical cryptography and hence, the internet and the current information security infrastructure in the coming decade. This is mainly due to the operational realizations of quantum algorithms such as Grover and Shor, to which the current classical encryption protocols are vulnerable. Blockchains, i.e., blockchain data structures and their data, rely heavily on classical cryptography. One approach to secure blockchains is to attempt to achieve conceptual information-theoretic security under certain assumptions by defining blockchains on quantum technologies. There have been two major conceptualizations of blockchains data structures on quantum registers: the time-entangled Greenberger-Horne-Zeilinger (GHZ) state blockchain and the quantum hypergraph blockchain. We conceptualize a new quantum blockchain framework combining features of both these schemes to achieve the conceptual information-theoretic protection against undetected measurement attack (physics-based disturbance detectability) of the time-entangled GHZ blockchain and the scalability and efficiency of the quantum hypergraph blockchain in the proposed quantum blockchain data structure and framework. In this work, we propose a novel quantum blockchain architecture that integrates temporal GHZ entanglement with phase encoding inspired by the quantum hypergraph blockchain. The proposed design combines the conceptual information-theoretic tamper sensitivity/resistance of temporal entanglement with improved encoding efficiency, offering a unified conceptual framework for scalable and secure quantum blockchains.
arXiv:2508.21790v2 Announce Type: replace-cross Abstract: Classical First-Passage-Time Distributions (FPTDs) have been extensively studied both theoretically and experimentally. Their quantum counterparts-Quantum First-Passage-Time Distributions (QFPTDs)-remain largely unexplored and have deep implications for both fundamental physics and the development of emerging quantum technologies. We measure the first QFPTDs using a motional mode of a single trapped ion. We develop a novel composite-phase laser pulse sequence to perform tunable stroboscopic single-shot projective measurements of the motional state of a trapped ion. We measure QFPTDs of the ion energy when coupled to electric-field noise. The measurement protocol developed here is broadly applicable to other quantum systems and provides a powerful method for exploring a broad range of QFPTD phenomena. With these results we open a new field of experimental investigations of QFPT processes with potential future relevance to quantum search algorithms, unraveling connections between classical and quantum dynamics, and study of the quantum measurement problem.
arXiv:2605.19181v2 Announce Type: replace Abstract: We study Conditional Value-at-Risk (CVaR) variants of two canonical sequential decision problems: Pandora's box and the prophet inequality. For Pandora's box, the risk-aware problem retains an exact Weitzman-style index solution after a one-dimensional variational reduction. For the prophet inequality, the picture is different: for every CVaR level \(\alpha\in(0,1)\), no positive constant approximation guarantee can hold without distributional structure, in sharp contrast with the risk-neutral case \(\alpha=1\), and we characterize the tight instance-dependent guarantee. Already in two-item hard instances, the prophet's CVaR benchmark can be made arbitrarily large while every online policy's CVaR remains bounded. This impossibility is due to the nature of CVaR objective: it measures only the worst \(\alpha\)-fraction of outcomes, so any compromise an online policy makes to preserve the chance of a large payoff in the upper \((1-\alpha)\)-fraction does not help its CVaR. It turns out that additional distributional structure restores a uniform result: under continuous reward distributions satisfying a recentered increasing-failure-rate-average (IFRA) condition, a threshold policy achieves an explicit constant bound.
arXiv:2606.18958v1 Announce Type: new Abstract: Cluster-scale full-stack simulation is essential for evaluating distributed software stacks and emerging hardware components before deployment. Such simulation must achieve both full-stack fidelity for the unmodified production stack and the simulation performance required for iterative configuration exploration. However, no existing method achieves both. We present LiveStack, an OS-level approach to cluster-scale full-stack simulation built on top of the Linux virtualization stack. LiveStack comprises four subsystems: simulation-oriented scheduling, live memory hierarchy management, simulation-aware IPC, and distributed simulation orchestration. Together, they coordinate live and modeled components under shared simulated time while controlling interference among co-located live hosts. These mechanisms point toward simulation-native OS support, where simulation control and orchestration become core OS responsibilities.
arXiv:2606.01139v3 Announce Type: replace Abstract: Agent skills are procedural artifacts that enable LLM agents to execute workflows, verify constraints, and recover from failures. Existing self-evolving methods refine skills using accumulated trajectories. However, they struggle in cold-start settings, where only an initial, imperfect skill is available. Consequently, skill construction defaults to expert authoring or one-shot LLM generation. Expert-authored skills are costly and may not align with how LLM agents actually execute tasks, while one-shot generated skills can be syntactically well formed yet behaviorally weak. To bridge this gap, we propose SkillRevise, an execution-grounded framework designed to iteratively refine these initial skills. SkillRevise diagnoses skill defects from execution evidence, retrieves relevant repair principles from a general memory, and applies execution-anchored edits. By re-executing candidates, it retains the first verifier-passing skill within the revision budget and falls back to empirical utility only when no candidate succeeds. Evaluated across three benchmarks and five LLMs, SkillRevise substantially outperforms one-shot baselines, improving the base agent's success rate on SkillsBench from 36.05% to 61.63%. Furthermore, the revised skills transfer across both executors and task environments, suggesting that SkillRevise captures reusable procedural knowledge beyond any single executor.