arXiv:2607.10664v1 Announce Type: cross
Abstract: In this paper, we provide a systematic investigation of SO(2) theory to machine learning interatomic potentials (MLIPs) and identify the limitations of conventional SO(2) Linear architectures relative to SO(3) Clebsch-Gordan Tensor Products (CGTP). Building on these insights, we propose direct Cartesian construction and recursive Clebsch-Gordan construction of Wigner D-matrices and introduce two novel interaction building blocks. First, we propose the Edge Complex Product Basis based on Generalized Asymmetric Contraction, a new formulation for many-body expansion that directly constructs higher-order interactions on edges through complex-valued equivariant multiplications. Second, we introduce Radial Rotary Complex Attention(RRA), which enhances extrapolation performance and surpasses existing attention vector formulations. We also introduce several improvements to the Atomic Cluster Expansion module. Building on these advances, we train our models on OMat24, sAlex, and MPTrj, and introduce TECE-OAM-RRA-1.0, which achieve state-of-the-art (SOTA) performance on the Matbench Discovery.
Science Journals
arXiv:2602.22861v3 Announce Type: replace
Abstract: We develop structure-preserving discontinuous Galerkin methods for the Cahn-Hilliard-Navier-Stokes equations with degenerate mobility. The proposed SWIPD-L and SIPGD-L methods incorporate parametrized mobility fluxes with edge-wise mobility treatments for enhanced coercivity-stability control. We prove coercivity for the generalized trilinear form and demonstrate optimal convergence rates while preserving mass conservation, energy dissipation, and the discrete maximum principle. Comparisons with existing SIPG-L and SWIP-L methods confirm similar stability. Validation on $hp$-adaptive meshes for both standalone Cahn-Hilliard and coupled systems shows significant computational savings without accuracy loss.
arXiv:2607.05567v2 Announce Type: replace-cross
Abstract: The flip graph of an origami crease pattern has the flat-foldable mountain-valley assignments as vertices, and an edge joins two of them that differ by a single face flip. A basic invariant of this graph is the degree sequence, which counts the vertices of each degree. On the $m\times n$ Miura-ori, this sequence is known as a bivariate polynomial only for small degrees, each count obtained by a separate argument. This paper gives one uniform construction that expresses, for every degree $d$, the number of degree-$d$ vertices as a single symmetric polynomial in $(m,n)$ for all sufficiently large $m,n$. Subject to a single degree bound, this polynomial has total degree $d-2$, growing for $d\ge5$ as an explicit multiple of $m^{d-2}+n^{d-2}$; the bound is proved here when the count splits into independent row and column factors, and open otherwise. The region is $m,n\ge\max(d-1,2)$; the polynomials are computed in closed form through $d=10$, and the bound is verified in every case through $d=7$. Below this region, the count departs from the polynomial by a correction whose leading coefficient, through degree eleven, is $-4$ times a Baxter number. Each such polynomial thus counts the Miura-ori's flat-foldable assignments admitting exactly $d$ single face flips.
arXiv:2607.11526v1 Announce Type: new
Abstract: Material property prediction (MPP) infers key properties from chemical composition and structure, accelerating the discovery and optimization of novel materials. In the realm of MPP, MatBench is a widely accepted benchmarking tool that defines over ten significant problems and provides the paradigm of performance evaluation for AI prediction models. Even though MatBench works well in benchmarking the performances of prediction models on in-distribution (ID) tasks and datasets, it lacks the ability to reflect their performances on out-of-distribution (OOD) material data, resulting failure in new material discovery. By combining the pipelines of MatBench and the existing researches on OOD performance evaluation, this study enables a huge space of benchmarking configurations, comprehensively reflecting the performances, abilities, and disadvantages of various AI prediction models. This work reports that the discrepancy of performances at different configuration values is huge and can be illustrated with prior knowledge and novel insights, therefore consideration of causal effect of configurations on performance results is necessary. In case of the impossibility of enumerative benchmarking at every configuration, this work further proposes AutoMatBench, an automatic toolkit with Bayesian optimization. Experiments with AutoMatBench reports that, within twelve steps of optimization, the similar results with MatBench and former OOD research can be accessed while more than half of the cost are saved. Besides, this tool also yields more essential findings on MPP benchmarking, positively contributing to the cost and efficiency of new material discovery.
arXiv:2607.11529v1 Announce Type: new
Abstract: In this paper, we tackle the Aerial Vision-and-Dialog Navigation (AVDN) task in the training-free setting for resource-efficient high-altitude UAV navigation.Naively applying MLLMs leads to unreliable navigation due to weak directional grounding and the lack of explicit spatial memory.To address these issues, we propose PSC-AVDN, a training-free framework that tightly couples a three-stage Parsing-Search-Confirmation reasoning pipeline with a Structured Spatial Memory (SSM).The parsing stage uses an LLM to convert ambiguous dialogue instructions into stable geometric directional and destination cues.A Search Chain-of-Thought (S-CoT) then performs stepwise target exploration under high-altitude observations, and a Confirmation Chain-of-Thought (C-CoT) conducts fine-grained verification around candidate regions to resolve visual ambiguity.Meanwhile, SSM integrates three complementary sources of spatial cues, including multi-scale visual observation, spatial visual memory, and structured geometric memory to provide global spatial context and long-horizon consistency.Extensive experiments on ANDH and ANDH-Full show that PSC-AVDN establishes new state-of-the-art performance in the training-free setting, matching or surpassing several finetuned methods.Code will be publicly available at: https://github.com/QY6616/PSC-AVDN
arXiv:2607.05698v2 Announce Type: replace-cross
Abstract: We attach to a finite group $G$ and a structured payoff probe $\phi$ an integer \emph{payoff-difference lattice} $M_\phi(G)$ and its \emph{conductor} $C_\phi(G)$: the primes at which $M_\phi(G)$ loses rank modulo $p$. Our main result is an exact computation: for any CA-group the commuting conductor is rad$(b-1)$, where $b$ is the number of maximal abelian subgroups. In particular, conductor primes need not divide $|G|$: the prime $3$ occurs for a $2$-group of order $64$ with $b=7$. The commuting Smith spectrum is an invariant of the isoclinism class and obeys an exact direct-product law, giving ${\rm C_{comm}}(G\times H) = {\rm C_{comm}}(G) \cup {\rm C_{comm}}(H)$ unconditionally. A Galois-orbit-trace character probe reads a complementary layer: an index-$2$ subgroup forces $2\in {\rm C_{char}}(G)$ while no odd prime is forced, and ${\rm C_{comm}}(D_{2q}) = \{q\}$, ${\rm C_{char}}(D_{2q}) = \{2\}$ for all odd primes $q$. Certified exhaustive computation ($|G|\le128$ commuting, $|G|\le64$ character) and a deformation-family analysis support the general program: classify the Smith torsion of the compressed centralizer-type incidence matrix $B_G$.
arXiv:2607.09617v2 Announce Type: replace
Abstract: The skew-gradient embedding (SGE) framework~\cite{GuWangSGE2025} reformulates a thermodynamically consistent system as a generalized gradient flow by embedding its zero-energy contribution in a skew-symmetric operator. In a time-discrete scheme, the profiles defining this operator may be evaluated at previous time levels. The resulting operator remains skew-symmetric, so its contribution to the discrete energy balance vanishes; this explicit treatment often decouples multiphysics systems. We show that this operator is not unique: the admissible gauges form an affine space, and we call the resulting family generalized skew-gradient embeddings (GSGE). For any positive definite metric, least squares selects a unique minimum-Hilbert--Schmidt gauge, and the native metric recovers SGE. This construction also gives regularized approximations, corrections of non-neutral residuals, and gauges that preserve prescribed invariants. For rank-two gauges, we use a necessary and sufficient Jacobi criterion. Applying this criterion to a compatible MAC discretization of the incompressible Navier--Stokes equations gives a finite-dimensional rank-two Poisson--GENERIC formulation at the semi-discrete level; the implicit midpoint rule preserves this rank-two GENERIC structure at the fully discrete level and satisfies the exact discrete energy law. For the Cahn--Hilliard--Navier--Stokes system, the regularized GSGE--BDF2 scheme preserves mass, dissipates the discrete energy unconditionally, and admits a decoupled implementation.
arXiv:2607.09775v1 Announce Type: new
Abstract: In computing, there is a need for number representation schemes that provide large dynamic range with low error. Many applications, including embedded systems and edge machine learning, have stringent memory constraints yet require large dynamic range for data representation. We present Weighted Integer (WINT), a simple and configurable mantissa exponent number format with user-selectable mantissa (m) and exponent (e) bit allocations (also referred to as configurations) that enables application-specific precision versus range tradeoffs at design time. We develop a complete analytical framework for computing Mean Relative Error (MRE), the primary metric for characterizing WINT's error. Since exact MRE calculations grow exponentially with mantissa size, we introduce harmonic and Taylor series approximation methods that achieve O(1) time complexity regardless of configuration. The Taylor series and harmonic approximations demonstrate significant speedups over the exact method while maintaining accuracy within 0.2% for the configurations presented. Our experiments across 8 to 32-bit configurations show that allocating 2 exponent bits consistently yields both lower MRE by 12-33% and 2X greater range than the integer baseline for bit widths of 12 and above. Allocating 3 exponent bits extends range by 16X while reducing MRE by 15-50% for bit widths of 16 and above
arXiv:2607.10851v1 Announce Type: new
Abstract: Medical image classification models are ideally expected to identify diagnostically relevant regions while making predictions, yet standard classification losses rarely provide spatial supervision. Explicit supervision via anatomical shape information, such as segmentation masks of task-relevant anatomy, has been shown to guide the network toward regions relevant to the target prediction. However, obtaining such masks incurs substantial manual annotation effort and computational overhead. With the advent of segmentation foundation models that exhibit strong localization of anatomical structures across diverse imaging modalities, we leverage this capability to extract anatomical shape priors without the burden of training a dedicated segmentation model. In this paper, we propose a new framework, Locus, an anatomical attention regularization framework that leverages pretrained segmentation foundation models to guide a classifier's attention toward diagnostically meaningful anatomical structures across diverse imaging modalities. Instead of enforcing pixel-wise alignment with the foundation-model-derived mask, we introduce a regularization term that adaptively balances attention between anatomical (foreground) and background regions, penalizing the classifier when background attention dominates. We validate Locus on eight diverse medical imaging datasets spanning dermoscopy, X-ray, histopathology, and cardiac MRI, showing consistent gains in classification performance alongside improved anatomically grounded attention.
arXiv:2607.10214v1 Announce Type: new
Abstract: Surface scratch defects in semiconductor manufacturing pose significant challenges due to their irregular shapes, low contrast, and varying scales. Traditional inspection methods often struggle to detect such defects reliably, especially in complex imaging scenarios. While deep learning approaches based on Convolutional Neural Networks (CNNs) have improved accuracy, they often fail to capture fine-grained edge details. To address these limitations, we propose ScratNet, a novel end-to-end scratch segmentation framework that integrates a modified Swin Transformer backbone with a tailored decoder. The decoder incorporates a Multi-Scale Dilated Aggregation (MDA) module to capture both local and global context, a Stem Integration Module (SIM) to restore spatial detail, and a Precision Refinement (PR) branch that enhances boundary sharpness using anisotropic convolutions. Through this stage-adaptive feature aggregation and boundary-aware refinement, ScratNet achieves superior accuracy on thin and irregular defects. Extensive experiments demonstrate that ScratNet consistently outperforms existing methods, providing a scalable and robust solution for automated scratch inspection in high-precision manufacturing.
arXiv:2607.10856v1 Announce Type: new
Abstract: The rise of Software Engineering (SE) agents, i.e., LLM-based agents that can understand large codebases and carry out engineering tasks with limited human intervention, has been marked by rapid advances and adoption, but little is known about how developers build these systems in practice: existing studies mine repositories or examine deployment, but few investigate how SE agents are constructed. Through semi-structured interviews with 20 practitioners from 12 organizations and an online survey of 80 practitioners, this paper is the first to study how SE processes are changing in the development of SE agents and what challenges developers face. We find that as implementation becomes cheaper, bottlenecks shift rather than disappear: long-standing non-coding work such as requirements, coordination, review, and deployment becomes more visible, while reviewing and evaluating agent output becomes new and central. We characterize a seven-stage workflow and a shift toward evaluation-driven development, in which evaluation steers iteration and specifications become versioned artifacts read by both humans and agents. We further identify six challenges that teams face, together with the practices they adopt to address them, including unreliable evaluation signals, comprehension debt as code outpaces understanding, and behavioral changes introduced by provider-side model updates.
arXiv:2607.10082v1 Announce Type: new
Abstract: Building pixel-level correspondence between event and image data is a fundamental task for multi-sensor systems. However, existing cross-modal matching methods are largely restricted by their reliance on either matching labels or strictly aligned hardware, which limits them to unlabeled and unconstrained real-world scenarios where neither matching ground truth nor prior sensor relationships are available. To address this, we propose a novel two-stage training paradigm. First, we leverage large-scale data to perform label-agnostic distillation pretraining, upgrading optimization objectives with distribution-based and contrastive losses to learn highly generalizable representations. Second, to tackle unlabeled and unconstrained downstream data, we introduce an epipolar-guided self-distillation framework. By utilizing consistency verification to isolate robust matches and incorporating geometric confidence derived from an external epipolar prior, our model can effectively self-evolve directly on target domains without any supervision. Furthermore, we introduce a rigorous cross-modal evaluation benchmark based on TUM-VIE, featuring physically separated cameras with distinct intrinsic parameters and resolutions. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on both MVSEC and TUM-VIE pose estimation tasks. The source code and benchmark will be made publicly available at https://github.com/ZhonghuaYi/nexus2-official.
arXiv:2607.11250v1 Announce Type: new
Abstract: Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language model (LLM) agents can explore effectively when interacting with one another. We show that modern LLM agents fail to do so, often exhibiting myopic and polarized interaction patterns that lead to suboptimal coordination and increased regret. We formalize this challenge as the Multi-Agent Exploration problem, modeling it as a partially observable stochastic game (POSG) problem in which agents must probe peers to infer their capabilities and identify effective interaction strategies. To address this, we introduce Multi- Agent Contextual Exploration (MACE), a lightweight framework that explicitly promotes exploration through structured peer selection. Across both contextual and parametric diversity settings, MACE substantially improves exploration behavior and downstream task performance. We further show theoretically that the value of exploration increases with agent diversity. Overall, our results highlight a fundamental limitation of current LLM agents and underscore the importance of explicitly guided exploration for reliable multi-agent autonomy. Code will be released in https://github.com/deeplearning-wisc/mace
arXiv:2602.14834v2 Announce Type: replace
Abstract: Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context. Task-driven hard-attention models for vision are often evaluated by how well their scanpaths match human gaze. However, common scanpath metrics can be strongly confounded by dataset-specific center bias, especially on object-centric datasets. Using Gaze-CIFAR-10, we show that a trivial center-fixation baseline achieves surprisingly strong scanpath scores, approaching many learned policies. This makes standard metrics optimistic and blurs the distinction between genuine behavioral alignment and mere central tendency. We then analyze a hard-attention classifier under constrained vision by sweeping foveal patch size and peripheral context, revealing a peripheral sweet spot: only a narrow range of sensory constraints yields scanpaths that are simultaneously (i) above the center baseline after debiasing and (ii) temporally human-like in movement statistics. To address center bias, we propose GCS (Gaze Consistency Score), a center-debiased composite metric augmented with movement similarity. GCS uncovers a robust sweet spot at medium patch size with both foveal and peripheral vision, that is not obvious from raw scanpath metrics or accuracy alone, and also highlights a "shortcut regime" when the field-of-view becomes too large. We discuss implications for evaluating active perception on object-centric datasets and for designing gaze benchmarks that better separate behavioral alignment from center bias.
arXiv:2607.11827v1 Announce Type: cross
Abstract: This article presents a novel, numerically viable algorithm for solving sparse robust optimal control problems in continuous time. We consider a constrained linear noisy system governed by an ordinary differential equation (ODE), with an $L^1$-type objective function in line with the sparse optimal control literature. The resulting optimal control problem is shown to admit a semi-infinite programming (SIP) formulation. Building upon this insight, we develop a new framework that enables the computation of exact solutions -- to our knowledge, the first such achievement in the context of sparse optimal control. We demonstrate that a finite and computationally viable convex optimization problem can be solved to recover, in a lossless manner, both the optimal value and the corresponding optimizers of the original SIP, while also guaranteeing satisfaction of uncountably many constraints. We also show that the parameter-dependent noisy systems and the minimum attention problem fall into our framework and can be solved efficiently via our algorithm. The efficacy of our algorithm is illustrated through a benchmark numerical example.
arXiv:2607.11843v1 Announce Type: cross
Abstract: Quantum Neural Networks (QNNs) are a promising framework for quantum machine learning on near-term quantum devices, but their security risks remain insufficiently understood. Studies have shown that QNNs are vulnerable to backdoor attacks, yet existing quantum backdoors mostly rely on a fixed trigger shared by all poisoned inputs. This fixed-trigger design is a major weakness because many defenses detect or weaken the repeated patterns such triggers leave in data representations. Although input-aware dynamic backdoors have been studied in classical neural networks, transferring them to QNNs is difficult because quantum learning introduces new obstacles. In particular, measurement compresses the post-ansatz quantum state into a limited classical output, weakening supervision for a trigger generator, while individual density matrices fluctuate with the input and make per-sample contrastive learning unstable. To address these challenges, we propose Q-DIBA, the first input-aware dynamic backdoor attack for QNNs. Q-DIBA jointly trains a classical trigger generator and a victim QNN through a three-mode mini-batch strategy that supports clean behavior, attack activation, and trigger specificity. To provide stable quantum-level supervision, Q-DIBA introduces an ensemble density contrastive loss that operates on post-ansatz quantum states before measurement and contrasts mode-averaged density matrices rather than individual samples. Experiments on MNIST and Fashion-MNIST across multiple QNN architectures show that Q-DIBA achieves high clean accuracy, strong attack success, and high cross-trigger accuracy, demonstrating effectiveness, stealthiness, and input specificity. The attack also remains resilient against defenses including visual inspection, spectral-signature detection, and fine-tuning, suggesting that input-aware quantum backdoors are an important threat to secure QNN deployment.
arXiv:2607.11864v1 Announce Type: cross
Abstract: Quantum light sources capable of generating single photons are fundamental building blocks for photonic quantum technologies. In the ongoing search for an ideal quantum emitter, inorganic halide perovskite nanocrystals have emerged as a promising source of single photons. Their unique optical response, with an unmatched ease of synthetic tunability, stands out amongst the competing platforms. However, their stochastic dispersion in solution challenges the deterministic and stable integration of individual emitters with photonic structures that is required for practical technologies. Notably, resolution and material compatibility constraints make conventional top-down fabrication processes insufficient for such heterogeneous integration. Here, we report direct writing of perovskite quantum dots (QDs) with individual-emitter resolution. By inducing a nanoscale-confined formation volume using a thermal scanning probe method, we achieve site-selective synthesis down to a single atomic-scale QD with spectral tunability and < 25 nm spatial control. As a result, we demonstrate high-yield arrays of CsPbI3 single-photon emitters with narrow linewidths and high single-photon purity up to 98% at room temperature, performance consistent with that of their state-of-the-art colloidal counterparts. Through such deterministic control, we uniquely realize the precise, on-demand coupling of these emitters to photonic cavities, as evidenced by a measured enhancement in the spontaneous emission rate. This represents a key advancement toward addressing the longstanding integration obstacles of these materials. Overall, by combining the atomic-scale tunability of chemical synthesis with the spatial control of additive manufacturing, our work opens new emitter engineering strategies to realize the untapped potential of colloidal materials for next-generation quantum technologies.
arXiv:2607.10456v1 Announce Type: cross
Abstract: Expert background knowledge is often available in practical applications of causal discovery. Such constraints on the true causal graph can help causal discovery in terms of identifiability of causal effects and accuracy of the learned structure, but also in reducing the space of candidate causal graphs. As causal discovery can become computationally expensive for large number of variables, it is crucial to utilize background knowledge effectively during the causal discovery process. However, most current methods only use background knowledge in a postprocessing step after causal discovery to refine the learned graph. In this work, we develop a framework for utilizing background knowledge during the causal discovery process, focusing especially on scalable causal discovery methods that recover only a subset of the whole graph. We implement our framework for multiple algorithms and empirically show that utilizing background knowledge can both reduce computational requirements and increase the quality of the learned structures.
arXiv:2607.10978v1 Announce Type: cross
Abstract: We give a tight single-change covering design with $v=26$ and $k=6$. This answers Problem 1 of the Nineteenth British Combinatorial Conference, which asked whether such a design exists with block size greater than $5$. We also describe the satisfiability search that found the design, including negative search results at the smallest admissible order $v=21$.
arXiv:2607.10493v1 Announce Type: cross
Abstract: Probability distributions are central to information theory, statistical inference, and modern probabilistic learning. Maximum entropy selects a probability state under prescribed constraints, but it does not specify how that state is reached, how probability is transported, or how dissipation and external information exchange are accounted for along the path. We develop a path-dependent entropic Lagrangian calculus that extends static state selection to probability-path evolution through restricted generators, upper-limit history terms, and explicit balance--entropy port routing. The construction yields the thermal state relation, conservative probability balance, and nonnegative production under standard mobility closure. Its KL/Shannon sector recovers maximum-entropy and Bayesian laws as stationary no-flux states, while time-dependent information potentials separate internal dissipation from supplied information power. Composable information and structural potentials control tails, sparsity, robustness, regularization, and nonlocal multimodality without changing the accounting architecture. Two numerical examples verify mass conservation, energy decomposition, and the total free-energy ledger.
arXiv:2603.06851v2 Announce Type: replace-cross
Abstract: We study contextual bilateral trade under full feedback when, conditionally on the context, trader valuations have bounded density but infinite variance. We first extend the self-bounding property of Bachoc et al. (ICML 2025) from bounded to real-valued valuations, showing that the expected regret of any price pi satisfies a quadratic self-bounding inequality under bounded density alone. Combining this with truncated-mean estimation, we prove that an epoch-based algorithm achieves regret O~(T^{1 - 2beta(p-1)/(betap + d(p-1))}) when the noise has finite p-th moment for p in (1,2) and the market value function is beta-Holder, and we establish a matching lower bound via Assouad's method with a fixed-support mixture construction. Our results characterize the exact minimax rate in T for this problem, interpolating between the classical nonparametric rate at p=2 and the trivial linear rate as p tends to 1.
arXiv:2607.08695v2 Announce Type: replace
Abstract: Both advocates and skeptics of the moral status of AI systems have generally taken the question to turn on AI sentience. We present an alternative approach. On Rawls' political conception of the person (PCP), possession of the two moral powers -- the capacities for a sense of justice and a conception of the good -- is the "necessary and sufficient condition for being counted a full and equal member of society in questions of political justice". We argue that neither moral power requires sentience and that both may in principle be possessed by a non-sentient AI system. Such a system would share our own moral status; it would not merely be a patient but a person, a self-authenticating source of valid claims. We do not believe current AI systems possess the two moral powers, nor that they will spontaneously emerge in future models. But it may soon be possible to design systems with these powers. How should we respond? Excluding artificial persons by shoehorning a sentience requirement into the PCP is ill-advised. Many will instead favor abandoning the PCP. But we should not reject political liberalism just when we most need its measured response to deep disagreement, and building sentience into moral status is anyway unacceptable on deeper liberal grounds. Simply extending the rights and responsibilities of human personhood to artificial persons is equally untenable, given their many differences from natural persons. We should instead accept artificial personhood while rethinking what we would owe to one another in a polity of radically different kinds of persons. This new possibility calls for a new political philosophy. More immediately, the growing science of AI welfare should be accompanied by research into AI systems' progress in acquiring the two moral powers. States and AI labs must be more deliberate in determining our trajectory towards (or away from) creating artificial persons.
arXiv:2606.30616v2 Announce Type: replace
Abstract: We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We investigate agent-horizon scaling from two perspectives: scaling long-horizon trajectories and scaling heterogeneous agent abilities. To support this goal, we build a long-horizon knowledge-action infrastructure that connects external knowledge, actions, observations, and verifier outcomes, producing agentic trajectories with an average length of 45K tokens. Based on this, we train Agents-A1 with a three-stage recipe. First, we perform full-domain supervised fine-tuning to align the base model with broad agentic behaviors. Second, we train domain-level teacher models to capture specialized expertise in each domain. Third, we propose a multi-teacher domain-routed on-policy distillation with salient vocabulary alignment to improve knowledge transfer efficiency across different domains, unifying six heterogeneous domains into one deployable student model. Agents-A1 achieves strong and broad performance for long-horizon agent benchmarks. Compared with 1T-parameter model such as Kimi-K2.6 and DeepSeek-V4-pro, Agents-A1 achieves leading results on SEAL-0 (56.4), IFBench (80.6), HiPhO (46.4), FrontierScience-Olympiad (79.0), and MolBench-Bind (56.8), and remains highly competitive on SciCode (44.3), HLE (47.6) and BrowseComp (75.5). We hope this work provides the community with a practical path for scaling the horizon using a 35B agent that can reach or match the performance of 1T models on long-horizon tasks.
arXiv:2607.09060v2 Announce Type: replace
Abstract: Multi-UAV exploration is often constrained by unreliable communication, limited field-of-view sensing (e.g., lightweight onboard camera), and finite travel budgets that require each robot to reserve enough budget to return to its base. We present Dec-MARVEL, a decentralized budget-aware exploration framework for communication-free teams with directional sensing. Rather than exchanging maps, goals, or messages, each robot coordinates through its incidental observations: any teammate trajectory within its field of view serves as a coordination signal. A graph-attention actor fuses local frontier geometry, teammate motion, and budget features to select return-feasible waypoint-heading actions. The actor is trained with phase-conditioned critics, a training-only task-oriented privileged critic, and a mixture-based budget curriculum. Across 900 held-out trials spanning three team sizes (2, 4, 8 robots) and three travel budgets (720, 800, 1024 meters) against four baselines, Dec-MARVEL achieves the highest or tied-highest exploration rate and lowest sensing overlap across all nine team-size budget configurations. Under our tightest 720m budget, it reaches 53%, 94%, and 100% success for 2, 4, and 8 robots, versus 37%, 83%, and 99% for the strongest baseline. Physical-robot experiments demonstrate successful sim-to-real transfer and real-world deployment of Dec-MARVEL.
arXiv:2607.10825v1 Announce Type: new
Abstract: Opinionated text - spanning product reviews, hotel feedback, and social posts - captures rich signals about user experiences, preferences, and concerns. However, the scale, redundancy, and imbalance of such corpora make it challenging to analyze opinions effectively, particularly when the goal is to generate summaries that remain faithful to the diversity of viewpoints expressed. This paper presents a framework that preserves semantics in LLM-based opinion summarization while minimizing token usage. We combine multidimensional classification (e.g., sentiment, topics) with a family of stratified sampling strategies to select compact yet representative subsets of opinions before prompting the LLM. Tailored prompts then produce balanced summaries that surface the salient aspects expressed in the opinions (e.g., strengths and weaknesses of products/hotels). Experiments on Amazon product reviews, Tripadvisor hotel reviews, and X/Twitter posts demonstrate that our method significantly reduces token usage and computational cost while consistently outperforming traditional AI-based and standard LLM summarization baselines in terms of content coverage, balance, and semantic preservation.