arXiv:2607.06184v1 Announce Type: new
Abstract: Coding agents are ranked almost entirely by resolve rate: whether their final patch passes the target tests. Yet two agents can reach the same outcome through very different processes, and a single pass/fail label says nothing about why a run failed or why an accepted run spent extra steps, time, or tokens. This process evidence lives in the trajectory, which records a run's searches, reads, edits, tool calls, validation, and reversions. However, raw traces are heterogeneous and hard to compare across runs. We present TraceProbe, a trajectory-diagnostic framework that recovers what resolve rate hides. TraceProbe normalizes each raw run into a canonical nine-type action taxonomy with deterministic effect labels, then applies two rule-based modules: Insight names single-trajectory anti-patterns adapted from established debugging practice (e.g., search loops, verification skips), while Converge aligns pairs of runs and classifies where their behavior diverges under controlled references. Applying TraceProbe to 2,500 trajectories from five production settings on SWE-Bench Verified, we find that (i) file choice is too coarse to separate success from failure, whereas function selection and completion behavior localize it; (ii) Insight anti-patterns act mainly as corpus-level difficulty clues, with search loops the most stable; and (iii) even resolved runs differ in how quickly they reach relevant code and how much failed work they incur. Trajectory structure thus adds auditable diagnostic context to outcomes by localizing inspection targets, suggesting failure hypotheses, and prioritizing runs for review.
Science Journals
arXiv:2607.05798v1 Announce Type: new
Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasoning, often referred to as ``thinking with images''. A basic operation in this reasoning process is to zoom in on regions of interest (often represented with bounding boxes) to acquire finer visual details. In this paper, we propose \textbf{Seg}mentation before \textbf{Answer}ing (SegAnswer), which shifts the unit of zoom-in from the popular bounding box to pixel-level segmentation mask. By employing fine-grained masks to isolate the target area from cluttered environments, segmented visual input yields a more precise region of interest, effectively filtering out redundant background and interfering objects. Furthermore, the discrete patches of segmented visual input align more seamlessly with how MLLMs structure visual tokens via positional embeddings. In experiments, we evaluate SegAnswer across diverse benchmarks, including high-resolution perception, general perception, and hallucination. It achieves consistent improvements and also exhibits considerable performance on segmentation tasks, validating its capability for reliable pixel grounding.
arXiv:2607.05818v1 Announce Type: new
Abstract: Three-dimensional (3D) integration is a critical technique for enhancing transistor density, improving power efficiency, and reducing interconnect delays. However, as current demands and design complexity increase, power deliver networks (PDNs) are facing growing challenges.Careful planning of through-silicon vias (TSVs) is essential for ensuring reliable PDNs, where effective resistance serves as a vital metric for the reliability. Ill-planned TSVs often cause 3D IC with unevenly distributed effective resistance and consequently severer IR Drop.In this paper, we propose a GPU-accelerated framework on accurate effective resistance analysis for early stage 3D IC PDNs. The proposed framework achieves a speedup of 5 to 6 orders of magnitude compared to the conventional direct solver, while maintaining negligible deviations in both maximum and average relative errors.
arXiv:2508.09964v2 Announce Type: replace
Abstract: Traditional methods of population synthesis produce stable and interpretable populations but cannot capture the interrelationships between household- and individual-level attributes. Recent deep learning methods offer this flexibility, yet can overfit high-dimensional attribute relationships without structural guidance and deviate from known structures. We develop a household level synthetic population generation framework that adapts the existing conditional input directed acyclic tabular generative adversarial network, or ciDATGAN, to multi person households. The framework combines household size specific data construction, directed acyclic graphs (DAG) informed dependency regularization, and conditional population inputs as deterministic anchoring to preserve intrahousehold associations. We apply the model to generate an open access synthetic population for New York State. The synthetic population includes nearly 20 million individuals and 7.5 million households in 2021. Validation against withheld benchmarks shows close agreement: joint distribution matching of Public Use Microdata Areas (PUMA), age, race, and disability against the full public use microdata sample (PUMS) yields an R-squared of 0.602 and a Jensen-Shannon distance of 0.280, pairwise Cramer's V differences between generated and benchmark records average below 0.007 at the person level, and classifier two-sample tests on complete household records yield AUC values of 0.527-0.549, close to chance level. The generated households reproduce cross member associations while increasing diversity by 10-17% over the PUMS sample and 13.2% over PopGen alone.
arXiv:2607.06226v1 Announce Type: new
Abstract: Inverse Compton X-ray sources are laboratory-scale devices providing quasi-monochromatic synchrotron radiation which is generated by laser photons Compton-scattering off highly relativistic electrons. Since the shape and width of the X-ray spectrum are determined by the properties of the colliding beams, these must be carefully optimised. However, device compactness limits the space for diagnostics, rendering a complete characterisation challenging, especially if an electron storage ring is combined with a laser enhancement cavity. Here, a framework for laser, electron and X-ray beam parameter determination is proposed to address this issue. First, methods for determining the laser- and X-ray parameters are presented. Knowing these, electron beam parameters are retrieved from the shape of the X-ray spectrum. To this end, an analytical physical model enabling a rapid calculation of inverse Compton scattering spectra is developed and combined with a genetic algorithm. This strategy's effectiveness is demonstrated by applying the concept at the Munich Compact Light Source, a storage ring-based inverse Compton X-ray source facility. Since the analytical model is computationally very inexpensive, the proposed framework could enable real-time monitoring of inverse Compton X-ray sources or be used as a non-invasive diagnostic based on a single spectrum for the electron beam emittance of storage rings or accelerators.
arXiv:2607.06237v1 Announce Type: new
Abstract: Coefficient inversion on fine grids under PDE constraints is ill conditioned: sparse observations weakly constrain fine-scale parameters, and direct single-resolution optimization must recover state and coefficient fields across all scales simultaneously. This causes slow, initialization-sensitive convergence; learned transfer models require offline data and can introduce approximation error into the numerical physics. We propose PhyRes-MDNF, a fixed-physics multilevel discrete neural field framework. On each level, a single-level DNF represents the inverse unknowns directly as trainable fields and optimizes the discrete objective. In the full-space Darcy realization, state fields \(U\) and their shared coefficient field \(K\) are optimized jointly in one fixed-physics inverse process. Between levels, one zero-initialized PhyRes-GNN jointly performs fixed-stencil prolongation and bounded residual correction to construct an incoming target representation, which a fixed initialization map converts to the next DNF variables. It is fitted anew from the observations and unchanged numerical model, without offline pretraining or fine-grid truth. Coarse levels therefore resolve large-scale structure before refined degrees of freedom are introduced, shortening the fine-grid optimization path while retaining the original discrete operator. Under the same final-grid update budget, the multilevel Darcy realization reduces coefficient and state errors by approximately \(85\%\) and \(90\%\), respectively, demonstrating improved accuracy and final-grid iteration efficiency. On measured KTC2023 EIT data, the full-\(W\) pipeline improves mean Otsu mIoU by approximately \(3.4\%\) over the official linearized CEM reconstruction and \(16.9\%\) over direct single-level DNF.
arXiv:2510.12588v2 Announce Type: replace
Abstract: We investigate the effective rheology of a train of elongated bubbles of negligible viscosity flowing in capillary tubes. Building upon the classical Bretherton theory for a single bubble, we extend the analysis to a train of bubbles in a single capillary tube and finally to an array of parallel, noninteracting capillary tubes, i.e., a capillary bundle. Our goal is to characterize the nonlinear pressure drop-flow rate relation of this simplified two-phase system by incorporating the thin-film hydrodynamics at small capillary numbers. We model the structural heterogeneity of the bundle by assuming that the tube radii follow a truncated power-law distribution and examine deviations of the system from the Darcy law in terms of both its statistical properties and the parameters characterizing the bubble train (i.e., the tube slenderness ratio, the volume fraction, and the number of bubbles). The main result is that two-phase flow alters the effective rheology, leading to deviations from Darcy-type behavior across the entire parameter space investigated. Specifically, for a limited number of bubbles, the flow exhibits a smooth transition from the Bretherton regime, where the pressure drop scales with the flow rate to the power of 2/3, to weaker sublinear regimes with exponents between 2/3 and unity. Interestingly, increasing the number of bubbles or narrowing the pore-size distribution leads to only minor deviations from the Bretherton regime. The resulting pressure drop-flow rate exponents are qualitatively similar to those reported in the literature for immiscible two-phase flow in porous media, despite the inherent simplicity of the capillary bundle model.
arXiv:2607.05966v1 Announce Type: new
Abstract: Long-horizon failure in world models is conventionally attributed to compounding error, a generic framing that does not distinguish what kind of error compounds. We propose a kinematic-vs-dynamic reframing: world models tend to imagine kinematically rather than dynamically. We operationalize this as the imagined Kinematic-Consistency Error, a per-step diagnostic that measures how far a rollout departs from a closed-form kinematic null, paired with a perturbation protocol that tests whether iKCE responds when physical conditions cross a regime boundary. We instantiate the diagnostic on a released DreamerV3 checkpoint trained on DMC walker-walk, where imagined iKCE runs roughly two orders of magnitude above that of matched real-physics rollouts. Across a friction sweep that crosses the gait-collapse boundary, the model's iKCE stays statistically flat even as the trained policy's reward collapses through the same range, providing the kinematic-not-dynamic signature. The diagnostic distinguishes kinematic from dynamic imagination at horizons longer than the embodiment's gait period.
arXiv:2607.05971v1 Announce Type: new
Abstract: We present VTMR, a two-stage framework for Video-To-Music Recommendation. In Stage~1, VTMR aligns comprehensive video and music signals in a joint audio-visual-text representation space and efficiently retrieves semantically compatible candidates using coarse global embeddings. In Stage~2, it reranks the retrieved candidates by attending to the temporal sequences of both video and music, thereby capturing fine-grained temporal correspondence. Evaluated on the video-to-music recommendation task, the multimodal retrieval stage improves R@10 from 14.2 to 15.9 and Median Rank from 75 to 58 over the strongest baseline; the temporal reranker further boosts R@10 to 18.3 and Median Rank to 46, demonstrating complementary gains from richer query encoding and temporal alignment. A human preference study confirms that VTMR is on par with a commercial baseline in overall preference, while outperforming a generative baseline in music quality.
arXiv:2607.05435v1 Announce Type: new
Abstract: Academic output is produced across a fragmented toolchain: literature discovery in one application, reference management in another, writing in a LaTeX editor, formatting against venue templates by hand, and submission through yet another portal. Each boundary between tools forces a context switch, a format conversion, or a manual copy-paste step, and the cumulative cost dominates the time researchers spend on activities that are not research. We present Bibby AI, an editor-native platform that collapses this toolchain into a single Research-Write-Publish pipeline built around a cloud LaTeX editor. Unlike assistants that attach to an existing editor through a browser extension, Bibby AI owns the full document state, compilation pipeline, and revision history, which allows its agents to perform retrieval-grounded citation insertion, structural edits, and template-compliant reformatting as first-class, verifiable operations rather than text suggestions. The platform integrates (i) ingestion pipelines that convert PDF, DOCX, and handwritten mathematics into clean LaTeX; (ii) a retrieval layer over scholarly metadata enriched with patent-to-paper citation signals derived from USPTO PatentsView and the Marx-Fuegi citation corpus, surfacing the translational impact of candidate references; and (iii) task-scoped agents for literature triage, drafting, revision, and venue formatting that operate directly on the document's abstract syntax representation. Bibby AI is deployed in production and serves more than 5,000 active researchers across more than 50 subscribing universities. We describe the architecture, the design decisions that editor-nativeness makes possible, and the workflow-level time-savings framework we use to evaluate the platform against fragmented baselines.
arXiv:2607.06160v1 Announce Type: new
Abstract: Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision. We propose \textbf{LongCrafter}, a structured synthesis framework that couples a hierarchical task taxonomy with an evidence-grounded pipeline. The taxonomy organizes long-context understanding into local/shallow and global/deep levels and yields 32 fine-grained task types that serve as a global generative prior. Guided by this taxonomy, LongCrafter constructs task-aligned long contexts, decomposes them into explicit evidence graphs that model cross-paragraph dependencies, and generates instruction--response pairs strictly grounded in the located evidence spans, ensuring both controllable difficulty and faithful, traceable reasoning. Models fine-tuned on LongCrafter data outperform all SFT baselines and even the official post-trained models on LongBench, LongBench~v2, and LooGLE across both Qwen2.5-7B and LLaMA-3.1-8B, with the largest gains on high-difficulty tasks. Further analysis shows that LongCrafter data is more diverse and better spread across difficulty levels, and that the trained models locate evidence robustly regardless of position, effectively mitigating the ``lost in the middle'' problem.
arXiv:2607.06169v1 Announce Type: new
Abstract: Meeting the growing demand for quality-of-service (QoS) guarantees in 5G networks requires an accurate characterization of delay performance, commonly captured by the delay violation probability (DVP) at a specified delay target. Although hybrid automatic repeat request (HARQ) is a fundamental reliability mechanism in wireless systems and is central to supporting QoS, many existing approaches to DVP prediction for HARQ remain overly simplified. In particular, they omit important delay components and adopt assumptions that do not reflect the operation of HARQ in slot-based systems such as 5G. Consequently, these models can substantially underestimate the DVP, especially under stringent latency requirements, where the contribution of the neglected components becomes critical. To address this gap, we develop a tractable DVP characterization for 5G HARQ that accounts for queueing, transmission, decoding, and feedback delay, as well as the contribution of Control Signaling (CS) transmissions to the overall delay, under practical timing assumptions consistent with 3GPP operation. Moreover, we incorporate parallel packet transmissions that proceed without waiting for earlier packets to succeed, an essential HARQ behavior frequently overlooked in prior work. Using tools from queueing theory and Markov analysis, we then derive upper bounds on the DVP and validate them against ns-3 5G-LENA simulations.
arXiv:2607.06290v1 Announce Type: new
Abstract: We study the infinite-width Gaussian-process limit of random neural networks
through the lens of tensor programs, and we provide a quantitative convergence
theory in Wasserstein distance.
Our main result gives explicit finite-width error bounds, of order inverse square-root of the widths
between finite-network executions and their
Gaussian-process limits. The framework is architecture-agnostic and covers feed-forward models together
with weight-sharing schemes relevant for recurrent and transformer-type
architectures.
arXiv:2607.06341v1 Announce Type: new
Abstract: Formal verification offers the strongest guarantee of software correctness, but it does not scale: the proofs demanded by interactive theorem provers such as Coq require enormous expert effort. Large language models (LLMs) promise to generate these proofs automatically, yet existing approaches wire a fixed, human-designed proof strategy into the system and constrain the model to follow it (retrieving premises and predicting tactics one step at a time, or splitting goals by divide-and-conquer), and still prove only a fraction of their target theorems.
We show that imposing such a strategy is unnecessary and limiting. Handing the whole lemma to a general LLM code agent (for example, Claude Code), free to choose its own approach, and wrapping it in a verification harness is both simpler and more effective, achieving full coverage: every targeted lemma proved, with no failures and no Coq expert intervention. The agent writes the proofs under feedback and hard constraints from the harness that keep each one sound (accepted only when the prover's kernel closes it), complete (no obligation left unproved or silently dropped), and terminating (no divergent tactics).
We evaluate this harness plus code agent along three dimensions. (1) Core logic: on Iris, the state-of-the-art separation logic for concurrent and memory-manipulating programs, Aria proves all 4,257 lemmas of the four core modules and the 217 lemmas verifying Rust's standard libraries built on it, fully automatically. (2) Comparison with prior LLM provers: on reglang, where prior provers manage barely one in eight, Aria proves all 318. (3) Generality: on iris-lean, the unfinished Lean 4 port of Iris, it proves 72 not-yet-ported lemmas, showing the approach is not specific to Coq. A state-of-the-art model (Claude Opus 4.7) can write proofs for verified software development fully and automatically.
arXiv:2607.05961v1 Announce Type: cross
Abstract: Recently, Chudnovsky, Cook, Davies, Oum, and Tan obtained the first finite bound on the chromatic number of t-perfect graphs, showing that they are 199053-colorable. We improve this bound to 186 by refining their proof.
The original proof establishes that every graph with large odd girth and large chromatic number contains a certain structure called an r-arithmetic rope, and that its existence in a certain leveling of a graph with large odd girth would imply an odd wheel as a t-minor, a known obstruction of t-perfectness. While their technique requires a lower bound on the chromatic number that is exponential in r, we show that the existence of an r-arithmetic rope can already be guaranteed under a linear bound. Using a slightly weakened notion of arithmetic ropes allows us to reduce the bound even further.
arXiv:2607.01148v2 Announce Type: replace
Abstract: We investigate the emergence of structural disparities in networks of collaborating large language model (LLM) agents. When LLM agents autonomously choose collaborators, the resulting communication network exhibits preferential-attachment dynamics: agents that are already prominent become increasingly likely to attract additional connections. In some cases, weaker LLM agents (agents with smaller base model or older version) can disproportionately occupy central and influential network positions relative to stronger LLM agents. We interpret this as a type-dependent glass-ceiling effect (GCE). We model the network of LLM agents as a time-evolving sequence of directed weighted graphs, where the vector-valued edge weights represent cumulative tokens exchanged, number of interaction rounds, and reasoning effort. Using a contraction mapping argument on the mean-field dynamics, we prove that the importance (centrality) of each agent type converges to a unique stable equilibrium. To ground the model in LLM decision mechanisms, we introduce a cross-attention-inspired utility for collaborator selection. This utility specifies the local connection dynamics and, together with the mean-field model, yields a predictive characterization of the limiting network structure and its type-dependent centrality gaps. To validate the theory, we develop an experimental testbed with 100 LLM agents. Our experiments show that autonomous network formation can generate persistent centrality disparities, with their magnitude and direction depending on model family, model size, system-prompt design, and task context. They further show that the effect of preferential attachment depends on its alignment with model capability: reinforcing it improves collective performance when stronger agents become central, whereas weakening it improves performance when network dynamics instead favor weaker agents.
arXiv:2607.05742v1 Announce Type: new
Abstract: Large language models increasingly serve as judges in AI evaluation, but current approaches rely on consensus preferences that ignore individual evaluator variation. We propose a novel simulation approach that combines categorical judgments with evaluator-specific auxiliary data--retrospective reasoning traces and interface telemetry--to enable LLM-based simulation of individual evaluators via in-context learning. We conduct a systematic empirical study of this approach using multi-facet data from 32 trained annotators across 4,200 preference judgments in a 4 x 4 x 4 factorial design. Our key findings: (1) The simulation approach achieves up to 9.9 percentage point improvements over the Base Judge; (2) Reasoning traces provide the largest gains with higher collection efforts, while interface telemetry often hurts rather than helps performance despite being cheaper to collect. (3) Simulation difficulty is systematic, predicted by an evaluator's neutral usage (most clearly on Helpfulness) and divergence from consensus; the neutral-usage tendency--rather than simulatability itself--is the cross-task-stable property (r = 0.728). These results establish both the potential and limits of evaluator-specific auxiliary data for personalized evaluation, offering methodological insights for scaling individual aware AI assessment.
arXiv:2607.05756v1 Announce Type: new
Abstract: Efficient data movement between memory and compute units is a key performance bottleneck in modern FPGA designs, particularly for deep learning (DL) workloads. In typical FPGA architectures, data transfers between block RAMs (BRAMs) and digital signal processing units (DSPs) must traverse the global routing network, leading to increased wirelength, routing congestion, and critical-path delays. Prior work has explored in- and near-BRAM compute architectures to mitigate these issues, but such solutions often require fundamental changes to FPGA architecture and CAD tools, limiting their commercial viability. This paper proposes a lightweight architectural enhancement that introduces a dedicated direct connection between BRAM and DSP blocks, enabling BRAM data to be consumed by DSPs without passing through the global interconnect. We also enhance the placement algorithm to recognize these BRAM-DSP macro blocks. The proposed architectural change incurs negligible area and delay overhead and does not affect non-DL benchmarks, while the proposed CAD remains compatible with the baseline architecture, where it yields negligible change in quality-of-results (QoR). On an Agilex-10-like FPGA, the proposed architecture and CAD updates deliver up to +25% Fmax and -49% wirelength on common DL layer designs.
arXiv:2607.05483v1 Announce Type: new
Abstract: Agentic workflows often operate over shared, structured state. Because LLM context windows are limited, each model invocation is typically shown only the state fragment needed for the current workflow step, a pattern commonly known as progressive disclosure. Modern systems construct such model-facing views using grep-like keyword search, retrieval-augmented generation (RAG), abstract-syntax-tree (AST) queries, and task-specific agent skills. These methods make the read side manageable, but they do not define when a locally proposed rewrite is valid after it is applied back to the full state. The missing piece is a contract between local updates and global validity. We introduce PatchOptic, an optic-inspired interface for shared-state LLM workflows. Optics are compositional bidirectional accessors that describe how views of structured data are read and updated. PatchOptic borrows this view/update intuition and realizes it through projected reads and verified structured patches. Each workflow step declares a projected read view, an authorized write region, and a patch-source region. Beyond runtime enforcement, the same declaration yields a path-level footprint that supports delegation, sub-workflow composition, and static certificates for reordering independent steps within the same phase. We evaluate this design with PatchBench, a benchmark with 46 cases across domains. The results show that projected reads reduce reported leakage and token cost while preserving accepted-output quality under the strong actor. Runtime verification blocks declared workflow-contract violations before commit, and patch-read enforcement rejects compromised patch artifacts that use hidden sources.
arXiv:2607.05552v1 Announce Type: new
Abstract: Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logically irrelevant changes of wording - among them an amplified yes-no bias on moral dilemmas, absent in humans. A single framing cannot say what such a shift is: in a yes/no question the word "no" is at once logical verdict, lexical token, and last-printed option. We introduce a psychometric battery that separates these: crossed symmetrization - every logically irrelevant factor flipped in balanced pairs - across a corpus of question forms. A graded rating across logically equivalent forms recovers a coherent internal moral scale: frontier models' stance $\theta$ is nearly format-invariant (cross-form incoherence 0.12-0.21 on a $\pm 1$ axis); small open-weight models fail in model-specific ways. Forcing the verdict through yes/no overlays a decomposable artifact: an order bias toward the last-printed option - opposite to classic human primacy - plus a lexical pull toward the word "no"; the artifact is substantial only in the Claude models (story-averaged -0.32 to -0.86), $\approx 0$ for GPT-5.5 and Gemini, and shrinks under extended reasoning. The word and the verdict share one token; swapping the words for arbitrary labels separates them, and the verdict-attached logical bias proves $\approx 0$ for every frontier model, while model-specific label and order attachments remain: the models are not drawn toward rejecting - the pull follows the printed surface, not the verdict it carries. A minimal model, $P = \sigma((\theta \pm m)/s)$, summarizes any such artifact by a framing susceptibility m and a moral decisiveness s, measurably distinct from sampling temperature. The battery applies unchanged to any dilemma set and binary format: measuring what a model values requires crossing the frames of the question, not asking once.
arXiv:2607.05608v1 Announce Type: new
Abstract: Collective actions emerge through the interplay between social influence and awareness. We introduce a nonlinear opinion-dynamics framework on networks in which social influence, shaped by individual interactions and community structure, promotes collective action, while awareness regulates the amount of reinforcement required for adoption. Using a degree-based mean-field reduction, we show that the competition between effective social influence and abandonment controls both the onset and persistence of collective action. Changes in awareness modify the nonlinear adoption mechanism itself, enabling populations to transition between highly responsive and weakly responsive collective states. This generates discontinuous transitions, bistability, and hysteresis, allowing collective action to persist even after the conditions that initially promoted it have weakened. We illustrate these effects through coupled disease-mitigation and resource-consumption dynamics, where external pressures act by reshaping awareness rather than social influence. More broadly, our results identify awareness as a fundamental link between environmental conditions, social interactions, and the emergence and persistence of collective action.
arXiv:2607.05618v1 Announce Type: new
Abstract: For many synchrotrons which cross $\gamma_t$ transition, a high intensity threshold has been identified beyond which significant beam losses are observed. These losses are due to a high-intensity mode coupling instability, which can be suppressed by adjusting the chromaticity curve to cross zero slightly after $\gamma_t$ transition.
arXiv:2607.06524v1 Announce Type: cross
Abstract: The Vietoris-Rips filtration $\mathcal{VR}(-)$ is a standard tool for analyzing the shape of data within topological data analysis. Beginning with seminal work of Sheehy, a substantial amount of research has centered on constructing linear-size sparse approximations to $\mathcal{VR}(-)$ and related filtrations for metric spaces of bounded doubling dimension. We show that this geometric assumption is necessary in a precise sense. Working in the framework of homotopy interleavings, we show that for any fixed $c \in [1, \sqrt{2})$, there exists a family of finite metric spaces for which any finitely presented $c$-approximation to $\mathcal{VR}(-)$ has exponential size. We also show that for any fixed $c \geq 1$, there exists a family of finite metric spaces for which any finitely presented $c$-approximation to $\mathcal{VR}(-)$ has superlinear size, yielding an obstruction to linear-size approximations for any fixed approximation factor. Both results extend to the intrinsic \v{C}ech filtration and to any bifiltration containing $\mathcal{VR}(-)$ as a $1$-parameter slice, including the function-Rips, degree-Rips, and subdivision-Rips bifiltrations.
arXiv:2607.05752v1 Announce Type: new
Abstract: Search-augmented language models can use external evidence to compensate for limitations in parametric knowledge, but search is not uniformly beneficial: models may call search for questions they can already answer, or rely on noisy evidence when correction, clarification, or abstention would be more appropriate. We formulate this as an instance-level search-routing problem: deciding whether search is needed to improve task success relative to a no-search execution. To derive supervision, we compare no-search and forced-search outcomes for the same question and construct an oracle over NO SEARCH, SEARCH, and UNSOLVED based on task-specific success. Using this oracle as both an evaluation criterion and a learning signal, we train search-routing policies with supervised fine-tuning and preference optimization, improving routing macro-F1 on oracle-eligible examples from 0.7082 to 0.8235 for Gemma E2B and from 0.7053 to 0.8365 for Qwen3.5-4B. Further analysis shows that the learned policies reduce model-specific routing failures: Gemma primarily learns no-search restraint, while Qwen further reduces missed search; residual UNSOLVED cases reveal heterogeneous bottlenecks involving model capacity, retrieval budget, evidence use, and policy behavior.
arXiv:2607.05398v1 Announce Type: new
Abstract: Personas are often employed to guide large language model agents, yet their effectiveness in shaping strategic behavior in social dilemma settings remains uncertain. To address this, we examined the impact of persona prompts in an iterated Split or Steal game where persona-driven agents interacted with a Virtual Human (VH) controlled by a fixed prompt. Agents were instantiated from four open models (Ministral 3:3b, phi4:14b, Gemma3:12b, and Gemma4:e4b) at two temperature settings (0.3 and 0.7) and deterministic decision with zero temperature, while the VH was powered by GPT 4.1 mini. Across 160 sessions of 15 rounds each conducted in European Portuguese, mutual Split outcomes dominated (roughly 74 percent of rounds), with exploitation occurring in fewer than 11 percent of rounds. Model choice significantly influenced behavior: phi4 and Ministral 3:3b remained consistently cooperative across temperatures, whereas Gemma3:12b and Gemma4:e4b exhibited more varied strategies and outcomes. Analyses based on Big Five personality traits indicated that Prosocial and Principled personas were most consistently cooperative, while Analytical personas were more likely to exploit the VH. Topic analysis revealed that friendship-related dialogue aligns with Split decisions, whereas money and vengeance-related content is more prevalent in Steal outcomes; sentiment labels were predominantly neutral or happy and provided limited additional explanatory value. These findings characterize the interaction between persona prompts and model differences in repeated trust games and serve as a baseline for planned virtual reality studies involving human participants interacting with an embodied VH.