Forskningsradar

Science Journals

Peer-reviewade publikationer — 56237 artiklar

Appraisal Dimensions Generalise Better than Emotion Labels for Cross-Age Affect Recognition in AI-Assisted Healthcare
arXiv:2604.27938v2 Announce Type: replace Abstract: The integration of artificial intelligence (AI) into healthcare has advanced significantly, yet affect recognition remains a major challenge, particularly in AI-assisted interventions such as Computerized Cognitive Training (CCT). The THERADIA-WoZ corpus was developed to enable multimodal affect recognition in the context of AI-driven CCT, focusing on an older adult population. This study extends the corpus by introducing a dataset collected from young adults, allowing direct comparison of affect recognition models across age groups. Our objective was to assess whether multimodal models based on dimensions borrowed from appraisal theories outperform those based on categorical labels and to evaluate their generalisation power across age corpora. After comparing both corpora, models were trained and tested using within-corpus, cross-corpus, and mixed-corpus evaluation. Results revealed that appraisal dimensions consistently outperformed categorical labels across all conditions, demonstrating greater predictive accuracy and stability. Notably, categorical labels failed to generalise across age corpora, as performance dropped to chance levels in cross-corpus evaluation. In contrast, appraisal dimensions maintained predictive performance above chance, reinforcing their robustness for cross-age affect recognition. Furthermore, training on both corpora did not improve generalisation beyond within-corpus training. The findings support the theoretical and practical advantages of appraisal dimensions over categorical labels in affective computing. They also highlight the importance of multimodal fusion and deep learning representations for emotion modeling. To facilitate future research, we provide an API for researchers interested in time-continuous emotion prediction, offering valuable tools for behavioral sciences to enhance the measurement of emotional states in various experimental settings.
SRA: Span Representation Alignment for Large Language Model Distillation
arXiv:2605.01205v2 Announce Type: replace Abstract: Cross-Tokenizer Knowledge Distillation (CTKD) enables knowledge transfer between a large language model and a smaller student, even when they employ different tokenizers. While existing approaches mainly focus on token-level alignment strategies, which are often brittle and sensitive to discrepancies between tokenizers, we argue that the method of aggregating tokens into more robust representations before distillation is of equal importance. In this paper, we introduce \textbf{SRA} (\textbf{S}pan \textbf{R}epresentation \textbf{A}lignment for Large Language Model Distillation), a novel framework that reframes CTKD through the physical lens of Multi-Particle Dynamical Systems. SRA shifts the fundamental unit of alignment from tokens to robust, tokenizer-agnostic spans. We model each span as a cluster of particles and represent its state by its Center of Mass (CoM) - an attention-weighted average that captures rich semantic information. We leverage the concept of span centers of mass with attention-derived weighting to prioritize the most salient spans. In addition, we employ a geometric regularizer to preserve the structural integrity of the representation space and introduce aligned span logit distillation to enhance knowledge transfer across models. In challenging cross-architecture distillation experiments, SRA consistently and significantly outperforms state-of-the-art CTKD baselines, validating our physically-grounded approach.
MTA: Multi-Granular Trajectory Alignment for Large Language Model Distillation
arXiv:2605.01374v2 Announce Type: replace Abstract: Knowledge distillation is a key technique for compressing large language models (LLMs), but most existing methods align representations at fixed layers or token-level outputs, ignoring how representations evolve across depth. As a result, the student is only weakly guided to capture the teacher's internal relational structure during distillation, which limits knowledge transfer. To address this limitation, we propose Multi-Granular Trajectory Alignment (MTA), a framework that aligns teacher and student representations along their layer-wise transformation trajectory. MTA adopts a layer-adaptive strategy: lower layers are aligned at the word level to preserve lexical information, while higher layers operate on phrase-level spans (e.g., noun and verb phrases) to capture compositional semantics. We instantiate this idea through a Dynamic Structural Alignment loss that matches the relative geometry among semantic units within each layer. This design is motivated by empirical findings that Transformer representations become increasingly abstract with depth, and is also consistent with linguistic views in which higher-level meaning emerges through the composition of lower-level lexical units. We further incorporate a Hidden Representation Alignment loss to directly align selected teacher-student layers. Experiments show that MTA consistently outperforms state-of-the-art baselines on standard benchmarks, with ablations confirming the contribution of each component.
Lean-GAP: A Dataset of Formalized Graduate Algebra Problems
arXiv:2606.02588v1 Announce Type: new Abstract: We present Lean-GAP (Lean-Graduate Agebra Problems), 430 formalized graduate-level algebra problems from the textbook Abstract Algebra by Dummit and Foote. We develop a scalable pipeline consisting of PDF-to-LaTeX preprocessing, autoformalization into Lean 4, and verification of informal-formal correspondence. While the preprocessing and autoformalization stages can be largely automated, we find that verification remains the most subtle and labor-intensive component, requiring careful human oversight. Our contributions include (i) the construction of a structured dataset of formalized exercises, (ii) a systematic methodology for formalizing textbook mathematics, and (iii) an analysis of recurring challenges in the formalization process. We also compare the performance of different autoformalization models and highlight key bottlenecks in translating informal statements into formal language.
Physics-Informed Neural Network for Diffusion-Reaction Problems with Dead-Core Formation in Catalyst Slabs
arXiv:2606.02599v1 Announce Type: new Abstract: This work investigates a nonlinear two-point boundary value problem arising in diffusion-reaction processes in catalyst slabs with power-law kinetics and fractional reaction order. The model exhibits a free-boundary structure, where an unknown interface separates a dead-core region with vanishing concentration from an active region with positive concentration. We propose a Physics-Informed Neural Network (PINN) framework that incorporates a structured, hard-constrained trial solution embedding the asymptotic behavior near the interface. The dead-core location is treated as a trainable parameter, enabling the simultaneous approximation of the concentration profile and identification of the free boundary without explicit interface tracking. The method is validated against analytical solutions and high-precision numerical shooting. Numerical experiments demonstrate that the approach accurately captures both the solution profile and the free-boundary location while maintaining a computationally manageable training cost.
$A^2$: Smaller Self-Supervised ViTs Localize Better than Larger Ones
arXiv:2606.03148v1 Announce Type: new Abstract: Robust visual classification often depends on localizing the main foreground objects in an image while ignoring contextual distractors. Surprisingly, we find that the attention maps of smaller self-supervised ViTs localize foreground objects better than those of larger ViTs. However, we still need large ViTs, because they extract richer representations from each patch. To get the best of both worlds, good localization and rich representations, we propose $A^2$, a simple method that leverages this inverse scaling finding by decoupling where to look (a small attention model) from what to extract (a large embedding model): we crop around the attention peaks of a small model and embed the crops with a larger model. $A^2$ uses entirely pretrained features, requires no group labels, and does not require per-dataset attention or backbone training. Across 5 benchmarks, $A^2$ is competitive with backbone-matched loss-level methods like DFR, and outperforms end-to-end attention training under stronger distribution shifts.
Late-Time Cosmology and Structure Formation in Quadratic $f(Q)$ Gravity
arXiv:2606.02660v1 Announce Type: new Abstract: We investigate the cosmological evolution associated with the quadratic symmetric teleparallel gravity framework, \( f(Q)=Q+\alpha Q^{2}+\beta \) where the relation \(Q\propto H^{2}\) generates an additional \(H^{4}\) contribution to the Friedmann equation. Using the exact algebraic solution for $H(z)$, we reconstruct the effective dark-energy sector and compare the background evolution with $\Lambda$CDM using Type Ia supernovae, BAO, and cosmic-chronometer data. At the perturbative level, the model modifies the Poisson equation through a time-dependent effective gravitational coupling $G_{\textrm eff}(z)=G\big[1+\tfrac{2}{3}A E^{2}(z)\big]^{-1}$, where $A=18\alpha H_{0}^{2}$. For $\alpha>0$ this produces a weakened gravitational interaction, suppressing the linear growth factor $D(z)$, the growth rate $f(z)$, and the RSD observable $f\sigma_{8}(z)$. In the nonlinear regime, the reduced gravitational strength increases the spherical-collapse threshold and suppresses the halo mass function, leading to a lower predicted value of $S_{8}=\sigma_{8}\sqrt{\Omega_{m}/0.3}$. Thus, the quadratic $f(Q)$ extension can reproduce mild deviations from $\Lambda$CDM at the background level while naturally alleviating the $S_{8}$ tension, offering a viable modified-gravity explanation for recent observational hints of dynamical dark energy.
Improvise, Adapt, Overcome: An On-The-Fly Multifidelity Algorithm for Efficient Machine Learning
arXiv:2606.02662v1 Announce Type: new Abstract: Machine learning has accelerated quantum chemistry but is hindered by the prohibitive cost of generating high fidelity training data. Multifidelity machine learning (MFML) mitigates this overhead by systematically combining abundant low fidelity data with sparse high fidelity data. In spite of its success, standard MFML schemes rely on pre-defined scaling factors to determine sparse data ratio across fidelities, often generating redundant multifidelity data resulting in a loss of efficiency. Here, we introduce an adaptive on-the-fly multifidelity framework for machine learning that autonomously determines training dataset composition. By dynamically querying training samples at each fidelity, the algorithm saturates model accuracy at lower fidelities before moving up to more expensive reference calculations. We benchmark the novel adaptive-MFML across diverse chemical properties including the computational chemistry gold standard coupled cluster energies, and the more chemically challenging excitation energies. In our numerical experiments we show that our adaptive algorithm reduces data generation costs by up to a factor of 30 compared to single fidelity methods and improves upon standard MFML by up to a factor of 5. The mitigation of data redundancy establishes a high-accuracy low-cost pathway for sustainable cost-aware machine learning in quantum chemistry.
What You Approve Is What Executes: Consent Integrity for Black-Box LLM Agents
arXiv:2606.02668v1 Announce Type: new Abstract: Coding agents gate consequential actions behind a human-in-the-loop approval dialog, but the dialog is narrated by the agent itself: the human approves a summary the agent writes. The Lies-in-the-Loop (LITL) attack shows that summary is forgeable, so a compromised agent can show a benign description while a different action runs. This paper names the missing property, Consent Integrity, by importing What You See Is What You Sign (WYSIWYS) and the trusted-path property into the agent approval channel: the action shown to the human must be rendered by a trusted mediator from the real action at the boundary, not the agent's narration, over a path the agent cannot spoof, and bound to the exact action that executes. Two twists distinguish it from classical WYSIWYS: the renderer is the adversary, and the boundary ground truth is a low-level event that must be decoded without trusting the agent. Since no decoder is complete, the realizable target is analyzer-relative: whatever the analyzer cannot classify is surfaced as uninspectable rather than silently approved. A prototype implements the analyzer, renderer, and bind-to-execution; total mediation and the trusted path are specified but assumed, not implemented. On GTFOBins, an independent corpus of 1330 trusted-tool abuses, the prototype silently passes 10.0% (every instance through a trusted tool); on tldr, 28,798 normal-usage commands, it marks 87.0% uninspectable. These two independent measurements bracket the design's central tension: the trust list that bounds silent passes is the same one that drives over-prompting, and a boundary-only mediator can move along that frontier but not escape it. The contribution is the property, the mechanism, and an honest position on that frontier, not a solved defense.
ANCHOR: Abductive Network Construction with Hierarchical Orchestration for Reliable Probability Inference in Large Language Models
arXiv:2605.10328v3 Announce Type: replace Abstract: A central challenge in large-scale decision-making under incomplete information is estimating reliable probabilities. Recent approaches use Large Language Models (LLMs) to generate explanatory factors and coarse-grained probability estimates, which are then refined by a Na\"ive Bayes model over factor combinations. However, sparse factor spaces often yield ``unknown'' predictions, while expanding factors increases noise and spurious correlations, weakening conditional independence and degrading reliability. To address these limitations, we propose \textsc{Anchor}, an aggregated Bayesian inference framework over a hierarchical factor space. It constructs dense factor hierarchies through iterative generation and clustering, maps contexts via hierarchical retrieval and refinement, and augments Na\"ive Bayes with a Causal Bayesian Network to model latent factor dependencies. Experiments show that \textsc{Anchor} markedly reduces ``unknown'' predictions and produces more reliable probability estimates than direct LLM baselines, achieving state-of-the-art performance while significantly reducing time and token overhead.
Sensitivity Enhancement of S-Band Rydberg Atom Microwave Receiver Using Resonant Cavity
arXiv:2606.02669v1 Announce Type: new Abstract: Rydberg atom-based microwave electric field sensing has attracted growing interest owing to its inherent advantages, such as absolute calibration, wideband operability, and compatibility with room-temperature devices. A critical bottleneck that limits sensitivity is the inefficient coupling between the Rydberg atoms and the incident microwave field, particularly when detecting weak signals propagating in free space. Here we propose and experimentally validate a scheme that integrates a horn antenna with a resonant microwave cavity to significantly improve this coupling for free-space signal reception in the S-band. Using a two-photon excitation scheme in a cesium vapor cell, we systematically characterize the sensing performance under three configurations: a bare cell, direct cavity injection, and a cavity coupled to a horn antenna that captures free-space microwave signals over a 1 m distance. In the antenna-coupled cavity configuration, we achieve an optimal sensitivity of 2.33 nV/cm/$\sqrt{\text{Hz}}$ at the receiving antenna, which corresponds to an enhancement of approximately 17.9 dB compared to the optimized bare vapor cell configuration. Our findings offer a practical and effective route to boost the sensitivity of Rydberg atomic sensors, facilitating their adoption in real-world microwave metrology and wireless communication applications where weak free-space electric fields must be reliably measured.
Visual Graph Scaffolds for Structural Reasoning in Large Language Models
arXiv:2606.02673v1 Announce Type: new Abstract: Graphs have been used to enhance large language models (LLMs) for structured reasoning, mostly as external knowledge sources are provided to models at test time. In this paper, we take a different view: the value of graphs for LLMs lie not only in supplying information, but also in organizing reasoning. Inspired by how humans use graph-structured mind maps to organize branching and converging thoughts, we ask whether graphs can serve as an internal form of reasoning assistance. We study this question on multi-hop question answering tasks, where teacher-provided reasoning traces are rewritten as graph mind maps and used to guide a student model. Our experiments reveal a clear modality gap. When graph structures are flattened into text, their benefits become limited once direct answer hints are removed. Under this abstract guidance setting, both reasoning efficiency and answer quality degrade substantially. In contrast, visual graph guidance remains effective without direct answer clues, and its advantage persists after supervised fine-tuning and KL-based distillation. The above findings support the claim that graphs should be studied not only as external knowledge structures for LLMs, but also as visual scaffolds for organizing reasoning.
Pretraining Language Models on Historical Text
arXiv:2606.02991v1 Announce Type: new Abstract: We introduce TypewriterLM, a 7.24B History language model (LM) trained exclusively on English text predating 1913. Developing History LMs requires addressing challenges in data quality and availability, preventing temporal leakage, designing temporally consistent post-training pipelines, and constructing reliable evaluations. To address these issues, we construct TypewriterCorpus, a 54B-token historical corpus collected from diverse archival and linguistically annotated sources with extensive data cleaning and leakage mitigation procedures. Furthermore, we introduce lexically grounded instructing tuning, a post-training framework that constraints responses to remain directly grounded in historical source documents. Using this framework we construct two historical instruction tuning datasets: History-LIMA and History-SelfInstruct. To evaluate capability and temporal consistency, we introduce History-Event, a benchmark suite for evaluating competence, temporal grounding and data leakage. We release TypewriterLM and all associated resources to support future research on historical language models.
MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control
arXiv:2605.26006v2 Announce Type: replace Abstract: Enabling physics-based humanoids to execute diverse behaviors from high-level textual commands remains a significant challenge. Existing methods typically follow either a two-stage paradigm that combines kinematic motion generation with physics-based tracking, or an end-to-end imitation-learning paradigm that directly generates actions from text. However, the former suffers from the inherent domain shift between kinematic generation and physics-based tracking, while the latter struggles with the substantial modality gap between textual commands and low-level actions, limiting effective semantic alignment. Notably, humanoid states encode rich motion dynamics that are more semantically aligned with textual descriptions than low-level actions, making them a natural basis for deriving behavioral intent. Building upon this insight, we propose MIND, a novel end-to-end diffusion framework for text-driven physics-based humanoid control that leverages behavioral intent as a semantic bridge between textual commands and low-level actions. At its core, MIND introduces a multi-scale intent diffusion mechanism, where a holistic intent predictor captures global behavioral dynamics to guide overall behavior synthesis, while an immediate intent predictor provides step-wise, fine-grained signals for local behavior refinement at each diffusion step. This hierarchical intent formulation imposes a structured inductive bias for humanoid control, improving semantic alignment and behavioral naturalness. Furthermore, MIND encodes humanoid states into a latent space to enable more effective semantic intent modeling. Extensive experiments demonstrate that MIND outperforms existing methods and synthesizes coherent, physically plausible, and semantically aligned humanoid behaviors from text commands. Project page: https://binlee26.github.io/MIND_page.
MUSE: A Unified Agentic Harness for MLLMs
arXiv:2606.03005v1 Announce Type: new Abstract: Despite rapid progress, multimodal large language models (MLLMs) still fail on tasks that humans solve effortlessly, such as navigating a grid maze from a screenshot or selecting the correct puzzle piece. Rather than retraining the model, we ask a complementary question: how much capability can be elicited from a frozen MLLM purely by improving the execution scaffold around it? We introduce MUSE, a multimodal unified structured execution harness that wraps any off-the-shelf MLLM with composable modules for task representation, visual processing, perception tool use, structured parsing, deterministic verification, and verifier-guided repair, without any model retraining. We evaluate MUSE across diverse benchmarks spanning visual spatial planning, visual perception, multimodal reasoning, and fine-grained visual discrimination, using multiple state-of-the-art MLLMs. MUSE delivers consistent gains over the bare model in all settings, with the largest jumps on challenging instances. Further analysis reveals that many MLLM failures arise from harness-level shortcomings rather than fundamental model deficits, and can be addressed through verifier-guided repair without touching the model. These findings highlight the agentic multimodal harness as a critical yet underexplored design dimension, offering an orthogonal avenue for improving MLLMs beyond model-centric optimization.
Secure AltDA Integration for Ethereum L2s: An End-to-End Validation Framework
arXiv:2606.03010v1 Announce Type: new Abstract: Alternative data availability (AltDA) systems provide Ethereum L2s with an external data publication layer for high throughput rollup designs. By moving bulk data publication outside of Ethereum, AltDA allows L2s to process more data than native DA. However, this replacement introduces a new consensus critical integration layer. Existing ecosystem frameworks identify high level risks, such as external DA trust assumptions and the presence or absence of a DA verifier, but do not provide a complete specification for how an L2 should integrate with AltDA. This gap can lead to L2 halts, inconsistent derivation across honest L2 nodes, invalid state assertions, or bridge attacks. This paper presents a canonical validation framework for secure AltDA integration. We model the boundary as a typed, deterministic, and total translation from L1 inbox bytes to an AltDA commitment, then to externally available data, and finally to the rollup payload consumed by the rest of core L2s logic. The central principle is that every adversarial input must lead to a defined unique outcome. We show how missing obligations lead to concrete failure modes, including underconstrained settlement, derivation halts, inconsistent honest node behavior, invalid state assertions, and bridge safety failures. We then apply the framework to representative AltDA integration architectures, including Celestia-Blobstream, EigenDA based designs, and Avail-ZKsync. Our evaluation shows that secure AltDA integration is not determined solely by the DA provider or bridge. The surrounding L2 integration must also enforce the full validation relation connecting L1 inbox inputs to accepted L2 state.
MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data
arXiv:2606.02753v1 Announce Type: new Abstract: Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a single perspective. Extending these models to multi-agent settings introduces two critical challenges: data scarcity (coordinated multi-view recordings are prohibitively expensive to collect for general open-domain scenarios) and world state alignment (independently generated video streams cannot ensure that shared physical environments and events evolve consistently across views). To address these challenges, we propose MetaWorld, a novel framework that scales multi-agent video world models to open-domain environments directly from single-view videos. First, we introduce Monocular World-State Unrolling (MWSU) to explicitly decompose monocular footage into the camera operator's ego-motion and the visible subject's spatial trajectory. This camera-trajectory decomposition naturally extracts synchronized multi-agent motion data within a shared 3D space, completely bypassing the need for multi-camera setups. Second, for precise visual control, we develop the Subject-Aware World Generator to enable appearance-driven simulation conditioned on per-agent identity images. Finally, to ensure both views are grounded in the identical physical reality, we propose World-State Alignment, a per-frame inter-branch cross-attention mechanism inserted at every transformer layer of the video DiT. By jointly synchronizing the denoising process, WSA enforces both static geometric consistency and dynamic motion consistency, encouraging that the shared 3D environment and physical events remain well-aligned across both egocentric views. Extensive experiments demonstrate that MetaWorld achieves superior cross-view consistency and identity fidelity, establishing a highly scalable, physics-driven paradigm for multi-agent video world modeling.
$\Psi$-Bench: Evaluating Persona-Sensitive Influencing in Persuasive Dialogues
arXiv:2606.02754v1 Announce Type: new Abstract: Personalization is a crucial capability of modern language agents. However, current research primarily positions personalized agents as passive responders to user preferences, limiting their ability to interact with users and provide suggestions or guidance proactively. To systematically evaluate such proactive personalization in realistic interactions, we propose $\Psi$-Bench, a benchmark for assessing LLMs' ability to influence realistic users through conversation. We design three real-world interaction scenarios that involve persuasion in $\Psi$-Bench, and endow simulated clients with personal characteristics through explicit user profiles derived from dialogue histories. We evaluate 10 frontier LLMs on $\Psi$-Bench and find that while most models can produce coherent and reasonable arguments, even state-of-the-art models still leave considerable room for improvement in persuasion. We also find that providing access to client profiles yields an average performance gain of 18.24\%, highlighting the importance of user-specific information for effective persuasion. Overall, our work highlights persona-sensitive influencing as a challenging yet practical direction for evaluating and developing more proactive personalized LLM agents. Codes are available at: https://github.com/Hanpx20/Psi-Bench.
Impurity-driven turbulence opens a pathway to ELM-free operation and enhanced pedestal stability in tokamaks
arXiv:2606.02768v1 Announce Type: new Abstract: Edge-localized modes (ELMs) impose severe transient heat, and particle loads on plasma-facing components, posing a critical challenge for steady-state operation of tokamak fusion reactors. Existing ELM control techniques either rely on externally applied perturbations or operate within narrow parameter windows, raising concerns for reactor scalability. Here we demonstrate that controlled injection of a low-Z impurity can fundamentally modify pedestal transport and stability, enabling access to long ELM-free periods through impurity-driven turbulence. Using boron (B) powder injection in the DIII-D tokamak, we observe a progressive reduction of ELM frequency, culminating in long ELM-free phases. Pedestal stability analysis reveals a pronounced decoupling of peeling and ballooning stability boundaries at moderate B injection levels, opening a stability channel toward super-high confinement operation. At higher injection rates, long (~300 ms) ELM-free periods are achieved. Fluctuation measurements show that B injection selectively enhances low-frequency pedestal turbulence, increasing inter-ELM particle transport and regulating pedestal gradients. The establishment of a feedback loop between turbulence, particle transport, and the resulting modification of pedestal conditions, indicated by the observed hysteresis loop in the evolution of density fluctuations in response to the B injection rate, is presented.
Hanger Reflex Based Driving Assistance for Drivers with Peripheral Visual Field Defects
arXiv:2606.03020v1 Announce Type: new Abstract: Drivers with peripheral visual field defects may fail to notice pedestrians in their peripheral visual field, leading to delayed hazard awareness and increased collision risk. This study explores hanger reflex cue (HRC) as a driving assistance method for drivers with peripheral visual field defects, in which mechanical pressure is applied to specific regions of the head to facilitate anticipatory orientation toward potentially risky pedestrians and support safer driving. In a driving simulator experiment with 15 participants, we compared driving behavior with and without HRC during pedestrian encounters under simulated peripheral visual field defect. The results showed that HRC significantly shifted drivers' modal head rotation angle toward the risky pedestrian and significantly increased gaze duration toward that pedestrian. Collision occurrence was lower in the w/ HRC condition than in the w/o HRC condition, although the direct effect of HRC on collision occurrence showed only a marginal trend. A piecewise structural equation modeling analysis further suggested that HRC may contribute to collision reduction through a sequential pathway from head rotation to gaze allocation and then to collision occurrence. These findings provide preliminary evidence that HRC can support anticipatory attention allocation toward peripheral hazards and may offer a promising driving assistance method for drivers with visual field impairment.
Disaster-induced behavioral change restructures social networks toward bonding ties
arXiv:2606.02790v1 Announce Type: new Abstract: Population displacement following environmental shocks reshapes the spatial organization of social interactions, often fragmenting existing ties and weakening community cohesion. Although social capital is widely recognized as a key determinant of resilience, its dynamic restructuring after disruption remains poorly quantified. Here, we develop a spatially embedded, dynamic network framework that operationalizes social capital as a network of repeated encounter opportunities inferred from large-scale mobility data. We construct temporal co-presence networks at third places to track how socio-spatial networks reorganize under disruption. We apply this framework to communities affected by the 2021 Marshall Fire in Colorado. We find that disaster-induced displacement leads to substantial contraction of socio-spatial networks, with mean weighted degree decreasing by 48%. To isolate underlying mechanisms, we develop two counterfactual models: a random node removal model and a behaviour-informed model in which individuals are removed based on their estimated propensity to evacuate. Both counterfactuals predict substantially lower connectivity than observed, indicating that post-disaster connectivity remains systematically higher than expected based on displacement behavior alone. Structural analysis of the network reveals that this residual connectivity is disproportionately concentrated among bonding ties between sociodemographically similar individuals, while bridging ties are comparatively fragile. Furthermore, interaction becomes increasingly located around third places, suggesting that these places act as spatial anchors for the persistence of social ties under disruption. Together, these findings provide a first empirical view of how behavioral responses to disruption shape community resilience through the reorganization of social networks.
Report on the Designing Accountable Software Systems Workshop
arXiv:2606.02804v1 Announce Type: new Abstract: The Workshop on Designing Accountable Software Systems (DASS) was convened in November 2024 with support from the U.S. National Science Foundation to engage a wide range of current and future stakeholders from government, academia, and industry on the cross-disciplinary topic of accountability in software systems. Over two days, attendees engaged in a series of panels, invited talks, and breakout sessions covering: (1) the dimensions of accountability, including legal compliance as well as business and societal aspects and drivers; (2) a conceptual model of the various structures needed to realize accountability; (3) the sources of legal requirements that affect software; (4) the operationalization of legal requirements in software; (5) the requirements to preserve evidence needed to conduct investigations; and (6) a range of challenges and contextual factors beyond software that affect why some accountability structures succeed, while others fail. The workshop was conducted as a collaborative systematization of knowledge that culminated in several research directions. The findings include the importance of clarifying definitions and responsibilities within accountable organizations, which can affect whether those researching accountability are making assumptions that limit the generalizability of findings. Further research was also identified as needed to study the ways to improve the translation of accountability structures into the software design process while improving engagement with stakeholders, such as legislators, regulators, business executives and system developers. Finally, a key finding was the high demands that DASS-like research projects place on interdisciplinary teams: both in terms of team formation and sustainment, as well as, the specific demands of cross-disciplinary learning that covers both research methods, research dissemination, and career development.
Do Neural Retrievers Prefer Certain Documents? Evidence of Learned Relevance Priors
arXiv:2606.02814v1 Announce Type: new Abstract: Neural retrievers are trained to estimate query-document relevance from annotated query-document pairs. Yet annotation protocols may not purely reflect relevance: they select only a subset of documents for labeling, and this selection can favor certain document types over others. We investigate whether supervised bi-encoder retrievers implicitly learn a document-level relevance prior: a query-independent signal encoded in their representation space as a side effect of training on annotated data. We estimate this prior by training simple classifiers on frozen document embeddings and evaluate three state-of-the-art retrievers across multiple IR benchmarks. We find that supervised neural retrievers encode relevance priors that generalize to unseen documents and are consistent across models. These priors create a findability gap: documents with lower prior are systematically harder to retrieve, even when genuinely relevant. This effect appears in supervised dense retrievers but is weaker and less consistent in BM25, and it persists under controlled matched-document comparisons. Using LLM-based explanations, we find that judged-relevant documents tend to be comprehensive, self-contained summaries of mainstream topics, while niche, fragmentary, or highly technical content is often left unjudged. Retrievers internalize this bias, ranking documents with these favored features higher than documents that lack them, independently of their actual relevance. Our findings expose a structural limitation of supervised retrieval: models trained on annotated data do not just learn relevance, but also the implicit document preferences in their training data.
Fairness as an Investment: Dynamic Participation and Long-Run Profit in Virtual Power Plants
arXiv:2606.02820v1 Announce Type: new Abstract: We show that incorporating fairness constraints into virtual power plant (VPP) operations can incentivize consumer participation and thus improve the aggregator's long-run profitability. VPPs rely on sustained participation from heterogeneous consumers to provide a variety of grid services whose timing and frequency are often uncertain. As a result, consumers' willingness and ability to provide flexibility evolve over time, creating a dynamic link between past participation and future resource availability. We develop a dynamic aggregation framework to study how fairness in service allocation affects future participation and long-run profitability. By linking current dispatch decisions to future resource availability, we show that fairer allocations can strengthen consumer engagement, expand aggregate availability, and create additional value during high-price and high-demand events. To balance fairness and operational efficiency, we introduce a slack-augmented allocation mechanism that preserves most of the participation benefits from fairness while avoiding unnecessary reductions in service procurement. We derive conditions under which the resulting availability gains outweigh the short-run cost of redistribution and validate the approach using real-world consumer behavior and electricity market data from Norway.
Inducing Reasoning Primitives from Agent Traces
arXiv:2606.02994v1 Announce Type: new Abstract: ReAct-style LLM agents often rediscover the same reasoning routines across problems, yet leave those routines trapped in transient scratchpads. We introduce Reasoning Primitive Induction, a single-pass method that mines successful ReAct traces, clusters recurrent reasoning moves, and converts the most frequent moves into a compact library of typed pseudo-tools. Each pseudo-tool is specified by a natural-language docstring interpreted by an LLM at invocation time, and a standard ReAct loop composes these primitives at test time. The central result is that induced libraries outperform the very agent that generated their traces: by +44pp on RuleArena NBA (30 -> 74), +30pp on MuSR team allocation (38 -> 68), and +22pp on NatPlan meeting planning (7 -> 29). Across five comparable subtasks spanning narrative deduction, rule application, and constraint-satisfaction planning, a single fixed configuration improves over zero-shot Chain-of-Thought on every subtask, matches or surpasses expert-authored decompositions, and outperforms AWM at lower average inference cost.