Forskningsradar

Science Journals

Peer-reviewade publikationer — 56239 artiklar

Full-Field Calibration of Coupled Thermomechanical Material Models at Finite Strain
arXiv:2606.05465v1 Announce Type: new Abstract: Calibrating thermomechanical material models from experiments is challenging because deformation, temperature, and force responses are strongly coupled, while measurements are usually restricted to specimen surfaces. We present a full-field calibration framework for coupled finite-strain thermomechanical material models using boundary displacement, reaction-force data, and temperature. The forward model is formulated as a near-incompressible thermo-hyperelastic problem with thermomechanical coupling derived from a Helmholtz free energy, and the inverse problem is posed as a PDE-constrained optimization problem with weighted observation terms for the available data streams. Reduced gradients are computed with adjoint sensitivities that are obtained by automatic differentiation, enabling gradient-based calibration of nonlinear transient thermomechanical systems. The formulation is first verified on synthetic examples involving uniform thermal preconditioning and localized transient rod contact, where the ground-truth parameters are recovered from full-field measurements and force observations. The same workflow is then applied to experimental thermomechanical data by first calibrating a hyperelastic mechanical baseline from cyclic equibiaxial loading and subsequently identifying thermal expansion and directional shrinkage parameters from surface-temperature and boundary-force histories. The results demonstrate that coupled thermomechanical parameters can be inferred from experimentally accessible surface data without requiring volumetric observations.
Robust Scene Transfer for PointGoal Navigation via Privileged Sensor Guided Contrastive Learning
arXiv:2606.05506v1 Announce Type: new Abstract: We propose a sensor-guided adaptive contrastive learning framework for visual representation learning in PointGoal navigation. During training, privileged LiDAR sensing guides the contrastive objective through a geometry-aware similarity metric and adaptive temperature scaling, encouraging visual embeddings to capture navigation-relevant structure rather than scene-specific appearance. The resulting encoder is pretrained independently, frozen, and used as the perceptual backbone for reinforcement learning, decoupling representation learning from policy optimization. We further introduce a cross-stage domain mismatch between representation pretraining and policy learning to suppress environment-specific shortcuts and promote reliance on task-relevant features. Extensive experiments in high-fidelity simulation demonstrate that our approach significantly improves policy-level scene transfer across diverse indoor and outdoor environments. At deployment, the agent relies only on monocular RGB observations together with standard task-related inputs such as goal position and proprioceptive signals, without access to LiDAR or other privileged sensors. Our method outperforms large pretrained vision models and standard contrastive baselines under severe appearance and semantic shifts. We also release a multimodal dataset to support future research on privileged-guided visual representation learning for navigation. The code is available at:
The Role of Instructional Guidance in Generative AI-Assisted Learning: Empirical Evidence from Construction Engineering Education
arXiv:2606.05509v1 Announce Type: new Abstract: Generative artificial intelligence (AI) is increasingly used to support self-directed learning, yet student interaction with such systems often remains unstructured, limiting engagement in deeper cognitive processes. This study examines how instructional guidance shapes student and AI interaction in construction education. A five-step prompting framework grounded in Generative Learning Theory (GLT) is introduced to guide learner interaction during review activities. A controlled experiment compares three learning conditions: slide-based learning, unprompted AI-supported learning, and prompted AI-supported learning. Learning performance is assessed using multiple-choice and open-ended tasks, and user experience is measured using the User Experience Questionnaire (UEQ). Performance differences are concentrated on tasks requiring explanation and reasoning. The prompted condition achieves higher open-ended scores, with an improvement of approximately 2 or 3 points on a scale of 18 (p < 0.01), while no significant differences are observed in multiple-choice performance. The unprompted condition remains comparable to slide-based learning. These findings indicate that the effectiveness of AI-supported learning depends on how interaction is structured. The proposed framework provides a basis for integrating learning science principles into generative AI systems for construction education.
The Dignity-Centric Stack: A Commons-Governed, Horizontally Federated Architecture for Human-Dignity AI
arXiv:2606.06083v1 Announce Type: new Abstract: The human-dignity-centric digital social contract grounds personal data in human dignity, data personalism, and data sovereignty, and articulates six dimensions of data governance: technological oversight, automation limits, economic justice, political legitimacy, social cohesion, and legal guarantees. It presupposes, however, that enforcement falls to State regulators, licensed fiduciaries, and multi-stakeholder bodies embedded in existing legal systems. This paper asks whether its normative content can instead be realized not as rules imposed on the owners of the AI stack from without, but as a commons-governed infrastructure that any person, firm, or State may use and fund while its governance stays horizontal, polycentric, and subsidiary. We construct the Dignity Stack, a six-layer architecture mapping each dimension onto a layer of commons-governed AI infrastructure, with protocols drawn from the Liberation Stack framework and from the cooperative, mutualist, and libertarian-municipalist traditions. The commons is State-agnostic rather than anti-State, anarchist in its horizontal means but not in the abolition of the State. Its central device is a decoupling of capital from control, by which the stack functions as a shared civic battery, charged by many contributors yet steered by none in proportion to its charge. We prove that this defeats formal capture through votes or surplus, and show that structural capture, the leverage of a dominant supplier free to withdraw what it provides, is resisted only insofar as operational supply is polycentric and substitutable, a condition demanding at the lower layers and perhaps presently unattainable at chip fabrication. We conclude, with explicit attention to its limits, that commons-governed AI realizes the values the contract proclaims more faithfully than the regulation it presupposes.
Missing Data on Physics Exams: Demographic Patterns, Course-Level Predictions, and Implications for Equity
arXiv:2606.05473v1 Announce Type: new Abstract: In a previous quantitative retrospective study we showed that different demographic groups of students leave different numbers of problems blank on physics exams, leading to inequities in course outcomes. In that work we argued that there were good reasons to treat these blanks as missing data, rather than indicators of a lack of understanding. In this paper, we refine this analysis and show more detailed breakdowns uncollected test item responses by race/ethnicity and first generation college student status, coming to the same conclusion: test item responses are uncollected for students with different ethnic and racial backgrounds at different rates, and these patterns exist even for high-performing students. We also correct an error from our previous work, finding here that there is no significant gender difference in uncollected test item responses. Finally, we provide a more robust analysis of course level data illustrating that blanks are a variable controlled at the course level rather than the student level, providing more evidence for the use of a course deficit model (rather than a student deficit model) when examining equity disparities, and also suggesting that there are plausible means for instructors to minimize uncollected test item responses, and therefore eliminate the bias associated with this missing data. We provide some suggestions for faculty who want to have more equitable course outcomes.
SHIELDS: Automating OS Hardening with Iterative Multi-Agent Remediation
arXiv:2606.05476v1 Announce Type: new Abstract: Security misconfigurations remain a leading cause of OS-level compromise, and manually keeping systems compliant with standards like Defense Information Systems Agency (DISA) Security Technical Implementation Guides (STIGs) is a tedious and expensive process. Existing compliance automation tools can reduce some of this burden, but they depend on static, pre-written corrective actions. In this paper, we introduce SHIELDS, a multi-agent system that uses large language models (LLMs) to approach OS hardening as an iterative, feedback-driven process. Instead of applying fixed remediations, SHIELDS continuously proposes fixes and refines them based on feedback from target system execution and validation scans. We evaluate the system across multiple virtual machine configurations using six contemporary LLMs ranging from 20B to 400B parameters, and find that SHIELDS successfully remediates up to 73% of scan findings. Our results also suggest that success in this setting depends less on model size (parameter count) than on effective tool use and information gathering, paving a practical path toward reducing the burden of security compliance in environments where compute is limited or security and privacy needs drive local model use.
Analytic patch trees: branch interface inheritance and fractal dimension fields
arXiv:2606.06400v1 Announce Type: new Abstract: The extension of the analytic fractal curve trees of (2601.17490} to analytic surface patch trees reveals a new geometric structure: branch points are replaced by interface curves that transmit the full analytical state of parent patches to their children. These interfaces prove to be central in determining the topology of the surface patch trees, including for the conditions for self-similarity of the interfaces, the patches and thus the trees. We establish the analytic conditions for the integrability and well-posedness of the surface patch trees and introduce further restrictions for conformality. We demonstrate that patch trees have a natural foliation that slices the trees into one dimensional curve trees, each of which has their own Hausdorff dimension, jointly creating a smooth dimension field. We extend the two dimensional surface model to arbitrary dimensions $n$ where $n-1$ interface manifolds transport the $n$ field state of the parent patches to their child branches. We note that the balance or discrepancy between patch field dimension and the dimensions in which the branches may evolve, determine the analytical regime from essentially geometrical to essentially operational.
A q-Tsallis Safe Approximation for Chance-Constrained Programs
arXiv:2606.06401v1 Announce Type: new Abstract: Classical chance-constrained programs are solved by safe approximations based on the empirical CVaR, which uses a uniform measure over scenarios and systematically underweights tail events under heavy-tailed distributions. We introduce \emph{q-CCP}, a non-extensive safe approximation grounded in the Riemannian geometry of the Tsallis statistical manifold: the rank-based q-CVaR escort weights are the $g^{(q)}$-geodesic projection onto the tail simplex face, and the q-CCP feasible set is a Tsallis-divergence ball (Proposition~12). This geometric foundation yields three results. First, q-CCP is a provable strict tightening of CVaR-CCP for all $q > 1$ (Theorem~7). Second, the empirical violation ratio satisfies $\rho(q) = [1-(1-\varepsilon)^{q+1}]/\varepsilon$, independent of the tail index $\nu$ (Proposition~10). Third, the feasible-region volume cost is monotone increasing in $q$ and $\nu$ (Proposition~11), providing a data-adaptive safety knob. The formulation inherits convexity and coherence from the q-CVaR functional and admits an iterative LP reformulation converging in 2--3 iterations. Experiments on 15 Ibovespa equities confirm the theory (violation ratio $0.241$, $q^* = 1.50$); an M5 inventory newsvendor experiment generalises the method to supply chain ($q^* = 1.88$, cost premium $1.155\times$, zero OOS stockout violations).
Exploring LLMs for South Asian Music Understanding and Generation
arXiv:2606.05522v1 Announce Type: new Abstract: Recent advancements in Large Language Models (LLMs) have shown promising results in music understanding and generation tasks. However, existing works remain confined to Western tonal traditions, offering little insight into whether current LLMs can handle structurally distinct low-resource musical traditions. We present the first systematic evaluation of LLM competence in South Asian classical music, a tradition governed by raga, tala-based melodic constraints that impose fundamentally different structural principles from Western harmony-driven music. We ground our evaluation in Hindustani classical theory and Bengali classical forms, including Rabindra and Nazrul Sangeet -- representative low-resource traditions within South Asian classical music. For music understanding evaluation, we introduce a 504-question-answer benchmark spanning raga grammar, cultural knowledge, and symbolic notation reasoning, evaluating 33 LLMs where frontier models such as Gemini 2.5 Pro achieve 85-90% accuracy, while most open-source models remain in the 23-40% range. For music generation, we design a five-level controlled prompting framework and find that even the strongest model produces stylistically faithful outputs only 40% of the time. These results reveal that structural validity and stylistic faithfulness in music generation are distinct objectives and highlight an open challenge for culturally grounded music modeling.
CHASE: Adversarial Red-Blue Teaming for Improving LLM Safety using Reinforcement Learning
arXiv:2606.05523v1 Announce Type: new Abstract: Despite advances in safety alignment, prompt-rewriting attacks such as persona modulation, fictional framing and persuasion-based reformulation, can bypass safety filters even on frontier models. Existing defenses either rely on non-scalable human curation or white-box optimisation that overfits to specific model internals, leaving aligned models brittle against the very class of adaptive black-box adversaries they will face in deployment. To address this gap, we introduce CHASE (Co-evolutionary Hardening through Adversarial Safety-Escalation), a closed-loop red-blue teaming framework in which a black-box attacker and a safety-aligned defender co-evolve. The attacker is trained via Group Relative Policy Optimization (GRPO) under a multiplicative reward that jointly enforces bypass effectiveness and intent fidelity, while the defender is hardened on the harvested adversarial rewrites through a two-stage GRPO + rejection-sampled SFT pipeline balanced with benign data. Evaluated on BeaverTails and JailbreakBench against five held-out attack families (PAIR, TAP, AutoDAN, PAP, Translation), CHASE cuts mean StrongREJECT score by 43.2\% with 0\% false-refusal on benign prompts. Beyond the headline result, CHASE shows that template-free RL exploration recovers latent attack primitives that transfer across mechanistically distinct attack families, suggesting a path toward LLM safety hardening that generalises beyond the narrow distributions achieved thus far in adversarial training.
Convergence of a discrete-in-time Approximation to a Degenerate Parabolic-Hyperbolic System
arXiv:2606.05879v1 Announce Type: cross Abstract: In this paper we consider an implicit semi-discrete approximation of a degenerate reaction-cross-diffusion system. Due to the symmetry in the parabolic part, this system is known to preserve segregation of densities -- initially non-overlapping densities belonging to different species remain segregated for all times, which leads to internal layers between different species. We show that time-discrete approximations exist and converge to a weak solution, as the timestep goes to zero.
Revisiting Lexicon Evaluation in Unsupervised Word Discovery
arXiv:2606.06183v1 Announce Type: cross Abstract: Building a lexicon from discovered word-like units is a central goal in zero-resource speech processing. But do our evaluations provide a trustworthy indication of lexicon quality? A common metric, normalized edit distance, averages the phoneme edit distances between discovered units in each cluster. We show that this metric has an inherent bias toward the quality of large clusters, inhibiting fair evaluation. Moreover, it ignores how well true classes are distributed across clusters. Based on established theory in clustering literature, we propose two metrics that address these shortcomings: a modified metric that weighs cluster size when assessing within-cluster consistency, and an inverse metric that assesses how true words are spread across clusters. Through experiments on synthetic and real-world lexicons, we demonstrate that combined, these metrics are: (1) more closely correlated with how similar a lexicon is to the ground-truth distribution, and (2) more robust to biases that skew lexicon evaluations.
Advancing Digital Government: Integrating Open Source Software Enablement Indicators in Maturity Indexes
arXiv:2510.04603v2 Announce Type: replace Abstract: Context: Open Source Software (OSS) is a vital public good, included across most of modern software stacks, significantly impacting GDP and national tech growth, while supporting interoperability, sovereignty, and transparency. However, systematic measurement of governmental OSS adoption remain limited. Research Aim: This study contributes to digital government maturity indexes by analyzing policies and support actions leveraging OSS for software reuse and collaborative development across 16 digitally mature countries, and proposing potential indicators for said indexes. It examines OSS policy formation, stated goals, key actors, and support mechanisms. Methodology: A qualitative approach is used combining desk research of policy documents with semi-structured interviews of government representatives, producing detailed country reports. These are cross-analyzed, focusing on OSS policy promotion, rationale, and implementation support. Results: Policies facilitating OSS reuse are widespread, targeting both inbound acquisition and outbound sharing, and are predominantly governed by central public sector organizations. Policy goals include interoperability, digital sovereignty, transparency, and cost efficiency, with security framed both as a risk and strength. Implementation is supported by diverse Open Source Program Offices (OSPOs) at multiple government levels, which foster capacity building, resource pooling, and sustainable project governance. Indicators are synthesized and proposed across 14 areas covering policy incentives and design, and implementation and support. Conclusions: OSS is a strategic enabler for public sector digital transformation. Clear policy frameworks, coupled with institutional support such as OSPOs, are essential. International digital maturity frameworks should expand OSS indicators to better guide and assess government adoption and impact.
Bridging CAD and Data-Driven Design: Attributed Feature Graphs for Engineering Design
arXiv:2606.06405v1 Announce Type: new Abstract: Engineering design is an iterative, simulation-driven process where traditional workflows rely heavily on computationally expensive analyses such as finite element and computational fluid dynamics. Although data-driven methods have accelerated design evaluation and optimization, most existing geometric representations discard parametric and feature-level semantics, limiting their integration with CAD-driven design workflows and reducing model interpretability. To address this gap, this work introduces Attributed Feature Graphs (AFGs), a feature-based representation that encodes design features, such as extrusions, ribs, and pockets, as nodes and their geometric or dependency relations as directed edges. AFGs preserve design intent and parametric structure while remaining compatible with standard graph-based learning methods, enabling end-to-end learning directly on CAD-derived feature graphs. The paper demonstrates the proposed representation through a surrogate-modeling case study on the CarHoods10K automotive hood frame dataset, where a Graph Neural Network (GNN) is trained as an evaluation engine to predict performance metrics from AFG inputs. The learned model achieves competitive surrogate performance compared with traditional data-driven approaches, but with the added benefit that engineers can map predictions back to specific CAD features and interpret how individual design elements influence system behavior. Furthermore, because AFGs are built from native CAD features, engineers can directly edit the underlying geometry in the CAD environment and reevaluate the design through the same learned model.
Expected String Stability of Human-Led Vehicle Platoons under Stochastic Communication Delays (Full Version)
arXiv:2606.06406v1 Announce Type: new Abstract: This paper studies expected $\mathcal{L}_2$ string stability of event-triggered vehicle platoons in which a human driver leads a chain of cooperatively controlled autonomous followers under stochastic communication delays. The leader's driving behavior propagates through the string via vehicle-to-vehicle (V2V) communication, so human-induced disturbances must not amplify along the platoon. Unlike deterministic approaches based on worst-case delay bounds, we derive string-stability conditions depending on the full delay distribution through integral inequalities. The closed-loop platoon is modeled as a stochastic hybrid system capturing vehicle dynamics, communication events, and event-triggering. This framework certifies string stability even when delays exceed deterministic admissible bounds with nonzero probability. Results are evaluated under several delay distributions using the MATLAB HyEQ simulator.
Semantic Partial Grounding via LLMs
arXiv:2602.22067v2 Announce Type: replace Abstract: Grounding is a critical step in classical planning, yet it often becomes a computational bottleneck due to the exponential growth in grounded actions and atoms as task size increases. Recent advances in partial grounding have addressed this challenge by incrementally grounding only the most promising operators, guided by predictive models. However, these approaches primarily rely on relational features or learned embeddings and do not leverage the textual and structural cues present in PDDL descriptions. We propose SPG-LLM, which uses LLMs to analyze the domain and problem files to heuristically identify potentially irrelevant objects, actions, and predicates prior to grounding, significantly reducing the size of the grounded task. Across seven hard-to-ground benchmarks, SPG-LLM achieves faster grounding-often by orders of magnitude-while delivering comparable or better plan costs in some domains.
VITO: Vascular Geometry and Blood Flow Estimation Using Inverse Topology Optimization
arXiv:2606.05487v1 Announce Type: new Abstract: Computed Tomography Angiography (CTA) is widely used to reconstruct vascular geometry from projection measurements, with conventional approaches such as Filtered Back-Projection (FBP) and Iterative Reconstruction (IR) forming the clinical standard. Blood flow is subsequently estimated through Computational Fluid Dynamics (CFD) simulations, which require vascular geometry and boundary conditions to be specified a priori. Since the geometry is fixed prior to flow estimation, the recovery of unknown anatomical features (e.g., missing branches or stenoses) is precluded. In this work, we present a fluid-physics-constrained reconstruction framework that leverages topology optimization (TO) to jointly recover vascular geometry and blood velocity directly from time-resolved CTA sinograms. The formulation couples a steady incompressible flow model with a transient advection-diffusion contrast transport model, mapped to sinogram space through a differentiable projection operator. The recovered velocity fields provide hemodynamic information and can support downstream estimation of wall shear stress and flow distribution, without requiring a separate CFD pipeline. The proposed method is demonstrated on synthetic phantoms under varying sparsity and noise levels, and on representative projection data.
CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents
arXiv:2601.09923v3 Announce Type: replace Abstract: AI agents are vulnerable to prompt injection attacks, where malicious content hijacks agent behavior. Among proposed defenses, architectural isolation provides the strongest guarantees by strictly separating trusted task planning from untrusted environment observations. However, applying this design to Computer Use Agents (CUAs), which automate tasks by viewing screens and executing actions, presents a fundamental challenge. Current agents require continuous observation of UI state to determine each action, which conflicts with the isolation required for security. We resolve this tension by demonstrating that UI workflows, while dynamic, are structurally predictable. Single-shot planning, where a trusted planner emits upfront a complete branching plan covering all anticipated runtime states, provides control flow integrity guarantees against arbitrary instruction injections. We introduce NOVA (Navigating via Observation, Verification, and Action) to make this viable in the combinatorially large UI state space, where the plan can invoke a perception model to resolve runtime values such as UI coordinates. We evaluate our design on OSWorld, and retain up to 57% of the performance of frontier models while improving performance for smaller open-source models by up to 19%, demonstrating that rigorous security and utility can coexist in CUAs. Although upfront planning prevents instruction injections, we show that additional measures are needed to defend against \textbf{Branch Steering} attacks, where adversaries deceive the perception model into routing execution down attacker-preferred branches of the plan, such as redirecting the agent to a malicious website.
TinyML-Driven Cybersecurity for Autonomous Spacecraft: Latency-Accuracy Analysis for SPARTA RF and Cyber Threat Detection
arXiv:2606.05779v1 Announce Type: new Abstract: Autonomous spacecraft require rapid, lightweight, and reliable onboard detection of cyber-RF threats. Using the SPARTA attack model, we analyze the latency-accuracy trade-offs of TinyML-compatible classical models -- Random Forest, Logistic Regression, SVM, and MLP -- for detecting uplink jamming, Fake-NR spoofing, payload manipulation, ground-segment compromise, and unauthorized command injection. We present a physics-informed theoretical analysis of each model's computational complexity, VC dimension, Lipschitz continuity, and latency scaling, supported by empirical measurements on adversarial RF spectrograms generated via BandErasure, FakeNR, and NoiseBurst corruption modes. Results show that Logistic Regression achieves microsecond-level inference with only a 1\% accuracy drop relative to Random Forest, making it an effective TinyML baseline for onboard autonomy. The study also identifies opportunities for advancing spacecraft cybersecurity through richer feature encoders and multi-timescale learning architectures, building on recent progress in edge intelligence and trustworthy AI.
TAPO: Tool-Aware Policy Optimization via Credit Transfer for Multimodal Search Agents
arXiv:2606.05784v1 Announce Type: new Abstract: We identify and formally characterize credit misassignment as a systematic failure mode of GRPO in tool-augmented multimodal search agents: its uniform broadcast of trajectory-level advantages to all tokens causes valuable tool-use steps in failing trajectories to be penalized no differently from valueless ones. We further empirically quantify the scale of this phenomenon. Over half of failing trajectories and failing tool-use actions exhibit correctable credit misassignment, demonstrating that the wasted training signal is both substantial and structurally exploitable. Building on this insight, we propose Tool-Aware Policy Optimization (TAPO), which exploits the parameter-determinism property of information-acquisition tools: similar call parameters define equivalent information-acquisition actions and should therefore share comparable action credit. TAPO constructs counterfactual witnesses within the current training batch and compensates misassigned negative credit via confidence-gated conservative advantage correction. It requires no additional annotation, models, or sampling, and introduces negligible computational overhead. Across multiple multimodal search benchmarks, TAPO delivers consistent, plug-and-play improvements over strong baselines for three mainstream RL algorithms (GRPO, GSPO, and SAPO). Our code and models will be publicly released upon acceptance.
HDST-GNN: Heterogeneous Dynamic Spatiotemporal Graph Neural Networks for Multi-Object Tracking in UAV Aerial Imagery
arXiv:2606.05587v1 Announce Type: new Abstract: Multi-object tracking (MOT) from UAV imagery presents unique challenges: altitude varies across sequences, objects are small and densely packed, and frequent occlusion causes identity switches. Existing graph-based trackers assume fixed spatial context and treat all objects uniformly, ignoring the heterogeneous lifecycle states of detections, active tracklets, and lost targets. We propose HDST-GNN, a Heterogeneous Dynamic Spatiotemporal Graph Neural Network with three novel contributions. First, Altitude-Adaptive Edge Construction estimates a camera-altitude proxy from mean object area and adjusts the graph connectivity radius accordingly. Second, Heterogeneous Node Representation models detections (Type-D), confirmed tracklets (Type-T), and lost tracklets (Type-L) as distinct node types with dedicated projections and typed edge relations. Third, Occlusion-Gated Temporal Aggregation gates each node's attention contribution by its occlusion confidence, preventing occluded nodes from corrupting neighbour embeddings. HDST-GNN is trained end-to-end with a differentiable Sinkhorn head using joint cross-entropy and triplet loss. On VisDrone2019-MOT with oracle detections, HDST-GNN achieves 94.51% MOTA and 97.24% IDF1, outperforming SORT by +5.0 MOTA points and reducing identity switches by 81%. With real YOLOv8n detections, HDST-GNN reduces identity switches by 49% vs. SORT. Ablation studies confirm the independent contribution of each component.
Auditing Demonstration Curation Metrics: Action-Only Scorers Fail on the Structural Defects That Degrade Imitation Policies
arXiv:2606.05588v1 Announce Type: new Abstract: Imitation-learning policies inherit the quality of the demonstrations they are trained on, and a growing set of curation metrics promise to score and filter low-quality demonstrations automatically. These metrics are each validated on different data with different protocols, so it is unclear which of them actually identify the demonstrations that harm a policy. We build a controlled testbed in which demonstration defects are injected with known type, and audit seven curation metrics along two axes: how well each separates defective from clean demonstrations, and whether training a behavior-cloning policy on each metric's curated subset improves task success. We study two defect regimes. Subtle perturbations (correlated action noise, tremor, truncation) are detectable by multivariate outlier scoring and, once removed, recover the full downstream gap. Structural errors, where the demonstration executes a wrong action at a key moment, are invisible to every action-only metric we test, and two of them are inverted: they score defective demonstrations as higher quality and, used for curation, tend to leave the policy at or below the uncurated baseline rather than above it. Only metrics that examine the state trajectory detect structural errors, and even the best of them recovers just a third of the downstream gap. High detection accuracy does not guarantee downstream improvement. We release the testbed and all curation implementations.
Sub-Kolmogorov Intermittency and Multifractal Dissipation in Multiphase Turbulence
arXiv:2606.05788v1 Announce Type: new Abstract: Multiphase turbulence displays stronger intermittency than its single-phase counterpart, yet the origin and geometrical organization of its most intense small-scale fluctuations remain poorly understood. Using direct numerical simulations of the incompressible Navier--Stokes equations with surface tension, we show that the local dissipative cutoff broadens strongly in the presence of interfaces, with dissipative events extending deep into the sub-Kolmogorov range. These events are spatially concentrated around topology-changing interfacial regions, namely breakup and coalescence. A multifractal analysis of the dissipation field further reveals that, while the spectrum above the Kolmogorov length, $\eta_K$, remains close to the single-phase case except for the most singular tail, the near- and sub-Kolmogorov range develops a markedly broader singularity spectrum supported on sparse intense structures. Our results show that breakup and coalescence do not simply perturb turbulence locally, but imprint a distinct multifractal organization on dissipation in multiphase turbulence.
FORTE: FOL-guided Optimal Refinement for Text-audio rEtrieval
arXiv:2606.05812v1 Announce Type: new Abstract: Text-to-audio retrieval has made significant progress with shared embedding models such as CLAP and Pengi, yet they often struggle with fine-grained semantic alignment due to the inherent modality gap between text and audio. In this work, we propose FORTE, a unified framework that integrates structured logical reasoning with parameter-efficient cross-modal alignment to improve retrieval precision. Our approach first transforms queries into first-order logic and refines them via a constrained search that preserves semantic invariance while introducing discriminative attributes. The refined representation is then aligned with audio embeddings using a lightweight projection module, followed by a predicate-aware re-ranking step that enforces logical consistency at inference. Extensive experiments on AudioCaps and Clotho demonstrate consistent improvements over strong baselines, particularly in challenging fine-grained scenarios. Our results highlight the effectiveness of combining symbolic reasoning with representation learning for cross-modal retrieval.
Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Models
arXiv:2601.18383v2 Announce Type: replace Abstract: Large Reasoning Models (LRMs) excel at solving complex problems by explicitly generating a reasoning trace before deriving the final answer. However, these extended generations incur substantial memory footprint and computational overhead, bottlenecking LRMs' efficiency. This work uses attention maps to analyze the influence of reasoning traces and uncover an interesting phenomenon: only some decision-critical tokens in a reasoning trace steer the model toward the final answer, while the remaining tokens contribute negligibly. Building on this observation, we propose Dynamic Thinking-Token Selection (DynTS). This method identifies decision-critical tokens and retains only their associated Key-Value (KV) cache states during inference, evicting the remaining redundant entries to optimize efficiency.