Forskningsradar

Science Journals

Peer-reviewade publikationer — 56239 artiklar

Immunity to Increasing Condition Numbers of Linear Superiorization versus Linear Programming
arXiv:2407.18709v3 Announce Type: replace-cross Abstract: Given a family of linear constraints and a linear objective function one can consider whether to apply a Linear Programming (LP) algorithm or use a Linear Superiorization (LinSup) algorithm on this data. In the LP methodology one aims at finding an optimal point, i.e., a point that fulfills the constraints and has the minimal value of the objective function over these constraints. The Linear Superiorization approach considers the same data as in linear programming problems but instead of attempting to solve with linear programming methods it employs perturbation resilient feasibility-seeking algorithms that steer the iterations toward a feasible point with reduced (not necessarily minimal) objective function value. This aim of the superiorization method (SM) is less demanding than aiming to reach full-fledged constrained optimality and it places more importance on reaching feasibility than on reaching optimality. Previous studies (e.g., [1]) compared LP and LinSup in terms of their respective outputs and the resources they use. Here, we investigate classical LP approaches and LinSup in terms of their sensitivity to condition numbers of the system of linear constraints. Condition numbers are a measure for the impact of deviations in the input data on the output of a problem and, in particular, they describe the factor of error propagation when given wrong or erroneous data. Therefore, the ability of LP and LinSup to cope with increased condition numbers, thus with illposed problems, is an important matter to consider which was not studied until now. We investigate experimentally the advantages and disadvantages of both LP and LinSup on exemplary sets of data of problems of linear programming with multiple condition numbers and different problem dimensions.
Zeroing Diagonals, Conjugate Hollowization, and Characterizing Nondefinite Operators
arXiv:2508.00096v2 Announce Type: replace Abstract: We prove the conjecture by Damm and Fassbender that, for real traceless matrices $L,M$, there exists orthogonal $R$ such that $\mathrm{diag}(R^\top L R) = (0,...,0,0,0)$ and $\mathrm{diag}(R M R^\top) = (0,...,0,*,*)$. We also prove for any pair $L,M$ of complex Hermitian traceless matrices, there exists a unitary $U$ such that $\mathrm{diag}(U^* L U) =\mathrm{diag}(U M U^*) = (0,...,0)$. The claims comprise a corollary to our more general theorem for $L,M$ of arbitrary trace. We also discuss severe limitations upon generalizing our theorem to general complex $L,M$. By setting $L = M$, much is revealed concerning freedom and constraint involved in introducing 0s to the diagonal of a single operator. From this we prove a novel characterization of real traceless matrices and complex Hermitian traceless matrices, strengthening the seminal theorem by Fillmore that every complex square matrix is unitarily similar to a hollow matrix. Our results are contextualized in a characterization of nondefinite matrices as a more general environment for introducing 0s to the main diagonal.
Dissipative Vortex Binaries in Compact Fluid Domains with Geometric Corrections
arXiv:2604.23857v2 Announce Type: replace Abstract: We study a dissipative extension of vortex-binary motion in a doubly periodic fluid domain. The underlying conservative system admits an exact integrable reduction to a single complex relative coordinate. Dissipation is introduced via a minimal rotated-velocity (mutual-friction) term, as motivated by finite-temperature superfluid dynamics, converting the Hamiltonian evolution into a mixed symplectic--gradient flow with monotonic energy decay for quantized vortices. In the local regime, the dissipative binary remains analytically solvable and admits closed-form solutions, with systematic corrections arising from the toroidal geometry. Equal same-sign vortices execute outward spiraling motion, while equal opposite-sign pairs (dipoles) undergo finite-time collapse in the planar limit. On the torus, however, the dipole orientation is no longer invariant: the geometry induces a slow angular drift, even in regimes where planar dynamics would preserve alignment. For unequal opposite-sign pairs, dissipation induces coupled contraction and rotation, leading to a finite-time nonlinear chirp characterized by $\dot{\omega}\propto\omega^2$, in contrast with electromagnetic and gravitational inspirals where $\dot{\omega}\propto \omega^{3}$ and $\dot{\omega}\propto \omega^{11/3}$. These results highlight the interplay between Hamiltonian structure, dissipation, and geometry in periodic fluid systems.
video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding
arXiv:2606.24477v2 Announce Type: replace Abstract: Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA). A practical and efficient solution is a two-stage paradigm: first perform coarse video understanding to localize relevant segments, and then re-watch these segments at higher temporal or spatial fidelity. In this paper, we present video-SALMONN-R$^3$, the first end-to-end video-LLM that enables re-watch through reinforcement learning without relying on chain-of-thought (CoT) cold-start. This design removes the need for costly CoT data annotations and avoids CoT-based supervised fine-tuning (SFT), which can otherwise degrade the pretrained video understanding abilities. To address the mismatch between the reasoning-first behavior induced by re-watch and the answer-first tendency of pretrained video-LLMs, we propose a re-answer strategy, in which the model first produces a direct answer in the first watch and then refines it after re-watching. Finally, to improve question adherence during re-watching, we propose a re-ask mechanism that re-injects the query when revisiting localized segments. Experimental results show that video-SALMONN-R$^3$ consistently outperforms both the base model and the QA-SFT baseline, while surpassing prior re-watch-based approaches with significantly lower computational cost. Code, models, and data will be publicly released upon acceptance.
Instance Generation for Patient-to-room Assignment and Admission Scheduling Based on Real Hospital Data
arXiv:2507.03423v2 Announce Type: replace-cross Abstract: Developing algorithms for real-life problems that perform well in practice depends on the availability of realistic data for testing. Obtaining real-life data for optimization problems in health care, however, is often difficult, and such data typically cannot be published, which limits reproducibility by other researchers. This is especially true for patient-related problems because of data privacy policies such as the patient-to-room assignment problem. Therefore, artificially generated instances are commonly used. To improve the generation of realistic instances, we develop a configurable instance generator for the patient-to-room assignment problem and other patient-related problems, featuring an easy-to-use graphical user interface. The design of the generator is based on an extensive empirical analysis of real hospital data, which identifies relevant ward-specific patterns such as patients' age and length-of-stay distributions. Moreover, as randomly generated instances are often infeasible, we address this issue in two ways. We implement a dynamic programming approach in the generator to optionally enforce feasibility and extend existing results from the literature to derive new combinatorial insights into patient-to-room feasibility.
Forensic Schema for Psychological Manipulation in Cyber Fraud: LLM-Driven Victim Reports Analysis
arXiv:2607.07751v1 Announce Type: new Abstract: Existing cybercrime classification schemas capture contact metadata and financial transactions but omit the psychological manipulation techniques perpetrators employ. We present a forensic schema (four categories, 35 questions) adding 11 manipulation indicators and cryptocurrency evidence fields to established forensic foundations. Applied to 10,994 victim reports via large language model (LLM)-driven annotation and validated against two human annotators (mean LLM-human $\kappa = 0.69$, matching inter-annotator $\kappa = 0.68$), the schema revealed a statistically distinct manipulation profile for each major fraud type (Cramer's $V$ up to $0.790$). A rationale-based evidence audit nonetheless exposed a forensic detail gap: detection of manipulation techniques was reliable, but victim narratives varied widely in the actionable detail supporting each Yes answer, and blockchain-specific identifiers were nearly absent. These findings point to AI-assisted victim intake with schema-informed follow-up questions as the most direct way to close the gap. The tiered annotation strategy also provides a reusable template for LLM-based extraction from other forensic text domains.
Dynamic masking for boundary-aware velocity reconstruction in volumetric particle tracking with moving solids
arXiv:2606.25748v2 Announce Type: replace Abstract: Volumetric particle tracking velocimetry (PTV) produces scattered Lagrangian tracks that must be reconstructed on an Eulerian grid before velocity gradients, pressure, or hydrodynamic loads can be evaluated. This step is usually performed on a domain treated as entirely fluid. When a solid body lies within the measurement volume, its surface kinematics are not imposed and the reconstruction is weakest in the steep-gradient region next to the body. We introduce LE-DM (Lagrangian-to-Eulerian reconstruction with Dynamic Masking), a constrained reconstruction framework for moving solid boundaries. A time-dependent signed-distance function classifies grid nodes as open fluid, boundary shell, or solid interior. The particle data, incompressibility constraint, prescribed surface velocity, and regularization terms are then assembled on the masked domain within a single solve. The method requires only a signed-distance field and a surface velocity, allowing stationary walls, translating, rotating, multiple, and deforming bodies to be represented in the same formulation. LE-DM is assessed using an analytical oscillating sphere, synthetic tracks from a CFD rising-sphere simulation, and a refractive-index-matched tomographic-PTV experiment on a freely rising sphere. The surface kinematics are enforced to solver tolerance, while the bulk reconstruction remains unchanged where no body is present. In the analytical case, the first-cell error is reduced from 14\% to 3\% of the body speed. In the experiment, LE-DM recovers the independently measured surface velocity, whereas an all-fluid reconstruction does not. The result is a divergence-free, boundary-consistent velocity field for pressure and force estimation.
Improving RCT-Based Treatment Effect Estimation Under Covariate Mismatch via Calibrated Alignment
arXiv:2603.19186v3 Announce Type: replace Abstract: Randomized controlled trials (RCTs) are the gold standard for estimating treatment effects, yet they are often underpowered for detecting effect heterogeneity. Large observational studies (OS) can supplement RCTs for conditional average treatment effect (CATE) estimation, but a key barrier is covariate mismatch: the two sources measure different, only partially overlapping, covariates. We propose CALM (Calibrated ALignment under covariate Mismatch), which learns embeddings that map each source's features into a common representation space. OS outcome models are transferred to the RCT embedding space and calibrated using trial data, preserving causal identification from randomization. Finite-sample risk bounds decompose into alignment error, outcome-model complexity, and calibration complexity terms, making explicit when the learned embedding is accurate enough to reduce variance. We instantiate CALM in two forms: a closed-form linear version, CALM-Lin, and a neural representation-learning version, CALM-NN. Across 51 simulation settings, calibration-based linear methods are effectively tied in linear-CATE regimes, while CALM-NN wins all 22 nonlinear-CATE settings by wide margins. Moreover, on two real-data studies CALM-NN delivers the largest gains over the trial-only baseline.
A Theory of Composable Lingos for Protocol Dialects
arXiv:2603.19908v2 Announce Type: replace Abstract: Formal patterns are formally specified solutions to frequently occurring distributed system problems that are generic, executable, and come with strong qualitative and/or quantitative formal guarantees. A formal pattern is a generic system transformation which transforms a usually infinite class of systems in need of the pattern's solution into enhanced versions of such systems that solve the problem in question. In this paper we demonstrate the application of formal patterns to protocol dialects. Dialects are methods for hardening protocols so as to endow them with light-weight security, especially against easy attacks that can lead to more serious ones. A lingo is a dialect's key security component, because attackers are unable to ''speak'' the lingo. A lingo's ''talk'' changes all the time, becoming a moving target for attackers. In this paper we present several formal patterns for both lingos and dialects. Lingo formal patterns can make lingos stronger by both transforming them and by composing several lingos into a stronger lingo. Dialects themselves can be obtained by the application of a single dialect formal pattern, generic on both the chosen lingo and the chosen protocol.
Two-stage dispersion mechanism of clean spherical bubbles rising in a chain
arXiv:2604.00612v2 Announce Type: replace Abstract: Wake-induced lift is a key mechanism governing the initial destabilization of bubbles rising in a chain (Atasi et al., 2023). Moore's wake model predicts limited interfacial vorticity and a relatively slender, spatially confined wake for clean spherical bubbles, suggesting that wake-mediated interactions weaken as the inter-bubble spacing increases. However, we observed pronounced large-scale lateral dispersion and strong bubble frequency dependence in controlled experiments where bubble diameter and generation frequency were independently varied, even when the inter-bubble separation exceed the characteristic wake length. A reduced-order model incorporating pairwise wake-induced interactions captured the onset of bubble chain destabilization but systematically underpredicted the subsequent emergence of large-scale dispersion. We demonstrate that bubbles rising in a chain collectively generate a mean upward liquid flow that modifies the local shear field, enhancing the lateral migration through shear-induced lift. Incorporating this self-induced weak flow into the model quantitatively reproduced both the dispersion magnitude and its frequency dependence. These results suggest that the dispersion of bubbles rising in a chain involves a two-stage mechanism, with initial chain destabilization mediated by wake interactions, followed by flow modification arising from two-way coupling between bubbles and the liquid. This collective mechanism highlights the importance of self-induced mean flow effects in continuum descriptions of bubble flows.
Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?
arXiv:2606.26428v2 Announce Type: replace Abstract: Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out of reach. These tasks are contact-rich, making data collection for imitation learning difficult, and sparse-reward, making direct exploration with reinforcement learning (RL) intractable. Consequently, prior work has made progress by structuring the problem with specialized grippers, tool attachments, and environment fixtures. In this work, we argue that before a robot can perfect precise assembly, it must first learn to play. We further ask the question: what factors in the process of learning to play matter for precise assembly? We propose Play2Perfect, an RL framework for task-agnostic pretraining through play on diverse objects and goals, which is then perfected on precise assembly. The goal of play is to acquire reusable manipulation priors, such as grasping, in-hand reorientation and pose reaching. Finetuning then adapts this general prior to assembly, focusing exploration on the final contact-rich, high-precision interactions needed for success. We systematically study key design choices in play pretraining, including object diversity, training objective, trajectory diversity, and goal precision. We show that our prior is 33x more sample-efficient than RL training from scratch, even when provided with dense, multi-stage rewards. We demonstrate zero-shot sim-to-real transfer, achieving 60% success on tight insertions with only 0.5 mm contact clearance, and over 50% success on long-horizon multi-part assembly and screwing.
Homomorphism Indistinguishability Beyond Graphs: Relational Weisfeiler--Leman and Hypertree Width
arXiv:2607.07934v1 Announce Type: new Abstract: The Weisfeiler--Leman (WL) algorithm is one of the most influential heuristics for the graph isomorphism problem. The expressive power of WL has been extensively studied in the contexts of descriptive complexity, logics, graph neural networks, and the theory of homomorphism indistinguishabily. Notably, two graphs are indistinguishable by the $k$-dimensional WL algorithm if and only if they are indistinguishable by homomorphism-counts from graphs of treewidth at most $k$. An intrinsic question is to find a natural version of the WL algorithm for relational structures of higher arity admitting an equivalent characterisation via homomorphism indistinguishability along bounded generalised hypertree width (GHW). Scheidt and Schweikardt solved this for $k=1$ by defining the RCR algorithm and showing indistinguishability from $\alpha$-acyclic structures. In this work, we resolve this for all $k\ge1$: we develop $k$-RCR and show that two structures $\mathcal{A}$ and $\mathcal{B}$ are insdistinguishable by $k$-RCR if and only if they have the same homomorphism-counts from all structures $\mathcal{C}$ of generalised hypertreewidth $\le k$. Moreover, we introduce a ``fractional'' version of $k$-RCR and show that two structures are insdistinguishable by fractional $k$-RCR if and only if they have the same homomorphism-counts from all structures with (a variant of) fractional hypertreewidth at most $k$. Last, we develop $k$-HyperOWL, the first relational WL algorithm operating directly on a relational structure. We show that $k$-HyperOWL is as expressive as $k$-RCR and that, given a structure $\mathcal{A}$, $k$-HyperOWL can compute $t$ iterative refinements in time $O(t|\mathcal{A}|^{k+1})$. Moreover, the colouring produced by $k$-HyperOWL can be used as a constructive preprocessing routine for counting homomorphisms from structures of generalised hypertreewidth $\le k$.
A Multi-cluster Boundary Learning Method for Out-of-Scope Intent Detection via MiniLM Embedding
arXiv:2607.07974v1 Announce Type: new Abstract: Intent detection is a critical task that bridges human intents and system actions in human-machine interaction systems. However, there still exist challenges for detecting out-of-scope (OOS) intents. (i) The traditional methods view the OOS intent detection as a multi-class classification, then the detection accuracy decreases as the class number of the known intents increases; (ii) LLM-embedding methods require large parameters, that makes them difficult to train and practically deploy. Thus, this work proposes a multi-cluster boundary learning method to detect OOS intents via MiniLM embedding (i.e., all-MiniLM-L6-v2) in an one-class classification workflow. The method learns the boundaries of multi-cluster embeddings generated by MiniLM from the training utterances, and then rejects the out-of-domain utterances as OOS intents. Experiments are conducted on public CLINC150, StackOverflow and Banking77 datasets. The results show that the method achieves the state-of-the-art OOS intent detection performance compared the other baselines. Ablation studies are also conducted and the results show that the used MiniLM can better adapt to the workflow and utterance embedding requirements. The code is available at supplementary materials.
LAP: Simple Command-line Tools for Teaching Logic, Algorithms, and Proof in Computer Science
arXiv:2607.08000v1 Announce Type: new Abstract: The LAP toolset is a set of command line tools for teaching logic in computer science. It provides implementations of standard algorithms for propositional and first order logic, including conversions to various normal forms, propositional satisfiability algorithms such as DPLL, Tseytin's transformation, and equivalence checking. Significantly, LAP also supports a language for expressing a natural deduction derivation for propositional or first order logic. The tools can check the derivation, provide meaningful feedback if it is wrong, or display the derivation in a variety of formats. The toolset is written in Java and has no dependencies other than a Java Virtual Machine. The code has been designed to be easy to read and to illuminate the data definitions and algorithms.
Aleena: Alignment Agent for Research Software Engineering Collaborations
arXiv:2607.08043v1 Announce Type: new Abstract: Research software collaborations span meetings, informal chats, pull requests, and GitHub issues. A decision surfaced in a Slack thread, refined in a meeting, and implemented in a pull request can lose its original rationale across these artifacts, leaving domain researchers and research software engineers with divergent mental models of project intent, ownership, and scientific assumptions. We argue that alignment in research software engineering is a continuous lifecycle problem, and that agentic AI can support stakeholder alignment and project-state tracking without replacing human decision-making. We present Aleena, an open-source lifecycle alignment agent that uses GitHub as a shared collaboration surface, transforming multi-modal stakeholder interactions into structured project records that surface risks, track open questions, and preserve decision continuity. Grounded in university-based research software engineering center experiences, this paper presents the motivating problem, system design, prototype, and illustrative lifecycle scenarios for Aleena.
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
arXiv:2607.08057v1 Announce Type: new Abstract: Despite the rapid advancements of large language models (LLMs), LLM serving systems remain memory-intensive and costly. The key-value (KV) cache, which stores KV tensors during autoregressive decoding, is crucial for enabling low-latency, high-throughput LLM inference serving. In this survey, we focus on system-aware KV infrastructure for serving LLMs (abbreviated as sKis). We revisit recent work from a system behavior perspective, organizing existing efforts into three dimensions: execution and scheduling (temporal), placement and migration (spatial), and representation and retention (structural). Furthermore, we analyze cross-behavior co-design affinity and behavior-objective links, highlighting future opportunities. Our work systematizes a rapidly evolving area, providing a foundation for understanding and innovating KV cache designs in modern LLM serving infrastructure.
Modular Pretraining Enables Access Control
arXiv:2607.08077v1 Announce Type: new Abstract: AI developers face a dual-use dilemma. An AI capability that helps one user cure a disease can help another synthesize one. This dilemma could be resolved with access control, limiting dual-use AI capabilities to trusted deployments with a legitimate need. A gold standard for access control would be to serve separate models with different capabilities to different users. However, training and deploying multiple models is prohibitively expensive. To address this challenge, we propose gradient-routed auxiliary modules (GRAM), a pre-training method that adds modules to a neural network and selectively updates them to induce specialization. Ablating a module at inference time removes its capability from the network, approximating a model trained on filtered data. We evaluate GRAM on synthetic stories and realistic dual-use data spanning virology, cybersecurity, nuclear physics, and specialized code. These experiments show that GRAM disables targeted capabilities while preserving the rest, and resists their recovery under finetuning better than post-hoc unlearning. Most importantly, a Chinchilla-optimal scaling analysis from 50M to 5B parameters shows that the gap between data-filtered and full-data models widens with scale on removed capabilities but stays small on retained ones, and that GRAM closely tracks data filtering. GRAM's training cost is independent of the number of supported capability profiles, yielding a 5x reduction over data filtering in our 5-profile setting.
MLQENABLER: Enabling Secure Machine Learning Queries over Encrypted Database in Cloud Computing
arXiv:2607.08197v1 Announce Type: new Abstract: In cloud computing, the public cloud service providers (CSPs) can provide cloud storage as the primary service while providing additional machine learning (ML)-based services by using the clients' data in storage. This business model extends the border of cloud computing services and brings in new business growth possibilities. Although it is promising, the model also brings in security concerns since the public commercial cloud cannot be fully trusted. For example, the public commercial clouds may sell clients' sensitive data to the government or other companies. To address the security concerns, an immediate solution is to require clients to encrypt their datasets before outsourcing to the cloud. However, if a database is formally encrypted, then the database contains only pseudorandom numbers, making it impossible to enable ML over it. In this project, we propose MLQENABLER (ML Queries Enabler) scheme to enable secure ML queries over encrypted database in cloud storage. MLQENABLER employs an index-aid approach to achieve security and ML capability simultaneously. Our initial experiments show that MLQENABLER achieves an acceptable security level while incurring only a slight ML performance degradation.
MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks
arXiv:2605.21917v2 Announce Type: replace Abstract: Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happened, but when, where, why, and with what consequence, at a scale manual labelling cannot support. We present MAVEN (Multi-stage Agentic Video Event aNnotation), a multi-stage agentic pipeline that turns raw videos into multi-task training data with Chain-of-Thought (CoT) reasoning traces, organized around a designated Event of Focus. At its core, MAVEN synthesizes a Multi-Scale Spatio-Temporal Event Description (MSTED) from three complementary caption levels; this explicit intermediate serves as the sole input to downstream Q&A generation across multiple task formats. Crucially, MAVEN supports agent-driven domain adaptation: given a new video dataset and target question examples, the agent redesigns all prompts top-down without manual re-engineering. A hierarchical refinement loop further classifies annotation errors against a taxonomy, traces root causes to the originating pipeline stage, and applies targeted edits that rewrite prompts or modify the pipeline structure itself, iteratively improving data quality. We apply MAVEN to label over 5,300 traffic videos and fine-tune Cosmos-Reason2-8B on the resulting data. On a private CCTV evaluation set, fine-tuning surpasses both Gemini 2.5 Pro and 3.1 Flash, including a $+38.8$-point gain in MCQ accuracy over zero-shot. On AccidentBench, CCTV-only training lifts Cosmos-Reason2 by $+10.7$ MCQ points and matches Gemini 2.5 Pro despite seeing no dashcam videos; adding agent-adapted dashcam annotations narrows the gap to Gemini 3.1 Flash, and RL post-training pushes overall performance past both Gemini baselines. Qualitative results on warehouse surveillance and public safety videos further show the agentic workflow readily adapts the pipeline to new domains.
The Memory Wall of Green Software: Empirical Energy Evaluation of Memento Design Pattern
arXiv:2607.07944v1 Announce Type: new Abstract: As Green Software Engineering matures, energy efficiency has transitioned into a mission-critical non-functional requirement. While software design patterns ensure structural integrity, their inherent abstraction layers impose an implicit "metabolic cost" that often remains obscured during the design phase. This paper empirically investigates the energy dynamics of the Memento design pattern, contrasting a direct, unabstracted baseline against Classic full-snapshot and Differential delta-encoding strategies. Leveraging the RAPL interface for high-fidelity hardware telemetry, we quantify energy dissipation across state volumes scaling from 10 MB to 200 MB. Our empirical results expose a critical architectural trade-off: the Differential strategy minimizes memory traffic, yielding a maximum energy reduction of 65.8% for mid-scale states, but collides with a catastrophic "memory wall" at 200 MB. At this saturation point, algorithmic optimizations are completely neutralized by severe GC thrashing and non-linear power spikes. We synthesize these findings into evidence-based heuristics, providing architects with a robust framework to reconcile structural design quality with sustainable Green IT imperatives.
A Puck-informed mode-resolved phase-field fatigue framework for unidirectional composites
arXiv:2607.07977v1 Announce Type: new Abstract: Fatigue fracture in unidirectional fibre-reinforced composites is strongly mode dependent: transverse and off-axis cycling is governed by matrix and inter-fibre mechanisms, whereas fibre-aligned cycling activates a longitudinal channel with a higher fracture-energy scale and a different crack topology. Single-damage-variable models can fit global stiffness loss but cannot identify the active mechanism. This work proposes a Puck-informed, mode-resolved phase-field fatigue framework with separate channels for fibre-dominated and matrix/inter-fibre fatigue. Each channel has its own fatigue history, threshold, and resistance-degradation law. Fatigue does not directly degrade elastic stiffness; it lowers the fracture resistance of the active channel, while the corresponding phase field controls stiffness loss and crack-path evolution. The formulation is implemented in Abaqus/Standard using a compact UMAT-UEL architecture with one orthotropic mechanical routine and two scalar phase-field layers. Using one fixed IM7/8552 material and fatigue card, the model is verified through one-element tests, parameter sweeps, and centred-notch and open-hole tension cases at 0, 45, and 90 degrees under monotonic and cyclic loading. Without orientation- or geometry-specific tuning, the framework reproduces transverse matrix/inter-fibre cracking at 90 degrees, off-axis cracking at 45 degrees, and longitudinal matrix splitting with delayed fibre activation at 0 degrees. The fatigue lives follow the expected ordering: 45- and 90-degree cases fail within about 1,000 cycles, while 0-degree cases run out to 200,000 cycles without fibre cracking. Additional load, hole-size, mesh, length-scale, and cycle-block studies confirm consistent crack modes and converged trends. The study is a numerical verification and cross-geometry consistency assessment, not a calibrated experimental life-prediction claim.
Physics-informed neural networks for shock capturing in inviscid flows around an airfoil
arXiv:2607.08130v1 Announce Type: new Abstract: Physics-informed neural networks (PINNs) have shown remarkable prospects in solving forward and inverse problems involving partial differential equations (PDEs). However, PINNs still face challenges in solving fluid mechanics problems involving shocks, especially in steady inviscid flows around an airfoil, where they may even fail to capture shocks. In this study, we first point out that the reason PINNs fail to capture shocks is that the steady Euler equations used to construct the loss function impose weak constraints, which are difficult to correct the continuous function approximation preference of neural networks, causing gradient descent converges to a smooth local optimum. Based on this insight, we propose to strengthen the physical constraints by reconstructing steady shock capturing as temporal evolution that gradually converges to the steady state solution. The unsteady Euler equations constructed by introducing time derivative terms into the steady equations are used to constrain PINNs. The output of PINNs is no longer required to directly approximate a flow field with shocks by minimizing the residuals of the steady Euler equations. Instead, shocks gradually form under the guidance of the temporal evolution law of the flow field. This additional temporal penalty alleviates the tendency of PINNs to converge to a smooth local optimum. Since obtaining the steady state solution requires solving the unsteady Euler equations over a long time in the time dimension, while the capability of PINNs to solve such problems is poor, we introduce a PDE loss function that embeds the concept of pseudo time-stepping to avoid this issue. In addition, to further improve the shock capturing accuracy, we develop a simplified formulation of the Euler equations. By solving four forward problems involving different flow conditions and geometries, we validate the effectiveness of the proposed method.
Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems
arXiv:2607.08010v1 Announce Type: new Abstract: Production LLM agents often waste latency and reliability by regenerating code for the same procedural steps on every request. We replace this inference-time coding loop with an agentic tool-making pipeline that compiles repeated SOP steps into validated, versioned tools before deployment. The tool-maker grounds synthesis in the live environment as it collects execution traces, observes backend schemas and values, generates candidate tools, and repairs them against labeled cases. At runtime, the production agent calls these tools directly and falls back to code generation only when needed. We deploy the approach in a Fulfillment Center alarm-triage system, where an agent diagnoses alarms against a 44-node SOP over heterogeneous metric backends. In production, tool calls reduce p50 latency by 42%. On 1,500 historical alarms, they reduce end-to-end error rate by up to 53% by suppressing run-to-run variance in repeated steps. Because tools return compact structured verdicts, they also enable a simpler direct-call architecture, reducing p50 latency by a further 62% in a controlled ablation. Versioned tools also improve auditability and expose specification gaps and upstream data drift. Our results show that self-evolving agents can make industrial LLM systems faster, more reliable, and easier to operate.
Prismata: Confining Cross-Site Prompt Injection in Web Agents
arXiv:2607.08147v1 Announce Type: new Abstract: Autonomous web agents promise to automate everyday browsing tasks, but inherit one of the web's oldest attack surfaces. Cross-Site Scripting proved that mixing trusted and untrusted content is dangerous, even on benign pages. Agents resurface this risk by interpreting natural language as instructions, allowing third-party and user-generated content to hijack the agent via prompt injection. The core challenge is that deriving a task-specific security policy requires reasoning over page structure that is entangled with the attacker's content. We present Prismata, a defense enforcing contextual least privilege for web agents, constraining both what the agent sees and what it can do. Prismata's dynamic trust derivation produces permission labels for page content, with structural confinement guarantees, inspired by classical integrity models, that bound any labeling errors so that labels can only decrease in privilege and mislabelings are bounded. Prismata's mechanical confinement enforces these labels by redacting content and restricting agent capabilities. Importantly, these mechanisms require no developer annotations, so Prismata supports the long tail of websites. Across recent published web agent attacks, including adaptive variants, Prismata substantially reduces attack success while preserving benign task utility.
Deployment-Time Memorization in Foundation-Model Agents
arXiv:2606.10062v2 Announce Type: replace Abstract: Foundation-model agents are increasingly long-lived systems that remember users across interactions, making memorization an explicit deployment-time function rather than solely a property of model weights. Existing work addresses parametric memorization or audits fixed memory configurations, but does not characterize how memory-design choices jointly shape personalization utility, extraction risk, and deletion fidelity. We study this surface as deployment-time memorization, formulating agent memory as a privacy-utility frontier measured by Personalization Recall (PR) and Adversarial Extraction Rate (AER), and sweeping three memory-design knobs: summarization aggressiveness, retrieval breadth (k), and deletion mode. We further introduce the Forgetting Residue Score (FRS) to quantify whether deleted information remains recoverable from derived memory tiers. On LongMemEval, key-fact summarization reduces canary extraction by 76% on Gemma 3 12B and 64% on GPT-4o-mini while preserving nearly all personalization recall; critically, once content is compressed away, increasing k no longer restores leakage. The same compression, however, induces a deletion-fidelity failure: raw-only deletion leaves derived summary copies recoverable in approximately 20% of instances, and only full-pipeline purge or tombstone redaction drives worst-tier residue to zero. Together, these results establish that persistent agent memory must be evaluated as a first-class memorization mechanism -- assessed by what it helps agents recall, what it makes extractable, and what it can truly erase.