arXiv:2607.16257v1 Announce Type: new
Abstract: Reinforcement learning (RL) has become a widely adopted technique for improving large language models (LLMs) on complex tasks. Despite this progress, existing RL methods still face challenges in training agents with longer-horizon interactions. One major bottleneck is distinguishing the contribution of different actions in long-horizon interaction, leading to high optimization variance. To address this, we introduce a novel policy gradient method, Hindsight Policy Optimization (HPO), that projects both the current policy distribution and the hindsight distribution into an intent space and extracts low-variance learning signals from the Wasserstein distance between them. We theoretically and empirically show that aggregating semantically similar states and actions in the intent space yields a bounded-variance estimator and improves policy performance stably. Our code is available online.
Science Journals
arXiv:2607.17075v1 Announce Type: new
Abstract: The advent of LLMs has significantly changed the research on privacy policy and data compliance analysis by enabling tasks that previously required specialized, domain-specific tools. However, it remains unclear to what extent LLMs can truly replicate the diverse functionalities, and the wide range of methodologies and analysis offered by prior work. In this paper, we conduct the first systematic evaluation of whether off-the-shelf LLMs can replace specialized privacy analysis tools. We study six representative tools spanning three major functionalities: contradiction detection, regulatory compliance analysis, and privacy policy summarization and aggregation, and across three intermediate tasks: structured data extraction using tuples, Semantic Role Labeling (SRL) and manual privacy policy labeling. We compare the performance of two state-of-the-art LLMs (GPT-5.2 and Gemini-2.5 in various configurations) against the tools by directly prompting the models to perform corresponding functionalities and tasks on a custom dataset of 10 privacy policies, allowing us to assess whether off-the-shelf models can produce tool-specific functionalities without further engineering or domain-specific training, major limitations in prior work. Our results show that LLMs consistently match or exceed the capabilities of existing tools across the functionalities. In manual labeling of first-party collection entities, LLMs achieved an average precision of 81.8% and recall of 70.9%, while for labeling of third-party sharing entities, they achieved an average precision of 91.4% and recall of 70.8% compared to the OPP-115 dataset. Overall, our findings indicate that LLMs can effectively perform a broad range of functionalities and tasks in privacy policy and regulation analysis that previously required specialized tools.
arXiv:2607.16649v1 Announce Type: new
Abstract: Magnetic Resonance Imaging (MRI) is often acquired with anisotropic resolution to reduce scan time, producing stair-step artifacts along the through-plane direction. In through-plane MRI super-resolution, an efficiency-fidelity trade-off arises: feed-forward regressors are fast but oversmooth at large slice-thicknesses, while sampling-based methods improve fidelity at high inference cost. We propose DRIFT, a two-stage thickness-conditioned rectified flow framework for through-plane MRI super-resolution with continuous input slice-thickness. Stage 1 employs an Anatomical Projection Network (APN) to map low-resolution patches to a coarse high-resolution manifold, providing a deterministic anatomical initialization that shortens the residual transport of Stage 2 and stabilizes slice-wise refinement. Stage 2 refines details via rectified flow and introduces a Physics-Aware Difficulty (PAD) metric derived from slice-thickness induced through-plane bandwidth deficit to guide an Adaptive Integration Scheduler (AIS), allocating ODE steps by thickness. A Consistent Endpoint Trajectory Alignment (CETA) loss enforces thickness-consistent reconstructions. Experiments show that DRIFT outperforms super-resolution baselines while reducing inference cost. Code, models, and interactive demos are available at https://yoonseokchoi-ai.github.io/drift-eccv2026/.
arXiv:2607.16347v1 Announce Type: new
Abstract: We study the classical single-machine deadline problem $1 \mid\mid \sum U_j$, in which each task has a deadline and an execution requirement and the goal is to select as many on-time tasks as possible. The standard Moore-Hodgson algorithm processes tasks by deadline and may later delete a previously accepted task. We study the insertion-only shortest-job-first rule of Lin and Wang: process the tasks in nondecreasing execution requirement, and accept a task exactly when doing so preserves feasibility. We give a direct $O(n\log n)$-time implementation using a balanced augmented BST keyed by deadline. Unlike the previous $O(n\log n)$ implementation of this SJF rule, our implementation needs neither a pre\"emptive schedule nor an amortized analysis of interval changes.
Our analysis gives an explicit threshold form of the rule's lexicographic (\emph{lex-first}) optimality: for every threshold~$e$, its outputs maximize the number of selected tasks whose execution requirement is at most~$e$. The analysis also reveals additional combinatorial structure. After the shorter tasks have been greedily fixed, the feasible choices within a single execution-requirement tier form a nested matroid. These tier matroids assemble, as a direct sum, into an overall laminar matroid whose bases are exactly the greedy outputs. Finally, a flow network encoding the deadline-prefix constraints gives a polymatroid rank function for the underlying scheduling feasibility structure. This flow view also recovers the nested matroids that govern the equal-execution tiers.
arXiv:2607.17076v1 Announce Type: new
Abstract: Routing quantum keys over low-earth-orbit (LEO) satellite constellations is harder than classical routing: satellite handovers couple consecutive scheduling decisions, stochastic cloud cover can silently zero a ground link, and finite-key effects eliminate short, low-elevation passes entirely. We present SATLOCK, a handover-aware Quantum Key Distribution (QKD) routing framework that combines (i) a composite channel model incorporating atmospheric loss, pointing jitter, Markov cloud cover, decoy-state estimation, and finite-key correction; (ii) an integer linear program (ILP) giving a provable handover-aware throughput upper bound; and (iii) a decentralized deep Q-network (DQN) baseline for weather-adaptive online routing. We evaluate two contention regimes on a Walker constellation serving intercontinental demands. In low contention (16 satellites, 6 demands), the ILP delivers 1,311 Mbit while the strongest heuristics reach 95--96\% of ILP. In high contention (8 satellites, 12 demands), where handovers become binding, heuristics drop to 89.5\% of ILP. The DQN agent reaches 91.8\% and 84.6\% of ILP in the two regimes; it learns effective per-demand weather policies but is limited in aggregate by the lack of cross-demand coordination.
arXiv:2607.16654v1 Announce Type: new
Abstract: We present a combined experimental and numerical investigation of the preferential alignment of Kolmogorov-size, high-aspect-ratio fibers in turbulent channel flow at friction Reynolds numbers $\mathit{Re}_{\tau}=300$ and $550$. Time-resolved volumetric measurements in the TU Wien Turbulent Water Channel are used to simultaneously track fibers and surrounding tracer particles, enabling the reconstruction of fiber trajectories together with a coarse-grained estimate of the local velocity-gradient tensor (VGT). Complementary direct numerical simulations (DNS) of channel flow laden with prolate ellipsoids provide a reference point-particle description. The analysis focuses on the channel core, where the experimental data recover the canonical alignment of vorticity with the intermediate strain-rate eigenvector, thereby supporting the reliability of the reconstructed VGT. We show that fibers preferentially align with the local vorticity direction, while weaker but still non-random alignments are observed with the strain eigenvectors. By measuring finite-time deformation along fiber trajectories through the left Cauchy--Green tensor, we further show that the strongest alignment occurs with the leading principal direction of Lagrangian stretching. The comparison with DNS shows overall good agreement, while deviations at higher Reynolds number suggest increasing finite-size filtering effects.
arXiv:2607.17028v1 Announce Type: new
Abstract: Let $A:\mathbb{F}_2^n\to\mathbb{F}_2^m$ be a binary linear map with fixed coordinate bases, let $C_A=\ker A$, and let $\lambda_A(y)$ be the minimum Hamming weight of a preimage of the syndrome $y$. We define $\operatorname{Shat}_{q,s}(A)$ as the least common check support of a $q$-dimensional syndrome subspace whose every nonzero element has coset-leader weight at least $s$. It therefore distinguishes release of $q$ independent syndromes from release of a subspace with no easy linear combination. Deleting check coordinates $F$ releases $\ker A_{\bar{F}}/\ker A$, canonically isomorphic to $(\operatorname{im} A)[F]$.
Finiteness implies $R_q(C_A)\ge \mathsf{N}_2(q,s)$, where $\mathsf{N}_2(q,s)$ is the shortest length of a binary code of dimension $q$ and distance at least $s$; profile-Griesmer bounds independently control common check support. The hierarchy is coordinate-relabeling invariant but can change under a change of check basis. For the pair-repetition code $C_n=\{(x,x):x\in\mathbb{F}_2^n\}$, the standard realization $H_0=[I_n\ I_n]$ has $\operatorname{Shat}_{q,s}(H_0)=\mathsf{N}_2(q,s)$ whenever feasible. For every $q\ge 1$ and $s\ge 2$, with $n=\mathsf{N}_2(q,s)$, a row-equivalent realization of the same code has value $q$.
For a simplicial coboundary map $A=\delta_k$, check erasure is top-face erasure and the released quotient is emergent cohomology. At $s=1$ the hierarchy reduces to generalized Hamming weights and is Tutte-determined; for $s\ge 2$, even identical labeled cut codes can have different values.
arXiv:2607.17077v1 Announce Type: new
Abstract: Adversarial attacks against vision models like object detectors are often evaluated under limited conditions, leaving their performance under-characterized. Bridging simulation and differentiable rendering enables more robust, end-to-end evaluation of these adversarial attacks, yet there is no easy-to-use, unified system that offers a rich set of customizable configurations for adversarial attacks across multiple scenes, objects, environmental and lighting conditions, and camera trajectories. We present ALLUDE, which addresses these gaps, offering first-of-its-kind evaluation capabilities across Linux and Windows. We comprehensively demonstrate ALLUDE's evaluation breadth through a two-pronged strategy: (1) using Latin Hypercube Sampling, we draw a representative subset from 5,400 configurations spanning 10 scene-object pairs, 9 weather conditions, 4 optimizers, 5 camera trajectories, and 3 detection models; (2) we stress-test existing attacks (CAMOU, RAUCA, FCA) under diverse weather conditions and continuous camera trajectories, revealing degradation of attack success across every attack, exposing evaluation gaps in prior work. Through ALLUDE's end-to-end differentiable rendering, adversarial attacks can be optimized against shifting real-world deployment conditions. Our cross-platform code is open source.
arXiv:2607.16521v1 Announce Type: new
Abstract: Distributed Denial of Service (DDoS) attacks continue to pose significant threats to network availability and security. While many detection systems focus on binary classification (attack vs. benign), effective mitigation often requires identifying the specific type of DDoS attack. This paper introduces a robust intrusion detection framework centered around a high-accuracy, multi-class classification model designed to precisely identify various DDoS attack types. We propose an ensemble architecture integrating Long Short-Term Memory (LSTM), K-Nearest Neighbors (KNN), and Random Forest (RF) models, whose outputs are synthesized by a Logistic Regression meta-learner. This approach explicitly addresses the ambiguity often encountered when combining predictions from multiple independent classifiers. Evaluated on the CIC-DDoS2019 dataset, our proposed ensemble meta-learning model achieves 96% accuracy in the multi-class identification task, significantly outperforming a baseline chain model (combining individual binary classifiers), which reached 92% accuracy and suffered from high ambiguity. Furthermore, integration and testing within a Software-Defined Networking (SDN) environment using Mininet and the Ryu controller demonstrated the practical applicability of our model, achieving 93% accuracy in identifying DDoS types in the emulated network traffic. Our work highlights the value of meta-learning ensembles for nuanced DDoS threat identification, paving the way for more adaptive and effective defense mechanisms.
arXiv:2607.16523v1 Announce Type: new
Abstract: One of the main strengths of Constraint Programming is the ability to reduce the search space via propagation. However, propagation is a double-edged sword, with more pruning power coming at the price of larger computation time. For each problem constraint, the best propagator depends on the specific instance and may change at search time. In the literature, Machine Learning (ML) techniques and activity-based heuristics have been applied respectively for choosing (statically) the propagators for a batch of problems and to adapt (dynamically) the propagation strength. We propose to merge those efforts by using an oracle function, obtained via ML, to decide whether to run complex propagators for a target constraint. A combination of design choices makes the approach flexible and easy to embed in state-of-the-art solvers. In this paper, we focus on investigating the feasibility of building an oracle for the Energetic Reasoning propagator. Our experiments show that high prediction accuracy can be obtained, provide suggestions for classification features, and highlight important issues to address when building such an oracle.
arXiv:2607.17817v1 Announce Type: new
Abstract: Photonic memory underpins optical information processing, neuromorphic photonics, and photonic computing. Existing studies typically treat dispersive, nonlinear, and driven-dissipative memory as distinct physical phenomena, despite all being governed by the evolution of the optical field. This work proposes a unified phase-based framework in which photonic memory is interpreted as a dynamical phase, with transitions between memory regimes governed by optical phase evolution, Kerr nonlinearity, and the balance between delayed feedback and dissipation. Within this framework, dispersive memory arises from frequency-dependent phase accumulation, nonlinear memory emerges through intensity-dependent phase evolution leading to bistability and hysteresis, and driven-dissipative memory is established through attractor convergence and memory stabilization. The proposed framework is validated through theoretical analysis and numerical simulations, demonstrating a continuous progression from linear dispersive memory to nonlinear and ultimately driven-dissipative memory. Experimental validation is performed using a silicon photonic waveguide incorporating chirped Bragg gratings. Group delay measurements reveal distinct linear and nonlinear memory responses, while the reconstructed memory distribution demonstrates the coexistence of dispersive, nonlinear, and driven-dissipative memory within a single integrated photonic device. To the best of the author's knowledge, this constitutes the first experimental demonstration supporting the coexistence of all three photonic memory regimes in a single photonic platform. These results establish optical phase as the unifying physical quantity underlying photonic memory and provide a common framework for designing future neuromorphic photonic systems, reservoir computers, and integrated photonic processors.
arXiv:2607.17956v1 Announce Type: new
Abstract: Learning is increasingly introduced into visual-inertial odometry (VIO), ranging from learned feature front-ends to learning-dominant motion and geometry estimation. However, learning more of the pipeline does not necessarily improve robustness when deployment conditions differ from the training distribution. This work asks whether robust VIO under distribution shift truly requires deeper learned estimation, or whether learning can be confined to visual measurement generation. We propose a minimal-learning stereo VIO framework in which SEA-RAFT is used only to propose dense stereo correspondences and predict their uncertainty, while temporal tracking, geometric verification, and state estimation remain explicit. Dense flow is sampled at sparse feature locations, filtered using predicted uncertainty and stereo epipolar consistency, and incorporated into a sliding-window stereo-inertial estimator through uncertainty-weighted reprojection factors. The same uncertainty is further propagated through stereo triangulation for downstream anisotropic 3D Gaussian mapping. Experiments on EuRoC, VIODE, and 4Seasons demonstrate accurate and stable estimation under motion blur, dynamic scenes, illumination changes, and large indoor-to-outdoor distribution shifts. Ablations show that learned flow alone is insufficient: the gains arise from combining learned correspondence proposals with geometric verification and uncertainty-aware weighting. These results suggest that, for OOD-robust VIO, carefully integrated learned visual measurements can be more effective than learning a larger fraction of the estimation pipeline. Code and configs for the benchmark will be open-source upon acceptance. A supplementary video is available at https://drive.google.com/file/d/1EVRhOkhanmNXHbQS1Vr80FoEIAYOYOV2/view
arXiv:2607.16258v1 Announce Type: new
Abstract: The application of artificial intelligence methods in power electronic converter modeling is becoming increasingly widespread, but existing applications still face many challenges, such as difficulties in multi-time-scale hybrid analysis and the lack of physics-aware evaluation criteria and constraints, resulting in poor performance. This paper proposes a Neural Controlled Differential Equation (Neural CDE) framework for learning continuous-time surrogate models of grid-forming inverters for electromagnetic transient (EMT) simulation, which relaxes the constraint of fixed sampling rates and enables multi-time-scale control analysis. Then, an affine-control formulation with dual slow/fast pathways is proposed to capture the hierarchical and multiscale behavior of converter dynamics, and a physics-inspired regularization method is utilized to enhance stability and coherence. Evaluated on EMT-generated trajectories, the model accurately reproduces transient responses, preserves effective damping and the dominant oscillatory characteristics, and maintains bounded long-horizon rollouts. The results show that Neural CDE-based component modeling offers a physically consistent surrogate modeling approach for EMT-level simulation studies.
arXiv:2607.17031v1 Announce Type: new
Abstract: Security-constrained unit commitment (SCUC) couples binary commitment, economic dispatch, reserves, and network security over a multiperiod horizon, making an exact solution computationally expensive for realistic system sizes. This paper proposes a three-layer hybrid framework in which a Bernoulli hybrid soft actor-critic (HSAC) policy proposes hourly commitments, a quantum-sampled auxiliary channel augments the state, and a native SCUC mixed-integer linear program recovers dispatch and security variables after only a limited subset of commitment binaries is enforced. The method is therefore solver-compatible rather than an end-to-end replacement for exact optimization. We formalize the SCUC-to-reinforcement-learning interface, derive the temporal coverage induced by the fixed cap, and evaluate the 14- 57- and 118-bus benchmark cases. The results show stable, low-cost recovery in the 14-bus case, where the best recovered schedule attains the full-horizon optimum; a very low screen-rejection rate in the 57-bus case; and a clear coverage bottleneck in the 118-bus case once the enforcement cap no longer spans a complete commitment period. The study, therefore, identifies the amount of useful commitment information that reaches the recovery model, under an exploratory Bernoulli actor and a small enforcement cap, as the dominant limitation that governs scalability
arXiv:2607.16656v1 Announce Type: new
Abstract: Object removal aims to eliminate target objects specified by a mask while preserving visual consistency with the surrounding regions. Existing methods typically rely on contextual information from surrounding regions. However, in dense scenes where the surrounding regions contain instances visually similar to the removal target, such reliance often leads to semantic interference, resulting in incomplete removal. This problem arises from erroneous information propagation in the attention space, where masked queries tend to align with such instances due to global similarity matching in self-attention. To address this challenge, we propose a Diffusion-based Object Removal framework for dense Scenes, dubbed DORS, built upon a Dynamic Attention Routing mechanism comprising two complementary components: Instance-Filtered Attention (IFA), which suppresses misleading semantic information from similar instances through dynamically constructed mask-guided attention constraints, and Context-Guided Routing (CGR), which dynamically routes complementary scene information to maintain visual consistency. We further introduce DOR-Bench, a benchmark tailored for object removal in dense scenes. Extensive experiments demonstrate that DORS outperforms state-of-the-art methods, particularly in reducing incomplete removal and duplicate artifacts. The code will be available at https://github.com/httang1224/DORS.
arXiv:2607.17080v1 Announce Type: new
Abstract: Base-2 digital nets are practical high-dimensional integration rules: the sample budget $N=2^m$ can be chosen independently of the ambient dimension, and the generating matrices provide algebraic control of projections and Walsh-dual weights. They are therefore well suited to problems whose error is governed by weighted or low-dimensional projection structure. However, when one restricts attention to a smooth low-dimensional projected component, a low-dimensional cubature rule with a comparable number of nodes can be substantially more accurate than the projected digital-net points. This raises the question of whether low-dimensional cubature accuracy can be inserted into a high-dimensional digital-net rule without forming the full tensor product.
We answer this question by a simple coordinate embedding: read the leading $p$ binary digits of each coordinate as an index into $2^p$ equal-weight cubature nodes, and replace the coordinate by the indexed node. When a projection forms the full $p$-bit grid, the transformed rule coincides on that projection with the corresponding product cubature rule; small projected $t$-values provide sufficient conditions for such full-grid recovery. For general integrands, the error separates into the corresponding product cubature error and a residual digital-net term. Experiments with scrambled Sobol' nets in dimension $50$ illustrate this mechanism and show finite-budget improvements for the smooth low-order and coordinate-decaying test functions considered here.
arXiv:2607.16790v1 Announce Type: new
Abstract: Fine-grained offensive language detection organizes labels into a hierarchical structure, for which two modeling paradigms exist: cascaded decomposition and joint multi-task modeling. Prior work rarely provides a direct, controlled comparison of the two paradigms in terms of accuracy, parameter count, and inference latency, and rarely verifies whether a chosen class-imbalance handling strategy is actually optimal. This paper proposes a three-level cascaded detection system whose training strategy is customized per subtask, together with two verification mechanisms. First, a controlled ablation study determines the best class-imbalance handling strategy for each subtask. Second, a joint multi-task model with a shared encoder is trained as an architectural control, yielding real measurements along the dimensions of accuracy, parameter count, and inference latency. Experiments show that the cascaded system attains macro-F1 scores of 0.795, 0.716, and 0.557 on the three subtasks of the official test set. The ablation study reveals that configuring the loss function purely by imbalance-severity intuition is suboptimal; reconfiguring based on the ablation results improves both performance and stability. End-to-end cascade evaluation shows that roughly one-fifth of the errors in the cascade pipeline originate from the first-stage filter and cannot be corrected by subsequent stages. Relative to the joint multi-task model, the cascaded architecture achieves higher accuracy on all three subtasks, with a 7.1-point macro-F1 gain on the most severely imbalanced subtask, at the cost of three times the parameters and 1.67 times the inference latency. Together, these results establish an explicit, quantifiable trade-off between the accuracy advantage of cascaded architectures and their deployment cost.
arXiv:2607.17987v1 Announce Type: new
Abstract: Smart-contract vulnerabilities often arise from inconsistencies between business paths that should correspond to one another, such as single and batch entry points, direct and adapter-based flows, quote and execution paths, or inverse operations such as buy and sell. Existing analyzers are effective for many local syntactic and data-flow patterns, but they provide limited support for bugs whose oracle is relational: whether two semantically paired paths preserve compatible guards, state transitions, value flows, and failure behavior.
This paper introduces chiral analysis, a relational model that treats paired business paths as implicit specifications for each other. We formalize chiral relations as static analogues of metamorphic relations, derive obligations over guards, actors, state, value, ordering, failure behavior, and external interactions, and report a vulnerability when a violated obligation has security impact. We implement this idea in ChiralDetector, a Solidity prototype that extracts business paths, ranks candidate pairs with static facts, applies LLM-based semantic filtering and detection, and validates and deduplicates findings.
In a preliminary evaluation on the Phi protocol, ChiralDetector reduced 3,217 statically ranked path pairs to 1,643 semantic candidates, produced 101 deduplicated finding groups, and retained 44 strict-validator positives that manually collapsed to 13 effective unique issues. These include cross-art Merkle proof reuse, fee unit mismatches, public state-tracking helpers, and refund propagation gaps. The results suggest that chiral analysis can expose business-logic bug classes that are difficult to express as single-function rules while providing a structured way to control LLM cost and validator precision.
arXiv:2607.16259v1 Announce Type: new
Abstract: Pretrained models are typically ranked on multi-task leaderboards to assess their effectiveness across diverse tasks. Rank confidence intervals were recently introduced as a method to quantify the uncertainty in these rankings by aggregating pairwise hypothesis tests. In this work, we analyze the sources of uncertainty in the knowledge evaluation benchmark MMLU and show how hypothesis tests can be modified to account for their effects. We demonstrate that ranking variability across MMLU subjects is substantial and should be considered when comparing LLMs or identifying the top-performing models.
arXiv:2607.17033v1 Announce Type: new
Abstract: Forecasting the outcomes of transition-metal-catalyzed reactions is notoriously complex due to the interplay of diverse physical and chemical variables. A persistent computational bottleneck has been effectively merging broad electronic descriptors with the localized, three-dimensional geometry of the reactive site. To bridge this representation gap, we present ChemFusion, a hybrid neural network that fuses conventional electronic features with explicit 3D atomic coordinates. Using a cross-attention mechanism, the model enables global electronic states to dynamically attend to specific spatial constraints within un-pooled molecular point clouds. When benchmarked against a diverse library of cross-couplings, this approach delivers exceptional predictive performance, decisively surpassing traditional single-modality frameworks. Importantly, extracting the attention matrices reveals that the architecture autonomously learns to identify and penalize restrictive steric hindrances. This provides a physically grounded interpretability, demonstrating that spatially aware networks can navigate complex reaction sterics that standard statistical models typically miss.
arXiv:2607.16660v1 Announce Type: new
Abstract: The increasing adoption of Large Language Models (LLMs) as AI components in modern software systems introduces distinct security risks to the software supply chain. While many considerations and safety mechanisms are in place for components of the traditional software supply chain, the recent rapid adoption of AI components and platforms has overlooked these hard learned lessons. Selecting and integrating AI models without clear guidance on how these choices affect system security may leave applications vulnerable to threats, such as malicious components, data leakage, and unintended behavior. The goal of this study is to understand practitioners' decision making process and security considerations in selecting and integrating AI components through an exploratory semi-structured interview study. Toward this goal, we conducted semistructured interviews with 22 software developers, architects, and AI practitioners across diverse organizations about how they integrate AI components into their software.
Our analysis finds that practitioners' model selection is predominantly driven by functional criteria, including performance, accuracy, cost, and specific features, e.g., tool calling or multimodal support, while security is rarely considered as an evaluation criterion. We observe a consistent lack of security concern throughout the AI component integration process, with established software supply chain lessons overlooked or ignored. The industry is repeating the historically costly mistakes of early software dependency management, prioritizing rapid reuse and availability over security and provenance. We distill our findings into actionable recommendations for AI adopters, model providers, and researchers, advocating for a proactive, security-by-design approach that integrates security evaluation into component selection and sustains it throughout the software development lifecycle.
arXiv:2512.11695v2 Announce Type: replace
Abstract: Particle Image Velocimetry (PIV) is among the central modalities for measuring flow fields across laboratory, industrial and environmental setting. Traditional PIV approaches typically depend on tuning parameters specific to the imaging setup, making the performance sensitive to variations in illumination, flow conditions, and seeding density. Similarly, state-of-the-art machine learning methods for flow quantification are fragile outside their training set. In our experiments, we observed that flow quantification would improve if different tunings (or algorithms) were applied to different regions of the same image pair. Motivated by this observation, we thus pose flow quantification as a multi-estimator fusion problem: several heterogeneous algorithms process the same image pair in parallel, and their dense flow fields are treated as complementary estimates. To fuse them, we adopt a consensus framework based on the alternating direction method of multipliers, incorporating priors such as smoothness and incompressibility. We perform several numerical experiments to demonstrate the benefits of this approach. For instance, we achieve a decrease in end-point-error of up to 20% of a dense-inverse-search estimator at an inference rate of 60Hz, and we show how performance can be increased with outlier rejection. Our method is implemented in JAX and integrated into Flow Gym, enabling reproducible comparisons with the state of the art and systematic evaluation across different base algorithms. Finally, we demonstrate successful deployment of our method in the same real-world active-fluids-control setup of Terpin and D'Andrea [1], where a reinforcement-learning agent uses our flow estimates to learn to minimize drag (down by 36%) or maximize it (up to 32%) with only two minutes of real-world interaction. Hardware and software are made available at ActiveFluidControl.com.
arXiv:2607.16524v1 Announce Type: new
Abstract: Cooperative multi-agent RL systems routinely use team-averaged rewards, a feedback-attribution choice that gives each agent the team outcome regardless of its individual contribution. We ask whether this leaves a measurable signature, geometric or behavioral, on learned representations. We propose EffRank/$n$ (effective rank normalized by agent count) and $D_\text{act}$ (mean pairwise KL divergence between agents' action distributions) as low-overhead diagnostics for reward-attribution effects, then test them on competent MAPPO agents in SMACv2 \texttt{protoss\_5\_vs\_5}, where unit type is encoded in the observation. In an observation $\times$ reward-attribution comparison (unit type observed vs.\ masked; individual damage-contribution reward vs.\ shared team reward), geometry follows observation rather than reward. With unit type observed, shared and individual rewards have similar EffRank/$n$ ($0.31{\pm}0.03$ vs.\ $0.29{\pm}0.02$) and probe accuracy ($0.75{\pm}0.05$ vs.\ $0.73{\pm}0.05$, both $\gg 1/3$ chance), while $D_\text{act}$ leans higher under individual rewards ($1.23{\pm}0.06$ vs.\ $1.07{\pm}0.20$). Masking unit type cuts the above-chance probe signal by more than half, to $0.49$ in both reward arms. In short: individually rewarded agents are competent and separable by role, but on SMACv2 the observation explains the geometry and reward attribution shows up mainly in behavior. Thus geometric diagnostics must control for observed role information and test persistent roles that are not directly observed. EffRank/$n$ and $D_\text{act}$ add $<$5\% overhead.
arXiv:2607.17927v1 Announce Type: new
Abstract: AI-based test agents promise to accelerate software testing by shortening feedback loops in continuous development and improving scalability and maintainability. To realize these benefits, engineers must still be able to assess if agent outputs are useful, valid, and reliable, rather than treating them as credible because they come from a capable system. This paper argues that overreliance on AI in testing is both an agency problem, in which engineers may cede cognitive control over test design decisions, and an assurance problem, in which testing artifacts may be accepted as evidence without sufficient scrutiny. We develop this argument through three theoretical lenses: software testing as cognitive problem-solving, test agents as adaptively autonomous entities, and test design argumentation as a means of making generated tests reviewable. We propose a framework for collecting data on overreliance in test agent workflows and identify specific modes of overdependence. The goal is to support accelerated testing without weakening judgment or the assurance value of testing evidence.
arXiv:2607.16348v1 Announce Type: new
Abstract: Machine learning-based intrusion detection systems (IDSs) often suffer from class imbalance and vulnerability to adversarial attacks, leading to degraded detection performance and reduced robustness. This study proposes a TabTransformer framework augmented by the Boundary-Seeking Generative Adversarial Network (BGAN) for flow-based intrusion detection using the CICIDS2017 dataset. BGAN serves a dual purpose by generating synthetic minority-class samples to mitigate data imbalance and producing adversarial samples to evaluate model robustness. Experimental results demonstrate that BGAN augmentation improves TabTransformer's Macro-F1 score from 82.96% to 86.50%, with the largest class-wise improvement observed for Web_Attack (F1 score: 0.29 to 0.61). Robustness evaluation shows that all non-augmented models experienced a 100% Performance Drop Rate (PDR) under adversarial testing, whereas all BGAN-augmented models achieved negative PDR values, indicating improved resilience. Furthermore, the augmented TabTransformer maintained stable and low False Triggered Rate (FTR) values (1.51%-2.92%) across all noise levels, compared with the BGAN-augmented Decision Tree, which reached 49.09% under benign perturbations. These findings demonstrate that BGAN consistently enhances both class balance and adversarial robustness, while the proposed BGAN-TabTransformer framework provides an effective and adaptive intrusion detection solution for adversarial network environments.