arXiv:2605.24672v1 Announce Type: new
Abstract: Theoretical models of evaporating droplets predict Marangoni flows orders of magnitude faster than those observed experimentally. While this discrepancy is often attributed to surface contamination, the underlying mechanism by which contaminants weaken Marangoni stresses remains unclear. In this study, we compare particle image velocimetry (PIV) experiments with a coupled hydrodynamic and solute transport model to investigate the internal flow of evaporating aqueous droplets containing salt, glycerol, or ethanol. By analyzing both sessile and pendant droplets, we demonstrate that the flow is driven entirely by natural convection, in contrast to theoretical predictions that use surface-tension gradients. Remarkably, in some cases, the experimental surface velocity is found to be directed against the predicted surface-tension gradient. We further prove that standard contamination models, whether based on surfactants lowering the surface tension or on surface rheology, cannot account for this flow reversal. Our results therefore suggest that Marangoni stresses are not merely reduced by contaminants, but that their macroscopic manifestation is effectively suppressed altogether.
Science Journals
arXiv:2605.25915v1 Announce Type: cross
Abstract: We report numerical simulations of the dissipative Gross-Pitaevskii equation for a bulk region of thermal-counterflow turbulence. Quasistationary states are obtained over a range of forcing, damping, and healing-length parameters. The mutual-friction acceleration exhibits cubic scaling with the mean relative velocity between the superfluid and normal-fluid components, and the coefficient of this scaling is linked to the phenomenological damping parameter. The intervortex spacing follows the expected dimensional scaling in the weak-forcing regime. Comparison with a straight-vortex-line model suggests that the vortex-line orientations are nearly isotropic.
arXiv:2605.24712v1 Announce Type: new
Abstract: Federated learning (FL) enables privacy-preserving collaborative training across distributed edge devices, but real deployments involve heterogeneous clients with different processing power, memory capacity, and communication latency, which often increase round duration and system cost. This paper proposes a hardware-aware federated learning framework for emotion recognition on session-partitioned IEMOCAP that integrates hardware profiling, top-K client selection, and adaptive local epochs within a unified training loop. We compare the method against FedAvg, FedProx, and random top-K selection under a non-IID setup and show that, across 50 federated rounds and 5 independent trials, the proposed approach achieves competitive validation accuracy (0.352), reduces total training time by about 36.5% compared to FedAvg, and lowers cumulative communication cost by 40%.
arXiv:2605.24667v1 Announce Type: new
Abstract: Mean cross-entropy is the standard validation metric for language models, but it can fail to track model quality during training. We examine this in two common scenarios. First, in Qwen2.5-1.5B SFT on synthetic fact-learning, we find that mean CE rises substantially after the initial learning phase while held-out fact-recall accuracy remains near its peak. Second, we find that in top-K distillation on TinyStories, decreasing K improves median CE while worsening mean CE; the Top-5 student attains the highest LLM-judge score and crosses below its teacher on median CE, despite having the worst mean CE. In both cases, median CE correlates much more closely with task performance than does mean CE.
Analyzing how bulk and tail percentile CE move during training reveals that training reshapes the empirical per-token CE distribution. In top-K distillation, smaller K yields a distribution with more mass at both extremes, decreasing the median and increasing the mean. In Qwen SFT, the bulk saturates quickly while the tail extends in the latter half of training. In both, the task-evaluation metric appears more sensitive to the bulk than to the tail.
Practically, we recommend reporting a small set of percentile CE summaries alongside the mean, and using concordance among them as a tool to keep track of distribution reshaping, as well as a low-cost diagnostic for when mean and median CE disagree on model selection.
arXiv:2605.25748v1 Announce Type: new
Abstract: Trajectory prediction methods have demonstrated remarkable capabilities in capturing complex motion patterns. However, existing methods rely on global state assumptions, suffer from insufficient belief inference under partial observability, and lack cognitive behavioral constraints in prediction. These limitations severely compromise both deployment feasibility and physical plausibility in real-world settings. In this work, we propose FEP-Diff, an agent-centric trajectory prediction framework grounded in the Free Energy Principle, aimed at achieving cognitively plausible predictions under realistic constraints. Specifically, a dual-branch spatiotemporal encoder extracts ego-motion dynamics and social interaction cues from local observations. Building upon this, a goal-conditioned belief learner infers multimodal latent belief distributions optimized via a free-energy objective, with a social consistency constraint on the local neighborhood graph to promote cognitive alignment among neighboring agents. Finally, a residual diffusion trajectory generator is conditioned on the learned belief representations with token-level proxy conditioning, producing precise and diverse future predictions. Extensive experiments on five public benchmarks demonstrate that FEP-Diff consistently outperforms state-of-the-art methods under restricted observability. Code: https://anonymous.4open.science/r/FEP-Diff-8876.
arXiv:2605.25415v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used in academic peer review, yet their reliability, alignment with human judgment, and robustness to adversarial attacks remain poorly understood. We present a systematic benchmark of LLM-as-a-Reviewer on 898 papers stratified from NeurIPS and ICLR, evaluating 12 LLMs along three axes: rating calibration, divergence from human reviewers, and resistance to prompt injection embedded via an invisible font-mapping attack. We find that LLMs systematically overrate weaker submissions and diverge from humans in topical emphasis, under-flagging Clarity and over-flagging Reproducibility, while producing reviews two to three times longer with lower lexical diversity and a more standardized vocabulary. Prompt injection remains highly effective. Simple hidden instructions can promote low-scoring papers to acceptance-level ratings in a substantial fraction of cases, with effectiveness varying sharply across model families. While LLMs offer utility in structuring evaluations, their integration into peer review requires safeguards against both intrinsic biases and adversarial risks.
arXiv:2605.24714v1 Announce Type: new
Abstract: In this paper, we study age of information (AoI) optimization for status updating in an integrated sensing and communication (ISAC) system. We consider a discrete-time architecture in which a base station interacts with a physical environment and a remote monitor, and at each time slot can operate in one of three modes: sensing, communication, or joint sensing and communication. Each mode is unreliable and incurs a different operational cost. The objective is to minimize a discounted infinite-horizon cost that combines the AoI at the monitor with action-dependent sensing and communication costs. For the single source scenario, we formulate the problem as a Markov decision process with a two-dimensional AoI state and prove that the optimal stationary policy admits an ordered threshold structure in the AoI state space. Since the AoI evolves over an infinite space, we truncate the state space to reduce complexity and rigorously bound the resulting error. The analysis analytically determines the truncation size needed to keep the error below a given threshold. For the multi-source scenario, we formulate the scheduling problem as a restless multi-armed bandit. We develop both a Whittle index policy and an approximate Whittle index policy for scheduling under two different regimes, one where indexability is guaranteed, and one where it is not. Numerical results illustrate the structure of the optimal policy in the single-source case and show that the proposed approximate Whittle index policy performs comparably to the Whittle index policy in the indexable regime, while remaining effective beyond it.
arXiv:2605.24663v1 Announce Type: new
Abstract: This paper presents CyBOKClaw, an interpretable human-in-the-loop retrieval framework for mapping cybersecurity keywords or phrases (KWoPs) to the Cyber Security Body of Knowledge (CyBOK). Rather than treating the task as strict exact classification, the framework is designed as a top-k candidate generator for expert review. It combines query normalization, curated term expansion, concept-level boosts, topic-description enrichment, and domain-sensitive ranking rules. Because educational KWoPs are often broad, ambiguous, and only approximately aligned with CyBOK terminology, strict exact matching provides only a partial account of practical utility. We therefore evaluate the framework using both structural retrieval metrics and an expert-guided top-5 usefulness metric, ECA-5 (Exact or Closest Acceptable Match at top-5), which records whether the returned candidates contain at least one mapping that an expert would judge exact or accept as the nearest practical CyBOK placement. On the development dataset, CyBOKClaw achieves 64.73% EXA-5 (Exact Match at top-5), 84.18% structural semantic alignment, and 91.88% ECA-5; on the validation dataset, it achieves 81.19% EXA-5, 93.32% structural semantic alignment, and 98.00% ECA-5. These results show that expert-guided top-k usefulness provides a more faithful account of practical CyBOK mapping utility than exact structural matching alone, and that CyBOKClaw is effective as a CyBOK-specific expert-support retrieval system.
arXiv:2605.24722v1 Announce Type: new
Abstract: High degrees of disagreement among annotators can exist for ambiguous objects, e.g. in medical images, underscoring the challenges of establishing ground truth annotations in object detection tasks. Despite this, all existing object detectors implicitly require access to ground truth annotations for either training or evaluation. The fundamental questions we target are: How can we learn an object detector with multiple annotators' annotations but without objective ground truth annotations due to object ambiguity, and how can we enable the learned detector to express meaningful model predictive uncertainties in detecting ambiguous objects? To answer these questions, we present an interpretable approach to calibrate probabilistic object detectors, where the calibration goal is to align the class confidence and bounding box variance estimates to the annotators' annotation distribution. We introduce an efficient yet effective framework to calibrate probabilistic object detectors by designing four evaluation metrics to measure calibration errors regarding classification and localization, and proposing a train-time calibration and post-hoc calibrator, all without the need to access any ground truth. This framework is generalizable to many existing probabilistic object detectors, such as the YOLO families and two-stage detectors. Empirical results with real-world and synthetic datasets of medical and natural images demonstrate the superior performance of the proposed framework with three popular object detectors.
arXiv:2605.24725v1 Announce Type: new
Abstract: Dynamic models of power systems are critical for analyzing grid response to disturbances and blackouts, but the release of real-world dynamic models is hindered by privacy and cybersecurity concerns, as such models carry sensitive information about transmission, generation, and load parameters. We develop an algorithm for synthesizing dynamic grid models from real-world power grids balancing two objectives: the privacy of the source grid, quantitatively measured using the notion of differential privacy, and the fidelity of the synthesized model. The algorithm applies privacy-preserving noise to obfuscate the original grid parameters, but then optimizes the perturbed parameters to ensure that the resulting model dynamics are statistically consistent with those observed in the source grid. Application to the frequency dynamics of the IEEE 30-bus system reveals the inherent privacy-fidelity trade-off: stricter privacy requirements degrade modeling fidelity, yet optimization significantly improves the quality of the synthesized models.
Critical Hawkes Processes with Random Fertilities: Stationarity in Law Beyond Infinite Mean Activity
arXiv:2605.24595v1 Announce Type: new
Abstract: Genuinely critical dynamics have been proposed to organize many natural and social systems, yet exact criticality is usually thought to preclude stationarity because the mean activity diverges. I show that this conclusion is not generally valid for self-exciting Hawkes point processes. At criticality, stationarity in law is controlled not by the mean intensity, but by local finiteness of the infinite-past Poisson-cluster construction. The relevant object is the fixed-window hitting probability \(H_T(u)\), the probability that a cluster born at time \(-u\) contributes at least one event to a window of length \(T\). For memory tails \(\mathbb{P}(T>t)\sim t^{-\theta}\) and fertility tails \(\mathbb{P}(\kappa>x)\sim x^{-\gamma}\), I prove stationarity for \(1<\gamma<2\) and \(\theta>\gamma\) via a finite-mean-lifetime criterion. In the finite-memory, finite-variance regime, \(H_T(u)\) is asymptotically comparable to the cluster-survival probability, and the exact local-finiteness condition fails. A direct asymptotic analysis of \(H_T\) gives the sharper condition \(\theta>\gamma-1\) for stationarity to hold in the infinite-fertility-variance regime. Thus broad fertility fluctuations can stabilize critical Hawkes dynamics in law, producing locally finite stationary sample paths despite infinite mean activity.
arXiv:2605.24726v1 Announce Type: new
Abstract: High-resolution printed circuit board (PCB) inspection suffers from resolution collapse when full-board images are resized to standard detector inputs: micro-scale defects shrink to a few pixels and are missed. Tile-based inference preserves local detail but introduces boundary artefacts at tile edges, causing split detections and false negatives. We present a systematic comparison of five inference strategies evaluated on two high-resolution PCB defect datasets, PCB-Defect (230 images, 1704 annotations) and HRIPCB (693 images, 2 953 annotations), spanning six defect classes. We show that training-inference scale consistency is critical: a detector trained on full images collapses to mAP@50 = 0.01 under tile inference, while the same architecture trained on 640*640 tile crops achieves 0.72 and 0.94 on the two datasets respectively. We further exploited Topology-Aware Tile Merging (TA-TM), a training-free post-processing method that builds a tile-adjacency graph and adjusts boundary-sensitive detection scores using neighbour-tile agreement before global NMS. Across both datasets, adding 128 px tile overlap raises boundary-zone recall from ~26-63% to ~70-100%, TA-TM achieves the best mAP@50 on both benchmarks, and tile inference recovers 46-100% of small defects missed entirely by full-image methods. Results are consistent across datasets, confirming the generalizability of the proposed strategy. TA-TM requires no retraining and is architecture-agnostic, making it directly applicable to existing PCB inspection pipelines.
arXiv:2605.24874v1 Announce Type: new
Abstract: Distributed vertical power delivery (DVPD) architectures employ multiple parallel voltage regulators (VRs) to meet the high-power and high current density demands of modern high performance computing (HPC) systems. While full parallel activation maximizes efficiency near peak load, medium to light load operation leads to efficiency degradation when all VRs remain active due to persistent switching and gate drive losses. This work proposes a load aware power system activation framework targeted at the medium to light load regime, in which the number of active VRs scales proportionally with instantaneous load power. A spatially informed selection strategy determines which VRs are activated from the available pool, aligning regulator placement with localized power demand. This locality aware activation minimizes lateral redistribution currents within the power plane and reduces conduction losses and voltage drops. Simulation results on a representative DVPD system demonstrate 2x to 3x switching loss reduction relative to conventional full-parallel light load operation, while sustaining an approximately 87% efficiency plateau across the 5% to 30% load range. Output ripple constraints are preserved, with inductor current ripple maintained within 6% and output voltage ripple within 2%, ensuring regulation integrity while improving overall conversion efficiency.
arXiv:2603.17688v2 Announce Type: replace
Abstract: In even spatial dimensions, solutions of the wave equation violate Huygens' principle, producing a persistent wake tail inside the light cone rather than a sharply localized propagating front. This intrinsic tail complicates refocusing. Here, we examine how the wake-tail structure of the two-dimensional wave equation affects refocusing, using the analytically tractable example of a pulse generated by a source localized in both space and time. Two idealized concentration strategies are considered. A spatial mirror reflects the outgoing pulse and produces refocusing, but the redirected signal is broadened, with the wake tail preserving its causal ordering behind the propagating front. A second strategy employs a time mirror generated by abrupt temporal modulation of the phase velocity, producing temporal reflection and transmission. This mechanism introduces an anti-causal response of the wake-tail, reversing its temporal ordering in a time-reversal-like manner; however, the pulse still undergoes distortion and wake-tail contributions persist through secondary radiation at the refocus point. These results demonstrate the fundamental connection between Huygens' principle and wave concentration, showing that the wake-tail structure intrinsic to two-dimensional propagation imposes a fundamental limit on perfect refocusing, even under idealized conditions.
arXiv:2605.24731v1 Announce Type: new
Abstract: This paper presents a novel passivity-based semi-autonomous attitude control framework, with a particular focus on attitude kinematics defined on the special orthogonal group $SO(3)$. While human-robot interaction facilitates the successful execution of complex tasks, ensuring stability of human-in-the-loop systems on the $SO(3)$ manifold remains a largely unsolved challenge. We first propose a new control architecture in which a multi-robot system preserves invariance of the average information fed back to the human operator through so-called stealthy control, and the human intervention is mediated through a virtual leader, which is coupled with the robots via a passivity-based attitude synchronization law. We then rigorously prove closed-loop stability of the proposed human-in-the-loop system under the assumption that the human behaves as a passive system. To support this analysis, simulation studies are conducted to identify the human operator as a dynamical system, and to examine passivity properties of the identified model.
arXiv:2605.25279v1 Announce Type: new
Abstract: Greenhouse agriculture in the Mediterranean region faces significant automation challenges due to its unique structural and environmental constraints. These environments are characterized by extremely narrow aisles, heterogeneous terrains ranging from concrete to tilled soil and severe optical interference caused by polyethylene covers, which induce specular reflections and "ghost points" in depth sensors. While autonomous navigation is essential for digitizing agricultural tasks, traditional solutions often rely on expensive 3D LiDAR systems that are economically unscalable for most facilities. To address this, this paper presents GreenSeg, a robust perception framework for autonomous navigation using RGB-D sensing. The proposed method introduces a dual-layer validation strategy: a robust global plane fitting combined with a surface curvature filter for terrain adaptability, and a seed-point-based Region Growing constraint to ensure the spatial continuity of the navigable plane. Experimental validation was conducted using the AGRICOBIOT I platform across four diurnal scenarios with varying solar elevations. The results show that GreenSeg consistently outperforms benchmark segmentation methods, achieving peak improvements of 11.58% in mean Recall and 19.24% in mIoU during critical rotational maneuvers at the end of corridors. These findings confirm that the proposed algorithm enables stable and safe autonomous navigation in unstructured, dynamic agricultural environments that are subject to budget constraints and sensitive to lighting conditions.
arXiv:2605.25000v1 Announce Type: cross
Abstract: Sodium bismuth titanate (NBT) is a promising oxide-ion conductor,but its electrical conductivity is highly sensitive to small changes in A-site stoichiometry and processing history.This sensitivity can reduce sample-to-sample reproducibility.Here we examine how precursor mixing controls structural uniformity and ionic transport in Na0.52Bi0.47TiO3 ceramics.Dry grinding,wet grinding with ethanol,and ball milling were compared by X-ray diffraction,electron microscopy,energy-dispersive spectroscopy,Eu3+ photoluminescence excitation spectroscopy,and electrochemical impedance spectroscopy.All processed powders and ceramics form the perovskite NBT phase within the detection limit of XRD.However,the microstructure,surface A-site cation ratio,Eu3+ excitation spectra,and electrical response change strongly with the mixing route.Continuous monitoring of Eu3+ excitation spectra at different emission wavelengths reveals different distributions of local Eu3+ environments.Larger spectral-shape variations are consistent with lower structural uniformity and stronger local distortion.Dry-ground samples show higher bulk conductivity than wet-ground samples,whereas wet-ground samples show much lower grain-boundary resistance.At 600 \u2103,the dry-60 min sample reaches a bulk conductivity of 13.54 mS cm-1,while wet-30 min shows the highest grain-boundary conductivity of 13.72 mS cm-1.These results suggest a processing-driven trade-off between bulk defect generation and grain-boundary blocking.Based on this processing understanding,Ca was introduced at the A site in Na0.52Bi0.47-xCaxTiO3.The x=0.04 sample reaches 8.35 mS cm-1 at 500 \u2103 and 18.98 mS cm-1 at 600 \u2103.
arXiv:2605.24754v1 Announce Type: new
Abstract: Neural network weights are increasingly a bottleneck for deployment, yet most compression pipelines treat layers independently and overlook cross-layer redundancy induced by function-preserving symmetries. We propose Motion-Compensated Weight Compression (MCWC), a weight-only codec that aligns permutation-symmetric blocks (e.g., hidden units and attention heads) to maximize cross-layer correspondence, turning depth into a predictable sequence. In the aligned coordinate system, MCWC uses a lightweight layer-sequential predictor with periodic keyframes and encodes only quantized prediction residuals using a learned entropy model trained under a rate distortion objective. A simple decoder reconstructs deployable weights by entropy decoding, dequantization, predictor-driven reconstruction, and inverse alignment, enabling fast weight materialization for inference. Across Transformer language modeling and vision classification, MCWC improves the rate accuracy Pareto frontier over strong quantization and learned weight-codec baselines, while maintaining competitive decode time. Ablations confirm that alignment, prediction, entropy modeling, and keyframe scheduling are each necessary for the full gains. Our code is available via https://github.com/Ism-ail11/MCWC.
arXiv:2605.24756v1 Announce Type: new
Abstract: Language-model agents increasingly emit uncertainty signals throughout a trajectory, but existing agentic UQ evaluations often conflate ranking usefulness with probabilistic truthfulness. AUROC, AUPRC, risk-coverage, Trajectory ECE, and scalarized trajectory scores evaluate discrimination, binwise calibration, or collapsed summaries, but do not strictly elicit the full prefix-conditioned success-probability trace $q_t = P^{\pi}(Y=1 | H_t)$. Building on prequential proper scoring, we introduce the Trajectory Proper Score (TPS), a predictor-agnostic family of strictly proper trajectory-level scoring rules for any per-step uncertainty signal calibrated into a probability of eventual success. We prove that TPS strictly elicits the success-probability process under complete observation, within the chosen score family and weight schedule. We extend the construction to administratively censored trajectories by projecting the complete-data score onto the observable stopped prefix, yielding an exact $q_Z$-weighted reduced score and a tractable approximation when $q_Z$ is unestimated. We further show that common trajectory evaluators target weaker objects than the full prefix-conditioned probability process: Trajectory ECE is resolution-blind, while scalarized Trajectory Brier elicits only the collapsed scalar, not the full trace. Experiments on StrategyQA, Tau2-Bench, HotpotQA, and WebShop show that these theoretical distinctions are operationally visible: probability recalibration can substantially change TPS while leaving rank metrics nearly unchanged, and the tractable censored approximation can change the verdict relative to complete-only evaluation.
arXiv:2605.24758v1 Announce Type: new
Abstract: We propose several linear, fully decoupled numerical schemes with first- and second-order temporal accuracy for a novel Q-tensor-based two-phase hydrodynamic model describing the coupling of active nematic liquid crystal solutions with isotropic solid substrates. The model is derived from the generalized Onsager principle and includes nontrivial terms that contribute zero to the total free-energy dissipation. We prove that the proposed decoupled linear schemes are thermodynamically consistent at the discrete level. In the passive limit, the SGE-BDF1 and SGE-PDG schemes are unconditionally energy stable, while the SGE-BDF2 scheme is energy stable with respect to a modified energy under a standard boundedness assumption and a sufficiently large stabilization parameter. We perform extensive numerical simulations to investigate how activity and other model parameters affect active nematic fluid-solid interactions. Finally, we analyze the physical mechanisms underlying the observed behaviors, providing deeper insight into the dynamics of soft confined active nematic fluids.
arXiv:2605.25502v1 Announce Type: new
Abstract: Educational aspect-based sentiment analysis (ABSA) can support course improvement, but public aspect-labeled student feedback remains scarce because educational reviews are private, institution-specific, and expensive to annotate. This study introduces a controlled synthetic benchmark for educational ABSA built from 10,000 synthetic course reviews with explicit train-validation-test splits and a 20-aspect pedagogical schema spanning instructional quality, assessment and course management, learning demand, learning environment, and engagement. The corpus is generated with sampled target labels, sampled nuance attributes, and a realism-tuned prompt refined through a three-cycle judge-editor procedure. On the resulting benchmark, local baselines with TF-IDF, two-step transformers, and joint encoders show that the task is nontrivial; the strongest untuned model, BERT, reaches a held-out detection micro-F1 of 0.2760, while a modest lower-rate BERT schedule improves this to 0.2930. Full-test GPT-based inference with gpt-5.2 reaches 0.2519 micro-F1 in zero-shot mode and 0.2501 with retrieval-based few-shot prompting, placing batch inference above the classical baseline and close to the compact joint encoders. A conservative external evaluation on 2,829 mapped student-feedback reviews from Herath et al. yields a micro-F1 of 0.4593 for BERT on a 9-aspect overlap, indicating partial synthetic-to-real transfer. Realism and faithfulness analyses are reported as generator diagnostics that clarify how the benchmark was stabilized and where label noise remains. The study therefore contributes a synthetic educational ABSA corpus, a documented generation procedure, and a reproducible benchmark setting for a domain in which public labeled data remain difficult to obtain.
arXiv:2605.24765v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly applied to cybersecurity question answering (QA) for critical tasks such as incident response and vulnerability analysis. However, real-world operational contexts, including system logs and network configurations, inherently contain sensitive identifiers, e.g., IP addresses, host names, and user accounts. Processing this data with cloud-based models is often unsafe or infeasible in regulated environments. Furthermore, progress in privacy-preserving QA is hindered by the lack of annotated, context-rich datasets capable of jointly evaluating operational reasoning and privacy preservation. To address this gap, we introduce CYBERMASKQA, a privacy-aware QA benchmark covering key security domains. Unlike existing benchmarks that primarily test factual knowledge, CYBERMASKQA grounds questions in realistic organizational contexts with explicit causal dependencies among assets and privileges. Generated through a systematic pipeline, the dataset combines human-curated base scenarios with LLM-driven semantic expansion, annotating each instance with precise private entity labels to enable controlled information disclosure. Evaluations of QA accuracy and masking performance demonstrate the benchmark's utility for developing deployable, context-aware cybersecurity models and facilitating nuanced studies of privacy-utility trade-offs. Upon acceptance, we will release the dataset and the generation framework.
arXiv:2605.24770v1 Announce Type: new
Abstract: Muon is a recently developed matrix-aware optimizer that has shown strong results in transformer training, but its behavior in vision transformers (ViTs) is not yet well understood. We study Muon for ViT training, largely on ImageNet-100 and Pl@ntNet-300K, comparing against AdamW under standard vision recipes involving mixup, cutmix, smoothing, and random augmentation and erasing. Muon consistently outperforms AdamW, with especially large gains on long-tailed Pl@ntNet macro top-1. These gains are also recipe-dependent, where Muon benefits much more than AdamW from advanced and significant data augmentation techniques. To understand this interaction, we analyze the singular-value structure of matrix gradients throughout the ViT. Within Muon training runs, removing heavy data augmentation induces a late-training spectral concentration and mode collapse in gradient matrices, primarily in deep MLP-down blocks. Under a fixed "full" augmentation recipe, the clearest Muon-AdamW contrast appears instead in QKV gradients, where AdamW gradient energy remains concentrated in a much narrower basis while Muon spreads energy across substantially more singular modes. Muon in ViTs is therefore best understood as an optimizer-recipe interaction. Under a fixed recipe, Muon differs from AdamW most clearly in attention projections, where its gradients consist of a broader spectral basis. Within Muon, a full training recipe is important for preventing late spectral concentration and mode collapse in deep feedforward blocks. We further demonstrate efficacy in training ViTs on image segmentation and masked autoencoder models, where Muon outperforms AdamW in all settings considered.
arXiv:2605.26054v1 Announce Type: new
Abstract: Variable-order time-fractional wave equations provide a flexible model for wave phenomena with evolving memory effects and anomalous temporal dynamics. Their numerical approximation is challenging because the variable-order fractional derivative generates time-dependent history weights and therefore lacks the standard time-translation-invariant convolution structure of constant-order fractional operators. In this paper, we develop and analyze a fully discrete energy-based discontinuous Galerkin (DG) method for wave equations with a Caputo-type variable-order time-fractional derivative. The equation is reformulated as a reduced first-order-in-time system, discretized in space by an energy-based DG method, and advanced in time using a second-order approximation of the variable-order Caputo derivative at a specially chosen point in each time interval. The main analytical novelty is a cumulative weight-variation estimate for the variable-order memory weights, which requires only that the variable order $\alpha:[0,T] \rightarrow (0,1)$ be Lipschitz continuous. Based on this estimate, we establish energy stability of the fully discrete scheme and derive second-order temporal convergence together with energy-norm spatial error estimates. The analysis gives suboptimal convergence on general affine simplicial or tensor-product meshes and optimal convergence under additional Cartesian and flux assumptions. Numerical experiments in one and two dimensions validate the theoretical findings.
arXiv:2605.25377v1 Announce Type: new
Abstract: Large Vision-Language Models (LVLMs) have advanced multimodal understanding, yet their reliability is limited by hallucination, where generated content conflicts with visual facts. Existing mitigation methods either rely on costly external interventions, such as instruction tuning and retrieval, or use internal mechanisms that remain limited by flawed attention weights and entangled hidden representations. We propose Adversarial Orthogonal Disentanglement (AOD), a latent geometric framework for mitigating LVLM hallucinations. AOD learns a hallucination-related direction through a minimax objective: a classifier concentrates hallucination signals into the projected component, while an adversary removes them from the orthogonal residual space via a Gradient Reversal Layer. The learned direction enables a training-free dual-forward-pass contrastive decoding strategy that suppresses hallucinations while preserving general capabilities. Experiments on three LVLMs across four hallucination and four utility benchmarks show that AOD consistently outperforms strong baselines. It improves POPE accuracy by over 6\% on average, boosts AMBER by 6\%, and maintains strong performance on utility tasks such as MMMU. Further analysis shows robust transfer across datasets, suggesting that AOD captures general hallucination-related biases rather than dataset-specific artifacts. Our source code and datasets are available at https://github.com/Hunter-Wrynn/AOD.