arXiv:2607.10809v1 Announce Type: new
Abstract: A compact and high order HWENO scheme using ADER (Arbitrary high order using DERivatives) time discretization is developed for hyperbolic conservation laws on the triangular mesh, which is the extension of the work on the structured mesh (Luo et. al. (2024) \cite{luo2023}). The Lax-Wendroff procedure is employed to convert time derivatives to spatial derivatives. Thanks to this, the cell averages of the derivatives of the solution can be obtained by the time accurate solution as Gaussian points along the cell interfaces through the Green-Gauss theorem instead of by the evolution solution directly in the conventional HWENO methods. Comparing with the existing Runge-Kutta HWENO (RK-HWENO) method on the unstructured mesh (Zhao et. al. (2025) \cite{zhao2025}), the new method has the following advantages. Firstly, the RK-HWENO method must solve the additional equations for reconstructions and time advancing, which is avoided for the new method. Secondly, the HWENO reconstruction in the new method is performed once per time step and is different from the RK-HWENO method, in which the reconstruction is performed several times every time step. Because of these advantages the new method is more efficient than the RK-HWENO method with smaller numerical errors and less computational costs. Besides, comparing with the existing ADER-WENO methods \cite{dumbser20071,dumbser20072} under the same order of accuracy, the stencil of the new method is more compact since the both the function and its first derivative values are used in the reconstruction of the HWENO schemes. Numerical examples demonstrate that the new method can achieve the high order for smooth solutions both in space and time, keep non-oscillatory near discontinuities.
Science Journals
arXiv:2607.09934v1 Announce Type: new
Abstract: A multi-species Bhatnagar-Gross-Krook (BGK) model for gas mixtures is presented that achieves the correct species-wise relaxation of velocities, temperatures, and pressure tensors according to the Boltzmann collision integral, as well as the correct mixture Prandtl number, while retaining a single relaxation term per species. The model extends the ellipsoidal statistical BGK (ESBGK) model by introducing relative relaxation targets for each species, derived from the Variable Hard Sphere (VHS) production rates of the Grad 13 approximation. Three approaches for the species relaxation frequency are proposed and analyzed: a Grad 13-based per-species frequency, a mixture-averaged frequency, and an empirical harmonic mean of the two. The model is implemented in the particle-based code PICLas and verified against Direct Simulation Monte Carlo (DSMC) results for a range of test cases, including 0D reservoir relaxation, mass diffusion, supersonic Couette flow, and hypersonic flow around a 70{\deg} blunted cone for binary and ternary gas mixtures. Across all test cases, the proposed model reproduces the correct Prandtl number, species temperature, velocity relaxation rates and pressure tensor relaxation, with the empirical relaxation frequency consistently yielding the best agreement with DSMC.
arXiv:2607.10577v1 Announce Type: new
Abstract: Perhaps one of the most intriguing phenomena in time-varying-media photonics is the amplification of light in a photonic time crystal (PTC). However, studies to date have focused only on the PTC-based amplification of coherent light. In this work, we theoretically examine the PTC-based amplification of thermal radiation, specifically blackbody radiation. Such amplification is fundamentally intriguing because of the inherently stochastic nature of thermal radiation, and technologically relevant because of its ubiquity. For simplicity, and because of the experimental relevance of transmission lines, we consider a one-dimensional medium. To analyze the PTC-based amplification of blackbody radiation, we examine the spatial correlations and spatial spectra of the electromagnetic fields. We show that the initially blackbody radiation periodically converges to Gaussian spatial correlations and spectra, with gradually increasing amplitudes, coherence lengths, and both spatial- and wavenumber-domain purities. We further demonstrate that these asymptotic behaviors are governed by the momentum band structure of the PTC and can be understood using a rotating-wave approximation for the pseudo-Hermitian dynamics of an electromagnetic field in a PTC.
arXiv:2607.10077v1 Announce Type: new
Abstract: Tabular learning is still dominated by gradient-boosted decision trees (GBDTs), while recent deep learning approaches have become increasingly competitive. However, applying deep tabular models to large-scale datasets remains challenging, as large sample sizes, high feature dimensionality, or many target classes can introduce substantial computational cost. We propose TabLoRA, a parameter-efficient trainable neural ensemble for large-scale tabular learning. Instead of using fully independent ensemble backbones, TabLoRA shares a common backbone across predictors and introduces predictor-specific low-rank adaptations, enabling ensemble-style prediction without full parameter duplication. Across benchmarks, TabLoRA achieves a favorable balance between predictive performance and practical efficiency compared with GBDT methods and recent deep learning baselines under the same resource constraints. Memory analysis and ablation studies further show that the proposed design improves the feasibility of neural ensemble learning while preserving much of the benefit of full ensembles.
arXiv:2607.10081v1 Announce Type: new
Abstract: The means to execute and orchestrate software components has changed from human-written code to descriptive prose. In high performance computing, this transition is represented in application orchestration, workload management, and system monitoring and debugging, to name a few. The underlying means to enable descriptive definition of tasks is the use of the Large Language Model with associated tool functions and resources. A combination of a model with access to such resources, modeled in software, encompasses an autonomous framework. As fully automated and agentic frameworks are developed for science, it is important to assess reliability and strategies scoped to specific tasks. In this work, we assess the extent to which an agentic framework can optimize and run an HPC scaling study with a low latency network in Amazon Web Services, accurately transform HPC job specifications between workload managers, and design and run an entire biosciences workflow. We find that the framework completes all three tasks while surfacing task-specific failure modes. In the scaling study, agents deploy and optimize applications but monitor running jobs inefficiently, preferring conservative fixed waits over event subscriptions. In job translation, they convert specifications between Slurm and Flux with high accuracy, with processor-affinity flags the most common error. In the bioscience workflow, the agent reproduces an expert-written variant-calling pipeline almost exactly -- agreeing with the reference call set in 18 of 19 completed runs -- and reaches this result through many distinct yet functionally equivalent workflow implementations. This information is invaluable moving forward to developing multi-cluster setups with scheduling and transformation handled by agents.
arXiv:2607.10164v1 Announce Type: new
Abstract: Surface tension is central to many two-phase flows, making accurate numerical schemes essential for predicting its effects. The integral formulation introduced by Popinet and Zaleski (1999) provides a natural discretisation that conserves momentum locally and globally and extends directly to variable surface tension, including Marangoni flows. However, to the authors' knowledge, only two-dimensional formulations have been reported, mainly because robust implementation in three dimensions is challenging for interfaces with complex geometries.
This work presents the first three-dimensional integral surface tension scheme, implemented within a sharp front-tracking framework. The method is tested for static and translating spherical droplets, oscillating droplets, thermocapillary motion, and rising bubbles. Results are compared with analytical solutions, experimental data, and established approaches, including the continuous surface force (CSF) and smoothing-based methods.
The proposed scheme produces spurious velocities comparable to CSF, while providing greater accuracy in all other tests. The largest improvements occur for droplets oscillating at low Ohnesorge numbers, variable-surface-tension flows, and strongly deforming rising bubbles. For a thermocapillary-driven droplet, terminal-velocity errors are reduced by up to five orders of magnitude relative to smoothing-based methods. The predicted steady-state shapes of rising bubbles also agree substantially better with experiments, particularly at low Morton numbers.
arXiv:2607.10195v1 Announce Type: new
Abstract: We investigate the problem of computing the distribution function for the shortest and longest path lengths in a directed graph with random edge lengths. Specifically, when these lengths are uniformly distributed, the problem reduces to computing the volume of a polytope defined by the graph structure. We establish that the problem is $\#P$-hard, even under the restricted condition that the random edge lengths are identically and independently distributed (i.i.d.) according to any continuous probability distribution with certain natural conditions, the local uniformity. This hardness result applies broadly: while the uniform distribution provides an essential case for the reduction, other distributions -- such as exponential or normal -- are similarly hard because they contain uniform distributions in every arbitrarily small interval. Furthermore, we show that the problem is contained within $\mathrm{XP}$ with respect to the treewidth $k$ of the underlying undirected graph. For the specific case of i.i.d. uniform edge lengths, we present a novel dynamic programming algorithm that processes a tree decomposition by iteratively performing convolutions to propagate distribution functions. Our approach achieves a time complexity of $n^{O(k^2)}$ for any fixed treewidth $k$.
arXiv:2607.10335v1 Announce Type: new
Abstract: The integration of power converters is profoundly changing the power system dynamics and poses significant challenges for stability analysis. The dynamic interactions between the power grid and the heterogeneous converters are highly complex and difficult to analyze due to the curse of dimensionality. Moreover, system stability varies with the operating points, which are determined by the voltage magnitude, active power, and reactive power of each converter. This further complicates the analysis as it is difficult to enumerate and examine all the possible operating points. To tackle these challenges, this paper proposes a geometric decentralized stability certificate for power electronics (PE)-dominated power systems, which can simultaneously handle heterogeneous power converters and their variable operating points. The certificate can be checked in a decentralized and modular manner, and it is scalable for large-scale power systems. Our approach is developed based on the concept of Davis-Wielandt (DW) shell and its projections, which can effectively visualize the characteristics of high-dimensional complex matrices. We investigate how the projections of the DW shell vary with operating points and how this variation can guide the search for worst-case operating conditions. We further propose an efficient algorithm to compute the stability margin and construct the certified operating regions. The effectiveness of the proposed method is validated through case studies on single-converter and 54-converter wind power systems.
arXiv:2607.10666v1 Announce Type: new
Abstract: Deploying AI-based visual inspection in manufacturing is hard because requirements change often, new defect types appear, and large labeled datasets are rarely available. We propose answer-conditioned chain-of-thought (CoT) distillation for rapidly adapting small vision-language models (VLMs) to new industrial tasks using minimal labeled data. A frontier VLM receives each training image along with its correct label and generates a justified visual explanation. A 3B-parameter model is then fine-tuned on these reasoning-augmented examples via LoRA. By conditioning on correct answers, we ensure all training reasoning is directed toward the correct conclusion, which is critical because frontier models score as low as 24.1% on our hardest task. We validate on four industrial classification tasks spanning three image modalities using only 18 to 30 labeled images per task. Across 4 seeds per task (32 training runs), our method outperforms direct fine-tuning on all 16 seed-task combinations, with mean improvements of +1.7 to +4.4 percentage points. A controlled equal-budget experiment confirms the improvement comes from reasoning quality, not additional training steps. An unconditioned baseline demonstrates that with out answer-conditioning, wrong reasoning degrades performance by 17.8 percentage points. On weld radiograph classification, the fine-tuned 3B model outperforms GPT-4.1 by 10.0pp using just 24 training images.
arXiv:2607.10539v1 Announce Type: new
Abstract: Existing approaches to infer user traits and generate responses consistent with a persona rely on static prompting. They lack calibrated uncertainty, ignore sequential evidence, and drift during long interactions. We present \textbf{AI YOU}, a framework that continually updates a personality profile with 22 dimensions from conversation and embodies it in a personal digital twin. Practically, the system combines prompting, Bayesian updating, and conformal prediction for persona inference. A periodically refreshed memory anchor and cognitive memory with three layers preserve persona consistency over long interactions. Across the main results, AI YOU \emph{(i)} achieves conformal coverage ranging from 0.921 to 0.976, \emph{(ii)} improves uncertainty calibration and reasoning grounded in memory, and \emph{(iii)} enhances persona fidelity over static prompting in role playing over 100 turns while reducing trait drift, for most evaluated backbones under adversarial settings with multiple agents. The prototype \emph{AI YOU Town} initializes an imaginative twin world for future interaction. The online demo is available at \href{https://quinnnnnne-ai-you.hf.space/}{\mbox{\texttt{quinnnnnne-ai-you.hf.space}}}.
arXiv:2607.10674v1 Announce Type: new
Abstract: As AI code tools become integrated into programming environments, students increasingly describe intended behavior in natural language and rely on these tools to generate code, shifting emphasis from code writing to specification. Yet little is known about the comments students write as specifications in AI-assisted programming tasks. We analyze a four-year dataset of undergraduate programming submissions and reflections from tasks in which students wrote comments to guide code generation and refined solutions using test-case feedback. We introduce a taxonomy spanning three dimensions: comment type, code expression level, and code construct. Using automated classification, we examine how these dimensions vary across attempts and how students describe the process in their reflections. Our findings show that students mostly wrote natural-language What comments, shifted toward How comments for more procedural constructs, and focused more on verifying generated code than on repeatedly rewriting comments.
arXiv:2607.09740v1 Announce Type: new
Abstract: Safe motion planning in advanced driver-assistance systems and autonomous vehicles requires an accurate understanding of how the surrounding traffic scene is likely to evolve. However, many existing lane-change prediction methods remain centered on a single target vehicle, while multi-agent forecasting approaches often describe scene evolution only through future positions and provide limited explicit information about the maneuver associated with each vehicle. This study proposes a dynamic scene graph attention framework that predicts the lane-change intention and future trajectory of every relevant vehicle within a local traffic scene. The scene is represented as a time-varying interaction graph in which vehicles are modeled as nodes and their spatial and kinematic relationships are encoded through explicit edge features. Temporal graph-attention message passing captures evolving inter-vehicle dependencies and pre-maneuver cues, while an intention-guided decoder links each predicted maneuver to its corresponding future motion. A scene-level consistency objective further encourages compatible multi-vehicle futures. Experiments on the NGSIM I-80, NGSIM US-101, and highD datasets demonstrate consistent improvements over competing baselines. DSiGAT achieves intention prediction accuracies of 90.12% and 90.97% on NGSIM I-80 and US-101, respectively, and reduces trajectory RMSE by up to 52.94% relative to the strongest baseline. It also produces lower inter-agent collision rates and joint displacement errors, indicating more coherent scene-level predictions. Ablation, sensitivity, robustness, and qualitative analyses further validate the contribution of the proposed components and the effectiveness of the scene-focused formulation.
arXiv:2607.11422v1 Announce Type: new
Abstract: We study the truncated trapezoidal rule, a family of A-stable second order multistep methods parametrized by an integer $J \ge 2$, a compromise between BDF2 and the trapezoidal rule. We obtain a closed-form expression for the coefficients that minimize the principal error constant under the A-stability constraint and we derive an explicit formula for the corresponding principal error constant. The latter decreases to the optimal Dahlquist value $1/12$ as $J$ increases, and is strictly smaller than the BDF2 constant for every $J \ge 2$. We apply the truncated trapezoidal rule within the convolution quadrature framework. Its analyticity in a neighbourhood of the closed unit disk yields milder regularity and perturbation requirements than those of the trapezoidal rule, while the error constant can be made arbitrarily close to the optimal one by increasing $J$. Numerical experiments show that convolution quadrature based on the truncated trapezoidal rule remains stable under symbol perturbations, where the trapezoidal rule fails, while achieving a smaller error constant than BDF2.
arXiv:2607.10694v1 Announce Type: new
Abstract: We study the problem of optimal continual fine-tuning for a pre-trained Foundation Model deployed at a resource-limited device. At each time slot, a new batch of training data arrives, and the controller is faced with two options: either use the data to fine-tune the model and incur a compute cost, or do not fine-tune the model and discard the data. After the decision, the performance of the current model is measured in terms of an application-specific performance metric such as classification accuracy. Our objective is to learn an optimal policy that determines \emph{when to fine-tune the model} on a single task (e.g., sentiment analysis), under a finite compute budget. We formulate this online decision-making problem as a constrained Markov Decision Process, where the system state captures three essential aspects: (\textit{i}) model's performance, (\textit{ii}) computational budget, and (\textit{iii}) data distribution relevance to historic data encountered up to that point. The transition to the next state is stochastic and therefore, we propose a reinforcement learning-based method to solve this problem, namely the \emph{actor-critic} algorithm. We also consider the special case where the performance of fine-tuning for a given model can be predicted or estimated prior to decision; in this case the problem becomes a Dynamic Programming one. Experiments with a large pre-trained model on a widely-used text classification dataset demonstrate that our method consistently outperforms fine-tuning approaches with the same compute budget by more than $4\%$ in terms of accuracy and achieves $97\%$ of full-parameter fine-tuning accuracy while requiring only $25\%$ of the fine-tuning steps.
arXiv:2607.11874v1 Announce Type: new
Abstract: Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, then train policies via reinforcement learning (RL) to track them. But how does this recipe transfer to dexterous manipulation? The answer is not obvious, as manipulation involves complex, contact-rich dynamics and requires delicate regulation of contact modes and forces. We present REGRIND, a minimalist retargeting-guided RL pipeline that learns dexterous manipulation policies from a single human demonstration. REGRIND retargets human hand-object motion to a robot reference that preserves hand-object spatial and contact relationships, trains a residual RL policy in simulation to track object-centric keypoints along that reference, and transfers the resulting policy zero-shot to hardware with careful system identification. The resulting policies produce fluid, human-like behavior on two different multi-fingered hands across contact-rich tool-use tasks, including operating a pair of scissors and turning a screwdriver. Through systematic hardware experiments, we identify and analyze the key factors that govern sim-to-real transfer in dexterous manipulation, offering practical guidance for retargeting-based learning in contact-rich settings. Videos and code are available at https://yunhaifeng.com/REGRIND.
arXiv:2607.10565v1 Announce Type: new
Abstract: End-to-end motion planning has emerged as a promising paradigm in autonomous driving, directly mapping raw sensor data to control commands via deep neural networks. Despite its advantages, its large model size hinders deployment in resource-constrained platforms. In this paper, we present BucketKD, a bucket-based knowledge distillation framework that yields compact and safety-aware end-to-end planners. Compared to the state-of-the-art approach, which relies on simplified planning state representations, BucketKD discretizes critical environmental variables into adaptive buckets that capture richer scene semantics while preserving efficiency. In addition, we design a safety-aware waypoint attention mechanism that evaluates each waypoint's risk level by accounting for both obstacle proximity and relative motion through a time-to-collision (TTC) formulation widely used in transportation research. This enables the student model to better retain safety-critical behaviors during distillation. Extensive experiments in CARLA using the Bench2Drive dataset show that BucketKD significantly outperforms the state-of-the-art in both planning accuracy and safety while maintaining strong compression ratios.
arXiv:2607.09721v1 Announce Type: cross
Abstract: To help evaluate the mathematical skills of current AI systems, we present a set of formulas for fundamental mathematical constants. These problems are attractive for AI evaluation because they are concrete and can be checked numerically to arbitrary precision, yet proving them may require non-obvious mathematics. Mathematical constants such as $\pi$, $e$, Catalan's constant, and special values of the Riemann zeta function have fascinated mathematicians for centuries. The search for formulas evaluating mathematical constants has produced some of the most beautiful mathematics in the field, especially in cases that yield irrationality proofs or fast convergence rates. Ramanujan's legacy is emblematic of this tradition. The list we provide contains two types of problems: formulas whose proofs are known to the authors but will remain encrypted for a short initial period; and formulas that are not yet proven. We are curious to see the achievements of AI in both cases.
arXiv:2607.09752v1 Announce Type: cross
Abstract: This paper presents a machine learning framework for data-driven inverse design of V-beam thermal sensors. The goal is to determine the optimal sensor geometry: beam inclination angle, beam length and beam width that achieves a target displacement under a given temperature. The design should also provide the geometry with minimum structure volume and minimum mechanical stress the sensor must support. This problem is ill-posed as for a given displacement there are multiple possible geometric configurations, causing direct regression methods to fail. We document a series of five exploratory trials that progressively revealed the nature of the problem culminating in a two-phase solution: a neural network forward model trained to map geometry and material constants to sensor responses, a gradient-descent inverse optimization over the frozen forward model, minimizing stress and volume simultaneously. The proposed pipeline utilizes a 3000-sample dataset and achieves a MAPE of 4.76% for predicting the displacement, more than 70% of predictions having MAPE of under 5%.
arXiv:2607.09772v1 Announce Type: new
Abstract: Autonomous driving systems require reliable safety validation before real-world deployment. However, large-scale road testing is costly, difffcult to reproduce, and inefffcient for exposing rare safety-critical scenarios. Conventional simulation improves repeatability, but an offfine simulator alone cannot continuously connect physical trafffc states, virtual reconstruction, algorithm evaluation, and scenario evolution. This paper proposes a risk-ffeld enhanced closed-loop digital twin framework for autonomous driving safety validation. The framework integrates physical data acquisition, data synchronization, virtual twin reconstruction, risk-aware scenario generation, autonomous driving algorithm evaluation, and safety analysis. A driving risk ffeld is introduced as a uniffed intermediate representation to describe obstacle, lane-departure, road-boundary, time-to-collision, and comfort-related risks around the ego vehicle. The risk ffeld ranks high-risk scenarios in the digital twin scenario library and provides dense safety guidance for reinforcement learning-based driving policies. A simulation-style evaluation protocol is designed to compare conventional reinforcement learning baselines, risk-penalty baselines, and the proposed risk-ffeld guided method. The study indicates that embedding explicit risk structure into digital twins can make autonomous driving validation more targeted, interpretable, and reusable, while its practical effectiveness remains bounded by model ffdelity, risk calibration, and sim-to-real transfer.
arXiv:2607.09782v1 Announce Type: new
Abstract: In evolutionary game dynamics, there exists a hypothesis, which states that, the dynamic structure of the game's steady -- state system is characterized by the linear superposition of eigenmanifolds, which depends specifically on the eigenvector structure at the Nash equilibrium and is ultimately governed by the game dynamics equations. This hypotheses has been supported widely in discrete strategy game. In continuous -- strategy game, using experimental data from human -- subject games, this paper finds that the hypothesis is supported in significant, too.
arXiv:2607.09767v1 Announce Type: cross
Abstract: The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech. It relies on content embeddings extracted from a frozen pretrained wav2vec2 encoder. These embeddings are decoded into an anonymized signal using vector quantization and a HiFi-GAN vocoder, both trained on LibriTTS without any waveform reconstruction loss or speaker embedding mapping. The training objective enforces that embeddings of the anonymized signal match those of the original one. While training, an auxiliary speaker classification branch with a gradient reversal layer is used to discard speakerspecific information. Results show that this straightforward embedding-based approach achieves very low WER (2.53) with an anonymization performance (EER 13.39) ranking within first level for VPC. Notably, emotions are partially preserved (UAR 43.91), even without a supporting training objective, while the anonymized voice is audible without reconstruction loss.
arXiv:2607.09826v1 Announce Type: new
Abstract: Dysgraphia is a specific learning disability that is prevalent among school-age children. It affects handwriting coherence, quality, fluency, and legibility, often hindering academic achievement and early learning development. This motor coordination disorder is typically diagnosed through subjective assessments based on clinician observation, which can be timeconsuming and prone to variability. In this paper, we introduce a deep learning-based framework for objective dysgraphia detection using online handwriting data captured via digitizing tablets. The proposed framework relies on two complementary branches: the first pipeline extracts both handcrafted and embedding-based kinematic features directly from raw temporal signals, while the second leverages image-based representations of the temporal signals generated using continuous wavelet transforms (CWT) and Gramian Angular Fields (GAF). The resulting features are then fused to leverage the complementary strengths of both representations. The four representations were evaluated separately and jointly using the publicly available DiaGraMo dataset, showing that the fusion of GAF, MOMENT, and hand-crafted kinematic features outperforms each individual representation, as well as other fusion schemes. These findings highlight the potential of the complementarity of image and signal based representations for more objective dysgraphia detection.
arXiv:2607.10749v1 Announce Type: new
Abstract: Reflections of water pose a significant challenge for computer vision systems, as standard deep learning models frequently confuse objects with their mirror images, producing spurious false positives and negatives in tasks such as object detection and semantic segmentation. As a result, detecting reflection axes in natural-water scenes is pivotal for reliable object detection and scene understanding. To mitigate this issue, we leverage the intrinsic imperfect reflective symmetry of water and introduce a Symmetry-Aware Water Reflection Detection Network, namely, SAWRD-Net, that couples dihedral group-equivariant convolutions with a matrix-decomposition decoder in an end-to-end framework. First, dihedral group convolutional layers extract geometry-consistent feature maps that explicitly encode both rotational and mirror symmetries. A Multi-scale Reflection Equivariant block then aggregates features across scales and employs a symmetric-attention mechanism to highlight reflection-relevant regions. The proposed matrix-decomposition decoder factorizes high-dimensional features into compact low-rank parameter and confidence spaces, after which the network directly regresses keypoints on the reflection axis. Then a robust principal component analysis fits the final axis. Evaluated on the largest available water reflection scene data set, SAWRD-Net achieves a true-positive rate of 0.890 against human annotations, outperforming all existing water reflection detectors.
arXiv:2607.09945v1 Announce Type: new
Abstract: Purpose: Motion compromises the utility of high-resolution 3D MRI, an established tool in quantitative neuroimaging research. Deep learning-based methods have shown promise for mitigating motion-induced artifacts, but their development typically requires simulated motion-corrupted data. Several open-source tools exist for this task, each implementing different algorithms. However, no scheme currently exists for evaluating the accuracy of these simulations, making it difficult for users to choose the most suitable tool. Developing such a scheme is the aim of this study. Methods: The essential ingredient of the desired scheme is a ground-truth reference simulation that does not suffer from sampling-induced error. To meet this requirement, the proposed scheme, APHABAMAS, leverages a digital phantom whose representations in both the image and Fourier domains can be expressed analytically under arbitrary rigid-body transformations. Results: APHABAMAS is used to quantify the sampling-induced errors of three existing simulation algorithms, establishing their first definitive accuracy-based ranking. Conclusions: APHABAMAS provides a rigorous tool for assessing the accuracy of high-resolution 3D MRI motion-artifact simulations. It allows the accuracy-based ranking of existing simulation algorithms to be established, thereby enabling informed selection of the most suitable algorithm for synthesizing motion-corrupted data.
arXiv:2607.09985v1 Announce Type: new
Abstract: Object pose estimation is a fundamental problem in 3D vision. Although recent state-of-the-art approaches achieve strong performance, they often overfit to existing benchmarks and exhibit limited generalization to novel categories and unseen scenes. We propose UniPose9D, a category-agnostic foundation model for 9D object pose estimation: given an instance mask/ROI and either an RGB-D observation or an RGB image with predicted depth, the model estimates rotation, translation, and metric size without category labels, CAD models, mean-shape priors, or reference views. Specifically, UniPose9D samples point pairs from the observed object geometry and uses DINOv2 and PointNet features to predict NOCS coordinates for each pair. To improve accuracy, we introduce a point-pair-based RANSAC N-hop Kabsch--Umeyama algorithm with an adaptive threshold. We further employ flow matching to address symmetric ambiguities and construct a large-scale training set by curating and aligning pose annotations from existing public datasets. Experiments across six datasets show that a single unified model can match or surpass specialist methods while generalizing to unseen objects and in-the-wild scenarios. Our code and model are available on https://github.com/qq456cvb/UniPose9D.