arXiv:2607.01205v1 Announce Type: new Abstract: We present Linkify, a framework for learning from interface-augmented assembly graphs to enable context-aware part retrieval in mechanical assemblies. While recent generative AI methods for CAD have focused largely on isolated parts or monolithic assemblies, the rich geometric information at the interfaces between parts, where function is realized, remains underexplored. We address this gap by recomputing high-fidelity interface geometry for the Fusion 360 Gallery Assembly dataset, correcting missing and erroneous contacts, and generating point-cloud representations of local contact regions. Using this data, we construct assembly graphs whose nodes encode part geometry and whose edges encode interface geometry via a pretrained point-cloud encoder. On top of this representation, we train a Graph Attention Network based on GATv2 to solve a masked part prediction task: given an assembly with one part held out, the model predicts the class of the missing component from a large vocabulary of geometrically clustered parts, thereby approximating a realistic part-retrieval scenario. Compared to non-graph baselines such as logistic regression and k-nearest neighbors operating on aggregated node features, Linkify achieves higher Top-K accuracy and F1 scores. Ablation studies on graph connectivity, edge attributes, and attention mechanisms demonstrate that accurate contact computation and dynamic attention over interfaces are critical for performance. Our corrected interface dataset and training pipeline, released publicly, provide a foundation for future interface-aware models for assembly retrieval, validation, and generative design.
Science Journals
arXiv:2607.01218v1 Announce Type: new Abstract: Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions. We formulate the \emph{state-prediction separation hypothesis}: disentangling the two roles yields better language modeling performance. We design a Transformer variant that uses two computation streams to separate the two functions, and conduct pretraining experiments across various scales. Our experiments show that state-prediction separation consistently offers better data and compute efficiencies, improving validation loss and outperforming standard Transformers by 2--3 percentage points on average on downstream tasks. We also conduct extensive empirical analysis that rules out potential confounders and demonstrates the fundamental difference in the gradients our design entails.
arXiv:2607.00740v1 Announce Type: new Abstract: Microservice availability is commonly assessed by fault injection and chaos experiments, but such experiments are costly, operationally risky, and difficult to repeat for every architectural change. Distributed tracing and deployment metadata provide cheaper evidence, yet they usually remain descriptive: they show which services interacted, not what endpoint-level availability property follows. This paper proposes a formal runtime availability model based on stochastic connectivity for resilience-oriented analysis of microservice endpoints. It treats endpoint availability under explicit fault scenarios as a measurable facet of microservice resilience, combining a typed service-dependency graph, a replication map, a probability measure over node and edge states, and request-specific success predicates. Its semantics separates computational failures of service replicas from communication failures of logical dependencies, showing that replication cannot compensate for bottleneck dependencies. The model can be reconstructed from traces and deployment artifacts, parameterized for architectural what-if analysis, and analyzed by Monte Carlo simulation before or alongside fault injection. We define the model, its trace-to-model construction, elementary semantic properties, and a synthetic adequacy study. The study matches closed-form oracle cases within sampling error and exposes boundaries caused by edge bottlenecks, correlated failures, missing traces, and time-dependent failures.
arXiv:2607.00760v1 Announce Type: new Abstract: Long-context LLM services now sustain prompts with hundreds of thousands to millions of tokens, making the key-value (KV) cache a first-order serving cost. Because the cache grows linearly with context length, it can exhaust GPU memory, force smaller batches, and reduce serving throughput. Prior KV cache compression techniques typically target only the sequence dimension or only the channel dimension, which leaves limited headroom as context windows scale. Compressing both dimensions promises higher memory reduction, but applying the two forms of compression directly leads to significant accuracy loss. This paper introduces MosaicKV, a dynamic two-D (dimensional) KV cache compression system for extremely long-context serving. MosaicKV uses dynamic two-D compression to address the accuracy challenge, exploiting the non-uniform importance distribution of elements within the KV cache. Instead of applying one compression pattern globally, MosaicKV identifies important elements for each KV vector and selects compression strategies at the granularity of KV cache segments. To address the performance challenge, where fine-grained sparsity and compression management overhead can offset the gains from compression, MosaicKV introduces compressed KV cache management. This mechanism uses underutilized GPU and CPU resources to maintain compressed KV caches and accelerate attention computation. Evaluation on an H800 GPU with multiple LLMs shows that MosaicKV delivers up to 16x attention speedup, 4.8x lower decode latency, and 7.3x higher throughput than the uncompressed baseline. At the same time, it reduces memory usage by 3x and incurs only 1.76% average accuracy loss on LongBench and RULER.
arXiv:2607.01234v1 Announce Type: new Abstract: In the last decade, Russia's strategic arsenal has pivoted towards a reliance on exotic nuclear-weapon delivery systems. One such system, the Burevestnik (NATO: 9M730) is claimed to be a nuclear-powered, nuclear-armed cruise missile capable of nearly indefinite flight. The air-breathing nuclear propulsion system used in this missile is unique, and its attributes are generally unfamiliar to both the aerospace and nuclear-security communities. To better understand the Burevestnik, and the potential of air-breathing nuclear propulsion systems generally, we have developed a nuclear-aircraft modeling toolkit capable of constraining the missile's performance characteristics. Using this framework, we conclude that the Burevestnik is a subsonic cruise missile system measuring $9.5 \pm 0.32$~m in length, with a $5.6 \pm 0.18$~m wingspan, likely powered by a direct-cycle nuclear turbojet (our calculations almost entirely exclude the possibility of a nuclear ramjet). Under these assumptions, our models predict a reactor thermal power of $4.3\pm 1.3$~MWth at cruise, with peak power demand during climb and terminal maneuvering exceeding $15$~MWth, which may be met with a supplemental chemical interburner. Monte Carlo simulations show that escaping neutrons will generate in excess of 5~TBq of gaseous radionuclides per MW-hr of flight, including isotopes such as $^{41}Ar$, $^{85m}Kr$, $^{83m}Kr$ and $^{14}C$, some of which may be detectable using existing monitoring networks.
arXiv:2607.00068v1 Announce Type: cross Abstract: The axion represents a strong candidate for weakly interacting dark matter. To date, high sensitivity lab based experiments and astrophysical observations have ruled out a substantial part of the axion mass and photon coupling parameter space. However, a challenge remains in searching for the presence of the axion in the higher mass range 0.01-1eV corresponding approximately to axion field oscillation at THz frequencies. This work investigates via numerical simulation the feasibility of a high sensitivity, lab-based axion sensor operating in this range, based on plasmonic electric field enhancement by a nanostructured metasurface, combined with heterodyne detection and quantum sensing via nitrogen-vacancy (NV) centers in diamond. Estimates of the sensor response to anomalous electromagnetic fields resulting from axion coupling are given using Ti/Au nanopillars on LiNb at axion mass corresponding to telecommunications wavelength ($\approx$0.8eV, 196 THz). Finally, the possibility of sensing in the lower axion mass $<$10$^{-2}$ to 10$^{-1}$eV range is explored using alternative materials, with CdTe as an example.
arXiv:2607.00413v1 Announce Type: cross Abstract: Spin correlations are among the most fundamental quantum observables in many-body systems, yet they remain difficult to access experimentally in relativistic heavy-ion collisions. Existing spin measurements, including hyperon polarization and vector-meson spin alignment, have revealed important single-particle spin phenomena, but genuine two-particle spin correlations in the produced hadronic system remain largely unexplored. Here we propose spin femtoscopy, a framework for accessing genuine two-particle spin correlations through spin-resolved femtoscopic measurements. The key principle is that different two-particle spin configurations can give rise to different femtoscopic correlation functions because of quantum statistics, spin-dependent final-state interactions. Using $\Lambda\Lambda$ pairs as a proof of principle, we exploit the self-analyzing weak decay of $\Lambda$ hyperons to construct spin-sensitive femtoscopic correlation functions with different singlet and triplet admixtures. We show that these observables provide experimental access to the spin-state populations of the pair and allow genuine spin correlations to be separated from spin-dependent femtoscopic mixing caused by quantum statistics and final-state interactions. This work extends femtoscopy from a probe of source geometry and final-state interactions to a framework for revealing the quantum spin structure of strongly interacting matter.
arXiv:2607.00284v1 Announce Type: cross Abstract: The fidelity of a quantum gate is sensitive to small deviations in the physical control parameters. Unfortunately, it is generally difficult to exactly model the implemented Hamiltonian for a set of user-defined parameters, necessitating on-device calibration. Here, we present an active learning framework based on Bayesian optimization with a Gaussian Process surrogate to find the optimal parameter set. We validate the technique through numerical calibration of the laser amplitude and frequencies that implement the trapped-ion M{\o}lmer S{\o}rensen gate. We show that a Gaussian process can model the Hamiltonian dynamics. The addition of active learning accelerates the discovery of the optimal parameter set with speed and final fidelity dependent on the quantum projection noise of the data. These results establish the utility of active learning and surrogate models for quantum calibration and control.
arXiv:2607.00328v1 Announce Type: cross Abstract: One of the traditional approaches for constructing approximate policies for dynamic assortment optimization problems is to use sampling-based inventory-agnostic policies. Such policies are called sampling-based, as they sample an assortment of products from a fixed distribution at each time period to offer to a customer of each type. Such policies are called inventory-agnostic, as the sampled assortments may include products without remaining inventories, so if a customer chooses a product without remaining inventories, then she leaves without a purchase. Inventory-agnostic nature of a policy is not a concern, because it is known that if the policy samples an assortment that includes products without remaining inventories, then dropping the products without remaining inventories does not degrade the performance. However, sampling-based nature of a policy is a concern, because sampling brings another source of uncertainty in the performance. In this paper, we give an algorithm to de-randomize any sampling-based inventory-agnostic policy, so the de-randomized policy offers a deterministic sequence of assortments within the support of the original policy without degrading the performance. Furthermore, we give a variation of our de-randomization algorithm that searches for a deterministic sequence of assortments beyond the support of the original policy. We show that we can implement the latter variation efficiently as long as we can solve the static assortment optimization problem under the choice model governing the choice process of the customers. As our crowning technical contribution, we study locally-optimal deterministic policies, where changing any single one of the assortments in the policy does not improve the total expected revenue. We show that any locally-optimal policy has a performance guarantee of 1/2 - epsilon when compared with the best sampling-based policy.
arXiv:2607.00459v1 Announce Type: cross Abstract: The Potential Field Source Surface (PFSS) extrapolation is a method for estimating the large scale coronal magnetic field from photospheric magnetograms. The source surface serves as the outer boundary of its solution domain, and is typically a spherical surface. An appropriate source surface radius ($R_{ss}$) enables more accurate identification of the coronal magnetic field topology and estimation of the open flux, thereby potentially enhancing the accuracy of space weather modeling. We prove the well-posedness of the PFSS forward problem and establish the existence and uniqueness of the optimal source surface by combining compactness of the admissible set with continuity of the objective functional. The objective functional is the mean squared error (MSE) between PFSS extrapolation and Parker Solar Probe (PSP) radial magnetic field measurements after Parker spiral backmapping and radial scaling for Encounters 1-19. The optimization algorithm is validated with an analytical solution, and Advanced Composition Explorer (ACE) in situ measurements are used as an independent cross-validation dataset. Additional evaluation metrics and Pareto analysis are used to identify the dominant metrics between open flux and polarity prediction accuracy. Our results show that the optimal $R_{ss}$ derived from the algorithm generally increase from solar minimum into the ascending phase of solar cycle 25. The optimized solution improves open flux agreement while preserving or improving polarity prediction accuracy relative to $2.5R_{s}$. The Pareto frontiers show a transition for dominant metrics from open flux during solar minimum to polarity prediction accuracy during the ascending phase.
arXiv:2607.00470v1 Announce Type: cross Abstract: We investigate a forecasting framework based on a simple discrete-time dynamic model with coefficients varying in time. The parameters of the model are recovered within a deep learning framework, which makes it possible to retain a transparent parametric structure while simultaneously accounting for complex and nonstationary patterns in the observed phenomenon. Our analysis covers two specifications of the noise process. Besides the standard Gaussian setting, we also consider Laplace-distributed noise, which can offer a more adequate description in the presence of heavier tails and sharper local fluctuations. For both cases, we formulate the predictive scheme of the model and analyze the associated uncertainty quantification, including the construction of prediction intervals. The results illustrate that a relatively simple model, when combined with time-dependent parameter estimation, can serve as a mathematically tractable and practically flexible tool for forecasting complex dynamics under different noise assumptions. The general model is stated for TVAR($p$), while the prediction-interval formulas and the numerical experiments are developed for the TVAR(1) case.
arXiv:2607.00668v1 Announce Type: cross Abstract: The chirality-induced spin selectivity (CISS) effect has been invoked to explain recent reports of differences in the time-resolved EPR signals between chiral and achiral molecules. However, the microscopic origin of these differences and their connection to CISS remains contested, particularly since these systems lack a metal interface. Here we introduce an intramolecular spinterface-like mechanism that naturally arises within donor-chiral bridge-acceptor (D--$\chi$B--A) complexes and quantitatively reproduces experimentally reported observed spin polarization in time-resolved EPR studies. In our two-electron Lindblad model, the photoexcited charge-transfer electron traversing the chiral bridge exchanges with the residual donor electron, which acts as a localized magnetic moment analogous to an induced magnetic moment on an electrode surface. The resulting through-bridge charge current produces an effective solenoidal field at the donor--bridge interface, breaking spin degeneracy and directional symmetry, thus enabling spin-selective transport without invoking intrinsic spin-orbit coupling on the bridge. We show that the interplay between this current-induced field, donor thermalization (which breaks time-reversal symmetry), and bridge spin mixing yields tens-of-percent polarization over realistic experimental conditions and charge-transfer time scales, matching reported CISS signatures in triads and DNA hairpins. By explicitly resolving the dependence on solenoidal coupling strength, temperature, and spin-mixing rates, the model identifies the regime in which internal spinterfaces can generate robust CISS-like spin filtering. These findings demonstrate that CISS-like signals in isolated D--$\chi$B--A complexes are fully compatible with a spinterface mechanism, providing a unified conceptual framework for interpreting both device-based and molecule-internal CISS platforms.
arXiv:2607.00844v1 Announce Type: cross Abstract: We consider the problem of designing input signals for an unknown linear time-invariant system in such a way that the resulting data, within a finite horizon, is suitable for identification with a desired accuracy. We consider both noise-free and noisy settings with $\ell_\infty$--bounded noise models. We will take into account general prior knowledge of the system parameters. Central in our study is the concept of universal inputs. An input is called universal for identification if, when applied to any system complying with the prior knowledge, it yields data suitable for accurate identification. We provide new methods for designing such universal inputs. Our results generalize the experiment design approach based on Willems et al.'s fundamental lemma that relies on persistently exciting inputs, and that is limited to prior knowledge on controllability. It turns out that for other types of prior knowledge, there exist universal inputs that outperform the persistently exciting ones, e.g., in terms of sample efficiency. Moreover, we investigate types of prior knowledge that enable experiment design for exact identification in the presence of noise.
arXiv:2607.00877v1 Announce Type: cross Abstract: Traditional variational Kalman filtering with unknown noise statistics suffers from inconsistent process covariance estimation and slow convergence speed, limiting its practical utility. To address these issues, we introduce a surrogate variable representing the process-noise-free state, which enables explicit modeling and inference of process noise statistics. In addition, we reformulate the conventional coordinate ascent variation inference (CAVI) as a marginalized maximum a posteriori problem, followed by a single-step hyperparameter fitting. This reformulation obviates the need for multiple inner iterations inherent to CAVI and decouples the design of the covariance tracking filters. Consequently, this architecture permits the deployment of higher-order filters for covariance tracking and enables sliding-window hyperparameter estimation. Notably, when this window encompasses all historical data, the covariance tracking estimator intrinsically operates as a zero-phase filter. Numerical simulations validate the theoretical framework, demonstrating the enhanced convergence speed and superior estimation accuracy compared with existing methods.
arXiv:2607.00884v1 Announce Type: cross Abstract: Searches for new physics at the LHC often look for localized excesses on smoothly falling background distributions. Several classes of background models have been considered, including polynomials and other parametric families; however, these approaches can require extensive analysis-specific development as datasets grow. In this work, we motivate the finite exponential mixture as a flexible semi-parametric class of functions for approximating falling distributions, drawing on results from extreme value theory. Using two published datasets ($n=28,619,185$ and $n=5,036$), we show that the exponential mixture performance is comparable to existing methods for both small and large datasets. Finally, in simulation studies ($n = 5,036$), we find that the finite exponential mixture exhibits small bias relative to the true statistical uncertainty while maintaining consistent nominal coverage in the bulk.
arXiv:2607.01057v1 Announce Type: cross Abstract: We study a broad class of graphical models whose independencies correspond to vertex separation in mixed graphs with directed, undirected, and bidirected edges, that are capable of encoding independence structures arising from feedback, latent and selection mechanisms. In particular, we introduce separable graphs, in which each missing edge implies the existence of a separating set for its endpoints, and essentially separable graphs, those graphs separation equivalent to a separable graph. We show that these models include many existing graph families used to define graphical models an provide several characterizations of separable graphs and essentially separable graphs. We also provide multiple characterizations of separation equivalence for separable graphs. One is a graphical characterization in terms of ordinary graph properties, extending earlier results for specific subfamilies Another is a separational characterization depending only on graph separation properties. Finally, we provide a canonical representation for the equivalence classes of essentially separable graphs and develop an algorithm that, under suitable assumptions, identifies the equivalence class of any essentially separable graph.
arXiv:2607.01089v1 Announce Type: cross Abstract: Active learning reduces labeling cost by querying the most informative unlabeled samples, but standard coreset methods ignore known data symmetries and can waste budget on transformed versions of the same instance. We propose GRINCO, a group-invariant coreset framework that performs acquisition in the quotient space induced by a transformation group, so that selection operates on orbits rather than raw samples. The method uses either canonical representatives or learned orbit-separating invariant embeddings to define practical quotient metrics, and combines quotient-space k-center selection with invariant training through an orbit-averaged loss. We further derive a generalization bound that relates excess orbit-averaged risk to quotient-space coverage, label uncertainty, and intra-orbit variability. Experiments on synthetic scale-invariant data and image benchmarks with rotation-induced redundancy show that GRINCO improves orbit coverage and achieves stronger label efficiency than conventional coreset baselines, especially when group-induced redundancy is substantial.
arXiv:2002.12459v2 Announce Type: replace Abstract: In the last few years, much effort has been devoted to developing join algorithms in order to achieve worst-case optimality for join queries over relational databases. Towards this end, the database community has had considerable success in developing succinct algorithms that achieve worst-case optimal runtime for full join queries, i.e the join is over all variables present in the input database. However, not much is known about join evaluation with {\em projections} beyond some simple techniques of pushing down the projection operator in the query execution plan. Such queries have a large number of applications in entity matching, graph analytics and searching over compressed graphs. In this paper, we study how a class of join queries with projections can be evaluated faster using worst-case optimal algorithms together with matrix multiplication. Crucially, our algorithms are parameterized by the output size of the final result, allowing for choice of the best execution strategy. We implement our algorithms as a subroutine and compare the performance with state-of-the-art techniques to show they can be improved upon by as much as 50x. More importantly, our experiments indicate that matrix multiplication is a useful operation that can help speed up join processing owing to highly optimized open source libraries that are also highly parallelizable.
arXiv:2106.14969v4 Announce Type: replace Abstract: In network design problems, such as compact routing, the goal is to route packets between nodes using the (approximated) shortest paths. A desirable property of these routes is a small number of hops, which makes them more reliable, and reduces the transmission costs. Following the overwhelming success of stochastic tree embeddings for algorithmic design, Haeupler, Hershkowitz, and Zuzic (STOC'21) studied hop-constrained Ramsey-type metric embeddings into trees. Specifically, embedding $f:G(V,E)\rightarrow T$ has Ramsey hop-distortion $(t,M,\beta,h)$ (here $t,\beta,h\ge1$ and $M\subseteq V$) if $\forall u,v\in M$, $d_G^{(\beta\cdot h)}(u,v)\le d_T(u,v)\le t\cdot d_G^{(h)}(u,v)$. $t$ is called the distortion, $\beta$ is called the hop-stretch, and $d_G^{(h)}(u,v)$ denotes the minimum weight of a $u-v$ path with at most $h$ hops. Haeupler {\em et al.} constructed embedding where $M$ contains $1-\epsilon$ fraction of the vertices and $\beta=t=O(\frac{\log^2 n}{\epsilon})$. They used their embedding to obtain multiple bicriteria approximation algorithms for hop-constrained network design problems. In this paper, we first improve the Ramsey-type embedding to obtain parameters $t=\beta=\frac{\tilde{O}(\log n)}{\epsilon}$, and generalize it to arbitrary distortion parameter $t$ (in the cost of reducing the size of $M$). This embedding immediately implies polynomial improvements for all the approximation algorithms from Haeupler {\em et al.}. Further, we construct hop-constrained clan embeddings (where each vertex has multiple copies), and use them to construct bicriteria approximation algorithms for the group Steiner tree problem, matching the state of the art of the non constrained version. Finally, we use our embedding results to construct hop constrained distance oracles, distance labeling, and most prominently, the first hop constrained compact routing scheme with provable guarantees.
arXiv:2403.06281v4 Announce Type: replace Abstract: Fuzzing has been widely used for testing embedded-system firmware in a fully rehosted environment without real peripherals, native system supports, or access to the source code and specifications. Some fuzzers emulate the MMIO behavior of missing peripherals based on the firmware binaries to boost code coverage. They emulate each individual MMIO read in the firmware with a fixed model. We find this ineffective when multiple MMIO reads collectively retrieve a data chunk, which greatly impedes the coverage growth. We propose ES-Fuzz to overcome this coverage bottleneck by adaptively modeling MMIO data chunks. ES-Fuzz runs alongside a given fuzzer and starts a new run when the fuzzer's coverage stagnates. In each run, it analyzes a high-coverage test case to infer new MMIO data-chunk models that unlock additional execution paths. We have implemented ES-Fuzz on Fuzzware and evaluated it on 24 popular firmware binaries. ES-Fuzz boosts Fuzzware's coverage by up to 68% in 11 -- and triggers additional bugs in 5 -- of them without degrading the coverage in the remainder. Its models describe a wide range of MMIO data chunks and the firmware's use of each across various contexts.
arXiv:2407.10887v4 Announce Type: replace Abstract: Growing concerns over the theft and misuse of Large Language Models (LLMs) underscore the need for effective fingerprinting to link a model to its original version and detect misuse. We define five essential properties for a successful fingerprint: Transparency, Efficiency, Persistence, Robustness, and Unforgeability. We present a novel fingerprinting framework that provides verifiable proof of ownership while preserving fingerprint integrity. Our approach makes two main contributions. First, a chain and hash technique that cryptographically binds fingerprint prompts to their responses, preventing collisions and enabling irrefutable ownership claims. Second, we address a realistic threat model in which instruction-tuned models' output distribution can be significantly altered through meta-prompts. By incorporating random padding and varied meta-prompt configurations during training, our method maintains robustness even under significant output style changes. Experiments show that our framework securely proves ownership, resists both benign transformations (e.g., fine-tuning) and adversarial fingerprint removal, and extends to fingerprinting LoRA adapters\footnote{We release our code at: https://github.com/microsoft/Chain-Hash.
arXiv:2505.07254v2 Announce Type: replace Abstract: Precise 3D state estimation in multi-object tracking (MOT) is critical for self-driving cars, particularly for objects occluded. Motion modeling in the Kalman filter with a constant motion assumption is widely used in MOT methods, but it neglects the continuous changes in objects' motion caused by traffic in urban environments. Although recent research introduces a multimodel Kalman filter that incorporates multiple motion models, these approaches incur significant computational overhead from the simultaneous processing of multiple models. To this end, this work introduces a motion-dynamics Kalman filter (MD-KF) that overcomes the constant-motion assumption while preserving the singularity of the motion model. MD-KF models the changes in objects' motion over successive measurements as Gaussian distributions, and adaptively adjusts a weighted motion model to account for these variations. MD-KF consistently outperforms constant and multimodel KF across multiple datasets with a significant reduction in computation latency compared to multimodel approaches. The proposed approach demonstrates its superiority in trajectory estimation during occlusion and state estimation stability for stationary objects.
arXiv:2505.19614v2 Announce Type: replace Abstract: Multimodal learning has seen remarkable progress, particularly with large-scale pre-training across various modalities. Most current approaches are built on the assumption of a deterministic one-to-one alignment between modalities. However, this oversimplifies real-world multimodal relationships, where their nature is inherently many-to-many. The many-to-many property, or multiplicity, is not a side-effect of noise or annotation error, but an inevitable outcome of intra-modal variability, representational asymmetry, and task-dependent ambiguity in multimodal tasks. We argue that multiplicity is a fundamental bottleneck that affects all stages of the multimodal learning pipeline: from data construction to model training and evaluation benchmarks. By formalizing its causes and consequences, we demonstrate how ignoring multiplicity leads to training uncertainty, unreliable evaluation, and degraded dataset quality. This position paper calls for new research directions on multimodal learning, including multiplicity-aware learning frameworks and dataset construction and evaluation protocols.
arXiv:2505.20857v2 Announce Type: replace Abstract: Motion retargeting for specific robot from existing motion datasets is one critical step in transferring motion patterns from human behaviors to and across various robots. However, inconsistencies in topological structure, geometrical parameters as well as joint correspondence make it difficult to handle diverse embodiments with a unified retargeting architecture. In this work, we propose a novel unified graph-conditioned diffusion-based motion generation framework for retargeting reference motions across diverse embodiments. The intrinsic characteristics of heterogeneous embodiments are represented with graph structure that effectively captures topological and geometrical features of different robots. Such a graph-based encoding further allows for knowledge exploitation at the joint level with a customized attention mechanisms developed in this work. For lacking ground truth motions of the desired embodiment, we utilize an energy-based guidance formulated as retargeting losses to train the diffusion model. As one of the first cross-embodiment motion retargeting methods in robotics, our experiments validate that the proposed model can retarget motions across heterogeneous embodiments in a unified manner. Moreover, it demonstrates a certain degree of generalization to both diverse skeletal structures and similar motion patterns.
arXiv:2506.10488v3 Announce Type: replace Abstract: In this work, we introduce the Sheet Music Benchmark (SMB), a dataset of six hundred and eighty-five pages specifically designed to benchmark Optical Music Recognition (OMR) research. SMB encompasses a diverse array of musical textures, including monophony, pianoform, quartet, and others, all encoded in Common Western Modern Notation using the Humdrum **kern format. Alongside SMB, we introduce the OMR Normalized Edit Distance (OMR-NED), a new metric tailored explicitly for evaluating OMR performance. OMR-NED builds upon the widely-used Symbol Error Rate (SER), offering a fine-grained and detailed error analysis that covers individual musical elements such as note heads, beams, pitches, accidentals, and other critical notation features. The resulting numeric score provided by OMR-NED facilitates clear comparisons, enabling researchers and end-users alike to identify optimal OMR approaches. Our work thus addresses a long-standing gap in OMR evaluation, and we support our contributions with baseline experiments using standardized SMB dataset splits for training and assessing state-of-the-art methods.