arXiv:2605.31289v2 Announce Type: replace
Abstract: Representation learning is a powerful tool for spatio-temporal abstraction within reinforcement learning (RL). Two well established approaches are through the successor representation (SR) and the default representation (DR). The SR encodes states by the future trajectories they induce, capturing information flow decoupled from reward. The DR builds on this by weighting trajectories with reward, integrating credit-assignment structure into the representation. Eigenvectors of both representations have been used to support a range of downstream tasks -- including option discovery, reward shaping, transfer learning, and exploration. We introduce a structurally distinct formulation: the terminal representation (TR). The TR encodes reward-weighted trajectories similarly to the DR, but can be learned as a lower-dimensionality object, and can be used directly for the mentioned applications without eigenvector computations. Eigendecomposition also imposes the assumption of symmetric transition dynamics, which the TR can bypass. In this work we develop the theoretical foundations of the TR: its derivation, convergence of two learning algorithms, its use for zero-shot compositionality, and equivalences between alternative reward formulations. We further show the TR is embedded in the top DR eigenvector, allowing it to capture the same underlying knowledge without eigendecomposition. Additionally, we provide empirical evidence of the TR as a viable alternative to existing representations in subsidiary applications, while requiring less computational overhead to learn, store, and use.
Science Journals
arXiv:2606.01131v2 Announce Type: replace
Abstract: Real-world asset tokenization is often presented as a mechanism for improving the liquidity of traditionally illiquid assets. However, on-chain representation and secondary-market liquidity are distinct outcomes. This paper examines whether tokenized real-world assets exhibit meaningful observed liquidity and identifies the token characteristics associated with higher market activity. Using token-level data from RWA.xyz and supplemental contract-level observations from Etherscan, the study constructs an Ethereum-based monthly panel of non-stablecoin real-world assets across three prominent categories: U.S. Treasury-backed tokens, gold-backed commodity tokens, and private-credit-related tokens. Liquidity is measured using turnover, active addresses, and an active-month indicator. The empirical design combines descriptive statistics, non-parametric group tests, and exploratory panel regressions suited to short and sparse token histories. The results show substantial heterogeneity across asset categories. Gold-backed tokens exhibit broader holder bases and more persistent on-chain activity than many Treasury and private-credit-related products, while outstanding asset value alone does not reliably predict observed liquidity. The paper contributes to the literature by developing a clearer empirical measurement framework for real-world-asset liquidity and showing that tokenization and liquidity should be analyzed as distinct outcomes.
arXiv:2605.12346v2 Announce Type: replace
Abstract: Dyadic Green's function is an important tool of computational photonics, giving deeper insights into light-matter interaction. We present an operator approach to the derivation of the dyadic Green's function of a generic anisotropic planarly-layered medium for both electric and magnetic fields. The resulting Green's function is expressed through the evolution operators (a kind of transfer matrices) of the comprising layers and the surface impedance tensors, the singular term being naturally separated from other terms. The operator approach to the Green's function simplifies both the conceptual understanding of the problem and the subsequent practical applications, some of which are demonstrated here. The proposed approach can be easily generalized to the case of spherical and cylindrical layers, as well as bi-anisotropic layered media. The obtained results can be applied in nanophotonics engineering problems.
arXiv:2606.01138v3 Announce Type: replace
Abstract: Agent-memory frameworks -- mem0, Letta/MemGPT, Cognee, Zep/Graphiti, MemoryOS, MemTensor -- each ship their own SDK, storage layout, and operational vocabulary. There is no shared wire format: every integration is bespoke, every migration rebuilds memory from scratch, and no framework ships a governance surface that lets a human review writes before they enter long-term storage. We present memorywire, a JSON-Schema 2020-12 wire format for five memory operations (remember, recall, forget, merge, expire) over four memory types (semantic, episodic, procedural, emotional), with a MemoryStore interface, a fan-out router, and an optional HITL governance channel. We describe an open-source reference implementation with five backend adapters (sqlite-vec, mem0, Letta, Cognee, pgvector); a microbenchmark on a 100-fact / 50-query labelled corpus (42 with non-empty gold ids + 8 no-match probes) achieving recall@5 = 1.000 on the 42 gold-id queries with ingest p50 = 37.8 ms and recall p50 = 40.6 ms; an adversarial-fusion experiment showing Reciprocal Rank Fusion holds recall@5 = 1.000 across a 1-of-N rank-0 injection sweep (K in {0, 5, ..., 50}) where max fusion collapses to 0.500 with 80% leak at K >= 5; and a 16-scenario cross-adapter conformance suite passing 68 of 80 cells with zero failures. The contribution is not a new algorithm; it is a packaging of established components (RRF, FSMs, STM/LTM consolidation, diff-and-approve workflows) into a venue-neutral protocol with an empirically validated reference, positioned to compose with the Model Context Protocol rather than compete with it. We further show that memorywire's provenance field is the strongest lever for recovering a poisoned store, evaluated with an external benchmark (PurgeBench).
arXiv:2606.04221v2 Announce Type: replace
Abstract: Hearing aids impose strict latency and power constraints that current DNN-based speech enhancement systems struggle to meet on embedded hardware. We characterize this gap by deploying both speech separation and denoising using the lightweight SuDoRM-RF++ architecture on the AMD-Xilinx Kria KV260, evaluated at FP32 and 16-bit fixed-point precision for each task. Across these configurations, first-sample latency tracks with on-chip parameter caching rather than arithmetic throughput, identifying data movement as the primary bottleneck. Precision reduction halves the model memory footprint without compromising objective speech quality. The fixed-point denoising accelerator achieves a first-sample latency of 9.7~ms, meeting the 10~ms clinical threshold, while speech separation reaches 16.0~ms. These measurements establish concrete resource requirements for embedded DNN-based speech enhancement and quantify the remaining gap to hearing aid deployment.
arXiv:2503.22998v2 Announce Type: replace
Abstract: Despite advancements in Graph Neural Networks (GNNs), adaptive attacks continue to challenge their robustness. Certified robustness via randomized smoothing offers provable guarantees but suffers from a severe accuracy-robustness trade-off, limiting its practical use. To bridge this gap, we introduce AuditVotes, the first framework that simultaneously achieves high clean accuracy and strong certified robustness. AuditVotes seamlessly integrates two novel components into the randomized smoothing pipeline: (1) graph rewiring augmentation, which denoises randomized graphs to recover data quality, and (2) conditional smoothing, which filters low-confidence votes to ensure prediction consistency. We establish a novel theoretical result, proving that certified robustness is preserved under arbitrary filtering functions. Designed for inductive learning, our framework generalizes to unseen nodes and applies broadly to other smoothing schemes, including de-randomized smoothing for graphs and Gaussian smoothing for images. Extensive experiments show AuditVotes delivers substantial gains: on Cora-ML under 20-edge attacks, it improves clean accuracy by 437.1% and certified accuracy by 409.3%, while maintaining comparable runtime to vanilla smoothing. As a widely applicable and efficient plug-in, AuditVotes offers higher accuracy and stronger guarantees, enabling the practical and certifiably robust GNNs in security-sensitive domains.
Holistic Fusion: Task- and Setup-Agnostic Robot Localization and State Estimation with Factor Graphs
arXiv:2504.06479v2 Announce Type: replace
Abstract: Seamless operation of mobile robots in challenging environments requires low-latency local motion estimation and accurate global localization. While most sensor-fusion approaches are designed for specific scenarios, this work introduces a flexible open-source solution for task- and setup-agnostic multimodal sensor fusion distinguished by its generality and usability. Holistic Fusion formulates sensor fusion as a combined estimation problem of i) the local and global robot state and ii) a (theoretically unlimited) number of dynamic variables, including automatic alignment of reference frames; this formulation fits countless real-world applications without conceptual modifications, offering a comprehensive solution beyond hard-coded/task-specific approaches. The proposed factor-graph formulation enables direct fusion of an arbitrary number of absolute, local, and landmark measurements expressed with respect to different frames by explicitly including them as states in the optimization and modeling their evolution as random walks. Moreover, local smoothness and consistency receive particular attention to prevent estimation jumps. Holistic Fusion enables low-latency and smooth online state estimation on typical robot hardware while simultaneously providing low-drift global localization at the IMU measurement rate. The efficacy of this released framework [1] is demonstrated in five real-world scenarios on three robotic platforms with distinct task requirements, highlighting the advantages of fusing multiple absolute measurement types [2]. [1] Code: https://github.com/leggedrobotics/holistic_fusion [2] Project: https://leggedrobotics.github.io/holistic_fusion
arXiv:2607.15989v1 Announce Type: new
Abstract: Understanding how cooperation persists despite the advantage of selfish behavior remains a central challenge in evolutionary dynamics. Classical models of public goods dilemmas predict dominance of defectors, yet natural and social systems often sustain cooperation. We study an eco-evolutionary public goods game on complex networks where cooperators and defectors diffuse at different rates. When the isolated system is in a defector-dominated coexistence regime, faster dispersal of defectors than cooperators leads to a symmetry-breaking transition that produces localized clusters of cooperators. In heterogeneous networks, nodes with higher connectivity become significantly more likely to exhibit cooperative dominance. A degree-based mean-field reduction supports this result by showing that network connectivity controls an effective coupling strength proportional to node degree, thereby producing a bifurcation that separates defector-dominated and cooperative states. We also address why not all hubs become cooperative by means of a multistability analysis. These results reveal how asymmetric mobility and heterogeneous connectivity jointly promote cooperation in structured populations.
arXiv:2607.15502v1 Announce Type: cross
Abstract: We determine the minimum order of a finite lattice set that contains a filled axis-parallel cube skeleton about every point of some $N$-point set of centers. For fixed integers $0\leq k<n$, the answer for $N$ centers is $N^{1-(n-k)/(2n^2)}$, up to constants depending on $n$ and $k$. Thornton proved every smaller exponent and gave a construction of this order; the endpoint lower bound was left open when $k\geq1$. Our proof combines a midpoint estimate, a labelled form of Shearer's projection inequality, and a strong induction that balances large and small radii without a dyadic pigeonhole loss. In particular, a lattice set containing a square boundary about each of $N$ centers has at least a constant times $N^{7/8}$ points.
arXiv:2510.25354v3 Announce Type: replace
Abstract: Hypergraphs provide a natural framework for modeling multiway interactions. We analyze a class of variational semi-supervised learning problems posed on random geometric hypergraphs and establish asymptotic consistency in the large-data limit. In particular, we identify scaling regimes that ensure well-posedness--yielding nontrivial label propagation rather than collapse to a constant labeling--and show that discrete minimizers converge, in the continuum, to solutions of a density-weighted p-Laplacian equation. We also propose Higher-Order Hypergraph Learning (HOHL), a multiscale regularization scheme based on powers of Laplacians associated with hypergraph-induced subgraphs. For geometric point clouds, we analyze an efficient multiscale Laplacian surrogate for HOHL and prove convergence to a higher-order Sobolev-type seminorm. Numerical experiments on standard benchmarks support the practical utility of the resulting higher-order regularization.
arXiv:2602.03676v2 Announce Type: replace
Abstract: We develop a statistical framework for wealth allocation in which equilibrium-like statistics follow from unbiased counting of admissible configurations rather than postulated exchange rules. Each agent is described by a value--wealth map $V_i(w)$, whose local resolution fixes the microscopic weight through a Jacobian relation. In a closed system, the microcanonical marginal and a reservoir expansion yield an emergent canonical distribution for the regular sector. Its partition sum gives a general condensation criterion: if this sector has finite wealth capacity, excess wealth concentrates on a small subset of agents. We extend the construction to open systems with variable wealth and agent number and to weak quasistatic driving. The global constraint determines an evolution equation for the common parameter $\lambda(t)$, while simultaneous changes in total wealth and value--wealth geometry produce a unified first-order response. The susceptibility $\chi_W=\sum_i\mathrm{Var}_i(w)$ equals the Fisher information of the joint canonical family, and Legendre duality gives $\mathrm{d}s^2=\chi_W\mathrm{d}\lambda^2=\chi_W^{-1}\mathrm{d}W^2=-\mathcal{S}''(W)\mathrm{d}W^2$. For power-law critical tails, finite capacity requires $p>2$; within this regime, $\chi_W$ diverges for $2<p\leq3$ and remains finite for $p>3$, while the critical boundary lies at finite Fisher--Rao distance. We also derive a qualified Cram'er--Rao duality and an open-system mixed-response relation. Contact-geometric, Airy-scaling, and stochastic-dynamical interpretations are identified only as conjectures or future work. The time-dependent results are quasistatic and do not determine microscopic relaxation times, while the information geometry describes the canonical family.
arXiv:2606.05062v2 Announce Type: replace
Abstract: We investigate the processes of local ordering for the partisan voter model on complex networks. In this model, agents hold a binary opinion and a fixed preference that biases updates toward alignment with their preferred opinion. We first study the dynamics on uncorrelated random networks and derive a pair approximation that resolves the densities of links connecting different classes of agents. The analytical predictions are in excellent agreement with Monte Carlo simulations. In this setting, partisan bias leaves the total stationary density of links connecting nodes in different sates unchanged and at the same value as in the standard voter model, but redistributes it among different categories of links. We then consider preference-dependent networks with homophilic and heterophilic attachment to analyze the competition between the global bias mechanism and the local effect of preference-based connectivity. In this case, structural correlations qualitatively modify the stationary state. We identify different regimes of local ordering in the space of parameters measuring the strength of the preference and the strength of the homophilic attachment. Our work clarifies the distinct roles of dynamical partisan bias and structural assortativity, and provides an analytical framework to study partisan opinion dynamics beyond mean-field theory.
arXiv:2606.07383v4 Announce Type: replace
Abstract: Vision-Language-Action (VLA) models have shown strong potential for robotic manipulation, but real-time deployment on edge hardware remains challenging. In this work, we identify VLM visual and context tokens as a major source of deployment latency: for GEMM-dominated projection operators, computation grows linearly with the number of input tokens when model dimensions are fixed. Motivated by this observation, we propose RhinoVLA, a deployment-oriented VLA model co-designed with the Huixi R1 edge SoC. RhinoVLA adopts a token-efficient Qwen3-VL backbone and a continuous Action Expert, reducing the VLM-side token and computation burden while preserving pretrained multimodal capability. To support cross-robot learning, RhinoVLA further introduces a unified interface that combines View Registry, 72D physical state-action slot space, and robotinstance LoRA, allowing heterogeneous robot observations and action schemas to be aligned under a shared policy. On the deployment side, RhinoVLA is optimized through hardware-aware compilation, mixed-precision execution, and parallel visual encoding. Experiments show that RhinoVLA achieves downstream performance comparable to {\pi}0.5 at a similar parameter scale, while reaching 11.69 Hz end-to-end inference on Huixi R1, meeting the 10 Hz real-time closedloop control target. The project will be open-sourced at https://github.com/HuixiAI/RhinoVLA.
arXiv:2606.07636v2 Announce Type: replace
Abstract: Long-form video editing over heterogeneous footage requires agents to coordinate source selection, multimodal analysis, timeline construction, narration and subtitle alignment, rendering, and revision while exposing intermediate state for inspection and repair. We present Crayotter, an open-source multimodal multi-agent demo system for prompt-driven long-form video editing. Crayotter organizes production around coverage-aware material preparation, artifact-grounded editing research, and tool-grounded timeline execution. Across these stages, retrieval reports, video analyses, editing blueprints, scheduler events, tool calls, intermediate renders, and final exports are treated as first-class artifacts rather than hidden transient state. The workbench supports local assets, agent-assisted retrieval, progress monitoring, artifact preview, failure diagnosis, interrupted-job resumption, and resource-aware asynchronous execution for long-running workflows. In a 23-theme evaluation, Crayotter achieves the highest human overall score (3.40/5) among the compared systems, with its largest margins in theme alignment, narrative coherence, and editing smoothness. These results show that long-horizon video editing agents can be made traceable, inspectable, and practically controllable through observable production artifacts. Code, traces, and examples are publicly available at https://github.com/idwts/Crayotter.
arXiv:2606.09421v3 Announce Type: replace
Abstract: Large language model agents increasingly rely on skills: reusable procedural documents encoding workflows, tool use, implementation patterns, validation checks, and domain rules. Skill rewriting is often treated as prompt compression, but shorter skills can make agents more expensive by removing sparse operational anchors that prevent exploration, debugging, and recovery. We study skill rewriting through this economic lens. Our controlled framework profiles skill structure, rewrites skills using information-preservation strategies, and evaluates the rewrites under fixed task instructions, environments, and verifiers. Experiments on SkillsBench reveal distinct quality--cost trade-offs across strategies: API/code anchoring, workflow guarding, and rule/formula anchoring benefit different task families, with no universally dominant template. In the main held-out evaluation, the learned policy reduces total cost by 7.0% and downstream agent-token cost by 6.0%; in frozen cross-model transfer, the corresponding reductions average 14.7% and 13.7%, while verifier quality is preserved. These results position skill design as cost-aware operational knowledge engineering rather than prompt compression. Resources: https://github.com/1Reminding/Skill_EE.
arXiv:2602.10936v2 Announce Type: replace
Abstract: We define trajectory predictive control (TPC) as a class of indirect data-driven predictive control (DDPC) methods that represent future outputs as linear in past inputs/outputs and future inputs. TPC unifies many DDPC variants with different predictor structures. We introduce a predictor with a state-space representation and show that with it, TPC inherits the mature theory of linear model predictive control. In numerical experiments, the state-space predictor outperforms existing predictors, especially for small training datasets.
arXiv:2602.13836v2 Announce Type: replace
Abstract: Speculative decoding has rapidly emerged as a leading approach for accelerating language model (LM) inference, as it offers substantial speedups while yielding identical outputs. This relies upon a small draft model, tasked with predicting the outputs of the target model. State-of-the-art speculative decoding methods use a draft model comprising a single decoder layer and output embedding matrix, with the latter dominating drafting time for the latest LMs. Recent work has sought to address this output distribution bottleneck by reducing the vocabulary of the draft model. While this can improve throughput, it compromises speculation effectiveness when the target token is out-of-vocabulary. In this paper, we argue for vocabulary speculation as an alternative to a reduced vocabulary. We propose SpecVocab, an efficient and effective method that selects a vocabulary subset per decoding step. Across a variety of tasks, we show that SpecVocab can achieve a higher acceptance length than state-of-the-art speculative decoding method, EAGLE-3. Notably, this yields up to an 8.1% increase in average throughput over EAGLE-3.
arXiv:2606.10931v3 Announce Type: replace
Abstract: Warning: This paper contains several toxic and offensive statements. Modern large language models (LLMs) are typically aligned through large-scale post-training to ensure fair and reliable behavior. In this work, we investigate how easily such guardrails can be broken by Group Relative Policy Optimization (GRPO). We show that one-shot GRPO training on a single biased example is sufficient to induce systematic bias, with stereotype-driven reasoning generalizing across attributes, categories, and benchmarks. We further find that models differ in their susceptibility based on the initial likelihood of producing biased outputs. Our results reveal a critical vulnerability in post-training: alignment can be overridden by a single example.
arXiv:2606.11781v3 Announce Type: replace
Abstract: We present a rotation-free magnetohydrodynamic dynamo driven by laminar thermal convection in a regular tetrahedral cavity. The tetrahedral boundaries organize the convective flow into a robust pattern of helical convection cells without global rotation or turbulence. Direct numerical simulations demonstrate exponential amplification of a weak seed magnetic field followed by nonlinear saturation, with the magnetic energy exceeding the kinetic energy. The velocity field develops $D_4$ dihedral symmetry, while the self-generated magnetic field exhibits a corresponding signed $D_4$ symmetry, including antisymmetry under $\pi$ rotations about the two horizontal axes. Analysis of the velocity and magnetic-field structures reveals a closed induction cycle sustained by geometry-induced helical convection. This system provides a conceptually simple setting for isolating and understanding the fundamental physical processes underlying magnetohydrodynamic dynamo action.
arXiv:2606.13669v3 Announce Type: replace
Abstract: Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration. Existing works often reduce papers to abstracts, surface mentions, and flat \texttt{cites} edges, omitting key entities, claims, evidence, mechanisms, and method lineages essential for scientific reasoning. To this end, we introduce \textbf{Agents-K1}, an end-to-end knowledge orchestration pipeline that converts raw documents into agent-native scientific knowledge graphs. Agents-K1 integrates three components under a unifying theoretical foundation: a multimodal parser whose five-module schema captures entities, multimodal evidence, citations, and typed inter-entity relations across the full paper rather than abstracts alone; a 4B information-extraction backbone trained with GRPO under a rule-based reward; and a graphanything CLI, a tri-source agent interface that unifies web search, multimodal graph retrieval, and cross-document traversal. On top of this, we process 2.46 million scientific papers across six subjects to produce \textbf{Scholar-KG}, of which we release a one-million-paper subset, and the full Scholar-KG is accessible via the SCP link below. The same pipeline can be extended to general-domain corpora and to schema-conformant data synthesis. Extensive experiments demonstrate that Agents-K1 achieves superior performance in scientific information extraction, knowledge graph construction, and multi-hop scientific reasoning.
arXiv:2606.15254v2 Announce Type: replace
Abstract: Biofilms are spatially structured microbial communities whose architecture, chemistry, mechanics, and cellular states evolve over time. Bulk assays and two-dimensional projections remain useful, but cannot alone resolve how these properties vary with depth or change during growth, treatment, dispersal, and regrowth. Imaging provides complementary routes to three-dimensional measurement: fluorescence microscopy supplies molecular, taxonomic, and functional specificity; optical coherence tomography resolves mesoscale architecture and dynamics; quantitative phase imaging and holotomography report refractive index and biomass-related changes; Raman methods provide chemical and metabolic contrast; and Brillouin microscopy probes mechanical response. We compare these modalities using four independent descriptors-contrast provenance, live volumetric capability, perturbation, and demonstrated biofilm use-and connect their signals to quantitative biological readouts. No single modality simultaneously maximizes spatial coverage, resolution, acquisition speed, molecular specificity, and low perturbation. Implementations from any contrast class can serve as a longitudinal backbone when perturbation is empirically controlled at the relevant spatial and temporal scale, while molecularly specific measurements remain indispensable for identifying species, molecules, and functional states. We therefore frame four-dimensional biofilm measurement as a validated measurement architecture that integrates a low-perturbation volumetric backbone with spatially registered, molecularly specific measurements acquired continuously or at predefined validation points. Achieving this integration will require compatible cultivation formats, controlled imaging dose, shared quantitative parameters, and robust cross-modality registration.
arXiv:2606.16073v2 Announce Type: replace
Abstract: Sampling from complex, unnormalized probability densities is a fundamental challenge in Bayesian inference and probabilistic modeling. While Markov chain Monte Carlo (MCMC) methods provide asymptotic guarantees, they often suffer from slow mixing and high computational costs due to fixed or manually tuned trajectory lengths. In this work, we propose a novel framework that treats trajectory termination as a learnable component of the sampling dynamics. By framing MCMC within the theory of non-acyclic generative flow networks (GFlowNets), we train state-dependent neural classifiers to decide when a trajectory has reached a high-density region and should terminate. We theoretically establish the connection between optimal classifiers and the target density via detailed balance conditions and introduce a multilevel training scheme to facilitate exploration in complex geometries. Experimental results across various benchmark densities demonstrate that our approach significantly reduces average trajectory lengths while improving mode coverage and mixing compared to standard MCMC baselines.
arXiv:2606.16127v2 Announce Type: replace
Abstract: The worldwide rise of authoritarianism and the growing role of Large Language Models (LLMs) in users' everyday lives raise the question of whether specific models exhibit or promote authoritarian attitudes. We introduce AuAu, a comprehensive benchmark for assessing the risk of authoritarian tendencies in LLM responses. AuAu combines three evaluation approaches: (i) psychometric questions from 15 human-validated instruments, (ii) vignettes probing intended behavior in concrete situations, and (iii) responses to realistic user prompts. Unlike prior work, AuAu measures not only overall authoritarian alignment but also its established sub-concepts: Authoritarian Aggression, Authoritarian Submission, and Conventionalism. Evaluating 17 models from China, the EU, Russia, and the USA, we find substantial authoritarian response rates on psychometric instruments across all models, though rates drop significantly on more realistic downstream tasks. Moreover, a simple authoritarian system prompt manipulates 15 of 17 models into promoting increased authoritarianism. Our results underscore the need for continued, systematic auditing of LLM-based AI systems to detect and mitigate authoritarian tendencies in their outputs.
arXiv:2606.16624v2 Announce Type: replace
Abstract: Plate and shell structures are widely used in engineering fields. Rapid response prediction for such structures under complex geometries, heterogeneous materials, and varying loads is important for engineering design, but conventional numerical methods usually require repeated modeling and solution when the physical configuration changes. To address this issue, this study proposes a geometry-aware variational physics-informed neural operator (GA-VINO) for Mindlin-Reissner plates. GA-VINO represents the plate geometry using boundary point clouds and incorporates a material encoder, a load encoder, and a scalar-parameter branch to handle spatially random material fields, spatially varying pressure loads, and sample-level uniform parameters. Through multi-branch point cloud encoding and cross-attention, GA-VINO fuses geometric, material, loading, and query point information, and predicts the transverse deflection and rotations at arbitrary query locations. Unlike conventional data-driven neural operators, GA-VINO requires no labeled solution data during training. Instead, it minimizes a variational physics-informed loss constructed from the discretized total potential energy of the Mindlin-Reissner plate. Compared with grid-based neural operators, GA-VINO directly processes irregular point clouds and allows different physical fields to be discretized on different point sets, avoiding forced interpolation onto a common grid. The method is validated on multiple examples involving different geometries, material fields, and load distributions. The results show that GA-VINO achieves promising accuracy in deflection, rotation, gradient-sensitive, and energy-based metrics, completes full-field inference for new samples within milliseconds, and exhibits promising cross-geometry generalization capability.
arXiv:2606.16900v2 Announce Type: replace
Abstract: Physical systems often exhibit heterogeneous mechanisms, where rapidly evolving dynamics coexist with persistent structures. Capturing such multiscale physical behavior remains challenging for existing neural operators, which typically rely on single dominant inductive bias and therefore couple distinct physical responses into a shared representation. We introduce the Unified Green's Function Framework across domains and propose the Factorized Neural Operators (FaNO), which decompose spectral representations into equivariant-inspired dynamic responses and invariant-inspired persistent responses, leading to better interpretability and generalization. Mechanistically, we show that the two operator branches spontaneously specialize into distinct physical roles that remain consistent across scales and domains: the equivariant-inspired branch captures rapidly varying transient dynamics, whereas the invariant-inspired branch extracts coherent persistent structures. This factorized mechanism of FaNO consistently improves prediction accuracy, parameter efficiency and cross-scale generalization across physical systems and domains. In particular, it maintains consistent predictions under long-horizon autoregressive rollout, cross-resolution extrapolation and physical-regime shifts. These findings suggest that scalable physical modeling may benefit from moving beyond single-inductive-bias formulations toward factorized operator representations that better reflect the heterogeneous organization of physical systems, accelerating the reliable deployment of machine learning for scientific computing and discovery.