arXiv:2607.06640v1 Announce Type: new
Abstract: A learned world model is usually judged by how faithfully it reconstructs its observations or predicts reward, as though quality were something the model simply has or lacks. But what a task actually needs from a model is narrower: the few predictive coordinates its queries depend on, which we call the closure. We show that how much of that closure a latent comes to represent is set not by the model's capacity or its observations but by the dimensionality of the objective it is trained against, and we measure this directly on a DreamerV3 stack in a controlled environment with known ground-truth closure. An aligned scalar value signal -- the objective at the heart of value equivalence -- installs only a one-dimensional projection of a closure that needs several dimensions: read through a single linear probe, the recoverable structure rises from R^2=0.10 to 0.76 as the scalar is replaced by the full objective. Sweeping the objective's dimensionality from one to four installs exactly that many predictive directions through an auxiliary head, and the same staircase appears -- at attenuated magnitude but the same rank -- through the model's own value head, so the dissociation is dimensional rather than an artifact of head form. Capacity-matched comparisons and in-situ pressure checks rule out the obvious alternatives. The law governs a regime, and we measure its boundary: on a companion closed-loop task whose structure is observable frame by frame, reconstruction installs that structure and the scalar objective suffices -- the objective decides what a latent represents exactly where cheaper training signals cannot already recover it. Value equivalence is thus not all-or-nothing but dimensional: the familiar single-reward objective is its rank-one corner, and a model installs as much of a task's structure as the objective it is asked to predict.
Science Journals
arXiv:2607.06687v1 Announce Type: new
Abstract: AI-mediated communication is increasingly being utilized to help facilitate interactions; however, in privacy sensitive domains, an AI mediator has the additional challenge of considering how to preserve privacy. In these contexts, a mediator may redact or withhold information, raising questions about how users perceive these interventions and whether explanations of system behavior can improve trust. In this work, we investigate how explanations of redaction operations can affect user trust in AI-mediated communication. We devise a scenario where a validated system removes sensitive content from messages and generates explanations of varying detail to communicate its decisions to recipients. We then conduct a user study with 180 participants that studies how user trust and preferences vary for cases with different amounts of redacted content and different levels of explanation detail. Our results show that participants believed our system was more effective at preserving privacy when explanations were provided (p<0.05, Cohen's d ~ 0.3). We also found that contextual factors had an impact; participants relied more on explanations and found them more helpful when the system performed extensive redactions (p<0.05, Cohen's f ~ 0.2). We also found that explanation preferences depended on individual differences as well, and factors such as age and baseline familiarity with AI affected user trust in our system. These findings highlight the importance and challenge of balancing transparency and privacy in AI-mediated communications and suggest that adaptive, context-aware explanations are essential for designing privacy-aware, trustworthy AI systems.
arXiv:2607.07313v1 Announce Type: cross
Abstract: Pairwise comparison (PC) via pairwise reciprocal matrices (PRMs) is central to the Analytic Hierarchy Process (AHP). Although the traditional eigenvector method is widely applied to derive priorities, its theoretical robustness in reflecting true priority vectors remains debated. Building upon a previous iteration of this study, this research develops the revised Least Penalty-Squared Prioritization (LPSP) optimization models, including the revised Least Product of Penalty and Direct Squares (LPPDS) and revised Weighted Squares (LPPWS), to minimize the revised Root Mean Penalty-Squared Variance (RMPSV) and the revised Root Mean Penalty-Weighted Square Variance (RMPSWV). However, solving these non-linear formulations is computationally complex for decision-makers. To overcome these limitations, this study proposes the Parallel Osprey Optimized Least Penalty-Squared Prioritization (POO-LPSP) method. By integrating an improved bio-inspired metaheuristic Parallel Osprey Optimization Algorithm (POOA), this framework efficiently solves complex LPSP models to minimize RMPSV and RMPSWV, thereby enhancing prioritization reliability. The practical utility and computational efficiency of the POO-LPSP method are validated through a numerical application focusing on a Generative AI (GAI) vendor selection problem. To extend, POO-LPSP can serve as a robust alternative to Saaty's Eigen system method for AHP applications.
arXiv:2607.07336v1 Announce Type: cross
Abstract: Large-scale combinatorial optimization remains demanding for classical heuristics, particularly when dense Quadratic Unconstrained Binary Optimization (QUBO) formulations induce large memory footprints, high CPU utilization, and long execution times. While near-term quantum processors cannot yet deliver unconditional quantum advantage, hybrid architectures can provide practical value by reducing the resource burden. This paper presents a resource-efficiency study of Hybrid Quantum Neighborhood Selection (HQNS), a framework that decomposes large dense QUBO instances into bounded-width quantum subproblems via stochastic frontier selection. We evaluate HQNS on the Maximum Diversity Subset Selection Problem (MDSSP), focusing on the trade-off between solution quality retention and resource consumption. Benchmarks up to N=1000 candidates show that HQNS preserves 99.9908% of the mean diversity score of an 11-restart parallel Simulated Annealing baseline, while reducing wall-clock time by 94.91%, peak CPU utilization by 64.68%, and peak memory usage by 88.61%. The QPU execution time remains bounded within a 6-7 second envelope across scales, indicating that the quantum component is decoupled from the global QUBO dimension when the frontier size is fixed. These results suggest that HQNS provides a resource-aware pathway for deploying hybrid quantum optimization in practical large-scale settings, serving as an efficient architecture for incorporating near-term quantum processors into classical optimization pipelines.
arXiv:2607.06782v1 Announce Type: new
Abstract: Global localization from 3D point clouds remains challenging under limited or asymmetric fields of view (FOV), which fail to provide the dense, symmetric coverage that place recognition methods assume. We present G-PROBE, a learning-free global localization framework that removes this assumption. A virtual sensor decomposition runs the same pipeline, by design, on configurations ranging from a narrow-FOV sensor to a panoramic or multi-sensor rig. The front-end enumerates cross-FOV branch ensembles that encode heading hypotheses for heading-invariant place recognition. A score-scale-invariant, tuning-free gamma-SGRT suppresses heading aliasing under partial FOV and provably becomes inert at symmetric 360 degrees. The back-end, CG-GICP, refines a coarse full-cloud GICP with a pass restricted to high-certainty co-observed points selected by a bird's-eye-view certainty map (a by-product of front-end scoring). This certainty coupling links descriptor evaluation to 6-DoF metric pose estimation without an external verification module. Evaluated on five LiDAR datasets and three modalities (mechanical, solid-state, FMCW), G-PROBE attains the highest learning-free multi-session F1 on average and is competitive in panoramic single-session settings. Where hand-crafted and zero-shot supervised baselines collapse under wide-to-narrow cross-sensor pairing, it remains usable end-to-end (up to 55.0% vs. no more than 6.8% success), and under FOV asymmetry (360 to 60 degrees) it retains about 54% Recall@1, about 18x the strongest learning-free baseline.
arXiv:2607.06815v1 Announce Type: new
Abstract: Autonomous negotiation agents are increasingly deployed in high-stakes settings such as insurance and procurement. While cryptographic techniques protect explicitly disclosed constraint values, they fail to address a subtler threat: behavioral privacy leakage, where an adversary infers private constraints from observable negotiation dynamics such as concession trajectories, timing, and convergence patterns. This paper investigates behavioral differential privacy in multi-round negotiation protocols. We design an adaptive stochastic negotiation policy that jointly guarantees $(\varepsilon, \delta)$-differential privacy, almost-sure convergence of the offer sequence (reaching agreement when the counterparty's reservation value permits), and high negotiation utility. Evaluated on 3,000 synthetic bilateral negotiations, our mechanism reduces adversarial inference accuracy by 43-50% while maintaining a negotiation success rate and utility above 90%, demonstrating that strong privacy guarantees can be achieved without significant loss of performance.
arXiv:2607.06946v1 Announce Type: new
Abstract: Magnetic fields generated by a current flowing through a U-shaped coil connecting two copper foils were measured using ultrafast proton radiography. Two $\sim$1.25 kJ, 1-ns laser pulses propagated through laser entrance holes in the front foil, and were focused to the back foil with an intensity of $\sim$3 $\times$ 10$^{16}$ W$/$cm$^{2}$. The intense laser-solid interaction induced a high voltage between the copper foils and generated a large current in the connecting coil. The proton data show $\sim$40-50 Tesla magnetic fields at the center of the coil $\sim$3-4 ns after laser irradiation. The experiments provide significant insight for future target designs that aim to develop a powerful source of external magnetic fields for various applications in high-energy-density science.
arXiv:2607.06964v1 Announce Type: new
Abstract: Bridging the gap between human pilot intent and autonomous flight operation is critical for real-world electric vertical takeoff and landing (eVTOL) aircraft deployment. Flight planning traditionally relies on classic algorithms that struggle to incorporate flexible human preferences. We present FRAMe, an End-to-End Large Language Model (LLM) Flight Planning tool with RAG-based Memory and Multi-modal Coach Agent. Our system integrates a planner LLM with a multi-modal coach agent and retrieval augmented generation (RAG)-based memory to generate flight plans that satisfy mission constraints while aligning with human flight operator preferences. We demonstrate the system in a range of real-world-inspired scenarios of varying difficulty levels. Across four LLMs, the full FRAMe system (RAG and coach) yields the highest validity for every planner (up to 93.8% aggregate, 99% on Easy scenarios for the strongest planner) and shifts preference-relevant metrics in the operator-favored direction where the metric has headroom. FRAMe signifies how advanced LLMs can be deployed for human-centric mission planning, translating natural language instructions into safe, efficient, and flexible flight routes. The code is available at: github.com/amin-tabrizian/FlightPlanningLLMs
arXiv:2607.06568v1 Announce Type: new
Abstract: We study constitutive conditions of hyperelastic potentials for incompressible material behavior in three dimensions. By means of a counterexample, we show that polyconvexity does not imply true-stress-true-strain monotonicity. Thus, polyconvexity alone is not strong enough to guarantee a physically reasonable response for idealized elasticity.
arXiv:2607.06901v1 Announce Type: new
Abstract: Virtual reality (VR) nature immersion is an increasingly popular field of research due to its potential to help people who do not have access to real nature. There are many questions surrounding how virtual forests can be designed to effectively reduce stress and restore attention. Many of these questions relate solely to visual aspects, but more recent literature has started exploring multisensory experiences. In these experiences, senses are treated as additive; however, certain results from the current literature may indicate that there are more complex, cross-sensory interactions occurring. For example, adding sound to visuals can increase stress reduction potential, but certain natural sounds can feel threatening if they are out of place within the virtual nature scene. Overall, cross-sensory interactions in VR nature environments (VNEs) are underexplored and challenge our current understanding of multisensory VNEs, and future explorations of these interactions are essential for designing optimal VNEs for stress reduction.
arXiv:2607.06917v1 Announce Type: new
Abstract: The projected linear system solver (PLSS), by incrementally appending columns to a random or deterministic sketching matrix, provides an attractive finite termination property for consistent linear systems. Nevertheless, a critical computational bottleneck of PLSS is accessing the whole coefficient matrix per iteration, making it prohibitive for extremely large-scale problems or applications with missing data. To alleviate this limitation, we propose a unified randomized PLSS (RPLSS) framework, built upon the tailored randomized row or column selection strategies that require only partial matrix information per iteration, for solving a general linear system, whether it is under- or overdetermined, and whether it is consistent or not. Within this framework, we develop a randomized Gaussian Kaczmarz method and its extended variant as row-action solvers, and randomized coordinate descent variants as column-action solvers. Theoretically, we prove that our methods inherit the finite termination property of PLSS, while achieving an exponential convergence rate, overcoming the sluggish convergence inherent in conventional randomized Kaczmarz and coordinate descent methods. Numerical experiments demonstrate the superiority of our method against state-of-the-art randomized methods, particularly in scenarios with large missing data.
arXiv:2607.07338v1 Announce Type: cross
Abstract: Nonlinear dynamics is ubiquitous in nature, ranging from chemical pattern formation to ocean circulation, yet its simulation on quantum computers is fundamentally limited by the unitary nature of quantum evolution. We propose the quantum Koopman method, a data-driven framework that embeds nonlinear dynamics into a learned linear representation and implements the resulting evolution using shallow quantum circuits. This method learns Koopman observables from trajectory data, projects the lifted dynamics onto a finite-dimensional subspace, and decomposes the corresponding non-unitary propagator into parallel spectral channels. We utilize the Koopman method on a superconducting processor to simulate three distinct nonlinear systems, comprising reaction-diffusion dynamics, fluid motion on a sphere, and satellite-derived observations of Gulf Stream currents, employing up to 32 parallel circuits of 10 qubits. These quantum simulations capture the dominant multiscale patterns and statistical signatures of the underlying dynamics, and reveal a transition from performance limited by hardware noise in weakly nonlinear systems to performance limited by finite-dimensional Koopman representations as nonlinear scale interactions increase. This transition identifies a practical boundary for quantum-amenable nonlinear dynamics, establishing a hardware-validated route for simulating moderately nonlinear dynamics on near-term quantum hardware.
arXiv:2607.07468v1 Announce Type: cross
Abstract: We study the recovery of sparse functions from finite, noisy, and indirect observations in the framework of statistical inverse learning. The unknown is modeled as an element of $\ell^1$, and observations are generated through a possibly nonlinear forward operator $A:\ell^1\to H$, where $H$ is a vector-valued reproducing kernel Hilbert space. We propose an $\ell^1$-regularized empirical risk minimizer and develop a theoretical analysis of its statistical properties.
Under mild assumptions, we establish almost-sure consistency and derive non-asymptotic high-probability convergence rates in both the prediction and $\ell^1$ reconstruction norms. The rates depend on the source smoothness parameter $r$, characterized by a variational source condition, and the effective dimension exponent $b$, describing the polynomial spectral decay of the covariance operator. We further prove matching minimax lower bounds, showing that the obtained convergence rates are optimal.
To relate the theory to practical sparsity models, we consider finitely smoothing operators of the form $A=G\circ S$, where $S$ is a synthesis operator, and show that approximation-space assumptions imply the required variational source conditions. In particular, we prove that membership in the approximation space $k_t$ is equivalent to polynomial decay of the best $n$-term approximation error.
Finally, we verify the assumptions for two representative inverse problems: reaction coefficient identification in elliptic PDEs and sparse computed tomography. For filtered Radon transforms, we derive explicit effective-dimension asymptotics, yielding concrete convergence rates for standard image models and sparsifying systems.
arXiv:2607.07691v1 Announce Type: cross
Abstract: The spectral condition number is a widely adopted measure of worst-case cost for quantum linear system solvers. Yet it can significantly overestimate the actual runtime for a typical problem instance. We present two quantum algorithms that produce the normalized solution $|x\rangle$ of linear system $Ax=| b \rangle$ to accuracy $\epsilon$ with complexity independent of the condition number $\kappa=\lVert A^{-1}\rVert$. We focus on the standard input model where $A$ is accessed through a block encoding and $| b \rangle$ is prepared by a unitary. But we also introduce an affine dilation model that encodes $A$ and $| b \rangle$ jointly, allowing further refinements of the query complexity. Our truncation-based solver makes an optimal number of queries to $| b \rangle$ and $\operatorname{\mathbf{O}}\left(\kappa_{\mathrm{eff}}\operatorname{polylog}\left(\frac{\kappa_{\mathrm{eff}}}{\epsilon}\right)\right)$ queries to $A$. We prove a family of upper bounds on the effective condition number, including $\kappa_{\mathrm{eff}}\leq\frac{\lVert(A^\dagger A)^{-t/2}|x\rangle\rVert^{1/t}}{\epsilon^{1/t}}$ for positive even integer $t$ and $\kappa_{\mathrm{eff}}\leq\frac{\lVert A^{-1\dagger}(A^\dagger A)^{-(t-1)/2}|x\rangle\rVert^{1/t}}{\epsilon^{1/t}}$ for positive odd $t$, overcoming the $\kappa$-barrier. Our filtering-based solver is extremely simple with a favorable runtime prefactor. In particular, the solver has query complexity $6\frac{\lVert A^{-1\dagger}|x\rangle\rVert}{\epsilon}\ln\left(\frac{1}{\epsilon}\right)$ to leading order when the solution norm is known. We then present a similarly simple solution norm estimator with the same asymptotic cost up to logarithmic factors. Our quantum linear system solvers thus substantially improve a recent algorithm of Li, enabling faster quantum linear system solving beyond the condition number.
arXiv:2607.07701v1 Announce Type: cross
Abstract: We specialize Agarwal's multi-level collective spontaneous-emission formalism to the four-level case by formulating it in the fully symmetric \SU(4) representation of $N$ identical atoms. In the irreducible representation $(N,0,0)$, the occupation-number basis forms a tetrahedral weight lattice on which the six embedded $\mathfrak{su}(2)$ transition subalgebras act as ladder operators. From these algebraic factors we obtain a compact Pauli-type population-rate equation and a closed-form expression for the total emitted intensity that apply to any combination of open dipole channels. The formalism is then specialized to the seven dipole-allowed four-level topologies -- tripod, inverted tripod, Y, inverted Y, double-$\Lambda$, closed cascade, and diamond -- and the resulting rate equations are solved numerically for atom numbers up to $N=50$. In every case the emitted intensity develops a delayed cooperative burst whose peak height obeys a power law $I_{\mathrm{peak}}=aN^{p}$ with topology-dependent parameters $(a,p)$; the fitted exponents lie in the range $1.81\lesssim p\lesssim 1.92$, indicating a superlinear. The \SU(4) tetrahedral flow and the seven configuration-dependent transients together provide a unified geometric picture of multi-channel collective dissipation in four-level atomic ensembles.
arXiv:2112.06339v3 Announce Type: replace
Abstract: We develop Boolean-valued domain theory and show how the lambda-calculus can be interpreted in using domain-valued random variables. We focus on the reflexive domain construction rather than the language and its semantics. The notion of equality has to be interpreted in the Boolean algebra and when we say that an equation is valid in the model we mean that its interpretation is the top element of the Boolean algebra.
arXiv:2204.08476v2 Announce Type: replace
Abstract: In recent years, with the increase of social investment in scientific research, the number of research results in various fields has increased significantly. Cross-disciplinary research results have gradually become an emerging frontier research direction. There is a certain dependence between a large number of research results. It is difficult to effectively analyze today's scientific research results when looking at a single research field in isolation. How to effectively use the huge number of scientific papers to help researchers becomes a challenge. This paper introduces the research status at home and abroad in terms of domain information mining and topic evolution law of scientific and technological papers from three aspects: the semantic feature representation learning of scientific and technological papers, the field information mining of scientific and technological papers, and the mining and prediction of research topic evolution rules of scientific and technological papers.
arXiv:2304.03388v2 Announce Type: replace
Abstract: Deep Neural Networks (DNNs) have become ubiquitous for their ability to solve problems across various domains, including computer vision, natural language processing, and speech recognition. However, as their adoption grows, they face a range of security threats, such as model stealing, architecture extraction, and manipulation, which can compromise their integrity, privacy, and functionality. Past works have relied on complex, fine-grained, and time-series analysis to launch DNN model extraction attacks. These approaches require extensive amounts of data, which are often challenging to acquire and analyze effectively. This paper introduces InferNet, an attack method that leverages simple, non-intrusive, and coarse-grained system-level information to identify the underlying DNN architecture of a victim's application. By analyzing GPU kernel calls, memory events, and system-level metrics, InferNet fingerprints the DNN and infers its architecture with very high accuracy. It can predict the architecture family (e.g., Inception vs. BERT), as well as the architecture variant (e.g., InceptionV1 vs. InceptionV3). The evaluation results demonstrate the effectiveness of InferNet across AI/ML frameworks (TensorFlow, PyTorch), different DNN types (vision, LLMs), and hardware platforms (NVIDIA Tesla T4, NVIDIA Quadro RTX 8000). The results show that InferNet achieves 100% model extraction accuracy using only a partial GPU profile under various attack settings.
arXiv:2307.02413v2 Announce Type: replace
Abstract: Network coordination across multiple domains is a complex task that requires seamless communication among network entities. Network operators aim to minimize costs while ensuring the requirements of user requests are met. Such efforts are highly challenging in decentralized environments with diverse network operators, where only partial knowledge of the complete network is available. Intent-driven multi-domain coordination offers various benefits, some inherent to IBN and others stemming from the standardization of the NBI. As standardization is still missing, there has not been a substantial initiative to develop tools that leverage this paradigm. MINDFul.jl is a Julia library that provides a means to accelerate research in this area, at both the architectural and algorithmic levels. It provides a stateful, modular representation of common metro/core IP-optical network equipment as well as the common intent operations. Finally, it introduces a novel modular IBN-over-SDN architecture and is part of a library ecosystem that facilitates event-based simulations with a hackable interface and offers visualization support.
arXiv:2308.13380v3 Announce Type: replace
Abstract: Is it possible to understand the intricacies of a dynamical system not solely from its input/output pattern, but also by observing the behavior of other systems within the same class? This central question drives the study presented in this paper. In response to this query, we introduce a novel paradigm for system identification, addressing two primary tasks: one-step-ahead prediction and multi-step simulation. Unlike conventional methods, we do not directly estimate a model for the specific system. Instead, we learn a meta model that represents a class of dynamical systems. This meta model is trained on a potentially infinite stream of synthetic data, generated by simulators whose settings are randomly extracted from a probability distribution. When provided with a context from a new system-specifically, an input/output sequence-the meta model implicitly discerns its dynamics, enabling predictions of its behavior. The proposed approach harnesses the power of Transformers, renowned for their \emph{in-context learning} capabilities. For one-step prediction, a GPT-like decoder-only architecture is utilized, whereas the simulation problem employs an encoder-decoder structure. Initial experimental results affirmatively answer our foundational question, opening doors to fresh research avenues in system identification.
arXiv:2407.11217v4 Announce Type: replace
Abstract: Clustering problems such as $k$-means and $k$-median are staples of unsupervised learning, and many algorithmic techniques have been developed to tackle their numerous aspects.
In this paper, we focus on the class of greedy approximation algorithm, that attracted less attention than local-search or primal-dual counterparts. In particular, we study the recursive greedy algorithm developed by Mettu and Plaxton [SIAM J. Comp 2003]. We provide a simplification of the algorithm, allowing for faster implementation, in graph metrics or in Euclidean space, where our algorithm matches or improves the state-of-the-art.
arXiv:2411.01683v3 Announce Type: replace
Abstract: Autonomous Vehicle (AV) perception systems require more than simply seeing, via e.g., object detection or scene segmentation. They need a holistic understanding of what is happening within the scene for safe interaction with other road users. Few datasets exist for the purpose of developing and training algorithms to comprehend the actions of other road users. This paper presents ROAD-Waymo, an extensive dataset for the development and benchmarking of techniques for agent, action, location and event detection in road scenes, provided as a layer upon the (US) Waymo Open dataset. Considerably larger and more challenging than any existing dataset (and encompassing multiple cities), it comes with 198k annotated video frames, 54k agent tubes, 3.9M bounding boxes and a total of 12.4M labels. The integrity of the dataset has been confirmed and enhanced via a novel annotation pipeline designed for automatically identifying violations of requirements specifically designed for this dataset. As ROAD-Waymo is compatible with the original (UK) ROAD dataset, it provides the opportunity to tackle domain adaptation between real-world road scenarios in different countries within a novel benchmark: ROAD++.
arXiv:2411.14038v3 Announce Type: replace
Abstract: Given a finite set $S$ of points in the plane and a finite set $\mathcal{D}$ of directions, a geometric spanning tree~$T$ of~$S$ is $\mathcal{D}$-monotone if every path in $T$ is monotone with respect to some direction in $\mathcal{D}$. We study the problem of computing, for a given point set $S$ and a given set $\mathcal{D}$ of directions, a minimum-length $\mathcal{D}$-monotone spanning tree of~$S$. We present a quadratic-time algorithm for two directions. More generally, we show that the problem belongs to the complexity class XP when parameterized by the number of directions. We further study, for a given positive integer $k$ and point set~$S$, the problem of finding a minimum-length $\mathcal{D}$-monotone spanning tree of $S$ over all possible sets~$\mathcal{D}$ of $k$ directions. We prove that this problem, too, is in XP when parameterized by~$k$, and present two algorithms that run in $O(n^2 \log n)$ and $O(n^6)$ time for $k=1$ and $k=2$, respectively, where $n$ is the number of points in~$S$.
Finally, in contrast to the classical Euclidean minimum spanning tree of a set of points, whose vertex degree is bounded by six, we show that for every even integer~$k$, there exists a point set~$S_k$ and a set $\mathcal{D}_k$ of $k$ directions such that any minimum-length $\mathcal{D}_k$-monotone spanning tree of $S_k$ has maximum vertex degree~$2k$.
arXiv:2501.02436v5 Announce Type: replace
Abstract: Advancements in artificial intelligence call for a deeper understanding of the fundamental mechanisms underlying deep learning. In this work, we propose a theoretical framework to analyze learning dynamics through the lens of dynamical systems theory. We redefine the notions of linearity and nonlinearity in neural networks by introducing two fundamental transformation units at the neuron level: order-preserving transformations and non-order-preserving transformations. Different transformation modes lead to distinct collective behaviors in weight vector organization, different modes of information extraction, and the emergence of qualitatively different learning phases. Transitions between these phases may occur during training, accounting for key phenomena such as grokking. To further characterize generalization and structural stability, we introduce the concept of attraction basins in both sample and weight spaces. The distribution of neurons with different transformation modes across layers, along with the structural characteristics of the two types of attraction basins, forms a set of core metrics for analyzing the performance of learning models. Hyperparameters such as depth, width, learning rate, and batch size act as control variables for fine-tuning these metrics. Our framework not only sheds light on the intrinsic advantages of deep learning, but also provides a novel perspective for optimizing network architectures and training strategies.
arXiv:2501.05711v4 Announce Type: replace
Abstract: Vision Language Models (VLMs) have achieved strong performance across a wide range of video understanding tasks. However, their viewpoint-invariant training limits their ability to infer egocentric properties, such as human-object interactions, from exocentric video observations. This limitation is particularly critical for applications such as Activities of Daily Living (ADL) monitoring, where understanding egocentric properties is essential but deploying wearable egocentric cameras is often impractical. We propose Ego2ExoVLM, a framework that enables VLMs to infer egocentric properties directly from exocentric videos by leveraging time-synchronized ego-exo video pairs during training. Our key insight is to treat the egocentric viewpoint as privileged supervision, providing rich interaction signals that are available only during training. Ego2ExoVLM consists of two complementary components: Ego2Exo Sequence Distillation, which transfers egocentric reasoning through a language-level sequence distillation objective, and Ego Adaptive Visual Tokens, which encourage the model to surface interaction-relevant cues within exocentric visual representations. To evaluate this capability, we introduce Ego-in-Exo Perception, a benchmark for assessing the understanding of egocentric properties from exocentric videos. We evaluate Ego2ExoVLM on 10 tasks spanning Ego-in-Exo Perception and existing ADL benchmarks, achieving state-of-the-art performance on the ADL-X benchmark suite and consistently outperforming strong baselines on our proposed benchmark. All code, models, and data will be released at https://github.com/dominickrei/EgoExo4ADL.