Forskningsradar

Science Journals

Peer-reviewade publikationer — 54780 artiklar

On the $\mathrm{In_{x}Ga_{1-x}As}$ channel noise in InP HEMTs from 4 K to 300 K
arXiv:2607.04757v1 Announce Type: cross Abstract: The InP high-electron-mobility transistor (HEMT) is indispensable for low-noise amplifiers (LNAs) in radio astronomy and quantum computing. The composition of the $\mathrm{In_{x}Ga_{1-x}As}$ channel in InP HEMT is known to influence the LNA noise performance. However, the various physical mechanisms responsible for noise generation are not fully characterized and understood. Here, we investigate the $\mathrm{In_{x}Ga_{1-x}As}$ channel noise from 4 K to 300 K for 100-nm gate-length InP HEMTs with channel indium content of 53\%, 60\% and 70\%. Channel noise was quantified by extracting the equivalent drain noise temperature $\mathit{T}_{d}$ using both on-wafer and LNA-based measurements, covering 40-300 K and 4-40 K, respectively. The 60\% indium channel InP HEMT exhibited the lowest channel noise across the full temperature range. The $\mathit{T}_{d}$ extracted from on-wafer characterization was found to obey a parabolic temperature dependence which predicted the $\mathit{T}_{d}$ at 4 K for all InP HEMTs in good agreement with LNA-based measurements. By expressing the channel noise as the sum of one thermal and one excess noise term, it was found that the former increased linearly with ambient temperature and dominated at 300 K. The channel noise at 4 K was determined by the excess noise term and exhibited a non-monotonic dependence on the channel indium content in the InP HEMT. The results suggest that the excess noise in the InP HEMT originates not only from temperature-independent shot noise but also from impact ionization and real-space transfer noise.
Non-asymptotic Convergence of Stochastic Gradient Descent in Score-based Generative Models
arXiv:2607.04775v1 Announce Type: cross Abstract: Score-based Generative Models (SGMs) have achieved impressive performance in data generation across a wide range of applications. While the statistical properties of their sampling procedures are increasingly well understood, the optimization dynamics underlying their training remain less explored. SGMs are typically trained by minimizing a weighted denoising scorematching objective, yet optimization guarantees with stochastic gradients remain limited. In this work, we study Stochastic Gradient Descent (SGD) for SGMs, contributing results in two complementary regimes. First, for general score parameterizations, we establish a non-convex convergence rate for SGD on the weighted denoising score-matching objective, with explicit dependence on the schedule-dependent weighting factors. Second, for overparameterized two-layer ReLU networks, we develop a Neural Tangent Kernel analysis tailored to diffusion training with stochastic gradients, yielding score-approximation error bounds along the SGD trajectory. Finally, our analysis quantifies the role of the reweighting factor in the score approximation error, providing theoretical guidance for weighting choices used in practice.
Context-Constrained Transfer Learning for Tabular Foundation Models via Data Distillation
arXiv:2607.04809v1 Announce Type: cross Abstract: Tabular Foundation Models (TFMs) have demonstrated strong empirical performance as black-box inference engines through in-context learning. However, their use in transfer learning is limited by two obstacles: strict context-size constraints and sensitivity to distribution shifts between source and target tasks. Directly pooling heterogeneous source data can therefore lead to negative transfer. To address these challenges, we propose Context-Constrained Transfer Learning via ANchoring and DIstillation (TL-ANDI), a posterior-aware distillation framework for TFMs. TL-ANDI constructs a compact source context by solving a budget-constrained optimal transport problem whose cost jointly measures target covariate coverage and posterior compatibility. The selected anchor samples are then equipped with locally distilled labels and combined with a residual calibration step using target data.
On a Boolean function without bold folding in the spectrum support and implications for greedy approaches to PDT depth
arXiv:2607.04806v1 Announce Type: new Abstract: We study Boolean functions and their Fourier spectrum supports in the context of parity decision trees (PDTs). Recently, H.~Hatami et al.~\cite{HHL+} constructed examples whose Fourier support \(\mathcal S\) satisfies $$ |(\mathcal S+\gamma_1)\cap(\mathcal S+\gamma_2)|=O(|\mathcal S|^{5/6}) $$ for all distinct \(\gamma_1,\gamma_2\), thereby refuting a natural greedy approach based on finding a single large folding direction. We strengthen this folding estimate by constructing an explicit infinite family of Boolean functions such that $$ |(\mathcal S+\gamma_1)\cap(\mathcal S+\gamma_2)|=O(|\mathcal S|^{1/2}) $$ for all distinct \(\gamma_1,\gamma_2\). The construction uses a special affine subspace partition, called an APLPS-partition, obtained from full linear spreads. In contrast with the probabilistic construction of \cite{HHL+}, our construction is explicit and has no background spectral components. We also discuss consequences for greedy approaches to PDT construction. Under the <<lazy>> assumption that the maximum-folding bound is inherited by all restrictions, the usual folding-counting argument cannot yield a PDT upper bound better than \(O(|\mathcal S|^{1/2})\), matching the known general upper bound. However, this inheritance assumption is false in general; hence our result refutes only this <<lazy>> maximum-folding approach, while a complete refutation of adaptive greedy strategies remains open.
Steering Optimisation Trajectories in Diffusion Representation Learning
arXiv:2607.05319v1 Announce Type: new Abstract: We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures. We trace this behaviour to optimisation dynamics; we analyse curves of image reconstruction against latent representation quality, revealing trajectories that organise around two distinct regimes early in training. Models in the reconstruction regime prioritise image fidelity early, whereas those in the disentanglement regime improve reconstruction and disentanglement more gradually. We hypothesise that this behaviour can be influenced by targeting shortcut pathways in the diffusion U-Net and controlling early noise-level exposure, thereby shaping the reconstruction-disentanglement trade-off during training. To steer optimisation toward stronger representations, we introduce SteeringDRL, combining gated residual U-Nets with a simple noise-level exposure curriculum for training. Across disentanglement benchmarks, SteeringDRL improves representation quality and reduces seed sensitivity. Our method further extends to spatial disentanglement in object-centric learning, improving segmentation quality on synthetic and real-world datasets.
Ranking the Impact of Contextual Specialization in Neural Speech Enhancement
arXiv:2607.04826v1 Announce Type: cross Abstract: We systematically investigate neural speech enhancement systems, ranging from very small ($\sim$10\,k parameters) to medium-large ($\sim$2-5\,M parameters), which specialize to acoustic conditions using contextual information such as speaker identity, noise type, speaker gender, spoken language, and SNR. By fine-tuning generalist models on specific data subsets, we find that specializing to a speaker's identity consistently yields the largest gains in estimated speech intelligibility and quality. In contrast, specializing to SNR, noise type, or gender offers only marginal benefits. Crucially, we show that a small model specialized to both a specific speaker and a specific noise type can match or exceed the performance of a generalist model ten times its size. Further, cross-lingual tests reveal that models specialized to a target language outperform multilingual generalists, suggesting that language is a salient feature for specialization. These findings highlight the potential of small, adaptive models for resource-constrained applications like hearing aids, which specialize on-the-fly to contextual information.
Parenclitic hypergraphs and their application in personalized cancer therapy
arXiv:2607.04938v1 Announce Type: cross Abstract: Understanding the differences between individual instances of the same complex system remains a central challenge, particularly in biological contexts. Parenclitic networks constitute a suitable means to detect deviations in correlations with respect to reference populations. Here, we introduce parenclitic hypergraphs, a general framework for identifying anomalies in higher-order correlations across arbitrary interaction orders. After validating the method on synthetic datasets and benchmark ones, we apply it to patient-derived cancer organoids, capturing temporal changes in gene expression between healthy and cancerous tissues as the disease progresses. Our approach not only reproduces known oncogenic signatures, but also reveals a previously unrecognized candidate therapeutic target. Since organoids are generated from individual patients, our method provides, for the first time, a viable protocol for personalized cancer therapy based on higher-order correlation patterns. These findings offer a novel, systems-level strategy for precision oncology grounded in complex systems theory.
Canonical quantization of neurons
arXiv:2607.05000v1 Announce Type: cross Abstract: Canonical quantization provides a systematic procedure for constructing quantum models from classical Hamiltonians. Here, we apply this principle to a fundamental computational primitive of machine learning: the neuron. Specifically, by viewing a neuron as a composition of an energy function and an activation function, we quantize this model by replacing the energy function with a quantum Hamiltonian and applying the activation function to it through matrix functional calculus. This results in an activation observable that can be measured on an input quantum state. We investigate the use of these quantized neurons for function approximation, where the objective is to learn an unknown observable from labeled quantum data. For this purpose, we develop hybrid quantum-classical algorithms for training and evaluation, including procedures for measuring the activation observable and estimating gradients of the squared loss error. Our algorithms for gradient estimation rely on basic primitives like classical random sampling, the Hadamard test, and Hamiltonian simulation, and those for measuring an activation observable rely on quantum algorithms known as the power of one qumode and Schroedingerization. Numerical experiments demonstrate that our quantized neurons exhibit enhanced expressive capabilities relative to corresponding classical neurons on representative learning tasks. Our work establishes canonical quantization as a principled framework for constructing quantum machine learning primitives and provides a foundation for developing neural architectures tailored to quantum data.
Routing Anonymity and Identifiability of Noisy Quantum Hardware
arXiv:2607.05281v1 Announce Type: cross Abstract: Present-day quantum computing is cloud-based, where a user submits a circuit to a service provider's proprietary backend hardware. While providers may wish to hide implementation details, scheduling choices, or even which physical device was used, noisy finite-shot outputs can carry backend-specific fingerprints: information imprinted in the classical output distribution that can reveal the backend identity. So far, such fingerprints have mostly been studied from a benchmarking perspective, with limited attention to privacy considerations for users and providers. This work develops the first formal framework for backend identifiability and its privacy implications. We introduce a backend-identifiability game and use it to formalise routing anonymity as a security notion for quantum cloud services. We show that backend identifiability is a hypothesis-testing problem and prove that, under passive i.i.d. access to a single backend, routing anonymity decays exponentially at the Chernoff rate. We also establish a utility-anonymity trade-off, imposing fundamental limits on how much backend-specific information can be removed from classical outputs without degrading their usefulness. In addition, we observe that, for noisy quantum hardware, identifying fingerprints are inherently an intermediate-depth phenomenon, and establish a depth principle using Pauli-transfer-matrix tools. We complement the theory with experiments on Amazon Braket on AWS, using ion-trap and superconducting quantum processors. We observe 87-90% classification between superconducting backends and 96-100% classification across physical platforms, and find that identifiability can survive natural forms of post-processing. Overall, these results establish routing anonymity as a distinct security requirement for quantum cloud computing, and provide a framework for quantifying and controlling the utility-anonymity trade-off.
Conflict-Based Lazy Search for Fast Multi-Manipulator Planning
arXiv:2607.04124v1 Announce Type: new Abstract: Employing multiple manipulators can boost efficiency and accomplish tasks that a single manipulator cannot do. However, real-time planning for multiple manipulators in a cluttered workspace still poses significant challenges for planning algorithms. This article proposes a new planning algorithm called Conflict-Based Lazy Search (CBLS) for multimanipulator planning. CBLS is built on Conflict-Based Search (CBS), an efficient multiagent pathfinding (MAPF) algorithm that has shown an order of magnitude speedup over previous approaches [1], [2]. CBS addresses MAPF by solving many single-agent pathfinding (SAPF) problems. Thus, its planning time directly depends on the efficiency of the SAPF algorithm adopted. Our CBLS algorithm enhances CBS with precomputation and lazy search. First, a lazily evaluated graph with controlled sparsity is precomputed for a single manipulator. Second, we propose the Lazy Edged-based A* (LEA*) for efficient SAPF. Since edge evaluation is the computational bottleneck of manipulator planning, LEA* uses lazy search and an edge queue to reduce the number of edge evaluations. We show that LEA* is optimally vertex efficient and has improved edge efficiency compared to A*. We apply the proposed CBLS to multi-manipulator planning problems and show its superior performance by comparing it with CBS and a sampling-based algorithm, namely, RRT-Connect.
MRMS: A Multi-Resolution Memory Substrate for Long-Lived AI Agents
arXiv:2607.04617v1 Announce Type: new Abstract: Long-lived AI agents require continuity across interactions, but continuity cannot be obtained by simply extending the prompt window. An agent must preserve useful prior experience, retrieve it selectively, distinguish personal context from external evidence, and revise memory when the underlying situation changes. We propose an architectural memory substrate organized along two orthogonal axes: a representational axis spanning structured records, vector representations, and graph relations; and a temporal axis spanning short-term traces, medium-term abstractions, and long-term semantic commitments. Its key design constraint is synchronized structured-vector-graph memory: structured records govern eligibility, vector representations support recall, and graph relations adjudicate support, contradiction, and supersession before gated context projection. Its central claim is that reliable personalization is a memory design problem: useful memory is structured, selectively exposed, continuously consolidated, and epistemically labeled rather than stored as undifferentiated conversation history. Beyond the framework, we instantiate MRMS as a lightweight prototype implementing structured records, vector retrieval, temporal policies, and graph-based revision. The prototype exercises the core substrate mechanisms through pre-generation memory selection, revision, boundary enforcement, and evidence attribution under controlled long-lived interaction scenarios with explicit evidence requirements.
Semi-Markovian switching in a fluctuating harmonic trap: An age-structured formulation
arXiv:2607.05173v1 Announce Type: cross Abstract: We study a Brownian particle in a harmonic trap whose stiffness switches between two values with arbitrary waiting-time statistics, generating semi-Markovian dynamics. To treat the resulting temporal memory, we formulate the problem in an enlarged age-structured state space, restoring Markovianity and yielding a local Fokker--Planck description. Within this framework, we derive exact steady-state integral equations for the spatial and birth distributions and obtain exact expressions for stationary moments, injected power, and potential energy. In the second part of the paper, we analyze the stochastic-resetting limit, corresponding to a particle alternately released and trapped. By representing the stationary spatial distribution as a superposition of Gaussian states with fluctuating variance, the problem can be reformulated as a switching process in variance space. This yields exact integral equations for the variance distributions and leads to a simplified description amenable to direct analytical treatment.
Data-driven atomistic modelling of hybrid halide perovskite passivation
arXiv:2607.05321v1 Announce Type: cross Abstract: Molecular passivation of surface defects is key to improving the optoelectronic performance of hybrid halide perovskite materials, but the underlying atomistic mechanisms are incompletely understood. While machine-learned interatomic potentials are now widely used to simulate complex molecular and crystalline systems, their application to experimentally-realistic scenarios - such as molecules coordinating to perovskite surfaces - is still far from trivial. Here, we describe a multistep training pipeline, resembling continuous fine-tuning used for large language models, to underpin atomistic modelling and computational experiments in this domain. Our protocol involves two components: (i) a large, curated, and open dataset of diverse metal and hybrid halide perovskite structures ('hyP-26'); and (ii) a small, specialised dataset for an amino-silane molecule passivating the surface, providing highly specific information for fine-tuning. We apply this approach to explore collective behaviour at a mixed-composition halide perovskite surface passivated with a varying coverage of amino-silane molecules, revealing an evolution of interactions with increasing molecular surface coverage.
Necklaces and Lyndon words in colexicographic order
arXiv:2607.05324v1 Announce Type: cross Abstract: We present the first constant-amortized-time algorithms for generating all length-$n$ necklaces and Lyndon words over a $k$-letter alphabet in colexicographic order, for arbitrary $k\geq 2$. Our approach introduces a novel class of words called \emph{quasinecklaces}, which serve as an easily generated superset of necklaces through which all necklaces can be efficiently identified. We derive a formula for the number $Q_k(n)$ of length-$n$ quasinecklaces and show that $Q_k(n)$ is proportional to the number of length-$n$ necklaces, which is the key property needed to achieve constant amortized time. We also apply our results to efficiently generate a well-known de Bruijn sequence and efficiently generate necklaces and Lyndon words subject to a weight constraint.
Fitted Occupancy-Ratio Evaluation without Bellman Completeness
arXiv:2607.05375v1 Announce Type: cross Abstract: Occupancy ratios correct distribution shift in offline reinforcement learning and are central to off-policy evaluation. Existing primal-dual and minimax methods typically estimate these ratios by enforcing occupancy-balance moments over a critic class. We propose fitted occupancy-ratio evaluation (FORE), a fitted fixed-point method that characterizes the discounted occupancy ratio through an adjoint Bellman recursion. At each iteration, FORE solves a single-level density-ratio objective on one-step-transition data, thereby projecting the adjoint Bellman image onto a log-ratio class in Kullback--Leibler (KL) divergence. Unlike analyses of fitted Q-evaluation, which typically require value-function realizability together with Bellman completeness or projected-operator stability, our central approximation condition is just realizability of the discounted occupancy ratio itself. Under this condition, the population KL-projected recursion contracts in relative entropy toward the true ratio by virtue of the adjoint Bellman operator being a KL-contraction. For the empirical recursion, we establish finite-sample regret bounds that yield convergence in KL up to log-ratio approximation error and a statistical error governed by the complexity of the ratio hypothesis class. The fitted ratio supports direct value estimation by reward reweighting, occupancy-weighted fitted Q-evaluation, and doubly robust estimation that combines the fitted ratio with a fitted Q-function. Together, these results identify discounted occupancy-ratio realizability as a sufficient condition for offline policy evaluation without any completeness assumptions.
Querying and Repairing Inconsistent Prioritized Knowledge Bases: Complexity Analysis and Links with Abstract Argumentation
arXiv:2003.05746v5 Announce Type: replace Abstract: In this paper, we explore the issue of inconsistency handling over prioritized knowledge bases (KBs), which consist of an ontology, a set of facts, and a priority relation between conflicting facts. In the database setting, a closely related scenario has been studied and led to the definition of three different notions of optimal repairs (global, Pareto, and completion) of a prioritized inconsistent database. After transferring the notions of globally-, Pareto- and completion-optimal repairs to our setting, we study the data complexity of the core reasoning tasks: query entailment under inconsistency-tolerant semantics based upon optimal repairs, existence of a unique optimal repair, and enumeration of all optimal repairs. Our results provide a nearly complete picture of the data complexity of these tasks for ontologies formulated in common DL-Lite dialects. The second contribution of our work is to clarify the relationship between optimal repairs and different notions of extensions for (set-based) argumentation frameworks. Among our results, we show that Pareto-optimal repairs correspond precisely to stable extensions (and often also to preferred extensions), and we propose a novel semantics for prioritized KBs which is inspired by grounded extensions and enjoys favourable computational properties. Our study also yields some results of independent interest concerning preference-based argumentation frameworks.
Structure-Guided Self-Supervised Matching for One-Shot Medical Landmark Detection
arXiv:2203.01687v3 Announce Type: replace Abstract: Medical landmark detection usually requires accurate expert annotations, which are laborious and difficult to scale across anatomical regions. In this work, we study an extreme annotation-efficient setting where only a single annotated template image is available. We propose SGB-Match, a structure-guided coarse-to-fine self-supervised matching framework for one-shot medical landmark detection. The framework first learns dense anatomical correspondence from unlabeled augmented image pairs, and then transfers the landmark definition from the annotated template to each target image through feature matching. Different from standard contrastive correspondence learning, where negative candidates are penalized by a structure-agnostic rule, we introduce a structure-guided bias into the contrastive objective. The bias is constructed from relative distance and edge-aware anatomical cues, and explicitly reweights the negative gradients: nearby structure-relevant candidates are weakly repelled, while distant or structure-irrelevant negatives are strongly suppressed. As a result, the learned feature space better preserves local anatomical structures around template landmarks and reduces confusing responses from repeated textures. We further adopt a global-to-local design, where a global encoder provides coarse landmark localization and a local encoder refines the prediction in a cropped region. Extensive experiments on four 2D radiological landmark datasets demonstrate that SGB-Match achieves strong one-shot performance across both public and newly collected datasets and consistently benefits from both structure-guided bias and two-stage refinement.
Representation Recycling for Streaming Video Analysis
arXiv:2204.13492v5 Announce Type: replace Abstract: We present StreamDEQ, a method that aims to infer frame-wise representations on videos with minimal per-frame computation. Conventional deep networks perform feature extraction from scratch at each frame in the absence of ad-hoc solutions. We instead aim to build streaming recognition models that can natively exploit temporal smoothness between consecutive video frames. We observe that the recently emerging implicit layer models provide a convenient foundation to construct such models, as they define representations as the fixed points of shallow networks, which need to be estimated using iterative methods. Our main insight is to distribute the inference iterations over the temporal axis by using the most recent representation as a starting point at each frame. This scheme effectively recycles the recent inference computations and greatly reduces the required processing time. Through extensive experimental analysis, we show that StreamDEQ is able to recover near-optimal representations within a few frames and maintain an up-to-date representation throughout the video duration. Our experiments on video semantic segmentation, video object detection, and human pose estimation in videos show that StreamDEQ achieves on-par accuracy with the baseline while providing 2-4x higher throughput.
On homomorphism related parameters of oriented triangle-free planar graphs
arXiv:2208.10188v2 Announce Type: replace Abstract: The first major contribution of this work is proving that the oriented relative clique number of oriented triangle-free planar graphs is $10$, which completely answers and closes an open problem posed by Sopena (Discrete Mathematics 2016) in the most recent survey on oriented colorings. The second major contribution of the paper is to prove that if all oriented triangle-free planar graphs admit a homomorphism to a particular oriented graph $\overrightarrow{T}$, then its underlying graph $T$ must have minimum degree at least $10$. This result implies that, for the family of oriented triangle-free planar graphs, the lower bounds of the parameters oriented chromatic number, pushable chromatic number, $2$-dipath $L(p,1)$-labeling span, and oriented $L(p,1)$-labeling span are at least $11$, $6$, $p+8$, and $2p+8$, respectively, where $p \geq 1$. That is, we are able to obtain improved lower bounds of a number of other parameters restricted to the family of oriented triangle-free planar graphs using our second major contribution.
Restricted Bernoulli Matrix Factorization: Balancing the trade-off between prediction accuracy and coverage in classification based collaborative filtering
arXiv:2210.10619v3 Announce Type: replace Abstract: Reliability measures associated with the prediction of the machine learning models are critical to strengthening user confidence in artificial intelligence. Therefore, those models that provide not only predictions, but also reliability, enjoy greater popularity. In the field of recommender systems, reliability is crucial, since users tend to prefer those recommendations that are sure to interest them, that is, high predictions with high reliabilities. In this paper, we propose Restricted Bernoulli Matrix Factorization (ResBeMF), a new algorithm aimed at enhancing the performance of classification-based collaborative filtering. This model is based on a collection of restricted matrix factorizations that jointly generate, for each user-item pair, a full probability distribution over the possible rating scores. To prove its effectiveness, the proposed model has been compared to other existing solutions in the literature in terms of prediction quality (Mean Absolute Error and accuracy scores), prediction quantity (coverage score) and recommendation quality (Mean Average Precision score). The experimental results demonstrate that the proposed model provides a good balance in terms of the quality measures used compared to other recommendation models.
Interpretable factorization of clinical questionnaires to identify latent factors of psychopathology
arXiv:2312.07762v4 Announce Type: replace Abstract: Psychiatry research seeks to understand the manifestations of psychopathology in behavior, as measured in questionnaire data, by identifying a small number of latent factors that explain them. While factor analysis is the canonical tool for this purpose, the resulting factors may not be interpretable, and may also be subject to confounding variables. Moreover, missing data are common, and explicit imputation is often required. To overcome these limitations, we introduce Interpretability Constrained Questionnaire Factorization (ICQF), a non-negative matrix factorization method with regularization tailored for questionnaire data. Our method aims to promote factor interpretability and solution stability. We provide an optimization procedure with theoretical convergence guarantees, and an automated procedure to determine latent dimensionality accurately. We validate these procedures using realistic synthetic data. We demonstrate the effectiveness of our method in a widely used general-purpose questionnaire, in two independent datasets (the Healthy Brain Network and Adolescent Brain Cognitive Development studies). Specifically, we show that ICQF preserves diagnostic information across a range of disorders, outperforming competing methods for smaller dataset sizes, and improves interpretability, as assessed by our clinical research collaborators and co-authors. This suggests that the regularization in our method matches domain characteristics, in addition to satisfying qualitative desiderata.
AppAgent: Multimodal Agents as Smartphone Users
arXiv:2312.13771v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have led to the creation of intelligent agents capable of performing complex tasks. This paper introduces a novel LLM-based multimodal agent framework designed to operate smartphone applications. Our framework enables the agent to operate smartphone applications through a simplified action space, mimicking human-like interactions such as tapping and swiping. This novel approach bypasses the need for system back-end access, thereby broadening its applicability across diverse apps. Central to our agent's functionality is its innovative learning method. The agent learns to navigate and use new apps either through autonomous exploration or by observing human demonstrations. This process generates a knowledge base that the agent refers to for executing complex tasks across different applications. To demonstrate the practicality of our agent, we conducted extensive testing over 50 tasks in 10 different applications, including social media, email, maps, shopping, and sophisticated image editing tools. The results affirm our agent's proficiency in handling a diverse array of high-level tasks.
Learning to Visually Connect Actions and their Effects
arXiv:2401.10805v4 Announce Type: replace Abstract: We introduce the novel concept of visually Connecting Actions and Their Effects (CATE) in video understanding. CATE can have applications in areas like task planning and learning from demonstration. We identify and explore two different aspects of the concept of CATE: Action Selection (AS) and Effect-Affinity Assessment (EAA), where video understanding models connect actions and effects at semantic and fine-grained levels, respectively. We design various baseline models for AS and EAA. Despite the intuitive nature of the task, we observe that models struggle, and humans outperform them by a large margin. Our experiments show that in solving AS and EAA, models learn intuitive properties like object tracking and pose encoding without explicit supervision. We demonstrate that CATE can be an effective self-supervised task for learning video representations from unlabeled videos. The study aims to showcase the fundamental nature and versatility of CATE, with the hope of inspiring advanced formulations and models.
TERC: A Transfer Entropy Redundancy Criterion for State Variable Selection in Reinforcement Learning
arXiv:2401.11512v3 Announce Type: replace Abstract: Identifying the most suitable variables to represent the state is a fundamental challenge in Reinforcement Learning (RL). These variables must efficiently capture the information necessary for making optimal decisions. In order to address this problem, in this paper, we introduce the Transfer Entropy Redundancy Criterion (TERC), an information-theoretic criterion, which determines if there is entropy transferred from observable state variables to actions during training. We define an algorithm based on TERC that provably excludes variables from the observable state that do not affect the agent's policy during learning. This yields compact state representations that reduce inference time by up to 2.6 times. Our approach is policy-dependent, making it agnostic to the underlying learning algorithm. The efficiency gains we demonstrate arise at retraining and inference time on the reduced state. Our method improves both retraining and inference efficiency. We demonstrate its effectiveness across three distinct algorithm classes, namely tabular Q-learning, Actor-Critic, and Proximal Policy Optimization (PPO), evaluated in a range of environments. Furthermore, to highlight the differences between the proposed methodology and the current state-of-the-art feature selection approaches, we present a series of controlled experiments on synthetic data, before generalizing to real-world decision-making tasks. We also introduce a representation of the problem that compactly captures the transfer of information from observable state variables to actions as Bayesian networks.