Forskningsradar

Science Journals

Peer-reviewade publikationer — 58665 artiklar

Hybrid quantum floating-point method for sharp arithmetic
arXiv:2607.06040v1 Announce Type: cross Abstract: There are several possible ways to encode random variables in a quantum state. The basis encoding of bit strings has paramount importance because it allows to load the values of a random variable through the superposition of corresponding basis states, and to then exploit quantum parallelism in processing algorithms. The basis encoding offers a natural way to represent an unsigned integer random variable, and extends to signed integers, as well as to fixed-point and floating-point variables. Each quantum representation of fractional numbers, however, involves a trade-off between accuracy and depth of manipulation circuits. Here, an efficient hybrid quantum-classical representation of quantum floating points is introduced. It combines a quantum register containing the values, with a classical register storing global information about the variable, namely the range and approximation tolerances. The sum and product operations are defined, in such a way as to ensure they are performed without overflow. By taking advantage of the stored classical information, the precision degradation that occurs due to rounding after repeated data manipulations, can be significantly reduced compared to known strategies. Ad hoc examples show up to around $90\%$ reduction in approximation, compared to previous techniques, after repeated additions. The method finds application in many algorithms of practical relevance and constitutes a significant advance in the design of arithmetic circuits with low depth and high accuracy.
Using Tanner Spectral Reduction to Improve Multi-Layer Optical Lattice Routing for Hypergraph-Product and Bivariate Bicycle qLDPC Codes
arXiv:2607.06177v1 Announce Type: cross Abstract: We characterize the Tanner graph spectrum of hypergraph-product (HGP) / lifted-product (LP) codes and bivariate-bicycle (BB) codes, informing qubit routing for three-dimensional reconfigurable qubit architectures. Syndrome-extraction routing depth on HGP/LP Tanner graphs reduces to a single SVD on the base parity-check matrix, using a spectral ratio $\beta_\text{HGP} = (1 + \beta_\text{base})/2$ where $\beta_\text{base} = \sigma_2(H)/\sigma_1(H)$ for the base parity-check matrix, and a diameter identity $D_T = 2 D_\text{base}$ where $D_\text{base}$ is the base Tanner graph diameter. Fourier spectral reduction reveals that the BB Tanner graph spectrum equals the union, over the $l \times m$ grid of characters of $\mathbb{Z}_l \times \mathbb{Z}_m$, of the singular values of a single $2 \times 2$ symbol matrix built from the two defining polynomials. This reduces spectral analysis from an $O((lm)^3)$ diagonalization of the $4lm$-node Tanner graph to $lm$ independent $2 \times 2$ SVDs. These results compose into a multi-layer three-dimensional AOL routing protocol with one-time setup cost $T_\text{Valiant} = O(\log N)$ atom rearrangements amortizable over a memory experiment of $R$ rounds. For a Tanner graph chromatic index $\chi'$ and $L_\text{layers}$ stacked AOL planes, the per-syndrome-cycle depth is $\lceil \chi'/L_\text{layers} \rceil$ AOL pattern activations with no atom motion, an $8\times$ step-count reduction at $L_\text{layers} \geq \chi' = 8$. Contingent on multi-layer AOL hardware, this yields an estimated $\sim50-300\times$ per-cycle wall-clock advantage over a single-layer AOD baseline (degrading to $\sim5-100\times$ under AOD-crosstalk overhead), reducing to equality in the single-layer limit. This paper therefore presents a route toward practical routing improvement for future quantum hardware incorporating multi-layer reconfigurable qubit architectures.
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios
arXiv:2607.06179v1 Announce Type: cross Abstract: There are some datasets of varying scales for audio classification (AC) applied to different tasks. However, annotated data is limited for most scenarios, such as domestic environments. To address this challenge, we propose an $\textbf{A}$utomatic $\textbf{A}$udio $\textbf{A}$nnotation Pipeline--TriA Pipeline, which can efficiently convert audio from various scenarios into high-quality training data with audio event annotations. A TriA dataset was constructed with the TriA Pipeline, over 2130 hours of audio covering 431 audio classes. Furthermore, we partitioned a prior-knowledge-guided subset (TriA$_{\mathrm{GK}}$) from TriA and conduct comparative experiments on three domestic AC tasks. Comparing the result on manually annotated data only and that on manually annotated data combines TriA$_{\mathrm{GK}}$, TriA$_{\mathrm{GK}}$ could achieve average relative gains of 3.97% in accuracy and 3.35% in Macro-F1, validating the effectiveness of TriA$_{\mathrm{GK}}$ and the TriA Pipeline.
The Cost of Lunar South-Polar Geometry, and Surface Beacons as the Efficient Fix: A Dilution-of-Precision Analysis
arXiv:2607.06212v1 Announce Type: cross Abstract: Lunar PNT architectures, NASA's Lunar Augmented Navigation Service (LANS), ESA's Moonlight, and allied concepts, place a small number of satellites in elliptical lunar frozen orbits (ELFO) to serve the south-polar region prioritized for exploration. We report a result that reframes the design trade: for a user at the lunar south pole, the satellite count needed to reach good geometry is roughly double what is currently planned, because the visible satellites cluster into a small solid angle overhead and dilution of precision is limited by their angular spread rather than their number. In a time-averaged simulation, orbit-only ELFO constellations of the planned size (4 to 6 satellites) give a south-polar median geometric DOP (GDOP) of 16 to 21, far worse than the GDOP of about 6 routine for terrestrial GNSS, and the constellation must grow to about 12 satellites before the median GDOP crosses 6. We then show that a small number of surface ranging beacons, a configuration absent from the lunar PNT literature, reaches the same geometric quality far more cheaply by supplying the near-horizon diversity the overhead cluster lacks: three beacons on elevated terrain around a -80 deg latitude user cut the median GDOP from 16.2 to 1.6, a factor of about 10, moving the user from 15% to 100% of the time below GDOP 6, geometry a purely orbital solution reaches only near a 24-satellite fleet. Because there is no atmospheric refraction, surface-to-surface line of sight is bounded by the geometric horizon, so beacon siting on crater rims and elevated terrain is itself a design variable. Surface-beacon augmentation is the lowest-cost, highest-leverage improvement available to lunar south-polar PNT, deployable on assets already planned for the region. The geometry engine is Validated against an independent DOP computation; the constellation and beacon scenario are Modelled.
From Pixels to Portraits: A Comprehensive Survey of Talking Head Generation Techniques and Applications
arXiv:2308.16041v2 Announce Type: replace Abstract: Talking head generation has progressed rapidly from landmark- and GAN-based facial animation to diffusion models, neural rendering, 3D-aware avatars, and foundation-model-assisted systems. This progress has enabled increasingly realistic audio-, image-, and video-driven talking heads, but it has also made the field difficult to navigate because methods differ substantially in their inputs, assumptions, controllability, temporal stability, computational cost, and risks of misuse. This survey provides a critical review of talking head generation techniques, organizing the literature into four broad families: image-driven, audio-driven, video-driven, and 3D/neural-rendering-based approaches. For each family, we discuss the underlying technical ideas, representative methods, strengths, limitations, datasets, and evaluation practices. Beyond cataloguing prior work, we analyse the persistent gap between commonly reported quantitative metrics and perceptual quality, and compare publicly available models in terms of inference time, memory requirements, and human-rated visual quality. We also examine emerging trends, including diffusion-based generation, 3D-aware representation learning, controllable emotional expression, real-time deployment, and the growing importance of provenance, watermarking, and deepfake detection. Finally, we identify open challenges around robust evaluation, identity preservation, lip synchronisation, temporal consistency, demographic fairness, computational efficiency, and responsible use. This review aims to provide researchers and practitioners with a structured and up-to-date map of the talking head generation landscape, while highlighting the technical and societal questions that should shape future work.
i-EXAM: Instructable and Explainable Attack Connectivity Graph Modeler
arXiv:2607.05888v1 Announce Type: new Abstract: i-EXAM is a planning-powered tool that helps system administrators to create security profiles of complex networks and perform what-if analyses to identify network hardening strategies. It leverages planning compilation that provides soundness and completeness guarantees to identify attack paths, evaluate security metrics, generate diverse hardening strategies, and explain these strategies in natural language using Large Language Models.
Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning
arXiv:2607.05773v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making. We introduce AgenticAI-Supervisor, an API and UI-driven RL Gym environment that decouples environment creation from scalable execution. By moving to verifiable execution outcomes, the platform generates high-fidelity traces and applies multi-dimensional reward shaping. Critically, our framework mitigates reward hacking through rigorous internal state validation and testing. This work provides a first look at our platform's core capabilities through a Customer Support Agent case study demonstrating a consistent closed-loop feedback for model optimization. Future work will focus on advanced features such as Computer Use, Tool Use, automated "stumping", and edge-case generation.
A behavioral principle underlying attacker-defender interactions in soccer
arXiv:2607.05845v1 Announce Type: new Abstract: Soccer is widely popular for its simple rules and complex yet coordinated play that unfolds on the pitch. Nevertheless, the fundamental mechanisms governing such play are not well understood: what shapes player interactions on the pitch? What short-term goals guide players' decisions about their movements over the next few seconds? We address these questions by focusing on one-on-one settings in open play, in which the attacker, in possession of the ball and typically dribbling, faces a defender aiming to stop or delay the attacker's actions over a short period. Here we develop a mathematical model of attacker-defender interactions and analyze 306 professional soccer games. Synthesizing the large-scale dataset with an analysis of the model reveals a simple behavioral principle that may underlie these interactions: the defender seeks to minimize their future relative speed to the attacker, whereas the attacker initiates their movements to preempt the defender's objective. This principle, relative-speed minimization, provides a consistent and unified account of the empirical data. Since our framework depends little on soccer-specific details, this principle may govern diverse pursuit-evasion scenarios as well as other invasion team sports.
Chains and Antichains inside Many-One Degrees and Variants
arXiv:2607.06218v1 Announce Type: cross Abstract: The relations between many-one degrees and one-one degrees have been studied since the beginning of recursion theory; early results from the 1960s include that many-one degrees always have a largest one-one degree and either that one-one degree is the only one-one degree inside the many-one degree or every countable linear order is noneffectively embeddable into the structure of one-one degrees inside the given many-one degree. Furthermore, the greatest recursive many-one degree is a special case, as it allows to embed ascending infinite chains but not descending infinite chains, all other many-one degrees fall into the two cases mentioned above. It remained open whether infinite antichains can always be embedded when the many-one degree is nonrecursive and nonirreducible; Odifreddi stated in a survey 1981 and in his book Classical Recursion Theory in the year 1989 this question explicitly as an open problem. D\"egtev had already in 1976 constructed antichains of one-one degrees inside all nonrecursive and nonirreducible recursively enumerable many-one degrees and Batyrshin generalised the result to all nonrecursive and nonirreducible limit-recursive many-one degrees. In 2026, Cintioli showed that there is a measure $1$ class of sets whose many-one degrees contain infinite antichains of one-one degrees. This class contains all rigid many-one degrees. The present work generalises Batyrshin's result to all nonrecursive and nonirreducible many-one degrees and solves therefore Odifreddi's open problem. The present work also proposes to deepen the study of reducibilities between one-one and many-one in recursion theory in order to get a more complete and detailed picture for the structures inside many-one degrees. In particular it studies finite-one and bounded finite-one reducibilities where the first was introduced by Maslova in the 1970ies.
Does Financial Trading Smooth Non-Convex Markets?
arXiv:2607.06316v1 Announce Type: cross Abstract: In non-convex markets, a competitive equilibrium may fail to exist. This turns out to be an important issue in real-world non-convex auction markets, such as electricity markets, as it complicates pricing and requires the auctioneer to resort to out-of-market discriminatory side payments to sustain an equilibrium. We investigate whether the introduction of convex financial trading induces a smoothing effect, mitigating the issues arising from non-convexities. We develop a two-stage non-convex market model (a forward market followed by a spot market) in which convex financial traders participate in the forward market. Our model predicts that financial trading reduces the magnitude of side payments required to support the cleared allocation. To test the prediction of our model, we examine the introduction of a transaction fee on financial traders in 2020 by PJM, the US's largest electricity market. We show that the substantial decline in financial trading volume caused by this policy coincided with a significant increase in side payments, in line with our theoretical predictions.
Differentially private quantum sensor networks
arXiv:2607.06521v1 Announce Type: cross Abstract: Quantum sensing is a promising technology capable of demonstrating clear advantage over comparable classical techniques for precise measurement. One application of quantum sensing is in function estimation, which can be done using a network of entangled quantum sensors, allowing for measurements with greater optimal sensitivity than unentangled sensing protocols. In cases where quantum sensor networks will be used to measure data that should remain private (e.g., biomedical data), it is imperative that these protocols include a privacy mechanism to hide sensitive information. In this work, we show that entangled sensor networks are vulnerable to certain privacy-violating attacks. To mitigate these attacks, we introduce secure sensing protocols endowed with differential privacy. We reconcile differential privacy with retaining Heisenberg-limited scaling, and introduce several protocols achieving varying balances between the two. We show that our main protocol, an $n$-node network sensing protocol that injects noise directly into the sensing Hamiltonian, exhibits a tradeoff between the desirable $O(1/n^2)$ Heisenberg scaling of the mean-squared error of the function estimate and the level of privacy attainable. Under assumptions on the network (a common source of randomness and a constant fraction of honest parties), we show that this protocol is locally implementable and achieves $(O(1), \delta)$-differential privacy for arbitrarily small $\delta$ while retaining Heisenberg scaling of the mean-squared error. We prove that our protocols are resilient to attacks by broad classes of classical and quantum adversaries, and find advantages in the privacy-utility tradeoff when using quantum techniques.
Probability-turbulence divergence: A tunable allotaxonometric instrument for comparing heavy-tailed categorical distributions
arXiv:2008.13078v4 Announce Type: replace Abstract: Real-world complex systems often comprise many distinct types of elements as well as many more types of networked interactions between elements. When the relative abundances of types can be measured well, we often observe heavy-tailed categorical distributions for type frequencies. For the comparison of type frequency distributions of two systems or a system with itself at different time points in time -- a facet of allotaxonometry -- a great range of probability divergences are available. Here, we introduce and explore `probability-turbulence divergence', a tunable, straightforward, and interpretable instrument for comparing normalizable categorical frequency distributions. We model probability-turbulence divergence (PTD) after rank-turbulence divergence (RTD). While probability-turbulence divergence is more limited in application than rank-turbulence divergence, it is more sensitive to changes in type frequency. We build allotaxonographs to display probability turbulence, incorporating a way to visually accommodate zero probabilities for `exclusive types' which are types that appear in only one system. We explore comparisons of example distributions taken from literature, social media, and ecology. We show how probability-turbulence divergence either explicitly or functionally generalizes many existing kinds of distances and measures, including, as special cases, $L^{(p)}$ norms, the S{\o}rensen-Dice coefficient (the $F_{1}$ statistic), and the Hellinger distance. We discuss similarities with the generalized entropies of R{\'e}nyi and Tsallis, and the diversity indices (or Hill numbers) from ecology. We close with thoughts on open problems concerning the optimization of the tuning of rank- and probability-turbulence divergence.
Deciding Conjugacy of a Rational Relation
arXiv:2307.06777v4 Announce Type: replace Abstract: The study of rational relations is fundamental to the study of formal languages and automata theory. A rational relation is conjugate if each pair of words in the relation is conjugate (or cyclic shifts of each other). The notion of conjugacy has been central in addressing many important algorithmic questions about rational relations. We address the problem of checking whether a rational relation is conjugate and show that it is decidable. Towards our decision procedure, we establish a new result that is of independent interest to word combinatorics. We identify a necessary and sufficient condition for the set of pairs given by $(a_0,b_0) G_1^* (a_1,b_1) \cdots G_k^*(a_k,b_k), k \geq 0$ to be conjugate, where $G_i$ is a (not necessarily rational) conjugate relation and $a_i, b_i$ are arbitrary words. This is similar to, and a nontrivial generalisation of, a characterisation given by Lyndon and Sch\"utzenberger in 1962 for the conjugacy of a pair of words. Furthermore, our condition can be evaluated in polynomial time, yielding a PTIME procedure for deciding the conjugacy of a rational relation given as a sumfree expression. Since any arbitrary rational expression can be expressed as a sum of sumfree expressions (with an exponential blow-up), decidability of conjugacy of rational relations follows.
A General Theory of Liquidity Provisioning for Prediction Markets
arXiv:2311.08725v3 Announce Type: replace Abstract: Liquidity provisioning in automated market makers is the practice of recruiting third-party liquidity providers (LPs) to contribute assets to the market in exchange for fees skimmed off of trades. This paper introduces a general framework for liquidity provisioning in cost function prediction markets. Our most general protocol allows LPs to submit or update an arbitrary cost function that specifies their liquidity over the entire price space. We show that our protocol encapsulates several notions of running market makers in parallel, which we prove to be equivalent. We also recover existing protocols from decentralized finance as special cases. In our protocol, liquidity can be expressed as a matrix-valued function, which we argue is necessary with three or more securities. Due to this inherent multidimensionality, the design of trading fees with three or more securities is nontrivial: we show that natural axioms on the design of these fees are incompatible.
Replication in Visual Diffusion Models: A Survey and Outlook
arXiv:2408.00001v2 Announce Type: replace Abstract: Visual diffusion models have revolutionized the field of creative AI, producing high-quality and diverse content. However, they inevitably memorize training images or videos, subsequently replicating their concepts, content, or styles during inference. This phenomenon raises significant concerns about privacy, security, and copyright within generated outputs. In this survey, we provide the first comprehensive review of replication in visual diffusion models, marking a novel contribution to the field by systematically categorizing the existing studies into unveiling, understanding, and mitigating this phenomenon. Specifically, unveiling mainly refers to the methods used to detect replication instances. Understanding involves analyzing the underlying mechanisms and factors that contribute to this phenomenon. Mitigation focuses on developing strategies to reduce or eliminate replication. Beyond these aspects, we also review papers focusing on its real-world influence. For instance, in the context of healthcare, replication is critically worrying due to privacy concerns related to patient data. Finally, the paper concludes with a discussion of the ongoing challenges, such as the difficulty in detecting and benchmarking replication, and outlines future directions including the development of more robust mitigation techniques. By synthesizing insights from diverse studies, this paper aims to equip researchers and practitioners with a deeper understanding at the intersection between AI technology and social good. We release this project at https://github.com/WangWenhao0716/Awesome-Diffusion-Replication.
MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning
arXiv:2409.06067v3 Announce Type: replace Abstract: Previous studies on federated learning (FL) often encounter performance degradation due to data heterogeneity among different clients. In light of the recent advances in multimodal large language models (MLLMs), such as GPT-4v and LLaVA, which demonstrate their exceptional proficiency in multimodal tasks, such as image captioning and multimodal question answering. We introduce a novel federated learning framework, named Multimodal Large Language Model Assisted Federated Learning (MLLM-LLaVA-FL), which employs powerful MLLMs at the server end to address the heterogeneous and long-tailed challenges. Owing to the advanced cross-modality representation capabilities and the extensive open-vocabulary prior knowledge of MLLMs, our framework is adept at harnessing the extensive, yet previously underexploited, open-source data accessible from websites and powerful server-side computational resources. Hence, the MLLM-LLaVA-FL not only enhances the performance but also avoids increasing the risk of privacy leakage and the computational burden on local devices, distinguishing it from prior methodologies. Our framework has three key stages. Initially, we conduct global visual-text pretraining of the model. This pretraining is facilitated by utilizing the extensive open-source data available online, with the assistance of MLLMs. Subsequently, the pretrained model is distributed among various clients for local training. Finally, once the locally trained models are transmitted back to the server, a global alignment is carried out under the supervision of MLLMs to further enhance the performance. Experimental evaluations on established benchmarks, show that our framework delivers promising performance in the typical scenarios with data heterogeneity and long-tail distribution across different clients in FL.
External quantum efficiency above 100% in photovoltaic cells due to the technique raising efficiencies for all kinds of solar cells
arXiv:2409.20066v3 Announce Type: replace Abstract: A V-shaped module (VSM) photovoltaic technique, which breaks traditional concepts, has been proven by European and American scientists to enable power-conversion-efficiency to increase by 50% for thin-film solar cells. Furthermore, the VSM technique raises power-conversion-efficiency for all kinds of solar cells (J. Environ. Sci. Eng. A 12, 214 (2023). https://doi.org/10.17265/2162-5298/2023.06.002). Mysterious dark energy is thought to result in the surprising VSM effect. The VSM approach opens up new avenues to study photovoltaic physics. In this study, the external quantum efficiency (EQE) of commercial polycrystalline silicon solar cells in the VSM was investigated, which exhibits a surprising phenomenon of EQE above 100%. In theory, non-infrared light incident into a solar cell can cause infrared emission. The VSM could trap the emitted infrared photons and lead to extra photoexcited carriers. The easy-to-reproduce VSM effect is thus explainable. The energy of emitted infrared photons is considered as the so-called dark energy. This study provides new clues and evidences to unravel the mystery of the surprising VSM effect. The VSM technique could also be used to develop photodetectors for infrared and ultraviolet light respectively.
Shaping Collaborations with Algorithms: How Agency and Heterogeneity Criteria Influence Team Formation and Outcomes
arXiv:2410.00346v2 Announce Type: replace Abstract: Across professional, scientific, entrepreneurial, and workplace collaboration platforms, algorithms increasingly shape how individuals find and connect with collaborators. These systems create tensions between user agency and organizational values: Should algorithms organize individuals directly in line with organizational goals, allow individuals to choose freely, or nudge choices toward those goals while preserving agency? This study examines how team formation algorithms that vary in user agency and incorporate organizational values--specifically, promoting teams with different expertise and backgrounds--influence collaborator selection, team composition, team processes, and team outcomes. We conducted a 2 x 2 between-subjects laboratory experiment using a team-formation recommendation system, manipulating user agency (assignment vs. choice) and heterogeneity criteria (included vs. not included). Across four conditions, 332 participants either selected collaborators through the system or were assigned to teams by the system, and then worked as members of 83 teams. Results show that modest differences in algorithm design can systematically reshape team composition and collaboration decisions, often without users fully perceiving the system's influence. While allowing user agency reinforced homophily, nudging by reordering recommendations based on heterogeneity criteria increased the selection of different collaborators and produced teams that performed better than those formed through unconstrained choice. Nevertheless, nudging operated without users' awareness, raising questions about transparency and autonomy. Our findings demonstrate that algorithms embedded in collaboration platforms constitute a distinct mode of algorithmic governance, where resolving tensions between user agency and organizational values raises questions about transparency, access, and control over collaboration.
Few-Medoids: An Embarrassingly Simple Coreset Selection Method for Few-Shot Knowledge Distillation
arXiv:2607.05891v1 Announce Type: new Abstract: Coreset selection aims to identify a small and highly representative subset of a massive dataset for efficient model training. The problem remains challenging even in the few-shot knowledge distillation (KD) setup, where a full-scale pre-trained teacher informs the student network. Typical sample selection strategies often struggle to surpass the random selection baseline. In this paper, we showcase few-medoids, an embarrassingly simple coreset selection strategy that chooses the samples closest to the centroid (average image) of each class. We present extensive KD experiments on four datasets, covering a wide range of image classification problems, and three teacher-student model pairs, comprising both convolutional and transformer networks. Although the proposed method is embarrassingly simple, our empirical results indicate that few-medoids is able to consistently surpass the random selection baseline, as well as the other coreset selection strategies. We therefore consider that few-medoids can be used as a drop-in replacement for commonly-used baselines (e.g. herding or k-center Greedy), in future research on coreset selection. To reproduce the reported results, we publicly release our code at https://github.com/CemilAndreiDilmac/Few-Shot-KD-Coreset.
Faster Exponential-Time Approximate Counting via Bounded Self-Reductions
arXiv:2607.06393v1 Announce Type: new Abstract: We give faster exponential-time randomised approximation algorithms for counting problems where polynomial-time approximation is unavailable and exact exponential-time counting remains expensive. For general \(n\)-vertex graphs, our independent-set counter runs in \(O^{\ast}(1.1869^{n})\) time, improving the previous \(O^{\ast}(1.2041^{n})\) general-graph bound. For \(n\)-variable \#\textsc{2-SAT}, we obtain an \(O^{\ast}(1.2373^{n})\)-time approximation algorithm, narrowly below Wahlstr{\"o}m's currently cited \(O^{\ast}(1.2377^{n})\) variable-parameter exact bound. The new algorithmic point is to take the square root after decomposition. For a single bounded unweighted self-reduction with \(f(x)\) positive leaves and recursion-compatible upper bound \(b(x)\), an enumerate-or-sample estimator gives an \((\varepsilon,\delta)\)-approximation in \[ O^{\ast}\!\left(\sqrt{b(x)}\,\varepsilon^{-2}\log \tfrac1\delta\right) \] time. After preprocessing decomposes an input into many bounded cores, the combined estimator pays \[ O^{\ast}\!\left(\sqrt{\sum_i b_i(x_i)}\,\varepsilon^{-2}\log \tfrac1\delta\right), \] rather than estimating the cores separately at cost \(\sum_i \sqrt{b_i(x_i)}\). The same conversion improves the bases for counting maximal cliques, minimal separators, and perfect matchings in subcubic graphs. Bounded unweighted self-reductions provide the formal language; at the level of counting classes, the resulting unweighted formulation has the same Karp closure as TotP. With explicit recursion-tree access, the framework yields black-box quantum speed-ups.
Trust-free Personalized Decentralized Learning
arXiv:2410.11378v3 Announce Type: replace Abstract: Personalized collaborative learning in federated settings faces a critical trade-off between customization and participant trust. Existing approaches typically rely on centralized coordinators or trusted peer groups, limiting their applicability in open, trust-averse environments. While recent decentralized methods explore anonymous knowledge sharing, they often lack global scalability and robust mechanisms against malicious peers. To bridge this gap, we propose TPFed, a \textit{Trust-free Personalized Decentralized Federated Learning} framework. TPFed replaces central aggregators with a blockchain-based bulletin board, enabling participants to dynamically select global communication partners based on Locality-Sensitive Hashing (LSH) and peer ranking. Crucially, we introduce an ``all-in-one'' knowledge distillation protocol that simultaneously handles knowledge transfer, model quality evaluation, and similarity verification via a public reference dataset. This design ensures secure, globally personalized collaboration without exposing local models or data. Extensive experiments demonstrate that TPFed significantly outperforms traditional federated baselines in both learning accuracy and system robustness against adversarial attacks.
SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation
arXiv:2607.05994v1 Announce Type: new Abstract: Human-Object Interaction (HOI) video generation aims to synthesize realistic videos of humans manipulating diverse objects, serving as a promising avenue for AI-driven live streaming e-commerce. A primary obstacle in this domain lies in the complexity of modeling fine-grained physical dynamics and the intricate spatial-temporal coordination between human hands and objects. Existing approaches to this problem typically rely on dense temporal guidance, e.g., frame-wise hand-object pose sequences, to strictly control the interaction process. However, such dense guidance incurs high annotation costs and affects motion synthesis diversity. To overcome these limitations, we introduce SparseCtrl-HOI, a novel sparse temporal control framework for HOI video generation. It requires only a few keyframes that capture interaction states at designated timestamps. Specifically, we employ a Time-Controlled Rotary Positional Embedding (TiRoPE) mechanism to temporally anchor these keyframes while preserving their spatial integrity. Subsequently, to govern the dynamics across intermediate frames, we propose a Motion Prior Injection Module that leverages Multimodal Large Language Models (MLLMs) to extract high-level motion priors. This empowers the model to hallucinate logically and physically plausible transitions. Furthermore, we build SparseHOI-5K, a high-quality and richly annotated dataset for HOI video generation with sparse temporal control. Comprehensive evaluations confirm that our method substantially reduces annotation overhead while synthesizing superior live-streaming e-commerce videos. Both our code and dataset are publicly available at https://mpi-lab.github.io/SparseCtrl-HOI.
Transferring Natural Language Datasets Between Languages Using Large Language Models for Modern Decision Support and Sci-Tech Analytical Systems
arXiv:2410.14074v2 Announce Type: replace Abstract: The decision-making process to rule R&D relies on information related to current trends in particular research areas. In this work, we investigated how one can use large language models (LLMs) to transfer the dataset and its annotation from one language to another. This is crucial since sharing knowledge between different languages could boost certain underresourced directions in the target language, saving lots of effort in data annotation or quick prototyping. We experiment with English and Russian pairs, translating the DEFT (Definition Extraction from Texts) corpus. This corpus contains three layers of annotation dedicated to term-definition pair mining, which is a rare annotation type for Russian. The presence of such a dataset is beneficial for the natural language processing methods of trend analysis in science since the terms and definitions are the basic blocks of any scientific field. We provide a pipeline for the annotation transfer using LLMs. In the end, we train the BERT-based models on the translated dataset to establish a baseline.
Tuned Reverse Distillation: Enhancing Multimodal Industrial Anomaly Detection with Crossmodal Tuners
arXiv:2412.08949v4 Announce Type: replace Abstract: Knowledge distillation (KD) has been widely studied in unsupervised image Anomaly Detection (AD), but its application to unsupervised multimodal AD remains underexplored. Existing KD-based methods for multimodal AD that use fused multimodal features to obtain teacher representations face challenges. Anomalies that only exist in one modality may not be effectively captured in the fused teacher features, leading to detection failures. Besides, these methods do not fully leverage the rich intra- and inter-modality information that are critical for effective anomaly detection. In this paper, we propose Tuned Reverse Distillation (TRD) based on Multi-branch design to realize Multimodal Industrial AD. By assigning independent branches to each modality, our method enables finer detection of anomalies within each modality. Furthermore, we enhance the interaction between modalities during the distillation process by designing two Crossmodal Tuners including Crossmodal Filter and Amplifier. With the idea of crossmodal mapping, the student network is allowed to better learn normal features while anomalies in all modalities are ensured to be effectively detected. Experimental verifications on multimodal AD datasets demonstrate that our method achieves state-of-the-art performance in multimodal anomaly detection and localization. Code is available at https://github.com/hito2448/TRD.
Leveraging Metamemory Agent for Enhanced Data-Free Code Generation in Large Language Models
arXiv:2501.07892v3 Announce Type: replace Abstract: Large language models (LLMs) have shown strong performance in automated code generation, with few-shot prompting widely used for its simplicity and effectiveness. However, few-shot methods depend on curated or manually crafted reference examples, limiting their applicability in data-free coding scenarios such as real-world data-free coding scenarios and benchmarks without training sets. Existing methods that generate reference examples via recitation or analogy cannot guarantee their authenticity or accuracy. Inspired by human metamemory, we propose a novel metamemory agent to enhance one-time code generation in data-free coding scenarios. The agent guides LLMs to recall relevant prior knowledge, evaluate confidence in recalled information, and selectively exploit reliable content for problem solving. This agent removes the need for external reference examples, improves the authenticity and accuracy of recalled knowledge, and adaptively tailors the recall\&evaluation process to each task. Extensive experiments demonstrate that the proposed metamemory agent significantly improves one-time code generation quality across data-free coding scenarios. The AI contribution is the metamemory agent, which makes self-recalled examples reliable through confidence evaluation and selection; the engineering application is data-free automated code generation, validated on eight public benchmarks.