arXiv:2607.09912v1 Announce Type: cross
Abstract: We report a heterodyne detection scheme for position readout of an optomechanical system, in particular an optically levitated particle, implemented via digital In-phase and Quadrature demodulation on a field-programmable gate array. Compared to the standard homodyne approach, the proposed method offers three key advantages: it remains robust in the presence of strong parasitic back-reflected fields that would otherwise prevent stable phase locking; it produces a signal linearly proportional to the particle displacement, eliminating phase-wrapping distortion; and its calibration factor is intrinsically immune to drifts in the optical power of the local oscillator or scattered field. We experimentally demonstrate and quantify all three advantages through simultaneous homodyne and heterodyne measurements on the same trapped particle. The proposed method can be used in any optomechanical system based on phase readout.
Science Journals
arXiv:2607.09883v1 Announce Type: new
Abstract: Ocular biometric systems face sophisticated presentation attacks, including high-resolution video replays and real-time generative deepfakes, which easily bypass static liveness checks. Current Presentation Attack Detection (PAD) frameworks typically rely on isolated physiological metrics, such as gaze tracking or the Pupillary Light Reflex (PLR), which can be spoofed independently. This paper proposes a Spatio-Luminance Sensor Fusion protocol, which introduces a dual-stream challenge-response framework for ocular liveness verification by uniting these metrics into a simultaneous authentication challenge. By generating a randomized, time-varying visual stimulus that fluctuates in both spatial trajectory and luminance intensity, we construct a mathematically coupled state-space likelihood model, termed the Synchronization Matrix, to evaluate the continuous cross-correlation between the expected biological latencies of smooth pursuit tracking and pupillary constriction. Using Monte Carlo simulation grounded in literature-derived latency distributions, we demonstrate theoretical separability between genuine and simulated attack conditions, and show that a multi-round challenge design improves the detection of generative deepfakes when a non-zero rendering-latency gap exists. This work provides a simulation-supported theoretical framework for next-generation dynamic spoofing defense in ocular and iris biometrics; human-subject validation is identified as necessary future work before deployment claims can be made.
arXiv:2607.11686v1 Announce Type: new
Abstract: Low-cost unmanned ground vehicles are often used in indoor places like warehouses, inspection corridors, and farm rows, where painted floor lines guide the robot. Line following is useful because it only needs one camera and little computing power, but it can fail when the line is blocked or turns sharply and goes out of view. Sensor-rich platforms tolerate this through hardware redundancy (LiDAR, GPS, multiple cameras), but camera-only systems must recover at runtime with no additional infrastructure. This paper presents a lightweight, two-stage recovery approach that restores guideline tracking without LiDAR, GPS, or a GPU. When the line is lost, the robot first turns in place while slowly relaxing its color checks and waiting for confirmation across multiple frames (Stage 1). If the line is still not found, monocular visual odometry moves the robot back to saved breadcrumb positions before it tries again (Stage 2). The system uses a depth-gated HSV line tracker, a YOLOv8n obstacle detector, and a visual odometry breadcrumb mapper, and it runs at 20 Hz on CPU-only hardware. The controller embeds a complete MAPE-K loop within a single 50 ms control tick, with no external adaptation manager required. The approach is evaluated across 119 fault-injected episodes on three Webots simulation courses. The method was successful in 86.6% of cases, with a median recovery time of 3.26 seconds. These results demonstrate that reliable visual recovery is feasible on camera-only UGVs within practical cost and computational limits.
arXiv:2607.10318v1 Announce Type: new
Abstract: This volume of the Electronic Proceedings in Theoretical Computer Science (EPTCS) includes the contributed papers presented at the 21st International Workshop on Logical Frameworks and Meta-Languages: Theory and Practice (LFMTP 2026), in Lisbon, Portugal, on July 24th, 2026, at the Federated Logic Conference (FLoC 2026) as a satellite event of the 11th International Conference on Formal Structures for Computation and Deduction (FSCD 2026). The program committee for this edition of LFMTP was chaired by Olivier Hermant and Sophie Tourret. More information about LFMTP can be found on https://lfmtp.org.
arXiv:2607.11880v1 Announce Type: cross
Abstract: Open questions in collisionless plasma dissipation can be addressed using space-based observations in different astrophysical environments, with implications for both astrophysical and laboratory plasma systems. We study a low-$\beta$, highly imbalanced, sub-Alfv\'enic stream observed by Parker Solar Probe (PSP) to identify and distinguish between signatures of stochastic heating (SH) and resonant heating (RH) by parallel ion cyclotron waves (ICWs). Prior work studying this stream (Bowen et al., 2025) showed that the SH rate, accounting for intermittency, matched the amplitude of the local energy transfer (LET) rate while the RH rate did not. This comparison relied on a number of assumptions regarding the nature of the diffusive process, and the calculation of the LET rate. We introduce a novel technique of inverting the proton guiding center equation to empirically measure velocity-space diffusion coefficients using three-dimensional proton velocity distribution functions (VDFs), from the ion electrostatic analyzer (SPANi) on PSP. Measured diffusion coefficients are used to determine phase-space heating rates, leading to a calculation of a fully kinetic heating rate independent of assumptions made in prior work. We show that scale-dependent analytic expressions for SH via non-coherent fluctuations match the empirical measurements from PSP data, provided that we account for intermittency in the heating calculation. In contrast, the derived heating rates for SH that accounts for the effects of the helicity barrier, and heating rates for RH via $\parallel$-ICWs do not peak in the same region of velocity-space as the empirical measurements, nor reach the required magnitude. Our approach provides novel methodology to uniquely identify and constrain heating processes in collisionless plasmas, and shows evidence of a Fokker-Planck like diffusive process in the near-Sun solar wind.
arXiv:2109.00586v2 Announce Type: replace
Abstract: The light absorption of [001] grown single-crystalline silicon wafers can be enhanced by chemical etching, e.g., with potassium hydroxide, resulting in a pyramid-like surface texture. Alongside advantageous photon harvesting in solar cells, the surface roughness leads to drawbacks when measuring diffusion behaviour of dopants in the heterogeneous structure. In this paper, we employ experimental and simulated scanning transmission electron beam induced current in combination with simulation of boron diffusion in a self-consistent framework to trace the dopant distribution underneath the pyramid-like surface texture. In order to account for surface recombination, an effective model projecting the system along the electron beam propagation direction is used in the EBIC simulation enabling a comparison to entire two-dimensional experimental maps. We find a good agreement between simulated and experimental data and thoroughly discuss how EBIC can be used in future experiments to quantify weak electric fields.
arXiv:2112.06362v5 Announce Type: replace
Abstract: We consider the problem of scheduling in multi-class, parallel-server queuing systems with uncertain rewards from job-server assignments. In this scenario, jobs incur holding costs while awaiting completion, and job-server assignments yield observable stochastic rewards with unknown mean values. The mean rewards for job-server assignments are assumed to follow a bilinear model with respect to features that characterize jobs and servers. Our objective is to minimize regret by maximizing the cumulative reward of job-server assignments over a time horizon, while keeping the total job holding cost bounded to ensure the stability of the queueing system. This problem is motivated by applications requiring resource allocation in network systems.
A central challenge is to control the tradeoff between reward maximization and fair allocation for the stability of the underlying queuing system (i.e., maximizing network throughput). To address this challenge, we propose a scheduling algorithm based on a weighted proportional fair criteria augmented with marginal costs for reward maximization, incorporating a bandit algorithm tailored for bilinear rewards. Our algorithm admits a regret--queue length tradeoff. For any fixed control parameter $V>0$, it ensures a uniform expected queue length and time-average holding-cost bounds. For a target horizon $T$, choosing $V_T=\Theta(\sqrt{IT})$ at initialization yields $\widetilde O((\sqrt I+d^2)\sqrt T+1/\delta)$ regret. Under this regret-optimized tuning, the corresponding expected queue length and time-average holding-cost bounds remain uniform over the execution time and scales as $O(\sqrt{IT}+1/\delta)$ and $O(\sqrt{IT}/\delta)$, respectively.
arXiv:2607.09885v1 Announce Type: new
Abstract: We present Index-1.9B, a series of open small language models developed at Bilibili. The series comprises four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embedding parameters pre-trained on 2.8 trillion predominantly Chinese and English tokens; Index-1.9B-Pure, a control variant trained with an identical recipe but with all instruction-like data strictly filtered from the corpus; Index-1.9B-Chat, aligned from the base model with supervised fine-tuning and direct preference optimization; and Index-1.9B-Character, which augments the chat model with retrieval-augmented generation for few-shot role-playing customization. Pre-training employs a Warmup-Stable-Decay learning-rate schedule in which the concentration of curated data is raised substantially during the decay phase, together with a Norm-Head output layer that stabilizes training under large learning rates. On a suite of standard benchmarks covering examination, reasoning, mathematics, and code, Index-1.9B-Base attains an average score of 64.92, competitive with or exceeding open models of several times its size. We further report controlled studies on model depth, learning-rate magnitude and scheduling, the interaction between learning-rate decay and data quality, and the effect of including instruction data during pre-training, and we document an unexplained surge in benchmark performance midway through the constant-learning-rate phase. All models, together with evaluation code, are released at https://github.com/bilibili/Index-1.9B.
arXiv:2607.09888v1 Announce Type: new
Abstract: Automated product recognition is a cornerstone of modern retail intelligence; however, accurately matching real-world, in-store images against extensive corporate catalogs remains a major scalability bottleneck for large-scale applications. In this work, we address this challenge by reformulating the task as an embedding-based cross-domain retrieval problem rather than a standard closed-set classification task. Specifically, we define the objective as retrieving the most corresponding catalog reference image for a given real-world product query crop from an expansive inventory. To bridge the severe domain gap between pristine studio packshots and noisy in-store queries, we introduce a novel catalog-to-real multi-stage contrastive learning paradigm (Cat2Real). This framework fine-tunes a vision backbone by systematically exploiting both item-level and image-level similarities to drive targeted hard negative mining. Extensive empirical evaluations demonstrate that our paradigm scales seamlessly to unseen products and categories, yielding outstanding zero-shot generalization performance even in the complete absence of real-world training images for novel inventory.
arXiv:2607.09891v1 Announce Type: new
Abstract: Audio deepfake detection models determine whether speech is genuine or artificially generated, but high overall accuracy can mask substantial performance disparities across demographic groups. In this work, we investigate gender bias in audio deepfake detection using the ASVspoof5 dataset. We use ASVspoof5 under a controlled custom split designed to isolate gender-composition effects. We train attack-specific models on nine training sets with different gender compositions, ranging from female-only to male-only. We use a ResNet18 classifier with LogSpectrogram and WavLM-Base+ features, and we evaluated six post-hoc threshold calibration methods. Experimental results show that training data composition strongly predicts bias direction, with the underrepresented gender performing worse at test time. WavLM-Base+ features are shown to produce gender performance gaps 3.0 to 4.3 times larger than LogSpectrogram under identical training conditions, and balanced training is found to reduce LogSpectrogram bias but leave WavLM bias largely intact. Moreover, all six calibration strategies, including Oracle calibration with full test-set label access, leave the Equal Error Rate gap unchanged at 1.317 pp, confirming that threshold adjustment cannot correct underlying score distribution disparities. Overall, these findings suggest that gender fairness in audio deepfake detection must be addressed at training time, as post-hoc methods can only partially mitigate the resulting disparities
arXiv:2204.04887v3 Announce Type: replace
Abstract: Since the era of big data, the Internet has been flooded with all kinds of information. Browsing information through the Internet has become an integral part of people's daily life. Unlike news data and social data on the Internet, cross-media science and technology information data has different characteristics. This data has become an important basis for researchers and scholars to track current hot spots and explore future directions of technology development. As the volume of science and technology information data becomes richer, traditional science and technology information retrieval systems, which support only unimodal data retrieval and use outdated keyword-matching models, can no longer meet the daily retrieval needs of science and technology scholars. Therefore, in view of this research background, it is of profound practical significance to study cross-media science and technology information data retrieval systems based on deep semantic features, in line with domestic and international technology-development trends.
arXiv:2607.09896v1 Announce Type: new
Abstract: Mitigating recoil events and minimizing optically induced heating are central challenges in the precise control and cooling of macroscopic particles. To overcome this, we propose trapping resonant dielectric particles for applications in ultra-high vacuum (UHV) levitodynamics. Contrary to other approaches, where suppressing the parasitic resonant scattering was achieved in a standing wave geometry, here we propose a single beam geometry in a dark trap regime. As a promising material platform, we focus on a class of transition-metal dichalcogenide (TMD) particles with high polarizability, characterized by refractive indices in the range $3.7$-$4.8$ and densities up to $9.3~\mathrm{g\,cm^{-3}}$. Using full Mie theory, we identify a range of TMD particle radii that support stable axial and radial magnetic quadrupole trapping in a bottle-beam configuration. We predict that for WS$_2$ particles with a mass of $0.5 \times 10^{12}\,\mathrm{amu}$, one can expect suppression of the scattering rate relative to the mechanical frequency down to $\Gamma/\Omega \simeq 0.02$. This corresponds to a coherence time extended by approximately three orders of magnitude compared with silica particles of the same mass trapped in conventional bright optical traps at UHV. Combined with significantly reduced internal heating, remaining well below the melting point of the material, dark trapping of resonant TMD macroscopic particles emerges as a promising platform for exploring quantum physics with large masses.
arXiv:2404.03578v3 Announce Type: replace
Abstract: The sim-to-real gap, which represents the disparity between training and testing environments, poses a significant challenge in reinforcement learning (RL). A promising approach to addressing this challenge is distributionally robust RL, often framed as a robust Markov decision process (RMDP). In this framework, the objective is to find a robust policy that achieves good performance under the worst-case scenario among all environments within a pre-specified uncertainty set centered around the training environment. Unlike previous work, which relies on a generative model or a pre-collected offline dataset enjoying good coverage of the deployment environment, we tackle robust RL via interactive data collection, where the learner interacts with the training environment only and refines the policy through trial and error. In this robust RL paradigm, two main challenges emerge: managing distributional robustness while striking a balance between exploration and exploitation during data collection. Initially, we establish that sample-efficient learning without additional assumptions is unattainable owing to the curse of support shift; i.e., the potential disjointedness of the distributional supports between the training and testing environments. To circumvent such a hardness result, we introduce the vanishing minimal value assumption to RMDPs with a total-variation (TV) distance robust set, postulating that the minimal value of the optimal robust value function is zero. We prove that such an assumption effectively eliminates support shift pathologies for RMDPs with a TV distance robust set, and present an algorithm with near-optimal sample complexity. To demonstrate the breadth of our framework, we extend our algorithm and theory to new robust set formulations and robust Markov games. To illustrate the operational relevance, we apply our algorithm to data-driven robust inventory control.
arXiv:2607.10390v1 Announce Type: new
Abstract: Although large language models (LLMs) have shown promising potential in news summarization tasks, their performance on long-document summarization remains challenging as their length often exceeds the input limits. As the agent investment, which provide possibility to improve the inherent capabilities of LLMs. To enhance the effectiveness of long-document summarization based on LLMs, this paper proposes an expert-editor stepwise questioning multi-agent method, in which the expert and the editor guide another agent to refine the summary by posing questions on different aspects of the content and providing targeted clues for revision. We conducted experiments on two representative long-document scientific datasets and evaluated the results through widely recognized automatic metrics. The results demonstrated the effectiveness of our method.
arXiv:2607.09683v1 Announce Type: new
Abstract: This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rotation with Beta Lloyd-Max and QJL, through a statistical validation methodology that separates systematic codec differences from implementation variance. Key findings reveal that while eigenbasis-based methods fail on heavy-tailed data due to covariance instability, they excel in structured regimes, with the effective semantic dimension ($d_{eff}$) adapting to calibration budgets rather than true data rank. (this is an abstract of the abstract thank you )
arXiv:2607.09983v1 Announce Type: new
Abstract: We consider covering and partitioning a simple polygon into pieces which either have unit geodesic radius or unit geodesic diameter, using the $\ell_2$-metric for distances. There is no known method for finding an exact solution to these problems, even when the input size is constant, and the problem is known to be NP-hard in the case of polygons with holes. With this in mind, we instead devote our attention to developing simple approximation algorithms that run in polynomial time. For the radius problem, we present the first known approximation algorithms for both covering and partitioning, achieving a factor of 9. For the diameter problem, we are only able to give a positive result for the partition version of the problem, where we improve upon a complicated 72-approximation from Abrahamsen and Rasmussen [SODA '25], achieving a simple 15-approximation.
arXiv:2607.09902v1 Announce Type: new
Abstract: Agentic coding tools are increasingly used to make autonomous repository-level changes to real-world projects. Prior work has largely evaluated these contributions at the pre-merge stage, through outcomes such as pull request acceptance and review effort. Far less is known about what happens to agentic code post-merge. Yet merge success alone does not reveal whether a contribution will remain stable or require bug fixes and other corrective maintenance downstream. We conduct a longitudinal empirical analysis of agentic and human contributions across 182 repositories, tracking their post-merge fate over time, characterizing the intent of subsequent modifications, and analyzing the defects and vulnerabilities they introduce. While the overall maintenance rates are similar, agentic contributions require significantly higher rates of corrective maintenance and introduce more security weaknesses and dependency vulnerabilities. We also find statistically significant evidence that agentic maintenance burden is associated with repository characteristics. In particular, each 10 percentage-point increase in a project's no-review rate is associated with roughly a 6% increase in agentic maintenance burden on average. As coding agents become pervasive in software development, our findings highlight the need to evaluate and design agentic tools not only to produce mergeable changes, but to produce contributions that remain secure and maintainable.
arXiv:2607.09908v1 Announce Type: new
Abstract: Recommender systems increasingly face a choice among heterogeneous agents -- collaborative filters, sequential models, content-based retrievers, and LLM-based rerankers -- yet no single agent is uniformly best. We study this choice as task-aware agent ranking under cost constraints using RouteRec, a framework that compares request-level hard selection with item-level learned aggregation over four traditional recommender agents and one LLM reranker agent. On MovieLens-1M, the full quality oracle has substantial headroom (HR@10 = 0.584), confirming that useful cross-agent signal exists. Under a leakage-free 5-fold out-of-fold protocol, however, hard selection remains below BM25 (0.223 vs. 0.254), and selective LLM escalation does not improve it. The same protocol yields a different outcome for learned aggregation: its cheap-only variant matches BM25 in HR and has a higher NDCG point estimate (0.123 vs. 0.114), while gated all-agent aggregation reaches HR@10 = 0.295 with 70.2\% LLM calls. The resulting lesson is not that routing is solved, but that request-level selection of one complete agent list is too coarse for this sparse fixed-candidate setting; item-level aggregation is the more promising action space.
arXiv:2607.09921v1 Announce Type: new
Abstract: We present a language-model forecasting system for merger arbitrage, a specialized high-stakes financial setting in which the task is to predict the outcome of announced M\&A deals. Unlike prior work on judgmental forecasting with LLMs, which has focused on broad mixed-topic benchmarks and short context such as news snippets, we study a setting that requires long-context reasoning over hundreds of pages of technical documents. Our system combines expert-guided context engineering with finetuning on hindsight-guided reasoning traces derived from historical deals. Given an announced deal, it outputs a probability distribution over three mutually exclusive outcomes: closing at announced terms, a higher bid, or deal termination. On an out-of-sample set of more than 400 large deals spanning 42 countries, our finetuned system achieves the best performance of any method we evaluate, reducing class-balanced Brier score to 0.151. This is 24\% below calibrated market-implied probabilities, 19\% below XGBoost, and 25-42\% below frontier language models. These results, together with ablation studies, show that LLM-based forecasting can succeed in specialized, long-context financial workflows, with hindsight-based supervision and expert-designed context playing a critical role.
arXiv:2607.09970v1 Announce Type: new
Abstract: Voice phishing (vishing) attacks have traditionally been limited by the need for human operators. The rapid emergence of high-quality AI voice synthesis and large language models (LLMs) reduces this bottleneck and enables scalable, automated scams. In this paper, we conduct a large-scale survey experiment (N=4100) and qualitative interviews (N=12) to assess U.S. adults' susceptibility to AI-powered voice phishing attacks. Participants were exposed to audio recordings or transcripts of scam scenarios generated using leading voice models such as Llama Full Duplex (Llama FD), Sesame, Gemini, OAI AVM, Play.AI, and ElevenLabs and the corresponding human baselines. The results show high compliance rates. Up to 36% of participants would or might comply with phishing requests in the "relative-in-distress" category. Overall compliance rate across all five scam categories was 16.5%, a striking figure given the low cost and high scalability of AI-automated voice phishing. Caller persuasiveness was the strongest predictor of compliance and certain models (most notably Sesame) achieved ratings comparable to human voices, or sometimes even slightly surpassing them. Our economic analysis suggests that while human-operated vishing is unprofitable at US wages, AI-powered vishing appears to be economically viable for several models. The primary risk of present-day AI-enabled vishing thus lies in the economics of automation rather than novel or "superhuman" persuasive techniques, though these cannot be ruled out for future systems. This raises significant concerns for the design of AI systems, consumer protection, and model release policies.
arXiv:2607.09978v1 Announce Type: new
Abstract: Here, we present a platform built on our inverted Graph Transformer Network, IMPRESSION-G2, which can accurately and rapidly reconstruct molecular bonding directly from experimental nuclear magnetic resonance (NMR) spectroscopic information. It comprises three interconnected stages: a one-shot model that predicts bond connectivity between atoms; a structure-correction stage that corrects the predicted structures by removing uncertain bonds and iteratively reassigning them; noise-augmented multi-shot prediction, generating an ensemble of candidate structures, which are ranked to identify the best-fit structure. By integrating a range of $^{1}$H and $^{13}$C NMR data, including two-dimensional (2D) experiments such as COSY, HSQC, and HMBC, the inverse-IMPRESSION platform correctly identifies the structures of 77.8% of molecules with up to 30 heavy atoms (H, C, N, O and F) using simulated NMR data, and 10 of 19 (53%) molecules using experimental NMR data. The experimental structures solved have molecular weights of up to 480 Da and are representative of the complex structures in synthetic and natural products that routinely challenge chemists. The inverse-IMPRESSION framework thus provides the first effective approach for automated molecular structure elucidation using graph-based machine learning on experimental data.
arXiv:2607.10100v1 Announce Type: new
Abstract: This paper aims to analyze a numerical scheme for semilinear subdiffusion problems with singular initial data beyond the $L^\infty$ framework. The main difficulty lies in the stronger singular behavior of the nonlinear term compared with previous analyses. Since the singular initial datum is too rough to guarantee a uniform $L^\infty$ bound for the solution, the usual Lipschitz framework in the base space is no longer sufficient. The analysis must instead be carried out in weaker fractional Sobolev-type spaces, where nonlinear composition is more delicate and the term $f(u(t))$ may exhibit an amplified singularity relative to that of $u(t)$. To overcome this difficulty, we exploit the smoothing properties of the subdiffusion solution operators and formulate suitable nonlinear assumptions in fractional operator spaces. These smoothing estimates allow part of the singularity to be transferred from the nonlinear term to the solution operators, where it can be controlled. Under these assumptions, we establish well-posedness and regularity results for the mild solution and derive a pointwise-in-time error estimate for the exponential convolution quadrature method. Numerical experiments confirm the predicted convergence rates.
arXiv:2607.11501v1 Announce Type: new
Abstract: Accurately modeling and understanding player experience is crucial for designing engaging puzzle games. To achieve this, a common approach involves collecting diverse user data to train predictive playtesting models that mimic player behavior. However, existing data-driven methods often lack the ability to capture the full range of player strategies and require extensive feature engineering and network architecture modeling. This limitation becomes particularly evident when new game mechanics or features are introduced, which necessitate continual adjustments to the models. To addrss these challenges, we propose a more generalized representation that reduces - or even eliminates - the need for ongoing feature-engineering maintenance. Specifically, we investigate two general-purpose network architectures: (a) a transformer-based model (BERT) and (b) a graph attention model (GAT), both of which are designed to effectively capture the relational structure of Candy Crush Saga (CCS) game boards. Our experiments compare these approaches to Convolutional Neural Networks (CNN) baselines, revealing better performance on challenging board configurations and underscoring the benefits of our generalizable representation.
arXiv:2607.10001v1 Announce Type: new
Abstract: We present the design of a language extension that allows the use of ellipses ("...") in patterns and expressions, which facilitates function definitions on lists that are more succinct and direct than the standard recursive ones. The semantics of ellipsis notation is defined via a program translation that is based on computing least general generalizations of the expressions on the boundary of ellipses. We show that ellipsis notation applies to a wide spectrum of function definitions. Ellipsis notation can also support the teaching of functional programming, and its inherently iterative nature may make it especially helpful to those with an imperative programming background. Like list comprehensions, ellipsis notation provides an attractive tool for functional languages that adds to their syntactic variety and enriches their expressiveness.
arXiv:2410.19553v2 Announce Type: replace
Abstract: This paper explores the impact of occlusions in video action detection. We facilitate this study by introducing five new benchmark datasets namely O-UCF and O-JHMDB consisting of synthetically controlled static/dynamic occlusions, OVIS-UCF and OVIS-JHMDB consisting of occlusions with realistic motions and Real-OUCF for occlusions in realistic-world scenarios. We formally confirm an intuitive expectation: existing models suffer a lot as occlusion severity is increased and exhibit different behaviours when occluders are static vs when they are moving. We discover several intriguing phenomenon emerging in neural nets: 1) transformers can naturally outperform CNN models which might have even used occlusion as a form of data augmentation during training 2) incorporating symbolic-components like capsules to such backbones allows them to bind to occluders never even seen during training and 3) Islands of agreement can emerge in realistic images/videos without instance-level supervision, distillation or contrastive-based objectives2(eg. video-textual training). Such emergent properties allow us to derive simple yet effective training recipes which lead to robust occlusion models inductively satisfying the first two stages of the binding mechanism (grouping/segregation). Models leveraging these recipes outperform existing video action-detectors under occlusion by 32.3% on O-UCF, 32.7% on O-JHMDB & 2.6% on Real-OUCF in terms of the vMAP metric. The code for this work has been released at https://github.com/rajatmodi62/OccludedActionBenchmark.