arXiv:2606.14224v1 Announce Type: new Abstract: The advent of photonic integrated circuits (PICs) will allow the replacement of the large aperture of an optical telescope by a dense array of small apertures combined interferometically. The light coming from aperture pairs can be combined by a PIC in order to extract interferogram characteristics known as complex visibilities, from which the observed object can then be reconstructed. In such a compact interferometric imager, the optical components dedicated to image formation in a regular telescope are no longer necessary. In particular, such a concept is relevant for space missions where weight and size are critical. In this communication, we study such an instrument concept, focusing on signal-to-noise considerations. We recall the design basis for the field and the spatial resolution, and we show that the spectral resolution must be no less than the field to resolution ratio. Then, we analyze the signal-to-noise ratio of this concept, assuming that each spatial frequency is recorded only once, and compare the signal-to-noise ratio with that of a monolithic telescope. We perform the comparison in Fourier space for an identical number of recorded photons. We show that the noise propagation of the interferometric imager is identical to that of a monolithic telescope that would have a flat Modulation Transfer Function with a level roughly given by the ratio of the small apertures' diameter to the maximum baseline. We conclude that the noise propagation in low and medium spatial frequencies is unfavorable for the interferometric imager.
Science Journals
arXiv:2606.14225v1 Announce Type: new Abstract: The paper studies the existence of \emph{polynomial} measures of dependence between two random variables: polynomial functions of the joint distribution that (i) vanish on independence and (ii) cannot increase under post-processing of either variable (the data processing inequality, DPI). Mutual information satisfies both properties but is transcendental in the joint distribution, making it impossible to estimate without bias from finitely many samples. A polynomial alternative would admit an exact finite-sample unbiased estimator, with the polynomial degree controlling the required sample size. The main result is negative: in the asymmetric setting where the variables have different alphabet sizes (\(|X| > |Y|\)), no nonzero polynomial can simultaneously satisfy DPI on the larger-alphabet side and vanish on independence. In the symmetric case \(|X| = |Y| = n\), we establish a structural result: every such polynomial is divisible by \((\det U)^2\), where \(U\) denotes the \(n \times n\) joint distribution matrix. Consequently, any nontrivial candidate must have degree at least \(2n\). The determinant-based mutual information attains this lower bound. These algebraic results have direct consequences for \emph{multi-task peer prediction}, a mechanism-design problem in which a principal incentivizes honest reports from agents who observe correlated signals but no ground truth. Every such mechanism running on \(\ell\) independent tasks gives rise to a polynomial measure of dependence of degree at most \(\ell\), so our lower bounds on polynomial degree translate directly into lower bounds on the number of tasks: in the asymmetric case no finite-task mechanism exists on the larger-alphabet side at all, while the symmetric case requires at least \(\ell \geq 2n\) tasks.
arXiv:2606.14227v1 Announce Type: new Abstract: Non-trivial alignments between vorticity and the strain-rate tensor play an important role in the evolution of velocity gradients and the energy cascade in isotropic turbulence. Here we explore how alignments between the fluctuating and mean density gradients impact the mechanisms governing the turbulent kinetic energy (TKE) and available potential energy (APE) across scales in stably stratified turbulence. This is motivated by analytical results that demonstrate a connection between them, and is conducted using direct numerical simulations (DNS) of statistically stationary, stably stratified turbulence for $Pr = 1, 7, 50$ in the strongly stratified regime. After demonstrating how the gradient field alignments depend on scale and $Pr$, we show that the alignments are intimately connected to the reversal of the buoyancy flux at small-scales, and that regions of strong alignment and misalignment correspond to regions where the horizontal TKE inter-scale flux becomes weak. The same is also true of the APE flux, except that at larger scales, regions of strong alignment are associated with an upscale APE flux. The TKE and APE dissipation rates, and the mixing coefficient also show a strong dependence on the alignment, especially for $Pr=1$. Finally, we explore the connection between the local alignment and stability of the flow, and we find a non-trivial relationship, with regions of strong alignment surprisingly occurring most often in stable regions. This demonstrates that the dynamical significance of the alignments on the flow energetics cannot be understood through a simple connection between the local alignments and local stability of the flow.
arXiv:2606.14297v1 Announce Type: new Abstract: Developing accurate crowd-counting models for Hajj pilgrimage scenes remains challenging because domain-specific annotated images are scarce and data collection during large gatherings raises privacy concerns. To address these limitations, this paper proposes Pix2Pix-Hybrid (P2P-H), a hybrid conditional GAN for structure-guided Hajj crowd-image synthesis and data augmentation. P2P-H builds on Pix2Pix and employs a U-Net generator conditioned on eight input channels that jointly encode structural cues (edges and grayscale) and contextual attributes (crowd density and time of day). To capture detailed textures in dense scenes, the framework integrates two multi-scale PatchGAN discriminators operating at different resolutions. The training procedure combines adversarial, perceptual, and feature-matching objectives with adaptive data augmentation and stabilization strategies. The model was trained on 993 real Hajj frames collected from 60 publicly available video sources, with conditioning attributes derived automatically to reduce manual labeling effort. Using this framework, we constructed CrowdH, a synthetic dataset of 10,000 high-resolution Hajj crowd images. Experimental results show that P2P-H improves structure-preserving conditional synthesis quality compared with Pix2Pix and StyleGAN2-ADA baselines and shows favorable transfer to other crowd datasets. To assess downstream utility, we further constructed CrowdH-Mix-469, an annotated mixed real-synthetic dataset comprising 384 real Hajj images and 85 selected synthetic images,and evaluated five crowd-counting models under real-only and real-plus-synthetic training. The selected synthetic data reduced MAE across all five models, with the strongest gain observed for CSRNet.
arXiv:2606.14520v1 Announce Type: new Abstract: The open ocean lacks systematic in situ wind observations, and satellite scatterometer calibration depends on collocated surface measurements largely absent away from coastlines. Compact wave-sensing drifters already retrieve the open-ocean wind vector, but at modest accuracy -- about 1-2 m/s in speed and with unreliable direction at low wind. We show that a compact freely drifting GNSS/IMU drifter (MELODI) improves both components at a small fraction of the cost of moored platforms. Wind speed is read from the full shape of the measured wave acceleration spectrum rather than a single equilibrium-range level. A supervised model (Wind Inversion using Tikhonov Regularization, WITR -- a regularized regression with a gated residual correction), trained against scatterometer winds and evaluated by leave-one-buoy-out cross-validation (ASCAT MetOp-B/C, HY-2B/C; ERA5 used during feature development), reaches an RMSE of 0.90 m/s and defines the empirical skill ceiling of the feature set; it serves both directly as a wind retrieval and as a teacher model for distillation. Symbolic regression then distills this teacher into compact, interpretable closed-form laws suited to onboard implementation -- a spectrum-only law (RMSE about 1.0 m/s) and a motion-enhanced reduced drag law (0.93-0.94 m/s) -- that reproduce the teacher's skill within the present validation uncertainty. Wind direction is recovered independently from IMU-derived directional wave moments in the wind-sea band, with a mean absolute error of 9.4 degrees against scatterometer winds. These accuracies are scatterometer-consistent (comparable to published scatterometer-buoy differences) and close to a factor-of-two improvement over the traditional single-band Toba inversion; winds are retrieved plausibly up to 18 m/s, with higher winds flagged as lower confidence.
arXiv:2606.13946v1 Announce Type: new Abstract: This paper studies an inverse problem in nonlinear optical encryption. We examine chosen-plaintext attacks (CPA) on a nonlinear optical encryption strategy that integrates double random phase encryption (DRPE) into a nonlinear optical propagation model to enhance the security of the combined system. We first demonstrate that the system's phase information can be decoded from carefully designed differential CPA data. We then demonstrate that the strength of the optical device's nonlinearity can also be recovered from CPA data, indicating that including this parameter as an additional security key does not enhance protection against CPA attacks, although numerical simulations show that strong nonlinearity still poses significant challenges for CPA attacks. Finally, we provide a stability analysis to demonstrate that small errors in decoded security keys result in only small errors in the decrypted text, even though the encryption process is nonlinear.
arXiv:2606.14235v1 Announce Type: new Abstract: Variational Inference (VI) is a fundamental inference technique in Bayesian machine learning for approximating complex posterior distributions. Traditional VI often relies on the mean-field factorization, which can inadequately capture true posterior complexity. Recent advancements have leveraged neural networks to model implicit distributions, offering increased flexibility. However, the practical constraints of neural network architectures still produces inaccuracies. In this paper, we propose a method called Implicit Variational Rejection Sampling (IVRS), which integrates implicit distributions with rejection sampling to improve the posterior approximation. Our method uses neural networks to construct implicit proposal distributions, and rejection sampling with a discriminator network that estimates the density ratio between the implicit proposal and the true posterior for refining the approximation. Towards this end, we introduce the Implicit Resampling Evidence Lower Bound (IR-ELBO) as a metric to characterize the resampled distribution's quality and derive a tighter variational lower bound. Experimental results demonstrate that our method outperforms traditional variational inference techniques.
arXiv:2606.14237v1 Announce Type: new Abstract: Accurate and robust localization is a fundamental requirement for service and inspection robots, particularly in feature-sparse indoor environments where traditional systems struggle due to a lack of distinct landmarks. While prior maps can enhance robustness, precise and compact maps capturing real-world details are often unavailable for new or frequently changing environments. This paper presents BIM-Loc, a novel discrepancy-aware LiDAR-based localization method that directly integrates Building Information Models (BIM) from the design phase. BIM-Loc simultaneously estimates trajectories aligned with the BIM coordinate system and identifies discrepancies between real-world observations and the as-designed BIM in an online fashion. Our core contributions include: (1) a novel multi-hit ray casting strategy for efficient BIM-point data association and projection of 3D observations into 2D texture space; (2) a pose graph optimization framework with BIM-integrated factors that enforces consistency among odometry, sequential scans, and BIM structures; and (3) a hierarchical Bayesian inference module that incrementally updates a continuous 2D surface representation for discrepancy detection, propagating updates from the pixel to the structure level. Extensive evaluations in both simulation and real-world applications demonstrate that BIM-Loc significantly outperforms state-of-the-art map-based methods in localization accuracy and robustness.
arXiv:2606.14240v1 Announce Type: new Abstract: Affordance reasoning, the inference of an object's action possibilities from its physical properties (e.g., shape and material), is fundamental to human physical understanding and increasingly critical for Large Language Models (LLMs). However, existing affordance benchmarks largely expose explicit object identities in the evaluation setup, allowing models to rely on memorized object-affordance mappings rather than reasoning over physical properties. To address this gap, we introduce Affordance20Q, a novel affordance reasoning benchmark formulated as a 20-Questions game without exposing the object's identity. In each game, the model identifies a hidden object's affordance from a candidate set by asking yes/no questions about its physical properties. Affordance20Q comprises 1,009 games over 454 objects and 59 affordances, all manually filtered, refined, and annotated. We conduct comprehensive experiments with 15 state-of-the-art LLMs and find a substantial gap (~20 points) compared to human performance. A KL-based information-gain (IG) analysis further shows that models fail to ask discriminating questions as the game progresses. To close the gap, we develop KB-Anchored Rule Induction (KARI), a pipeline based on LLMs that generates affordance rules grounded in evidence from knowledge bases (KBs). KARI improves open-source LLMs by up to 15.2 points, while the limited coverage of KBs hinders further gains. We release all our code and data at https://github.com/1171-jpg/Affordance20Q.git
arXiv:2606.14242v1 Announce Type: new Abstract: Neutral gas transport directly affects the ionization source, propellant utilization, and low-frequency discharge oscillations in Hall thrusters. High-fidelity particle-based neutral models or DSMC methods can describe rarefied gas transport, but they are computationally expensive; in contrast, reduced neutral-continuity models are cheaper but require a closure for the neutral velocity or face-normal flux. Under a low-pressure collisionless approximation, this work adopts a free-molecular preprocessing strategy to provide a reference density field and the mean-velocity or face-normal-flux closure used by the reduced neutral-continuity equation in a manner consistent with the underlying transport model.On this basis, a particle-based free-molecular faceflux preprocessor is used as a stochastic reference, and an SN-DFEM deterministic free-molecular preprocessor is proposed to generate the corresponding reference density, velocity moments, and face-normal fluxes within a unified free-molecular transport framework. Results show that the SN-DFEM preprocessor preserves the main neutral-density and velocity structures and reduces the statistical error in face-flux closure by about three orders of magnitude in the baseline continuity-recovery test. A prescribed moderate ionization-loss case further demonstrates the extension of the framework to free-molecular preprocessing with volumetric neutral removal.
arXiv:2606.14243v1 Announce Type: new Abstract: Knowledge injection aims to equip large language models (LLMs) with external, domain-specific, or time-sensitive knowledge. Existing approaches typically face a trade-off between flexibility and integration: retrieval-augmented generation keeps knowledge outside the model but only provides prompt-level augmentation, whereas post-training based methods encode new knowledge into shared parameters but may introduce catastrophic forgetting, knowledge conflict, and costly updates. In this paper, we propose Decoupled Mixture-of-Experts (DMoE), a modular architecture for parametric knowledge injection that decouples both experts and the router from the base model. DMoE converts external knowledge corpora into independently updatable expert modules and uses a lightweight uncertainty-aware router to activate relevant experts only when the base model lacks sufficient knowledge during generation. To support efficient auto-regressive inference, DMoE attaches experts only to the final-layer feed-forward network, preserving KV-cache reuse while enabling parameter-level knowledge augmentation. Experiments on knowledge-intensive benchmarks show that DMoE consistently improves answer quality over retrieval and adapter-based baselines.
arXiv:2606.14245v1 Announce Type: new Abstract: Drug-target interaction (DTI) and affinity (DTA) predictors increasingly achieve strong benchmark scores, yet their internal use of sequence, fingerprint, and graph features often remains opaque. We present an interpretability audit of BridgeDPI architecture on three different datasets including Gao, Human, and C.elegans. This study combines gradient-based attributions -- integrated gradients, saliency, layer-wise relevance propagation, SmoothGrad, and SmoothGrad-IG -- with feature-wise occlusion ablation and strict intersection consensus across methods to reduce single-explainer bias. We summarize sensitivity and signed effects at raw inputs, at the bridge similarity scaffold, and through the graph convolution, including edge-level sensitivities and targeted edge removals. The results show that explainability is most informative when treated as model criticism: it reveals modality dominance, padding and special-token artifacts, dataset-dependent cooperative versus suppressive effects across layers, and chemistry-consistent fragment and composition motifs where methods agree. These analyses do not substitute for structural or experimental ground truth, yet they can provide testable hypotheses for downstream validation in computational drug discovery pipelines. More broadly, applying modern XAI to contemporary DTI/DTA models is still an early pass over the rich structure implicit in trained weights and data -- yet even this first layer of scrutiny already helps researchers relate predictions to drug- and target-side representations and to prioritize external validation.
arXiv:2412.05828v3 Announce Type: replace Abstract: Optimization problems in communication networks and information systems often contain coupled multiplicative or fractional terms, such as sum-of-products, sum-of-ratios, and logarithmic product-ratio structures. These problems are generally non-convex and difficult to solve, which motivates the development of tractable transformation and approximation techniques. In this paper, we propose an inequality-based transform framework for handling multiplicative and fractional terms involving an arbitrary number of coupled functions. The proposed framework is built upon the harmonic-mean, geometric-mean, arithmetic-mean, and quadratic-mean inequalities, and yields lower-bound and upper-bound surrogates for product-type terms. We derive the corresponding auxiliary-variable updates in closed form and show that the constructed surrogates are tight and first-order consistent at the current iterate. Based on these properties, we develop a class of successive approximation (SA) methods for sum-of-products/ratios minimization and maximization problems. When the transformed surrogate is convex for minimization or concave for maximization, the proposed method reduces to a standard successive convex approximation (SCA) method. When such convexity or concavity is not guaranteed, we further develop gradient-based SA variants and establish their sublinear convergence to an $\epsilon$-stationary point under standard smoothness and boundedness assumptions. We also discuss extensions to logarithmic product-ratio objectives and non-convex constraints. Numerical studies and application examples, including transmit-energy minimization, age-of-information minimization, semantic utility maximization, reliability-aware routing, cooperative edge caching, and product-loss learning, demonstrate the versatility and effectiveness of the proposed transform framework.
arXiv:2606.14504v1 Announce Type: new Abstract: Physical adversarial attacks on vision systems are typically studied through scene manipulation, such as adversarial patches or projections, where the adversary controls what the camera observes. Camera-side attacks using stickers or auxiliary optics have also been explored, but they treat attacks as image-space perturbations from designed patterns. This misses how physical imperfections interact with scene-dependent lighting and optics. We identify a threat: passive lens-side damage that is persistent yet trigger-conditioned, producing optical artifacts that bias geometric inference under particular visual conditions. We instantiate this threat through Scratch-induced Lens Adversarial Streak Hijacking SLASH, a physical-world attack caused by small scratches on a camera lens or protective cover. Scratches interact with bright light sources and specular reflections to create structured streak artifacts that distort depth cues. Since the perturbation is fixed in the optical path but triggered by the scene, it is both persistent and selective. We formulate the attack in optical space, model the scratch pattern as a trigger-conditioned optical channel, and optimize one fixed configuration across diverse viewing conditions. We evaluate SLASH on monocular depth estimation and monocular 3D object detection in digital and real-world settings. Under the fixed-scratch constraint, directional depth shifts reach up to 32% relative error for monocular depth estimation, with consistent effects on monocular 3D object detection. Physical experiments confirm transfer to real camera recordings, inducing depth shifts above the model's natural prediction baseline. These findings reveal an attack surface where benign-looking hardware imperfections act as latent, scene-triggered adversarial mechanisms, challenging assumptions about physical robustness and motivating defenses for secure vision systems.
Security Threats and Their Impact on Blockchain Interoperability: Identification and Countermeasures
arXiv:2606.14554v1 Announce Type: new Abstract: Blockchain interoperability enables independent blockchain systems to communicate and exchange assets across heterogeneous networks. However, the lack of comprehensive security mechanisms remains a critical weakness -- one that attackers have already exploited to cause hundreds of millions of dollars in asset losses. This paper presents a systematic identification and classification of security threats facing interoperable blockchain systems, along with corresponding countermeasures for each. We organize threats into five categories: (1) core blockchain attacks, (2) network attacks, (3) interoperability-specific attacks, (4) social engineering, and (5) code vulnerabilities, with particular attention to smart contract weaknesses. For each identified threat, we analyze its attack surface and propose effective defensive strategies. The resulting taxonomy provides a structured foundation for designing and evaluating secure blockchain interoperability solutions.
arXiv:2606.14527v1 Announce Type: new Abstract: We report multi-angle reflectivity measurements in the extreme-ultraviolet (XUV) range for mono- and bilayer MoS$_2$ on a Si$_3$N$_4$ substrate. Using a single-sheet 2D conductivity model, we extract the complex optical response of the MoS$_2$ bilayer between 25 and 90 eV and derive an effective refractive index by introducing a thickness equal to the interlayer spacing. The MoS$_2$ monolayer response is consistently reproduced either by halving the 2D conductivity or the effective thickness, indicating a robust scaling with layer number. The resulting optical constants display a broad resonance at the Mo N$_{2,3}$ edge with no signatures of sharp core-exciton features despite the reduced dimensionality. First-principles calculations reproduce the experimental results and show that local-field (Hartree) effects dominate the XUV response, while screened-exchange (SEX) contributions remain weak and mainly induce spectral shifts. Our analysis demonstrates that excitonic effects play a minor role in the XUV optical response of atomically thin MoS$_2$, highlighting key differences with respect to the visible and infrared regimes, and calling for a reassessment of the use of Mo-based transition metal dichalcogenides in attosecond spectroscopy and XUV excitonics.
arXiv:2412.16317v2 Announce Type: replace Abstract: The Epstein zeta function generalises the classical Riemann zeta function to oscillatory lattice sums in higher dimensions and has recently emerged as a key tool in the simulation of long-range interacting classical and quantum many-body systems. Its computation and analytic properties are therefore of significant interest, yet a rigorous and comprehensive treatment has been lacking. We address this gap by introducing a superexponentially convergent algorithm, complete with error bounds, for computing the Epstein zeta function in any dimension with arbitrary real parameters. Our approach is accompanied by a detailed analysis of the analytic properties of the Epstein zeta function. We first present a concise reformulation of its meromorphic continuation, functional equation, and symmetries. We then establish, for the first time, its joint holomorphic continuation in all parameters and offer a complete characterization of the resulting complex singularity structure, which governs convergence rates in numerical algorithms based on the function. Recognizing that the function can be decomposed into power-law singularities and a regularised analytic part, we provide an algorithm for removing singularities without cancellation error. This facilitates the evaluation of integrals and enables fast precomputations through interpolation methods. We present the first high-performance implementation for arbitrary real arguments in EpsteinLib, a C library with Python and Julia bindings, and rigorously benchmark its performance and accuracy, achieving full-precision evaluation against known analytic results in dimensions 1,2,3,4,6, and 8 and against an arbitrary precision implementation across the entire parameter range. Finally, we apply our methods to the computation of quantum dispersion relations in spin systems and Casimir energies in three-dimensional geometries.
arXiv:2606.14252v1 Announce Type: new Abstract: As multi-drone fleets scale, zone assignment rapidly evolves into an intractable NP-hard combinatorial problem that overwhelms classical exhaustive search. While quantum optimization promises to shatter these classical bottlenecks, mapping complex spatial tasks from human intent to restricted quantum hardware remains a severe challenge. To bridge this gap, we present an end-to-end framework integrating a fine-tuned Large Language Model (LLM) front-end with a highly scalable, domain-specific quantum-classical backend. The front-end utilizes Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) to translate free-form natural language instructions into structurally robust Quadratic Unconstrained Binary Optimization (QUBO) constraints without false negatives. To overcome the strict qubit limits of near-term quantum devices, our framework features a novel constraint-preserving graph partitioner and a compressed separator-based dynamic programming (DP) merge. By structurally encoding constraints via W-state initialization and XY-mixers in Conditional Value-at-Risk Quantum Approximate Optimization (CVaR-QAOA), the pipeline stays highly compact. Empirical results demonstrate that this architecture circumvents classical scaling walls, recovering the global optimum on 100% of idealized oracle cases and 96.3% under real QAOA sampling, enabling natural-language-guided task allocation at previously intractable scales.
arXiv:2606.14255v1 Announce Type: new Abstract: Diffusion-based Vision-Language-Action (VLA) policies have demonstrated strong capability in modeling expressive and multimodal action distributions. However, their reliance on iterative sampling introduces substantial inference latency, which limits their applicability to reactive closed-loop robot manipulation. To address this limitation, we propose \texttt{ReactVLA}, a lightweight and low-latency VLA framework for real-time robotic manipulation. \texttt{ReactVLA} combines two complementary designs: (1) an improved Mean Flow (iMF) action generator that reduces expensive multi-step diffusion sampling to one-to-few-step action generation, and (2) Attention Residuals (AttnRes), a dynamic depth-wise feature routing mechanism that replaces uniform residual accumulation to better preserve task-relevant multimodal representations. We evaluate \texttt{ReactVLA} on large-scale simulation benchmarks, including LIBERO and RoboIMI, as well as real-world robotic manipulation tasks. Experimental results show that \texttt{ReactVLA} consistently outperforms similarly sized VLA baselines, including SmolVLA and $\pi_0$. On challenging precision manipulation tasks, \texttt{ReactVLA} achieves up to a 1.65$\times$ improvement in task performance while providing more than a 4$\times$ increase in inference speed compared with leading VLA models. Finally, it reduces real-world policy latency to below 38.6 ms, enabling fast reactive control on physical robot platforms. Please check out our project website at: https://game-loader.github.io/ReactVLA/.
arXiv:2606.14259v1 Announce Type: new Abstract: Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data aspects, architecture design, and optimization properties. Yet these explanations are often studied in isolation, leaving their relative importance unclear. In this work, we revisit these hypotheses through a controlled empirical study across vision, language, genomics, and graph tasks, spanning modern and classical architectures, and carefully designed training setups. Our results suggest that no single factor consistently explains the Adam--SGD gap. For instance, the Adam advantage can (1) persist under a uniform vocabulary distribution yet nearly disappear under a heavy-tailed one; (2) reverse in favor of SGD in softmax-attention models; and (3) become larger under soft architectural modifications, e.g., when ReLU is replaced by a GeLU nonlinearity. This suggests that the gap arises from nontrivial data and architecture interactions, rather than from a single common factor. Yet, we observe a pattern across our settings: a \emph{crossover batch size} at which the relative advantage shifts from SGD to Adam as the batch size scales. These empirical results are captured by our theoretical gap model, which predicts this batch-size-dependent crossover. Our perspective helps reconcile several existing hypotheses while offering practical insights across domains.
arXiv:2606.14267v1 Announce Type: new Abstract: Floor plans encapsulate compact spatial priors, enabling agents to navigate unseen scenes more efficiently. While prior work has explored floor plan-guided navigation, it has focused mainly on PointNav and a limited set of environments. To bridge this gap, we introduce FloVerse, a new task for floor plan-guided embodied navigation that unifies PointNav, ObjectNav, and ImageNav. To support FloVerse, we assemble FloVerse-1.6K, a large-scale dataset of 1.6K scenes from HM3D and Gibson 4+, paired with corresponding floor plans, comprising 240K expert trajectories and 12M RGBD frames. We further propose ThreeDiff, a two-stage imitation learning policy comprising a planner, a diffusion-based multimodal goal-reasoning module trained via masked-modality modeling, and a refiner, a depth-based trajectory-refinement module for safe execution. Extensive experiments demonstrate that (1) floor-plan priors improve navigation performance across all goal modalities, and (2) ThreeDiff implicitly captures spatial information from floor plans. These results underscore the effectiveness of spatial priors and validate our proposed unified approach for floor plan-guided embodied navigation.
arXiv:2606.14269v1 Announce Type: new Abstract: Fixed-cardinality retrieval injects a constant top-K chunks into the generator regardless of query complexity, causing over-retrieval for narrow queries and under-retrieval for compositional ones. We describe ScoreGate, a lightweight score-space decision mechanism that controls retrieval cardinality at inference time using two scores already produced by the standard pipeline: bi-encoder similarity s_i and cross-encoder reranker score r_i, with no additional model inference calls required. Its core insight is that cross-encoder affirmation can rescue semantically relevant chunks that bi-encoder retrieval ranks poorly due to vocabulary mismatch -- a failure mode unaddressed by fixed-K or single-score thresholding. On MS MARCO (200 dev queries), ScoreGate achieves MRR@10 = 0.401 with 35% fewer retained chunks than Standard Top-K. On an internal benchmark (n=300, Fleiss' kappa=0.87), ScoreGate observed zero false positives (95% CI [96.4%, 100%]) at 97.77-99.34% recall, with 34.8% fewer tokens per query and only 31ms added latency. Results on both MS MARCO and real-world production traffic suggest that adaptive retrieval cardinality can improve retrieval efficiency without degrading retrieval quality.
arXiv:2510.09585v4 Announce Type: replace Abstract: Community Notes (formerly known as Birdwatch) is the first large-scale crowdsourced content moderation initiative launched by X (formerly Twitter) in January 2021. As the Community Notes model gains momentum across other social media platforms, there is a growing need to assess its underlying dynamics and effectiveness. This paper provides a descriptive investigation of Community Notes during its first four years, examining its linguistic diversity, sourcing practices, Contributor activity, rating behaviour, and interaction networks. In addition, we release a curated dataset and accompanying source code to support future research, along with a review of prior research on Community Notes. We parsed Notes and ratings data from the first four years of the program and conducted language detection across all Notes. For English-language Notes, we extracted embedded URLs and identified discussion topics in each Note. Additionally, we constructed monthly interaction networks among the Contributors. Together, the descriptive analysis, dataset, code, and literature review provide a foundation for advancing research on Community Notes and community-based content moderation more broadly.
arXiv:2501.11842v5 Announce Type: replace Abstract: The intrinsic integration of Rydberg atomic receivers into wireless communication systems is proposed, by harnessing the principles of quantum physics in wireless communications. More particularly, we conceive a pair of Rydberg atomic receivers, one incorporates a local oscillator (LO), referred to as an LO-dressed receiver, while the other operates without an LO and is termed an LO-free receiver. The appropriate wireless model is developed for each configuration, elaborating on the receiver's responses to the radio frequency (RF) signal, on the potential noise sources, and on the signal-to-noise ratio (SNR) performance. The developed wireless model conforms to the classical RF framework, facilitating compatibility with established signal processing methodologies. Next, we investigate the associated distortion effects that might occur, specifically identifying the conditions under which distortion arises and demonstrating the boundaries of linear dynamic ranges. This provides critical insights into its practical implementations in wireless systems. Finally, extensive simulation results are provided for characterizing the performance of wireless systems, harnessing this pair of Rydberg atomic receivers. Our results demonstrate that LO-dressed systems achieve a significant SNR gain of approximately 40~50 dB over conventional RF receivers in the standard quantum limit regime. This SNR head-room translates into reduced symbol error rates, enabling efficient and reliable transmission with higher-order constellations.
arXiv:2502.07242v2 Announce Type: replace Abstract: This article explores nonlinear analogues of skew quasi-cyclic codes of index~$\ell$, i.e., $\mathbb{F}_{q^m}[X;\sigma]$-submodules of $\left(\mathbb{F}_{q^m}[X;\sigma]/(X^n - 1)\right)^\ell$. After introducing nonlinear skew quasi-cyclic codes, we then determine the module structure of these codes by using a two-fold iteration of the Smith normal form of matrices over skew polynomial rings. We show that actually a single use of the Smith normal form will suffice to determine the elementary divisors of the code. Along the way, we also describe duals of our codes with respect to appropriately chosen inner products.