arXiv:2607.09729v1 Announce Type: new
Abstract: In his 1996 doctoral thesis, Maurice Pagnucco created the first AGM-like abductive expansion operation. Taking his operation as a basis, as well as a taxonomy -- inspired by Atocha Aliseda -- responsible for highlighting and formalizing the main components of abductive reasoning, the main aim of this paper is to present a new paraconsistent AGM-like abductive expansion operation -- capable of assimilating contradictory explanatory hypotheses without trivialization and the consequent absurd epistemic state -- with its postulates and its transitively relational partial meet construction. To a large extent, the formal development presented in this paper was only made possible by the recent creation of the paraconsistent logic RCbr, an LFI (Logics of Formal Inconsistencies) that establishes properties especially relevant to belief revision contexts, in particular, the ability to be self-extensional -- i.e., to satisfy the replacement property. This is the first of two papers: the paraconsistent abductive expansion operation announced here -- which is part of a new system called AGMpabd -- despite bringing many interesting features, does not assign any relevant epistemic role to the paraconsistent operators of negation and consistency. Only in a second paper will an analogous paraconsistent abductive expansion operation -- which is part of another new system, AGMcircabd -- be enhanced in this direction. Nevertheless, to the best of my knowledge, the operation developed in this paper is the first of its kind in the AGM literature.
Science Journals
arXiv:2607.10969v1 Announce Type: new
Abstract: Given a large graph, how to generate a compact summary graph that is configurable by the user and supports multiple graph queries with either no loss or with high accuracy? The ever growing size of graph datasets makes the above question on graph summarization very pertinent. Although, there are several approaches, there does not exist a configurable graph summarization method that offers high compression along with support for multiple graph queries on the summary graph with high accuracy, and allows the user to configure the summarization based on: (1) lossless or lossy summarization, (2) amount of tolerable neighborhood loss, (3) the type of loss it can tolerate, in terms of false positive edges (i.e., extra edges), false negative edges (i.e., missing edges), or neither, in both the (a) reconstructed graph and the (b) query answers. To overcome these limitations, we propose a novel graph summarization framework CGS (Configurable Graph Summarizer) that builds upon the idea of aggregating nodes with common neighborhoods. The CGS framework consists of three summarization variants, CGS-E, CGS-I and CGS-U. While CGS-E is a lossless scheme, CGS-I and CGS-U are lossy schemes that allow reconstruction of the input graph with no false positive edges and no false negative edges, respectively. To bound the graph reconstruction loss, we introduce a user-specified parameter neighborhood loss tolerance threshold, that limits the maximum loss allowed in the neighborhood of each node. This allows graph reconstruction and neighborhood query evaluation with either no loss or with bounded loss guarantees. Empirical evaluation on several synthetic and real-world graphs shows that CGS offers superior summarization than the state-of-the-art methods, and can answer graph queries with fairly high accuracy and efficiency.
arXiv:2607.10970v1 Announce Type: new
Abstract: Federated learning distributes data among $n$ clients, making it vulnerable to malicious attacks and data heterogeneity, which together pose challenges for robust learning. To tackle this issue, centered clipping and Huber aggregators have been exploited for Byzantine robustness. In this paper, we first demonstrate their equivalence via convex conjugate theory, and show that they can yield biased solutions in the presence of outliers, leading to failure under high data heterogeneity and a substantial fraction of outliers. Next, we propose a new robust aggregation rule that utilizes the truncated-quadratic (TQ) loss, effectively mitigating the biases of existing methods, such as centered clipping and Huber aggregators. We show that our aggregator achieves order-optimal Byzantine-robust learning under nonconvex loss functions and heterogeneous data, ultimately enhancing the reliability of federated learning systems. Additionally, we provide a robust deviation estimation strategy for TQ, demonstrating its effectiveness. Furthermore, we show that TQ maintains robustness even when only an estimate of the number of Byzantine clients is available. Finally, experimental results on MNIST, Fashion-MNIST, and CIFAR-10, indicate that our aggregator provides better robustness performance than the competing techniques.
arXiv:2607.11181v1 Announce Type: new
Abstract: The (quasi)particles or structured wavepackets in parabolic potential exhibit well-known harmonic oscillations, typically described by the Lissajous equations. However, such conventional harmonic laws rely on a fundamental assumption that the different constituent components of the (quasi)particles or wavepackets do not interact. Here we challenge this paradigm, by taking advantage of intrinsic couplings among distinct constituents-specifically by leveraging nontrivial couplings between vortices and antivortices embedded in a spatially structured wavepacket. We demonstrate theoretically and experimentally abnormal motions by considering two different optical waveforms. For a vortex-antivortexcoupled dipole mode, we reveal counterintuitive propagation regimes, including periodic annihilation and regeneration of the dipole, its non-orbital motion and realization of a critical equilibrium state without nonlinearity. For a circular chain of vortices with an antivortex set at the center, we successfully tune the oscillation frequency of the overall configuration in the potential, thus disobeying the classical Lissajous trajectories, by precisely engineering the nonlocal vortex-antivortex couplings. Since the harmonic oscillations have been proven to be fundamental physical phenomena in distinct disciplines and led to numerous important applications, our demonstrations provide different opportunities to trigger considerable investigations and potential applications, by leveraging the underlying anomalous motions of the vortex-antivortex-coupled wavepackets in the parabolic potential.
arXiv:1708.04326v1 Announce Type: cross
Abstract: This paper evaluates existing and newly proposed answer selection methods based on pre-trained word embeddings. Word embeddings are highly effective in various natural language processing tasks and their integration into traditional information retrieval (IR) systems allows for the capture of semantic relatedness between questions and answers. Empirical results on three publicly available data sets show significant gains over traditional term frequency based approaches in both supervised and unsupervised settings. We show that combining these word embedding features with traditional learning-to-rank techniques can achieve similar performance to state-of-the-art neural networks trained for the answer selection task.
arXiv:2509.03758v5 Announce Type: replace
Abstract: We propose a data-driven interpolation framework for reconstructing real-valued functions on smooth manifolds from scattered pointwise observations. The method combines a Gaussian Nadaraya--Watson kernel interpolant with a Voronoi-adaptive bandwidth determined entirely by the geometry of the sampled data, yielding an explicit closed-form construction that requires neither training, iterative optimization, preprocessing, nor parameter tuning.
The proposed interpolant satisfies several theoretical properties. It reproduces the observed data exactly, enforces a vanishing intrinsic gradient at every sample point, and, in the dense-sampling limit, attenuates high-frequency oscillatory components through the geometric regularization induced by the adaptive bandwidth. Furthermore, the construction admits an interpretation in terms of minimizing a discrete total variation--type functional, establishing a natural connection with compressed sensing and sparsity-promoting regularization.
Unlike classical kernel interpolation methods employing a fixed global bandwidth, the proposed adaptive strategy automatically adjusts to the local sampling geometry through the Voronoi tessellation while preserving an explicit analytical formulation. Because the interpolant is available in closed form, the overall computational cost is entirely determined by the inference stage: evaluating the interpolant at a query point requires only the computation of Gaussian kernel weights and their weighted combination, resulting in linear complexity with respect to the number of sample points. In contrast to many data-driven interpolation approaches, no additional offline computational stage is required before inference.
arXiv:2607.10678v1 Announce Type: new
Abstract: Emotional intelligence enables humans to recognize emotions, infer their causes, reason about interventions, and modify their environment to achieve desired affective states. Despite recent advances in artificial intelligence (AI), current models remain largely limited to generating realistic content or performing semantic reasoning, with little capacity for understanding, predicting, and personalizing human emotional responses. Here we introduce Emotion-augmented geneRatiOn System (EROS), a hybrid AI framework that integrates symbolic reasoning with deep learning to enable personalized emotion augmentation through visual content. Leveraging large-scale image-emotion datasets, EROS discovers generalizable affective rules, identifies emotion-relevant image regions, and predicts context-aware visual modifications that preserve scene semantics while steering emotional responses toward desired targets. To account for individual variability, EROS incorporates an expandable memory bank that supports inference-time personalization without model fine-tuning, yielding interpretable emotional profiles and rapid adaptation to new users. Across extensive human psychophysics experiments, EROS elicits target emotional responses more effectively than state-of-the-art large multimodal models while adapting to individual affective preferences. Beyond affective computing, EROS provides a foundation for AI systems that can understand, reason about, and augment human cognitive states, with potential applications in mental health, adaptive media, education, and human-computer interaction.
arXiv:2607.10681v1 Announce Type: new
Abstract: In pre-LayerNorm looped transformers, LayerNorm inside the recurrent block acts as an implicit gain controller: by coupling the block's local Lipschitz constant inversely to the activation scale, it renders the recurrence Jacobian non-normal -- asymptotically contractive at every verified fixed point even where its operator norm exceeds 1 -- so the true stability budget is the spectral margin, not an operator-norm bound. That margin depletes as the carry $\rho \to 1$, and a minority of initializations never converge to a fixed point at all, so the diagonal carry constraint $\rho(\bar{A}) < 1$ is necessary but not sufficient for convergence of the full recurrence. Training experiments across six tasks, including a controlled ablation, reveal that the linear carry is not the depth-memory mechanism: gradient descent routes memory through the block's more expressive nonlinear recurrence and leaves the stability-constrained carry at rest -- the carry's role is stabilization, not memory. We characterize the boundary of this claim: on tasks with axis-aligned per-channel structure, gradient descent does recruit the carry. All results are derived analytically and verified in a from-scratch, CPU-scale implementation; verification at larger scale is needed.
arXiv:2607.10682v1 Announce Type: new
Abstract: Existing robotic perception is constrained by sensors that are either robot-mounted or permanently fixed in the environment, locking perception to a limited set of viewpoints. Yet as robots perform increasingly diverse tasks, the most informative viewpoint shifts from one task to the next-often somewhere onboard sensor and static infrastructure can not readily satisfy. To address this gap, we propose SensorPerch, a novel realization of active perception that decouples sensing from both the robot embodiment and the environment by treating sensors as independent physical entities that the robot can autonomously detach and re-attach within the environment. SensorPerch presents one realization of this paradigm: a lightweight, wireless, reconfigurable sensor platform that can perch on diverse surfaces, paired with a viewpoint-selection framework that determines task-optimal sensor placements. Together, these enable robots to construct task-relevant viewpoints on demand, independent of the robot's current position and available fixed infrastructure. We demonstrate the paradigm on two task classes: (i) object-coupled perception, where SensorPerch enables persistent object-state detection beyond the robot's current position, achieving successful event detection even when the robot is not nearby; and (ii) policy-coupled perception, where SensorPerch allows robots to construct diverse, policy-specific viewpoints for various policies, achieving success rates comparable to those obtained using oracle viewpoints.
arXiv:2603.25559v3 Announce Type: replace
Abstract: Non-fixed flexible antenna architectures, such as fluid antenna system (FAS), movable antenna (MA), and pinching antenna, have garnered significant interest in recent years. Among them, rotatable antenna (RA) has emerged as a promising technology for enhancing wireless communication and sensing performance through flexible antenna orientation/boresight rotation. By enabling mechanical or electronic boresight adjustment without altering physical antenna positions, RA introduces additional spatial degrees of freedom (DoFs) beyond conventional beamforming. In this paper, we provide a comprehensive tutorial on the fundamentals, architectures, and applications of RA-empowered wireless networks. Specifically, we begin by reviewing the historical evolution of RA-related technologies and clarifying the distinctive role of RA among flexible antenna architectures. Then, we establish a unified mathematical framework for RA-enabled systems, including general antenna/array rotation models, as well as channel models that cover near- and far-field propagation characteristics, wideband frequency selectivity, and polarization effects. Building upon this foundation, we investigate antenna/array rotation optimization in representative communication and sensing scenarios. Furthermore, we examine RA channel estimation/acquisition strategies encompassing orientation scheduling mechanisms and signal processing methods that exploit multi-view channel observations. Beyond theoretical modeling and algorithmic design, we discuss practical RA configurations and deployment strategies. We also present recent RA prototypes and experimental results that validate the practical performance gains enabled by antenna rotation. Finally, we highlight promising extensions of RA to emerging wireless paradigms and outline open challenges to inspire future research.
arXiv:2603.26599v2 Announce Type: replace
Abstract: Large-scale video diffusion models achieve impressive visual quality, yet often fail to preserve geometric consistency. Prior approaches improve consistency either by augmenting the generator with additional modules or applying geometry-aware alignment. However, architectural modifications can compromise the generalization of internet-scale pretrained models, while existing alignment methods are limited to static scenes and rely on RGB-space rewards that require repeated VAE decoding, incurring substantial compute overhead and failing to generalize to highly dynamic real-world scenes. To preserve the pretrained capacity while improving geometric consistency, we propose VGGRPO (Visual Geometry GRPO), a latent geometry-guided framework for geometry-aware video post-training. VGGRPO introduces a Latent Geometry Model (LGM) that stitches video diffusion latents to geometry foundation models, enabling direct decoding of scene geometry from the latent space. By constructing LGM from a geometry model with 4D reconstruction capability, VGGRPO naturally extends to dynamic scenes, overcoming the static-scene limitations of prior methods. Building on this, we perform latent-space Group Relative Policy Optimization with two complementary rewards: a camera motion smoothness reward that penalizes jittery trajectories, and a geometry reprojection consistency reward that enforces cross-view geometric coherence. Experiments on both static and dynamic benchmarks show that VGGRPO improves camera stability, geometry consistency, and overall quality while eliminating costly VAE decoding, making latent-space geometry-guided reinforcement an efficient and flexible approach to world-consistent video generation.
arXiv:2603.29042v2 Announce Type: replace
Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive. Highly performant English-focused models do not generalize across languages, while multilingual models underutilize pretrained representations. It also remains unclear how data scale, architecture, and training objective contribute to multilingual PR. We present PhoneticXEUS -- trained on large-scale multilingual data and achieving state-of-the-art performance on both multilingual (17.7% PFER) and accented English speech (10.6% PFER). Through controlled ablations with evaluations across 100+ languages under a unified scheme, we empirically establish our training recipe and quantify the impact of SSL representations, data scale, and loss objectives. In addition, we analyze error patterns across language families, accented speech, and articulatory features. All data and code are released openly at https://github.com/changelinglab/PhoneticXeus
arXiv:2604.00137v2 Announce Type: replace
Abstract: Tool-integrated LLMs retrieve information, perform computations, and take real-world actions, but their reliability depends on both tool-use accuracy and intrinsic tool accuracy, including tool correctness, stability, and safety. While prior work primarily emphasizes tool use, intrinsic tool accuracy remains underexamined. We introduce OpenTools, a community-driven and maintainable toolbox for discovering, using, evaluating, and contributing open-source tools. OpenTools standardizes tool interfaces, converts documented Python functions into reviewable bundles, supports maintainer-triggered evaluation, and combines non-executing risk inspection with optional advisory LLM review. A public web demo allows users to run tools and agents, inspect evidence, contribute tests, and submit tools for maintainer review, while MCP enables controlled access from external applications. Experiments show that community-contributed, task-specific tools yield relative gains of 6% to 22% over an existing toolbox across multiple agent architectures, highlighting the importance of intrinsic tool accuracy.
arXiv:2604.00878v2 Announce Type: replace
Abstract: Actor-level stance detection aims to determine an author expressed position toward specific geopolitical actors mentioned or implicated in a text. Although transformer-based models have achieved relatively good performance in stance classification, they typically rely on unified representations that may not sufficiently capture heterogeneous linguistic signals, such as contrastive discourse structures, framing cues, and salient lexical indicators. This motivates the need for adaptive architectures that explicitly model diverse stance-expressive patterns. In this paper, we propose StanceMoE, a context-enhanced Mixture-of-Experts (MoE) architecture built upon a fine-tuned BERT encoder for actor-level stance detection. Our model integrates six expert modules designed to capture complementary linguistic signals, including global semantic orientation, salient lexical cues, clause-level focus, phrase-level patterns, framing indicators, and contrast-driven discourse shifts. A context-aware gating mechanism dynamically weights expert contributions, enabling adaptive routing based on input characteristics. Experiments are conducted on the StanceNakba 2026 Subtask A dataset, comprising 1,401 annotated English texts where the target actor is implicit in the text. StanceMoE achieves a macro-F1 score of 94.26%, outperforming traditional baselines, and alternative BERT-based variants.
arXiv:2607.10866v1 Announce Type: new
Abstract: This work is concerned with the development of control strategies for the Direct Internal Recycling Loop (DIRL) system, which is an essential part of the Tritium Fuel Cycle (TFC) for the fueling of fusion reactors. As a first step, a control-oriented model is developed that describes dynamic behavior of DIRL with interactions between the torus reactor, buffer and fuel units, and recirculation streams. This model is used to evaluate the controllability, stability of the DIRL, and interactions between input and output variables. Moreover, the direct recycling of isotopes from the exhaust gases is discussed from a control perspective. It is observed that Gas Distribution and Storage (GDS) within the DIRL is associated with significant control challenges due to input-output interactions and competing process objectives. Three control strategies are developed and evaluated for the GDS of the European Demonstration fusion power plant (EU-DEMO): a Multiple Input Multiple Output (MIMO) control scheme, a redesigned GDS configuration with extended input variables enabling decentralised Single Input Single Output (SISO) control, and an extended-input MIMO control scheme addressing protium dilution. All strategies are assessed against three control objectives: maintaining GDS pressure around a prescribed set-point to ensure process safety; regulating the tritium-deuterium fuelling ratio for optimal reactor operation; and managing protium concentration to prevent fuelling dilution.
arXiv:2607.10995v1 Announce Type: new
Abstract: Recent generalizable 3D Gaussian Splatting models have advanced long-sequence novel view synthesis (NVS), but at the cost of substantial redundant computation. We identify that the redundancy can be mitigated based on two observations: (i) high-precision geometry is not strictly required for high-quality NVS; (ii) appearance learning is generally easier than geometry recovery. Motivated by these insights, we propose an asymmetric architecture that decouples geometry and appearance modeling. The geometry branch processes coarse-grained tokens with most of the parameters for multi-view reconstruction, while the appearance branch operates on fine-grained tokens to capture details using significantly fewer parameters. The two branches interact through bilateral connections, enabling mutual guidance for their respective tasks. This task-aware asymmetry reduces the computational redundancy and allocates the computation more judiciously, thereby increasing parameter efficiency and enabling smaller models to achieve strong performance. On 32-view 960P inputs, our model matches optimization-based methods while delivering nearly 800x speedup, and surpasses the zero-shot performance of state-of-the-art generalizable models with markedly fewer parameters and reduced training/inference overhead, achieving an overall efficiency improvement.
arXiv:2607.11337v1 Announce Type: new
Abstract: Montgomery's trick accelerates simultaneous modular inversion of $N$ inputs by amortizing a single shared inversion, but auxiliary multiplications for complement products are typically scheduled in a linear, serial form. We construct a maximally parallelizable data-flow graph (DFG) that computes all $\overline{x}$ complement~products by scheduling auxiliary multiplications into idle multiplier slots during accumulation of the product of all inputs, and that of the shared inversion. This scheduling ensures the post-inversion phase adds exactly one multiplication layer of latency regardless of $N$, yielding a critical path latency of $\lceil \log_2 N \rceil$ multiply layers, one inversion, and one final parallel multiply layer.
arXiv:2607.10869v1 Announce Type: new
Abstract: We study the population gradient flow of an infinitely wide two-layer neural network learning a misspecified single-index model in high dimension. The two layers are optimized jointly, with a perturbative parameter tuning the relative training speed between the first and second layer. This setting was considered by Berthier, Montanari and Zhou in \cite{berthier2024learning}, who conjectured a hierarchical learning scenario with explicit timescales as the second layer is trained faster than the first. In this paper, we prove that the constant and linear components of the hidden link function are indeed recovered within the predicted timescales, at sharp explicit thresholds. We then analyze the onset of learning of the quadratic component and show that the components learned at earlier stages continue to influence the dynamics in an essential way. Our proof is based on quantitative approximation results for singularly perturbed flows evolving near a manifold defined by integral constraints. At a phenomenological level, we also show that the empirical measure of the weights displays singular behaviour when reaching the quadratic component of the hidden link, with a small fraction of neurons growing significantly while the remaining ones rearrange to preserve the components already learned.
arXiv:2607.11312v1 Announce Type: new
Abstract: We introduce Skill Learning from Video Memory (SLVMBench), the first benchmark that jointly evaluates whether video large language models (video-LLMs) can learn skills from long video memory and apply them to real-time tasks. SLVMBench presents models with 2-3 hour video streams that contain a tutorial video embedded in a stream of arbitrary irrelevant videos, resembling real-world human learning practices. Video-LLMs are asked to apply the acquired skill to answer real-time questions about an ongoing video. Unlike long-video understanding benchmarks that emphasize passive comprehension and skill-learning benchmarks that rely on short, immediate demonstrations, SLVMBench tests the full pipeline of memorizing and extracting procedural knowledge, as well as transferring it to real-time tasks. Moreover, rigorous human annotations feature sub-second-level temporal calibration, manually engineered questions eliminating common-sense guessing, and collated tutorials to ensure coverage of the required skills. Evaluations on state-of-the-art proprietary and open-source video LLMs show that video-LLMs struggle substantially with learning and applying skill knowledge from videos. Moreover, performance degrades markedly when the skill knowledge is placed within a long video memory. These results reveal a key limitation of existing video LLMs and position SLVMBench as the first benchmark for studying real-time skill acquisition and application from long-context video memory.
arXiv:2607.10873v1 Announce Type: new
Abstract: Achieving optimal screw placement for orthopedic surgeries requires frequent alignment checks and multiple anatomical views under X-ray -- a process known as "fluoro-hunting" that increases radiation exposure to patients and surgical teams. This work introduces X-GuideAR, an augmented reality (AR) framework for identifying optimal X-ray views, aimed at reducing radiation exposure while ensuring accurate screw placement. To exemplify the benefits of X-GuideAR, we focus on S2 alar-iliac (S2AI) screw placement. Our system provides radiation-free guidance for view acquisition and drilling by generating synthetic X-ray previews that accelerate fluoro-hunting. Once the required anatomical views are identified using these previews, a real X-ray is acquired, and the preview of the drilling trajectory is augmented onto it, facilitating precise screw placement with minimal additional radiation. A preliminary study involving eight S2AI trajectories performed by an expert spine surgeon demonstrated a 62.3% reduction in the number of X-rays. Post-procedure evaluations showed that trajectories done with X-GuideAR supported an average safe screw diameter of 12.95 mm compared to 5.9 mm under the conventional workflow, suggesting improved bony containment and potential biomechanical benefit. X-GuideAR shows great potential to reduce radiation exposure and streamline S2AI screw placement, offering a promising direction toward safer and more efficient surgeries.
arXiv:2607.10998v1 Announce Type: new
Abstract: Precise Event Spotting (PES) requires distinguishing visually similar yet semantically distinct adjacent frames, making it fundamentally different from image classification and coarse action recognition. Although self-distillation methods such as DINO have shown strong representation learning ability in images, we find that directly applying them to PES is ineffective: without supervised guidance, subtle but crucial motion cues are often suppressed as noise, leading to representations that are insensitive to precise event boundaries. To address this, we propose Temporal Feature Distillation, a semi-supervised objective that aligns temporally informative backbone features, rather than projection-head outputs, to preserve motion-sensitive and boundary-aware cues for frame-level localization. A supervised warm-up with a ramp-up schedule further stabilizes training by ensuring that meaningful event cues are learned before unlabeled distillation begins. We also introduce Transformer Gate Shift, a multi-scale gated shifting module that injects motion-aware temporal information into Vision Transformers. Experiments on four fine-grained sports benchmarks show consistent improvements over fully supervised and semi-supervised baselines. Under 10\% supervision on FSPerf, our method improves mAP by 4.54 points over the strongest competing approach, and with only 80\% labeled data, it matches or surpasses the fully supervised 100\% baseline on two of the four datasets.
arXiv:2607.10867v1 Announce Type: cross
Abstract: GroupFunctions.jl is a Julia library for computing individual matrix elements of irreducible representations of U(d). These matrix elements, called group functions, can be evaluated symbolically or numerically. For SU(2), they reduce to the Wigner D-functions. The library computes these matrix elements in a carrier-space basis enumerated by Gelfand-Tsetlin patterns. It can also compute entire representation operators, construct input unitaries from parameterisations common in quantum optics, translate Gelfand-Tsetlin patterns into occupation-number kets, and compute the associated Schur functions. Results can be exported in a form compatible with Mathematica.
arXiv:2607.10336v1 Announce Type: new
Abstract: This letter presents PrismAD, a decoupled end-to-end autonomous driving framework based on a Semantic Mixture-of-Planners. Existing planners usually aggregate heterogeneous scene tokens into a coupled representation space, forcing a single planning branch to jointly model agent interaction, road geometry, and driving intention. Such coupling may weaken factor-specific reasoning and obscure the contribution of different planning cues. To address this limitation, PrismAD partitions scene tokens into interaction, geometry, and intent groups, and assigns them to independent planning experts with the same architecture but separate parameters. Each expert learns a specialized motion-planning representation, while a semantics-aware router adaptively aggregates expert predictions with separate routing weights for motion prediction and ego planning. Sparse top-$K$ activation with noisy gating is further introduced to improve routing robustness and reduce unnecessary expert computation. Extensive experiments on the nuScenes open-loop dataset and NeuroNCAP closed-loop benchmark demonstrate that PrismAD exhibits competitive performance. Our code will be released soon.
arXiv:2607.10698v1 Announce Type: new
Abstract: We study the modality gap in CLIP-style dual-encoder contrastive learning, where image and text embeddings remain misaligned despite being trained in a shared space. We argue that the gap is induced by a failure of the InfoNCE formulation with independent encoders. We conduct a uni-modal experiment with two independent encoders and identical initialization conditions and find that InfoNCE actively generates a gap at low temperatures. We provide a theoretical analysis of this phenomenon and show that the modality gap is indeed a mode-failure of InfoNCE, but only at low temperatures. We propose a simple modification called xNCE, which uses intermodal as well as intra-modality negative contrastive pairs. xNCE matches retrieval performance on MS-COCO while consistently reducing the gap even at low temperatures. Notably, xNCE improves zero-shot classification over the InfoNCE baseline across all benchmarks, whereas high-temperature InfoNCE and regularized InfoNCE both fail to do so, demonstrating that xNCE reduces the modality gap without sacrificing the discriminative geometry needed for transfer.
arXiv:2607.11186v1 Announce Type: new
Abstract: X-cut thin-film lithium tantalate (TFLT) offers a unique combination of third nonlinearity, electro-optic effects, and a high optical damage threshold. However, its strong Raman response has historically hindered broadband Kerr comb generation. Here, we leverage this inherent Raman response by engineering coupling-defined dissipation. This allows us to reconfigure the relative thresholds of Raman and Kerr processes without modifying the intrinsic microresonator dispersion. Through this coupling-engineered threshold control, we can deliberately access distinct comb states, ranging from pure Kerr combs to Raman-Kerr synergistic broadband combs. We demonstrate a Kerr comb spanning 450 nm and a Raman-Kerr comb spanning 650 nm, representing the broadest combs reported to date on X-cut TFLT platforms. Moreover, in strongly coupled devices, we show that a single near-infrared pump can generate visible emission across multiple bands (from violet to red) via cascaded second sum-frequency processes. Our work demonstrates that a strong Raman response can be transformed from a parasitic competitor into an enabling mechanism for achieving broader comb spectra and generating polychromatic visible light. This work establishes X-cut TFLT as a powerful monolithic platform for nonlinear light sources, electro-optic functions, and complex photonic systems.