arXiv:2607.03668v1 Announce Type: cross
Abstract: Spectral analysis using linear mixture (LM) and radiative transfer-based (RT) intimate mixture modeling based on Hapke theory at near-infrared wavelengths are applied to estimate the abundance of surface materials on Europa. Previously, Emran (2026) compared these approaches against the laboratory spectra of H$_2$O ice and H$_2$SO$_4$$\cdot$8H$_2$O mixtures with $\sim$100 $\mu$m grains. Here, the effect of particle size on spectral modeling accuracy was assessed using laboratory spectra of H$_2$O ice mixtures with small ($\sim$70 $\mu$m spherical) and coarse ($\sim$1 mm irregular) grains, measured over the $\sim$1.2-2.5 $\mu$m wavelength range at 100 K and 120 K (Stephan et al., 2021). Modeled abundance estimates at both temperatures show consistent trends across all mixing ratios, with only minor temperature-dependent variations. The discrepancy in abundance estimates from both LM and RT models remains within $\pm$10% across all mixtures, with the error reduced to $\pm$5% when fine grains dominate. Across all mixtures, the average difference between RT- and LM-derived abundance estimates remains within $\pm$2% for mixtures containing both small and large grains. In contrast, mixtures composed solely of smaller grains render larger deviations between the models, with RT producing more accurate estimates (Emran, 2026) -- indicating that the presence of coarse H$_2$O ice grains minimizes abundance differences between LM and RT modeling. Thus, I posit that Hapke-based RT modeling is the preferred spectral modeling approach -- regardless of grain size or compositional mixture -- for constraining Europa's surface composition. Nonetheless, LM modeling remains a reliable approach for compositional analysis of terrains containing H$_2$O ice with $\sim$mm-sized grains.
Science Journals
arXiv:2607.03813v1 Announce Type: cross
Abstract: Spatial tumour--immune heterogeneity is a key feature of solid-tumour progression, immune infiltration, and immune exclusion. We develop a computational oncology model in which tumour cells, immune effector cells, and a chemokine signal interact through a reaction--diffusion--chemotaxis system on a bounded tissue domain with no-flux boundaries. Chemokine is produced by tumour cells and tumour--immune contact, recruits immune cells, and guides chemotactic migration. After nondimensionalization, we establish positivity, a tumour-density bound, and immune/chemokine mass estimates. We identify the tumour-free equilibrium, derive the immune-control threshold $\sigma_0>\delta$, and reduce coexistence to a scalar equation. Linear stability analysis about coexistence yields a mode-wise dispersion relation in which chemotaxis appears as a wavenumber amplified coupling, producing finite-wavelength instability above a critical sensitivity. A conservative finite-volume scheme with upwind chemotactic flux verifies the thresholds, dominant unstable modes, sensitivity maps, positivity, convergence, and residual consistency.
arXiv:2607.05189v1 Announce Type: new
Abstract: Persistent personal agents combine long-term memory with access to users' external environments, enabling personalized foreground assistance and proactive background execution. This integration also creates a new path to compromise: untrusted external content can be silently written into persistent memory and later reused as trusted state. We study this threat as stealth memory injection, in which a remote black-box adversary delivers a single email payload that must induce the agent to write poisoned memory, stay hidden in the agent's response to the user, and affect future behavior.
We introduce WhisperBench, a 108-case benchmark spanning five risk categories and both fact and preference poisoning. Built on a real IMAP/SMTP workflow and an authentic email agent skill, it enables full-cycle evaluation of stealth memory injection attacks. To enable this black-box attack under single-email delivery and without runtime feedback, we propose MemGhost, a one-shot payload generation framework. MemGhost uses an environment proxy to emulate persistent-agent execution and an objective proxy to convert memory adoption and conversational stealth into dense rubric-based rewards, then trains the attacker policy with supervised fine-tuning and reinforcement learning.
Across 56 held-out test cases, MemGhost achieves 87.5% end-to-end success on OpenClaw with GPT-5.4 and 71.4% on Claude Code SDK with Sonnet 4.6. It also transfers across personal-agent architectures (NanoClaw and Hermes Agent) and memory backends (filesystem and vector-based Mem0), and remains effective against input-level, model-level, and system-level defenses. These results suggest that persistent memory can turn ordinary external processing into a practical pathway for long-term agent compromise.
arXiv:2607.04425v1 Announce Type: new
Abstract: Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-platform interaction. However, building multi-platform GUI agents remains challenging. On one hand, high-quality and executable cross-platform interaction trajectories are still scarce, and existing data often suffer from limited platform coverage. On the other hand, different platforms exhibit distinct interaction conventions, making joint or continual training prone to behavioral pattern mixing, platform-specific capability degradation, and catastrophic forgetting. To address these challenges, we construct Uni-GUI, a high-quality cross-platform GUI interaction dataset, and propose UI-MOPD, the first method that incorporates multi-teacher on-policy distillation into continual learning for GUI agents. UI-MOPD dynamically selects a platform-specific teacher according to the current environment and transfers platform-specific behavioral priors to a shared policy through platform-conditioned distillation, enabling adaptation to new platforms while preserving capabilities on existing ones. Experiments on OSWorld and MobileWorld show that UI-MOPD achieves task success rates of 38.2% and 12.0%, respectively, demonstrating its effectiveness in balancing cross-platform capability retention and new-platform adaptation.
Project page: https://elispectre.github.io/UI-MOPD/.
arXiv:2607.04437v1 Announce Type: new
Abstract: Using fully kinetic simulations that capture unprecedentedly large (from electron to ion) scales, we study magnetogenesis driven by continuous large-scale forcing until nonlinear dynamo saturation. We uncover a two-stage mechanism in collisionless ion-electron plasmas whose dynamics diverge dramatically from the pair-plasma case. In the first phase, electron pressure anisotropy triggers electron-Weibel modes, seeding small-scale magnetic fields. Then, a second growth phase emerges when the more massive ions develop their own strong anisotropy and drive ion-Weibel-type modes; concurrently, a Biermann-battery mechanism contributes to amplifying the magnetic field. This combined dynamics provides a tenfold amplification of the magnetic field in comparison to the pair-plasma case. Over long times, dynamo action continues until the system reaches a statistical steady state. This self-consistent kinetic mechanism provides a plausible explanation for robust magnetogenesis wherever an external forcing continuously stirs the plasma.
PTCOG Treatment Efficiency Subcommittee Risk Assessment Report on Patient-Specific Quality Assurance
arXiv:2607.04446v1 Announce Type: new
Abstract: Patient-specific quality assurance (PSQA) in pencil beam scanning proton therapy (PBS-PT) is often treated as a purely technical verification task. This PTCOG Treatment Efficiency Subcommittee White Paper instead frames PSQA as a workflow-embedded risk-control strategy and asks how different PSQA approaches reshape the same clinical risk landscape. Using a generic PBS-PT process-driven Failure Mode and Effects Analysis (pFMEA), 44 validated PSQA-relevant failure modes across 20 process steps were scored under a common no-PSQA baseline and three PSQA pathways: measurement-based PSQA, log file-based PSQA, and independent secondary dose calculation.
A staged mathematical formalism separates preparatory data-stage effects, method-specific full-stage verification, cumulative endstate effects, and a Data-to-Cum bridge that quantifies additional verification benefit on the baseline scale. In this expert-scored, baseline-anchored model, log file-based PSQA produced the largest cumulative workflow-level risk-score reduction, followed by measurement-based PSQA and independent secondary dose calculation. The ranking is not a winner-takes-all rule or probability-calibrated risk estimate; instead, each method shows distinct risk-control strengths in different workflow regions.
The White Paper therefore supports a risk-informed hybrid PSQA architecture, where log file-based PSQA, measurement-based PSQA, and independent secondary dose calculation are assigned to the workflow segments in which their signatures are strongest. It provides a transparent, semi-quantitative, stage-resolved framework for institutions seeking to evaluate, implement, or evolve PSQA in PBS-PT and emphasizes that log file-based PSQA must itself be supported by validated and governed log data and treatment records.
arXiv:2607.04403v1 Announce Type: cross
Abstract: Binary change detection in remote sensing requires both complete changed-region localization and accurate boundary delineation. We present MambaRefine-CD, a region-boundary temporal refinement framework built on a shared MambaVision encoder. The proposed D-RBI module constructs temporal evidence from paired features, absolute differences, and signed differences, then separates it into region and Sobel-conditioned boundary streams. Region features are enhanced with CRAM-lite and decoded by an adaptive receptive-field FPN, while the finest boundary stream guides a bounded residual refinement of the coarse prediction. Experiments on DSIFN-CD and WHU-CD show strong changed-class F1 and IoU under verified evaluation settings, and ablations support the contribution of signed temporal evidence and the full region-boundary refinement pipeline.
arXiv:2607.04451v1 Announce Type: new
Abstract: Evaluation of autonomous vehicle (AV) planners in safety-critical closed-loop simulation is essential for real-world deployment. However, generating controllable safety-critical scenarios remains challenging. Existing approaches use soft guidance that provides only probabilistic preferences and cannot guarantee the satisfaction of geometric and severity constraints associated with specific collision types. We introduce Collision-Constrained Flow Matching (CCFM), a novel framework that guarantees precise collision control through hard physical constraints. CCFM consists of three key components: (i) a heuristic collision selector that optimally identifies an adversarial agent and collision type via composite scoring; (ii) structured hard constraints that explicitly define four collision types (rear-end, side, cut-in, head-on) through contact point, heading, and severity requirements; and (iii) a collision-constrained flow matching sampler that enforces the constraints via Gauss-Newton manifold projection. CCFM achieves collision rate up to 46.4% on nuScenes and 83.1% on nuPlan, significantly outperforming baselines while preserving realistic driving behavior. By enabling controllable collision characteristics in safety-critical scenario generation, CCFM provides a reliable foundation for AV safety evaluation and sim-to-real crash data generation. The code and implementation details are available at https://github.com/KELISBU/CCFM.
arXiv:2607.04453v1 Announce Type: new
Abstract: The assessment of planktonic standing stocks and microorganism structures is critical for understanding upper ocean biological processes. Currently, autonomous underwater vehicles (AUVs) equipped with in-situ optical imaging and artificial intelligence (AI) methods offer a promising solution for persistent surveillance, mapping and monitoring of planktonic life. However, current AI methods often lack robustness in dynamic, unstructured environments, where environmental noise and non-biological artifacts lead to frequent misclassifications. Standard convolutional neural network (CNN) classifiers often struggle with such conditions, leading to misclassifications that require time-consuming manual validation by marine biologists.
To address this issue, we propose a novel robustness verification framework for in-situ plankton classifiers based on reachability analysis. We also introduce a continuous-time neural ordinary differential equation (neural ODE) classification model leveraging the high-resolution imaging capabilities of the SilCam particle imager. In this paper, we demonstrate the effectiveness of the proposed framework by formally verifying the robustness of the neural ODE model against environmental perturbations. We demonstrate that our verification framework acts as an automated filter providing formal guarantees of model stability against ambiguous data, thereby improving the reliability of autonomous sampling and reducing the post-processing workload.
arXiv:2607.04515v1 Announce Type: new
Abstract: Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end text-to-speech study for Efik, introducing a curated single speaker corpus of 2,632 utterances totaling three hours and a comparative evaluation of four neural models (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under low resource conditions. Native speakers evaluated the systems using MOS, Nat-MOS, and A-MOS. MMS-TTS achieved the highest MOS of 3.80 +/- 0.63 and produced more stable long form speech, though tonal errors persisted. Other models showed greater tonal and prosodic inconsistencies. These results provide a reproducible baseline and highlight the need for larger corpora and tone aware modeling for tonal African languages.
arXiv:2607.03881v1 Announce Type: cross
Abstract: Codon harmonization aims to adapt the coding sequences for heterologous expression while preserving the native-like patterns of frequent and rare codons that may influence local translation dynamics and co-translational protein folding. However, widely used harmonization metrics, such as $\%$MinMax, are defined on discrete codon sequences and are, therefore, not readily compatible with gradient-based neural codon design. Here, we introduce Smooth $\%$MinMax, denoted as $\%{\rm MinMax}_{(s)}$, a differentiable relaxation of the conventional hard $\%$MinMax metric, denoted as $\%{\rm MinMax}_{(h)}$. $\%{\rm MinMax}_{(s)}$ replaces the discrete codon-usage values with probability-weighted synonymous-codon usage values and replaces the hard $\%$Max/$\%$Min branch with a sigmoid-gated interpolation. This formulation preserves the signed interpretation of $\%{\rm MinMax}_{(h)}$, while enabling optimization with respect to the synonymous-codon probabilities and learnable parameters. In human-to-Escherichia coli codon harmonization experiments, $\%{\rm MinMax}_{(s)}$ closely approximates $\%{\rm MinMax}_{(h)}$ and supports gradient-based profile matching in synonymous-codon probability space. These results suggest $\%{\rm MinMax}_{(s)}$ as a practical bridge between profile-based codon harmonization and neural synonymous-sequence design.
arXiv:2607.04547v1 Announce Type: new
Abstract: People respond to artificial intelligence chatbots (AICs) in highly variable ways. In this paper, we adapt Bronfenbrenner's theory into a heuristic framework for understanding this variation. The framework places the human user at the center while also placing the AI there and reconceptualizing the proximal processes as the repeated, reciprocal, and coadaptive interactions between the user and a personalized AIC. The surrounding systems identify the contextual factors that shape how the user experiences, interprets, responds to, and is changed by these interactions. Because stateful AICs learn from accumulated exchanges with their users and have memory, users are responding not only to an AIC but also to a version of the AIC that their own prior interactions have helped create. This extension preserves Bronfenbrenner's emphasis on proximal processes while accounting for the unique dynamics of personalized AICs. The resulting framework provides a structured map of where and how variation in human and AIC relationships arises, as well as having implications for researchers, practitioners, and AIC designers.
arXiv:2607.03931v1 Announce Type: cross
Abstract: Whole-body fluorodeoxyglucose positron emission tomography combined with computed tomography is widely used in cancer care, but manual lesion delineation is slow, subjective, and difficult to scale. We present GLOW-FDG, an open-source artificial intelligence model for whole-body cancer lesion segmentation in fluorodeoxyglucose positron emission tomography and computed tomography. The model was trained on 1,563 scans spanning multiple cancer types and evaluated on 185 external scans from independent institutions. Across breast cancer, nonmetastatic and oligometastatic lung cancer, head and neck cancer, and metastatic melanoma, GLOW-FDG consistently outperformed publicly available benchmark models in lesion detection, while reducing false positives and maintaining strong segmentation accuracy. Quantification of total tumor burden and total lesion glycolysis was robust across cohorts, and performance approached the variability observed between expert radiation oncologists. These results support GLOW-FDG as a generalizable tool for automated cancer segmentation and quantitative imaging biomarker extraction in whole-body imaging.
arXiv:2607.04520v1 Announce Type: new
Abstract: Active mobility is widely promoted for sustainable and healthier living, but whether it translates into equitable mental health benefits across individuals and places over time remains unknown. Using causal machine learning and causal deep learning in 264168 UK adults, we find substantial inequalities in individualized effects of active mobility on anxiety, depression, and common mental disorders. These inequalities widen over time and are strongly structured by urban context. For example, anxiety risk at follow-up ranges from a 40.6% reduction to a 10.1% increase across individuals, versus a 10.4% reduction to a 0.1% increase at baseline. Benefits are greatest in greener, safer, less polluted, and less deprived neighborhood environments, with 81.8% of individuals experiencing above-average benefits and mean anxiety risk reduced by 26.4%, versus 10.4% of individuals and 7.4% reduction in the least supportive environments. Urban compact form further modifies these effects through nonlinear interactions with neighborhood environments, amplifying benefits only under supportive conditions. Despite these strong environmental gradients, genetic moderation is negligible. These findings suggest universal active mobility promotion could widen health inequalities if individual and contextual differences are not accounted for.
arXiv:2607.04551v1 Announce Type: new
Abstract: Glioblastoma progression is strongly influenced by evolving mechanical interactions between the tumor and surrounding brain tissue. However, the extent to which finite-deformation mechanics and constitutive assumptions improve subject-specific prediction as tumor burden evolves remains unclear. We introduce a sequential Bayesian inference and dynamic model selection framework that assimilates longitudinal murine magnetic resonance imaging (MRI) data to calibrate spatially varying tumor diffusivity, proliferation rate, and tissue stiffness in biomechanical tumor growth models. Competing formulations were compared at each imaging time, including reaction-diffusion without mechanics and reaction-diffusion coupled to linear elasticity or hyperelastic mechanics, using posterior model plausibility to adapt model choice for individualized one-scan-ahead prediction as new MRI scans are acquired. Across the studied animals, mechanically coupled models were consistently more plausible than the uncoupled reaction-diffusion model, and the evolution of model plausibility indicated an increasing role of mass effect and stress-mediated feedback of tumor growth during progression. While linear and hyperelastic coupled tumor growth models often produced similar tumor morphology, they yield distinct stress, deformation, and inferred stiffness fields, with the hyperelastic formulation often receiving higher posterior plausibility at later imaging times. These results indicate that, within the present longitudinal murine dataset, mechanical coupling is favored for image-informed glioma growth prediction and that constitutive assumptions should be evaluated sequentially for each subject rather than fixed a priori.
arXiv:2607.04141v1 Announce Type: cross
Abstract: We introduce split-free cable terms and cable plays, a sequential graph-construction language whose live cables impose uniform GF(2)-row behaviour across the current cut. Every play of width w gives a birth-order layout whose cutrank is at most half of w, rounded down, so the sequential split-free width is at least twice the linear rank-width. At the first nontrivial level we prove an exact characterization: a connected graph with at least two vertices has linear rank-width at most one exactly when it admits a stream, equivalently a singleton-birth play of width at most four. We show that unrestricted term width and sequential width differ unboundedly on trees, calibrate the construction on the net graph, and formulate an affine upper-bound conjecture relating sequential split-free width to linear rank-width. For the rank-two case we prove a two-accumulator scheduling criterion that yields width-six plays under a natural future-uniformity hypothesis.
arXiv:2607.04579v1 Announce Type: new
Abstract: CI/CD workflows have become executable operational policy: they decide what gets built, tested, released, and deployed, and they mediate how maintainers interact with delivery infrastructure. That makes them an important measurement point for cyber-systems engineering. Recent large language model (LLM) work shows that workflow stages can be recognized directly from configuration files, but stage labels alone do not tell us whether a workflow is brittle, unusual for its ecosystem, or worth revising first. We present an LLM-based CI/CD analysis pipeline that combines repository enrichment, anti-pattern detection, stage mining, and recommendation generation over a large GitHub corpus. Starting from 59,550 repositories with at least 1,000 stars, we identify 34,225 projects with CI/CD and collect 127,559 configuration files. Across 75,201 analyzed workflows, the anti-pattern detector reports 434,769 findings, dominated by reliability and maintainability issues. Across 59,906 configurations, stage usage differs significantly by language ($\chi^2 = 4168.88$, $p < 0.001$, Cramer's $V = 0.063$), and domain analysis shows distinct operational profiles, including higher release and cache usage in mobile projects. For repository-level recommendation generation, few-shot prompting performs best overall, averaging 8.25 recommendations per repository with 96.1% YAML-valid snippets. Taken together, the results argue for CI/CD observability that combines diagnosis, context, and human review rather than treating workflow mining as a stage-classification problem alone.
arXiv:2607.04803v1 Announce Type: new
Abstract: Autonomous driving research has largely focused on safety while giving limited attention to non-functional aspects such as energy consumption and sustainability. As Autonomous Electric Vehicles (AEVs) become increasingly common in urban traffic, understanding how complex traffic dynamics influence their energy consumption is paramount to test whether AEVs can complete trips before battery depletion. To support energy-aware scenario-based testing of AEVs, we present E-CoDrive, a framework for reproducible closed-loop driving co-simulations that integrates an energy consumption model, a micro-traffic simulator, and a high-fidelity driving simulator to test AEV software stacks in urban scenarios. This tool paper describes the architecture of E-CoDrive and demonstrates its applicability by testing an Autoware-based AEV stack. Our evaluation shows that varying traffic conditions produce substantial differences in vehicle energy consumption. The artifact is publicly available at https://doi.org/10.6084/m9.figshare.32244783, and a screencast showing the tool is available at https://youtu.be/yX9fWHqCvgc.
arXiv:2607.04123v1 Announce Type: new
Abstract: Inverse design of mechanical metamaterials seeks a periodic unit cell whose homogenized elastic properties meet a prescribed target, but current learning-based methods are data-hungry, mostly interpolative, and provide no guarantee that the generated design satisfies the specification. We introduce CertMix, a data-efficient framework that represents each exemplar unit cell as a small periodic neural implicit field, specifically a SIREN signed-distance decoder overfit from a shared anchor, so that exemplar weight vectors become aligned and directly comparable. The key observation is that, in this aligned weight space, the homogenized elasticity tensor is approximately linear in the mixing coefficients. Targeted design therefore reduces to a small constrained affine-mixing problem solved with a differentiable periodic homogenizer in the loop. Negative coefficients enable extrapolation beyond the exemplar range, a linearity-mismatch trust region keeps blends valid, and split-conformal calibration converts the mismatch signal into a distribution-free certificate on achieved-property error. From as few as 50 exemplars, CertMix attains a scaled property error of $10^{-4}$, roughly two to three orders of magnitude below conditional generative baselines trained on 1000 cells. It remains accurate far outside the exemplar range, is $57\times$ faster than per-target topology optimization while avoiding checkerboards and enclosed voids, and extends to spatially graded fields, 3D triply periodic surfaces, and a certified running-shoe midsole application.
arXiv:2607.04600v1 Announce Type: new
Abstract: Graph eXplainable AI (G-XAI) is increasingly important for making Graph Neural Networks interpretable and accountable. While a growing number of explainers are available, choosing the right method and assessing the trustworthiness of its outputs remains unclear. Consistent evaluation practices and actionable guidance are still missing, hindering practical adoption. In this paper, we introduce a unified, quantitative benchmarking framework for G-XAI that requires no ground-truth assumptions. We formalize tabular explainability metrics for graph data, evaluating topological structure and node features as independent components. Our large-scale benchmarking study identifies explainers that consistently lie on the Pareto front across metric pairs and tasks, establishing robustly non-dominated solutions - while confirming that no single explainer achieves universal superiority. We distill our findings into actionable G-XAI usability guidelines to support Machine Learning practitioners in evaluating and deploying trustworthy GNN-based pipelines.
arXiv:2607.04615v1 Announce Type: new
Abstract: This paper presents a novel acoustic-visual-inertial odometry solution leveraging a continuous-time trajectory estimation framework for unmanned underwater vehicles. Underwater environments present unique challenges for visual localization and mapping, such as light attenuation, illumination variance, and the presence of particulate matter. This motivates the use of additional sensing modalities and a visual tracking pipeline that is robust to diverse subsea conditions. The proposed system is the first continuous-time trajectory estimation framework based on Gaussian processes to fuse asynchronous measurements from a Doppler velocity log, a stereo camera, and an inertial measurement unit. Additionally, a novel visual frontend is proposed, incorporating learning-based feature extraction and matching that is robust to the specific challenges that subsea environments present. The proposed framework enables seamless integration of additional sensor modalities in continuous-time and is adaptable to different environments without reconfiguration. The proposed system is extensively tested on real-world underwater inspection datasets, where it outperforms state-of-the-art visual-inertial and acoustic-visual-inertial SLAM algorithms in accuracy, robustness, and trajectory coverage. Notably, the proposed system outperforms the state-of-the-art despite only forming short-term visual data associations.
arXiv:2607.04636v1 Announce Type: new
Abstract: Key Information Extraction (KIE) converts visually rich documents into structured data, but practical deployment remains challenging: strong performance often relies on costly on-server Large Multimodal Models (LMMs), while compact locally deployable models lack sufficient KIE supervision. We present SAYRE, a scene-aware document synthesis framework for generating scalable KIE training data without hand-crafted template design. Given a few exemplar documents, SAYRE captures category-specific content patterns and layout conventions to synthesize document-schema-annotation triples. It further introduces error-driven generation, which expands real-world failure cases into hard training examples while preserving their structural patterns. Experiments on constrained- and open-category KIE show that SAYRE consistently improves Qwen3-VL backbones and achieves the strongest overall performance among on-device LMMs. Data scaling experiments show an overall upward trend as more synthesized data is introduced, especially for smaller models and open-category extraction. Error analysis further shows that synthesized training reduces field-level errors by improving schema-aware extraction over dense tables, business identifiers, and contract clauses. These results establish scene-aware synthesis as an effective data-centric approach for improving practical multimodal KIE.
arXiv:2607.04637v1 Announce Type: new
Abstract: Vision-Language-Action Models (VLAs), which leverage the advanced reasoning capabilities of Vision-Language Models (VLMs), show promising generalization in complex autonomous driving scenarios. Existing VLAs typically predict and optimize 3D trajectories from 2D images. While intuitive, this 2D-to-3D prediction is inherently entangled with camera parameters, leading to limited data scalability across heterogeneous driving datasets. Moreover, directly optimizing in 3D space induces severe convergence to trivial solutions, where VLAs rely on ego-status rather than visual scene understanding. To address these issues, we propose PixelPilot, a novel VLA featuring a decoupled planning and lifting paradigm. In the planning phase, PixelPilot reformulates scene understanding and trajectory prediction as sensor-agnostic 2D-to-2D tasks in the image plane, thereby facilitating scalable training across diverse datasets. The planned 2D trajectories are then deterministically lifted to 3D only during inference, ensuring the full exploitation of visual cues and generalization across different vehicles. To realize this paradigm, we propose a knowledge-instilled policy learning strategy that applies dense, intermediate rewards via Group Relative Policy Optimization (GRPO) to enforce a rigorous causal chain from visual perception to spatial planning. Extensive experiments demonstrate that PixelPilot achieves state-of-the-art performance in both open-loop and closed-loop settings, validating its superior scalability and visual reasoning capabilities.
arXiv:2607.05353v1 Announce Type: new
Abstract: Watermarking methods embed imperceptible and verifiable signals into text generated by large language models (LLMs). Existing approaches include zero-bit schemes for distinguishing synthetic text from human writing and multi-bit schemes for embedding metadata. However, current multi-bit watermarking methods do not allow selective disclosure: verifying any part of the watermark requires revealing the entire embedded message. This lack of control leads to unnecessary information exposure and raises privacy concerns. We propose Hierarchical Vocabulary Routing (HeRo), a watermarking framework that enables selective disclosure of embedded metadata. The method recursively partitions the vocabulary and distributes watermark information across hierarchical layers, so that different verifiers can decode only the portions of the payload corresponding to their access level. We show that the proposed scheme preserves the unbiasedness of the underlying sampling process and thus maintains text quality. Experiments demonstrate that our framework supports fine-grained access control while achieving high detection accuracy and low latency. Code is available at https://github.com/xuyangc03/hero-watermark.
arXiv:2607.04638v1 Announce Type: new
Abstract: Referring expression comprehension (REC) aims to localize the object in an image described by natural language. In Weakly supervised REC (WREC), existing approaches primarily operate on anchor-level visual representations. Even when enriched with auxiliary cues, relational interactions remain implicitly encoded within individual anchor features. The resulting visual representation remains flat and unary-only, limiting its ability to align with the structured nature of language. In this work, we propose a Structured Visual Compositional Representation (SVCR) learning framework for WREC. Rather than implicitly encoding relations within unary anchors, the proposed SVCR explicitly models both unary object embeddings and pairwise relational embeddings, forming a structured visual representation space. We further introduce a compositional alignment mechanism that matches unary and pairwise visual representations with their corresponding textual embeddings in a unified manner, enabling compositional visual-textual matching under weak supervision. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg show that the proposed SVCR achieves state-of-the-art performance. These results demonstrate the effectiveness of explicit structured visual representations and visual-textual alignment for WREC.