arXiv:2607.01670v1 Announce Type: new Abstract: Day-ahead wind power forecasting is essential for cost-effective power-system operation. It is primarily driven by future meteorological conditions while retaining temporal dependencies in power generation. In practice, observed wind-farm power often entangles physically available power with local environmental effects and latent operational states, such as shutdowns and curtailment. Existing physical models provide useful constraints but adapt poorly across wind farms, whereas data-driven models can capture rich correlations but often conflate meteorological effects with state-induced deviations. In this study, we propose UniWind, a wind power forecasting model based on physics-informed state routing. UniWind first employs a Physical Prior Estimator to construct a site-calibrated physical prior by combining site-conditioned monotonic warping with a shared physical power curve. It further applies a physical upper-bound constraint to shape this prior as a soft envelope of available wind power generation. UniWind then proposes a Latent State Encoder to model operating-state embeddings and transforms the physical prior into final power forecasts through a State-aware Power Corrector, which uses knowledge-guided supervised state routing and bounded, state-specific expert correction. Full-shot and cross-farm zero-shot experiments on more than 20 real-world datasets demonstrate the accuracy and robustness of UniWind.
Science Journals
arXiv:2607.01671v1 Announce Type: new Abstract: Self-reference and solution independence are core properties underlying intractability. This paper establishes a finite combinatorial analogue of G\"odel's incompleteness theorems within Boolean $K$-SAT. While standard random $K$-SAT has assignment correlations that disrupt solution independence, we resolve this via a logarithmic-width ensemble ($K = O(\log N)$). Here, satisfying assignments converge to a Poisson distribution, letting unsatisfiable and uniquely satisfiable formulas coexist. By executing a single-clause substitution conditioned on the unique solution, we construct structurally irreducible SAT/UNSAT pairs that are indistinguishable via local evaluation. Using algorithmic information theory and Shannon channels, we prove that deductive pipelines restricted to a sublinear window suffer from an informational blind spot, forcing a descriptive lower bound of $K(\mathcal{A}) \geq \Omega(N^{1-\delta})$. This deficit forces any Resolution refutation of the UNSAT instance to utilize wide clauses ($w(\pi) \geq \Omega(N^{1-\delta})$), triggering an exponential proof-tree explosion ($S(\phi) \geq \exp(\Omega(N^{1-2\delta}))$). As $\delta \rightarrow 0^+$, this bound converges to the worst-case $2^N$ threshold, reframing the Strong Exponential Time Hypothesis (SETH) as a direct projection of G\"odel incompleteness onto finite computation. We diagnose the decades-long stagnation in complexity theory. Transitioning from Turing's class separation to a G\"odelian paradigm of instance indistinguishability, we introduce a multi-dimensional comparative framework that contrasts these two historical lineages across distinct perspectives. The self-referential hardness exhibits physical invariance: it precludes quantum shortcuts due to the necessity of global semantic analysis and delineates a scaling bottleneck for machine learning architectures operating on lossy, local compression.
arXiv:2607.01701v1 Announce Type: new Abstract: The rising demand for AI-generated videos is fueled by advances in large-scale Text-to-Video (T2V) models, trained on extensive datasets of video clips spanning diverse resolutions and durations. To address this data heterogeneity, current training methods often use a bucketing strategy that groups samples into discrete buckets for efficiency. However, this approach struggles to scale with compute and data volumes under static parallelism schemes, such as data and sequence parallelism, leading to significant workload imbalances and hardware under-utilization. In this paper, we present Arachne, a novel training framework for efficient T2V model training at scale. Arachne decomposes the training process into fine-grained computational units, called \textit{cascades}, orchestrating their distributed execution and synchronization across the cluster through coordinated spatial and temporal optimization. Our comprehensive evaluation demonstrates that Arachne reduces iteration time by up to 65\% over leading frameworks, exhibiting a positive scaling trend where its performance advantages amplify as training scale grows.
arXiv:2607.02074v1 Announce Type: new Abstract: Recent advancements in LiDAR-only 3D object detection have demonstrated improved detection accuracy over benchmark datasets. However, the adversarial robustness of these models remains untested. Very few adversarial robustness studies exist for LiDAR-only 3D object detection and unfortunately, even they are limited to legacy models. Moreover, there is a systemic gap in the existing evaluation frameworks that rely simply on mAP ignoring other structural and predictive factors. To fill this gap, we propose a holistic framework that evaluates adversarial robustness using two structural factors (point cloud density and point cloud localization) and three predictive factors (misclassification, localization error, distance from ego). Using this framework, we perform an empirical study and critical analysis on recent and legacy state-of-the-art models using adversarial attacks specifically designed for LiDAR-based models. Our key finding is that high-capacity, voxel-based detectors are more susceptible to structured coordinate perturbations than pillar-based detectors. Additionally, non-anchor-based detectors demonstrate poor adversarial robustness, which necessitates rethinking model training techniques. Overall, our results demonstrate that recent models are as vulnerable to adversarial attacks as their predecessors. Therefore, we argue that there is a need to improve the evaluation benchmarks for 3D object detection that not only reward architectural modifications for improving detection accuracy, but also evaluate whether the design choices improve adversarial robustness.
arXiv:2607.01253v1 Announce Type: new Abstract: Rationale. The diagnostic radiologist's role in 2035 will not look like it does today. Imaging AI is already changing how worklists are organized, how reports are generated, and which cases require a radiologist's attention. What remains genuinely contested is not whether the role changes but how. Approach. Three subject-matter experts (two radiologists and one health tech professional with more than 20 years of experience in medical imaging IT) independently authored 2035 job descriptions for the diagnostic radiologist using a shared template. Each author wrote from a distinct vantage point: one optimistic, one framed as a trade-off view incorporating workforce economics, and one structured around professional stratification. The three versions were published openly and subjected to a structured comparison across seven dimensions. Key findings. The three versions agree on direction but disagree on magnitude. All three describe a radiologist whose routine workload is AI-managed, who carries accountability for AI output, and who spends more time on complex cases and clinical collaboration than today's radiologist does. They diverge on headcount, career security, and whether the profession expands broadly, concentrates into a smaller well-compensated group, or stratifies into sharply differentiated tiers. Conclusion. AI won't eliminate the diagnostic radiologist. Whether it expands, concentrates, or stratifies the profession depends on choices health systems haven't made yet. The clinical argument for optimism is real. So is the economic argument for caution. Both can be true simultaneously. Keywords: radiology workforce; artificial intelligence; diagnostic radiology; job redesign; medical imaging IT; AI governance
arXiv:2509.15919v2 Announce Type: replace Abstract: Experiments have shown that strong coupling between molecular excitations and a mode of a Fabry--P\'erot cavity can significantly alter molecular properties, such as reaction rates and equilibrium constants. However, in spite of the large body of theoretical work, the mechanism behind this change is still not well understood. In order to make progress, we first take a step back and investigate the appropriateness of the Hamiltonian that most recent studies are based on. In particular, we investigate the dipole self-energy, which can be divided into in self terms and cross terms. While the self terms are an indispensable part of the Hamiltonian, the cross terms -- which have received attention as they seem to mediate distance-independent interactions between all molecules in the cavity -- are known to, under certain conditions, cancel exactly with the usually neglected intermolecular Coulombic interactions. In this work, we revisit how this cancellation comes about in free space and in a perfect cavity, clarifying that it can only be found when looking beyond the single-mode approximation and taking the full continuum of light modes into account. We also provide numerical evidence suggesting that this cancellation may extend to the case of an imperfect cavity, and show how the situation changes for a more realistic cavity in the framework of macroscopic QED. Finally, we discuss the implications of this cancellation for the single-mode Hamiltonian.
arXiv:2607.02075v1 Announce Type: new Abstract: We present HandsOnWorld, a framework for hand-controlled egocentric video generation that forgoes multi-view and marker-based motion capture, learning instead from unconstrained monocular video. Such generality is bottlenecked by the scarcity of scalable 3D hand annotations: large egocentric corpora lack finger-level labels, whereas precise hand datasets are confined to narrow, instrumented settings, limiting prior hand-controlled generators to restricted scene distributions. We instead annotate 3D hands directly on in-the-wild egocentric video through monocular reconstruction, introducing a protagonist-centered annotation pipeline that filters the reconstructions at the action-semantic, image-quality, and 3D-geometric levels to build EgoVid-Pro, a dataset of clean, protagonist-only hand trajectories spanning 103K clips and roughly 12M frames across diverse everyday scenes. To resolve the camera-hand entanglement induced by large ego-motion, we further propose the Pl\"{u}cker Hand Map, a 3D-aware control signal that extends Pl\"{u}cker-ray representations from camera rays to the hand surface, disentangling camera and hand motion at the representation level. Experiments show that \method surpasses prior hand-controlled generators in reconstruction fidelity and control accuracy, and generalizes to out-of-distribution everyday scenes beyond the laboratory datasets on which prior methods rely.
arXiv:2607.01669v1 Announce Type: new Abstract: This work investigates the effect of batch sampling strategies during training for text-to-audio music generation under low-data and small-scale model settings. This paper describes our approach and findings for the ICME 2026 Grand Challenge on Academic Text-to-Music Generation. Training data are clustered using either text embeddings or audio embeddings, and samples with similar characteristics are grouped within the same mini-batch to mitigate gradient interference. The effects of modality and cluster granularity on clustering are analyzed. Results show that clustering based on text embeddings achieves better performance on objective evaluation metrics than clustering based on audio embeddings. In addition, different cluster granularity leads to different behaviors across evaluation criteria: a moderate number of clusters performs best on objective metrics, while a larger number of clusters tends to exhibit music with more coherent structure in listening tests.
arXiv:2607.01277v1 Announce Type: new Abstract: Large language models (LLMs) can be induced to produce harmful content through multi turn strategies in which no single user message appears clearly unsafe. Existing runtime safeguards commonly evaluate prompts or responses as isolated messages, which limits their ability to recover ac-cumulated intent, verify asserted authority, or detect harmful objectives decomposed across a dialogue. This paper presents the Cognitive Firewall, a proactive runtime oversight framework that interposes an independent oversight model between a user and a protected target mod l. The framework decomposes safety assessment into four categorical gates: an intent gate that identi-fies the operational objective of a request, a zero trust context gate that treats claimed roles and permissions as unverified evidence, a consistency gate that detects escalation and decomposition across turns, and an output risk gate that inspects candidate responses before release. Gate decisions are combined through escalation rather than score averaging, allowing any confident danger signal to block an interaction while preserving an auditable rationale. Experiments on four jailbreak benchmarks and a benign safety test set show that the Cognitive Firewall substantially reduces attack success across single turn, multi turn, authority based, and human crafted attacks. It lowers attack success to 2 percent or below on three attack sets and to 14 percent on the most difficult human crafted set, while maintaining an 8 percent over refusal rate. These results indicate that decomposed, conversation level oversight can improve proactive containment and auditability for LLM safety.
arXiv:2607.01296v1 Announce Type: cross Abstract: What would it be like to be in a superposition of yesterday, today, and tomorrow? This question may seem at best entertaining, but it is necessary, and exploring it allows us to understand how exact irreversible clocks and change are possible, despite the Unruh-Wald and Hegerfeldt-Ruijsenaars no-go theorems forbidding them. Unruh and Wald (1989) proved that if energy is bounded from below, no observable can increase monotonically with the Schr\"odinger time parameter t. Perfectly monotonic clocks and irreversible observable changes (Hegerfeldt-Ruijsenaars, 1980) seem impossible. From the perspective of the Schr\"odinger time, the world appears in a superposition of different intrinsic clock states indicating different times and opposite time directions. This seems to directly contradict our daily experiences of time and change. I show that there is no contradiction: from an intrinsic perspective of the world, sharp irreversible changes do happen, because the macroscopic pointer states resolve the superposition of different times. Large-scale time-reversing or discontinuous transitions are not internally observable in the records. From the intrinsic perspective, an unbounded intrinsic-time translation generator plays the role of the Hamiltonian, generating only forward time evolution with respect to the intrinsic time, but not to the Schr\"odinger parameter t, which is thus not justified to play the role of time. This allows sharp time observables even if the external Hamiltonian is bounded from below. In addition, this leads to a stationary wavefunction of the universe satisfying a Wheeler-DeWitt-type equation, without assuming gravity.
arXiv:2607.01280v1 Announce Type: new Abstract: Programming-by-example systems infer programs from a small set of input-output examples. Robust PBE work usually models wrong examples as samples from a stochastic noise process and then minimizes an expected or empirical loss. This paper studies a different failure mode: an adversary who sees the synthesizer and chooses the examples whose corruption most damages the returned program. We formalize fixed-set worst-case corruption for finite PBE version spaces, implement exact-within-bounded-pool and heuristic corruption searches for a string-transformation DSL, and introduce version-space partition aggregation (VPA), a defense that synthesizes on disjoint example groups and votes by semantic signatures. The central claim is deliberately bounded and partly negative: low-margin PBE tasks have an adversarial robustness dimension that random-typo and noisy-PBE evaluations miss, while semantic partition aggregation helps only when the clean semantics keep a partition vote margin, which often fails on realistic tasks. Evidence from curated/generated DSL tasks, accepted public SyGuS PBE_SLIA slices, SYNTRA Playgol v2, and noisy-PBE objective baselines supports that boundary. One curated edit flips all 8 spike tasks while 200-trial typo, DSL-pool, and distance-matched random controls succeed on 10.3%, 11.0%, and 16.7%; generated margin-1 rows flip under budget 1 yet VPA recovers them; on public SyGuS the vote margin is near one, so an adaptive attacker drives VPA accuracy to zero; accepted public SyGuS slices move across exact-within-pool budget boundaries; and Playgol shows positive paired-bootstrap gaps against typo and same-pool random controls on the 141 accepted rows. A small exact-output prompt harness over 20 controlled margin-1 tasks shows the same qualitative clean-to-attacked pattern across local and API models, while it is treated as a scope check, not a broad LLM benchmark.
arXiv:2607.02105v1 Announce Type: new Abstract: In this paper, we introduce a weighted derivative histopolation framework on families of intervals. The degrees of freedom consist of one scalar normalization and weighted integral moments of the derivative over a prescribed family of subintervals. We prove that the resulting scheme is unisolvent on $\Pi_N$ when the interval family separates polynomials of degree at most $N-1$ through weighted moments and the normalization is nonzero on constants. Thus, the derivative moments determine the polynomial up to an additive constant, and the scalar normalization fixes this remaining degree of freedom. This gives a sharp criterion for the well-posedness of the interpolation problem and a complete characterization of the admissible scalar normalizations. We then show how admissible families of intervals can be constructed from a fixed grid. When the endpoints of the intervals belong to the grid, admissibility is reduced to the nonsingularity of an interval matrix associated with the family, which depends only on the representation of the intervals in terms of consecutive cells. For Jacobi weights, the associated data matrices have a natural block structure in Jacobi polynomial bases, and the reduced derivative matrix can be expressed in terms of shifted Jacobi moment matrices. We next study Chebyshev configurations in which this structure becomes explicit. For the four classical Chebyshev families, suitable polynomial bases lead to diagonal Gram matrices for the reduced derivative matrices. We show that this diagonal structure depends on the simultaneous choice of the weight, the basis, and the grid. Numerical experiments on equispaced and Chebyshev--Lobatto nodes show the behaviour of the method for different interval families and for different Jacobi parameters.
arXiv:2607.02115v1 Announce Type: new Abstract: For a recommender service, we view the customer journey as a chain of item recommendations: a useful item changes the user's state and therefore what should be retrieved next. Standard matrix-factorization retrieval ignores this -- it builds one user vector and returns the top-$K$ items by a static score, treating them as independent. We ask a narrow question: when is it worth planning over the user-state dynamics that fold-in induces? To answer it we propose casting top-$K$ retrieval as an MDP over the implicit-ALS posterior $(A^{-1},u)$, where an action is an item and the transition is a closed-form rank-one fold-in, and the trajectory reward combines a relevance similarity with a posterior-alignment term. Under the same fixed embeddings we compare static retrieval, one-step planning, and horizon-$K$ MCTS across five datasets and two protocols: a per-user leave-last-$n$ split and a stricter global time split. Dynamics-aware planning tends to overcome static retrieval on all datasets under leave-last-$n$, and the gains hold on MovieLens-1M and the VK-LSVD slices under the global time split. A single step of lookahead already captures most of the gain, so the lightweight planning layer turns static top-$K$ scoring into a short decision and improves retrieval over fixed collaborative-filtering embeddings, with no retraining and no change to the representation. These gains depend on measuring relevance with cosine rather than inner-product similarity, which is otherwise entangled with item popularity.
arXiv:2607.02030v1 Announce Type: new Abstract: We consider a mathematical model of a poro-visco-elastic medium subject to frictional contact with a rigid obstacle, and study its numerical approximation. This model couples the Biot equations and contact conditions in the form of normal compliance and Coulomb friction. The resulting variational problem consists of a linear partial differential equation coupled to a nonlinear variational inequality. We propose and analyze a fully discrete numerical scheme for this problem, using conformal finite elements in space and the implicit Euler method in time. Existence and uniqueness of the discrete solution is established, and stability and a priori error estimates are derived. A numerical experiment is performed in which numerical error estimates are computed and compared to the theoretical results.
arXiv:2607.01728v1 Announce Type: new Abstract: Visual regression testing (VRT) is a standard quality assurance step in modern software release pipelines. On every change, it re-renders user interface (UI) screenshots, compares each one against an approved baseline image, and routes any detected difference to a human reviewer who decides whether it is an intended update or an unintended regression. A widely used approach, especially in open-source and continuous-integration pipelines, is pixel-level comparison, which is semantically blind and treats rendering noise and genuine defects identically, producing large volumes of false positives that force developers and testers to spend substantial time and effort manually reviewing flagged differences at every release cycle. Industry tools apply machine learning to VRT, but lack public evaluation. More critically, no dataset or benchmark exists to support natural language descriptions of UI changes, a capability that tells testers what changed in words instead of leaving them to interpret a binary flag or a highlighted region. To address the gap, we propose a new task, Web UI Image Change Captioning (WUICC), which sits at the intersection of VRT and image difference captioning (IDC), and release WUICC-bench, its first dataset and benchmark for the task. We evaluate eleven representative IDC methods, together with two zero-shot general-purpose LLMs. We find that: (1) these methods tend to struggle in the Web UI domain due to its layout diversity, dense text, and fine-grained changes, and (2) yet the trained methods already suppress non-meaningful visual noise far more selectively than the pixel-level comparison VRT relies on, providing a solid foundation for future domain-specific research.
arXiv:2607.02040v1 Announce Type: new Abstract: This paper presents a four-channel prototype system for the geometric combining and coherent addition of tightly focused femtosecond laser radiation into a standing-wave field configuration. A stabilization system for beam pointing and relative phase of the four optical channels has been implemented, and its performance has been experimentally demonstrated. To characterize the standing-wave electromagnetic field distribution at the main focus of the system, an original measurement technique based on a fiber subwavelength optical probe has been employed. This work has been conducted in support of the exawatt-scale XCELS project.
Estimating unknown dynamics and cost as a bilinear system with Koopman-based Inverse Optimal Control
arXiv:2501.18318v2 Announce Type: replace Abstract: In this work, we address the challenge of approximating unknown system dynamics and cost functions through a Koopman-based Inverse Optimal Control (IOC) framework. Using optimal trajectories, a modified Extended Dynamic Mode Decomposition with control (EDMDc) constructs a bilinear control system in lifted coordinates. Pontryagin's Maximum Principle (PMP) conditions are then derived, revealing structural similarities to the inverse Linear Quadratic Regulator (LQR) problem. This allows tractable cost recovery without resorting to nonlinear IOC formulations. The bilinear representation also inherits the analytical advantages of linear systems. Simulation and robotic experiments validate the approach, showing accurate estimation of both dynamics and costs, and illustrating its potential for general control and modeling applications.
arXiv:2607.01293v1 Announce Type: new Abstract: We present RuleChef, a framework that uses large language models (LLMs) to generate executable rules for NLP tasks such as text classification, Named Entity Recognition (NER), or relation extraction. Rules are generated based on a task description and a set of labeled examples, then they are iteratively improved based both on additional examples and on human feedback overexisting rules. RuleChef can also be used to bootstrap rules using the observed input-output pairs from any existing model for a given task. LLMs are used only at learning time, synthesizing rules and iteratively patching them based on failures measured on a held-out split. The result of this process is a fast, deterministic, and inspectable rule system. Preliminary evaluation is performed on both classification and NER tasks. We release RuleChef as open-source software under an Apache 2.0
arXiv:2607.01304v1 Announce Type: new Abstract: ROS 2 (Robot Operating System 2) has emerged as the de facto standard for modern robot software development, with middleware implementations such as the Data Distribution Service (DDS) and Zenoh forming the core infrastructure for distributed robotic communication. Despite their architectural flexibility, these middleware systems exhibit structural limitations, particularly under dynamic and resource-constrained wireless environments. This paper presents a systematic survey of ROS 2 middleware and introduces a conceptual framework to examine its architectural limits through three structural dimensions required by distributed robotic systems, namely Space, Time, and State. We first provide a structured analysis of middleware architecture and operational dynamics, including discovery, data exchange, and state management mechanisms. Building on this foundation, we formalize Time as temporal predictability for control loops, Space as spatial abstraction from physical topology to enable modular deployment, and State as contextual continuity despite dynamic node participation and intermittent connectivity. Through a comprehensive review of existing implementations and prior studies, we organize middleware research according to the structural trade-offs that arise among these dimensions. Under constrained wireless conditions, spatial abstraction can obscure network variability and weaken temporal guarantees, while mechanisms that preserve state continuity introduce computational and network overhead that competes with time-critical communication. These interactions reveal structural trade-offs that characterize the practical limits of contemporary robot middleware. By synthesizing architectural patterns and identifying gaps in current modeling and analysis approaches, this survey outlines a principled research roadmap for robust and scalable robotic middleware architectures.
arXiv:2511.07688v2 Announce Type: replace Abstract: Light propagation through turbulence produces speckles, whose ensemble behavior is typically characterized by snapshot intensity statistics. Here, we track the spatiotemporal evolution of individual speckles and quantify fragmentation, localization, and persistence under different diffraction and turbulence scales. Beam fragmentation coincides with complete spatial decorrelation defined by the magnitude-squared coherence. Fragmentation occurs closer to the source for larger beams, which indicates that smaller beams are more robust to decoherence. Subsequently, speckles are both spatially localized and persistent over distances significantly longer than their associated Rayleigh length. The combination of localization and persistence impacts the statistics of light relevant to their long-distance signaling and sensing.
arXiv:2607.01733v1 Announce Type: new Abstract: Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We observe that as supervised ASR training data increases, the contribution of LLM priors becomes less evident, and simple speech-text joint training under-utilizes textual knowledge. We therefore propose Joint Speech-Text Interleaved Pretraining (JSTIP), an ASR-oriented pretraining strategy that constructs word-level and segment-level interleaved speech-text sequences within aligned pairs for speech-LLM architectures that accept continuous inputs. Experiments on 38k hours of ASR data show consistent entity accuracy improvement compared to ASR-only and joint speech-text training baselines. JSTIP achieves on-par entity recognition performance using domain transcription text compared to synthetic speech-text pairs, simplifying domain adaptation. Benefiting from textual pretraining and domain text data, JSTIP is competitive with open-source ASR and Speech-LLM systems in medical entity recognition. The zero-shot speech question answering behaviors further suggest that interleaving reduces the speech-text modality gap and preserves the LLM generative prior, which is likely the reason for the entity improvements on the ASR task.
arXiv:2607.02139v1 Announce Type: new Abstract: Zero-shot object counting (ZOC) aims to count instances of arbitrary object categories specified only through textual prompts. Recent training-free approaches leverage foundation models such as SAM to reformulate counting as a prompt-driven segmentation task, eliminating the need for costly counting-specific training data with point-level annotations. More recently, SAM3 introduced promptable concept segmentation, enabling the zero-shot segmentation of all instances corresponding to a text-defined concept. However, SAM3 struggles in densely populated scenes containing numerous small objects, where limited image resolution and insufficient attention to target-relevant regions often lead to missed instances and poor instance separation, hindering accurate object counting. To address this limitation, we propose AdaCount, a training-free framework for ZOC based on similarity-guided spatial and feature adaptation. AdaCount first estimates a prototype-driven similarity map that identifies target-relevant regions. This similarity map subsequently guides two complementary adaptations: (i) similarity-guided spatial warping, which reallocates image resolution toward target instances, and (ii) feature modulation, which amplifies target-relevant encoder representations. Together, these adaptations enable SAM3 to devote greater representational capacity to target-relevant regions while preserving global image context, without requiring any model retraining. Extensive experiments across six diverse counting benchmarks establish AdaCount as a new SOTA among training-free ZOC approaches.
arXiv:2607.01367v1 Announce Type: new Abstract: A proposed method for the control of groups of inspection spacecraft is Multi-Agent Reinforcement Learning (MARL). While MARL has already been employed for this purpose in previous work, the reward functions used focus on reaching a finite set of predetermined inspection points around the target. In this work, we study and develop a generalized reward function for the MARL inspection task informed by the analysis of 3D reconstructions of inspected objects in orbit. Because the reward function is generalized such that any number of images at arbitrary locations may evaluated, we also allow trained agents to have complete control over when images are collected. With this approach, we gather insights into best practices for not only the specific MARL inspection task, but also gain key takeaways informative to the broader inspection task outside of a MARL context.
arXiv:2607.01737v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have substantially advanced video understanding, yet long-form video QA remains challenging under fixed input token budgets, where uniform sampling can be inefficient for evidence localization. We propose ReQuest , an uncertainty-driven, question-adaptive keyframe selection pipeline that aligns question intent with relevant video content through selective computation. ReQuest integrates (i) a lightweight question-aware selector distilled from MLLM-generated supervision, (ii) Re-thinking Routing that triggers additional inference only when the model is uncertain with a length-adaptive criterion, and (iii) uncertainty-guided adaptive non-maximum suppression that selects temporally diverse frames while adjusting spacing based on question difficulty. As a plug-andplay method, ReQuest improves long-video QA without modifying or fine-tuning the underlying MLLM. Experiments on Video-MME, MLVU, and LongVideoBench demonstrate consistent accuracy gains with competitive computational cost, with particularly strong improvements in medium and long video regimes.
arXiv:2607.01739v1 Announce Type: new Abstract: Despite significant technological progress, the realization of fully autonomous berthing and unberthing remains a significant challenge. One of the primary obstacles is the complex, non-linear nature of low-speed ship dynamics, which are difficult to model and control and often necessitate equally complex maneuvering models and control systems. This study proposes a simplified approach to bridge this gap by modeling the ship dynamics in the form of a time-invariant, continuous-time linear state-space system. The model parameters are estimated through system identification using the Covariance Adaptation Strategy Evolution Strategy (CMA-ES) applied to full-scale maneuvering data. Validation results demonstrate a strong agreement between the model output and empirical data. This outcome demonstrates the significant potential of simplified models to effectively define the maneuvering motion of a ship at low speeds.