Forskningsradar

Science Journals

Peer-reviewade publikationer — 60797 artiklar

HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning
arXiv:2606.29126v2 Announce Type: replace Abstract: Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing protocols treat messages as flat dense vectors detached from the structure of the observations they summarize. This design overlooks an important source of inductive bias in many cooperative environments, where observations naturally follow a hierarchy such as groups and entities. We propose \textsc{HiComm}, a plug-in communication module that grounds messages in the sender's hierarchical observation. \textsc{HiComm} is receiver-driven: the receiver issues a query, and the hierarchy is resolved through a three-stage decoding process that first selects a group, then a sender, and then an entity within that group, returning the corresponding feature slice as the message. This converts communication from unstructured vector transmission into structured information retrieval over the sender's observation hierarchy. We instantiate this mechanism with Straight-Through Gumbel-Softmax for differentiable discrete selection and a lightweight shared projection design that attaches to standard MARL pipelines. Experiments across cooperative MARL tasks with different observation structures and coordination demands show that \textsc{HiComm} matches or outperforms representative learned communication baselines while reducing communication volume by up to $23\times$ per receiver per episode.
Mobile Base Station Positioning in Smart Ports Based on Kriged Sparse Measurements and Obstacle Inference
arXiv:2607.00709v1 Announce Type: new Abstract: Smart-port wireless networks suffer from dynamic radio blockage caused by container stacks and industrial structures, challenging efficient mobile integrated access and backhaul (MIAB) deployment. Existing approaches rely on obstacle maps, geometry information, or computationally intensive propagation models that limit adaptability. This paper presents DOCKING, a radio environment map (REM)-driven framework that converts sparse radio measurements into optimization-ready obstacle representations for MIAB deployment. The framework infers propagation-relevant obstacle abstractions from reconstructed REMs, eliminating the need for obstacle-geometry databases while relying only on known network parameters and sparse measurements. Reference signal received power (RSRP) and signal-to-interference-plus-noise ratio (SINR) observations are reconstructed using Ordinary Kriging (OKG), and dominant attenuation regions are approximated by compact cuboidal blockage models. The inferred geometry feeds a backhaul-aware optimization that determines MIAB placement, user equipment (UE) association, and backhaul selection. Under realistic smart-port conditions, REM reconstruction achieves prediction errors below 3 dB at the 90th percentile using only 15% spatial sampling, while obstacle characterization exceeds 85% true-positive coverage. Capacity gains reach 150% in sparse deployments, and a fast Genetic Algorithm converges within 5-15 s per network snapshot. A field campaign using real measurements validates the workflow, showing throughput trends consistent with optimization predictions. Results demonstrate that sparse radio measurements provide sufficient environmental awareness for practical obstacle-aware MIAB deployment in obstruction-prone industrial environments.
Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark
arXiv:2607.00710v1 Announce Type: new Abstract: Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what datasets contain rather than how to strategically design impactful ones. This is especially limiting for small and medium-sized labs and startups that cannot afford to misallocate scarce resources. We argue that impactful dataset creation begins with a diagnosis: whether a research question is blocked by a data problem or an evaluation problem, and proceeds by selecting the minimal data operator(s) that closes the resulting gap, recording new data only when no cheaper operator(s) suffices. We analyze the evolution of major autonomous driving (AD) datasets through this lens and distill a strategic framework spanning gap identification, operator choice, sensor suite design, and annotation strategy. We ground the framework in a running case study of our KITScenes dataset family. The datasets are available at: https://kitscenes.com/
Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption
arXiv:2607.00712v1 Announce Type: new Abstract: Autoregressive (AR) streaming models have emerged as a powerful paradigm for long video generation. However, the linearly growing Key-Value (KV) cache poses a significant bottleneck, leading to memory overload and degraded inference throughput. A common compression method is to drop redundant KV tokens, which often breaks long-range dependencies, resulting in temporal flickering and identity loss. In this paper, we propose Instance-Specific Parametric Absorption (ISPA), a novel framework that shifts the KV cache compression from discarding to distilling. The core idea is to transit a subset of layers from Full-Attention (F-Layers) to memory-efficient Local-Attention (L-Layers) by "absorbing" historical context into the model's weights. Specifically, during a brief warmup phase, ISPA monitors the output discrepancy between global and local attention. At the transition point, we solve a closed-form least-squares problem to compute an instance-specific weight modulation that compensates for the missing history. Experiments across architectures (1.3B to 14B) demonstrate that ISPA can remove up to 50\% of the KV cache with near-lossless visual quality. We hope this perspective encourages future work to explore parametric memory consolidation beyond external token-level cache management for streaming generative models.
Self-conditioned Flow Map Language Models via Fixed-point Flows
arXiv:2607.00714v1 Announce Type: new Abstract: Self-conditioning is a core technique that enhances continuous flow-based language models, where the model learns to denoise generated text by conditioning on its own denoising estimate. While empirically successful, its performance improvements are poorly understood. Moreover, there is growing interest in the use of few-step generators based on flow maps, for which how to leverage self-conditioning is unclear. Here, we show that flow language models with self-conditioning solve a fixed-point iteration that bootstraps the performance of the learned denoiser. We use this viewpoint to formulate fixed-point flows, a two-dimensional class of self-conditioned flows, where the first dimension represents the flow process and the second represents the fixed-point iteration. We show that fixed-point flows define valid flow maps, and show that they can be distilled from self-conditioned flow models by compressing both fixed-point iterations and the flow process, the former with fixed-point distillation and the latter with flow map distillation. Our resulting flow map language model, FMLM$^\star$, outperforms state-of-the-art self-conditioned models and few-step models in one- and few-step generation on OpenWebText. Code is available at https://github.com/Ugness/self-conditioned-fmlm.
Fractal-Fractional HIV Dynamics with Mittag-Leffler Kernel: Analysis, Stability, and Numerical Simulations
arXiv:2607.00061v1 Announce Type: new Abstract: In this paper, a fractal--fractional HIV model with the Mittag--Leffler kernel is proposed using the Atangana--Baleanu--Caputo operator to capture the memory and hereditary properties of the disease dynamics. The existence and uniqueness of the solutions are investigated using suitable analytical techniques, and the Hyers--Ulam stability analysis is carried out to verify the stability behavior of the proposed system. For the numerical simulations, the Newton polynomial approximation method together with the Atangana--Toufik numerical scheme is employed to obtain approximate solutions for different parameter settings. Furthermore, several visualization techniques, including sensitivity heatmap representation and tornado diagram analysis, are utilized to study the influence of model parameters on the HIV dynamics. The obtained numerical results demonstrate that the proposed fractal--fractional framework provides an effective and reliable approach for analyzing the transient and long-term behavior of HIV transmission dynamics.
Structure-preserving dynamical low-rank approximation for parametric elastic guided waves
arXiv:2606.30469v2 Announce Type: replace Abstract: Elastic guided waves are widely used in Structural Health Monitoring (SHM). In many-query settings, the computational cost of high-fidelity simulations motivates the use of projection-based reduced order modeling (ROM). However, the transport-dominated and dispersive nature of guided waves challenges static linear subspaces. In addition, preserving the Hamiltonian structure of the equations for energy conservation necessitates dedicated projection techniques. While the Dynamical Low Rank Approximation (DLRA) has proven effective for other wave equations, its application to elastic guided waves in SHM has remained unexplored. In this work, we introduce a structure-preserving parametric ROM framework that leverages the DLRA in an off-line/on-line strategy. During the off-line stage, a time-dependent symplectic reduced basis is constructed from training simulations. For a simplified class of parameter dependencies, we derive a closed-form solution of the nonlinear basis evolution equation. This analytical result yields a closed-form, energy-preserving reduced propagator during wave propagation, eliminating on-line time integration after the loading phase. We validate our approach on a 2D elasticity problem featuring dispersive guided waves interacting with a damage. The results demonstrate high compression ratios (rank $\sim 10-30$), low full field reconstruction errors ($\sim 10^{-3}-10^{-2}$), speedups of two to three orders of magnitude, and excellent long-time energy conservation.
3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement
arXiv:2606.30514v2 Announce Type: replace Abstract: Human image animation, which aims to generate a video of a reference subject following a provided action sequence, has received increasing research interest. With the development of diffusion-based/flow-based video foundation models, existing animation works have began to upgrade the guidance information from 2D skeleton/pose to 3D modeling conditions. Despite achieving reasonable results, these approaches face challenges in synthesizing trajectory-controllable human motion within natural scene under changed camera views. In this work, we present a scene-adaptive human image animation framework that controls both human motion and camera trajectories within a reconstructed 3D environment for video generation. To achieve this, we first develop a ground-adaptive 3D motion retargeting approach to enable user-friendly motion trajectory control adapting to the changes of elevations of ground and orientations automatically. Then we design a viewpoint-adaptive latent fusion mechanism to inject point-cloud geometric priors through scene-visibility masking into the generative process, providing precise guidance of viewpoint changes under camera control. Experiments on two standard human image animation benchmark datasets demonstrate remarkable improvements of our method over the state of the arts in related video generation metics. Project page: https://robinhood256100.github.io/web-disp
KGS-GCN: Kinematics-Driven Gaussian Splatting and Probabilistic Topology for Skeleton-Based Action Recognition
arXiv:2603.16943v2 Announce Type: replace Abstract: Skeleton-based action recognition is widely applied in sensor-based systems, including human-computer interaction and intelligent surveillance. However, typical sensors produce sparse and discrete joint coordinates, often leading to the loss of fine-grained spatiotemporal information during dynamic movements. Furthermore, predefined physical topologies restrict modeling potential long-range dependencies. To address these challenges, we propose KGS-GCN, which integrates kinematics-driven Gaussian splatting and probabilistic topology within a graph convolutional network. A Gaussian splatting module constructs anisotropic covariance matrices by extracting instantaneous joint velocity vectors, rendering sparse skeleton sequences into multi-view continuous heatmaps rich in spatiotemporal semantics. Additionally, a probabilistic topology construction strategy transcends physical connectivity limitations by utilizing the Bhattacharyya distance to quantify statistical correlations between joint Gaussian distributions, generating an adaptive prior adjacency matrix. Finally, the lightweight multi-view rendering branch and topological GCN backbone are unified through a visual context gating mechanism, enabling seamless fusion of continuous dynamic cues with structural priors while maintaining high computational efficiency, requiring only 1.4M parameters and 1.3 GFLOPs. Extensive experiments on multiple benchmark datasets demonstrate that KGS-GCN significantly enhances the modeling of complex spatiotemporal dynamics and achieves competitive performance at low computational cost, establishing an efficient paradigm for improving the perceptual robustness of low-fidelity sensor data.
Neural Signatures of Programming Expertise: Classifying Programmer Skill Levels Using EEG Data
arXiv:2606.30879v2 Announce Type: replace Abstract: Accurately assessing a programmer's skill level is critical for hiring, team composition, and performance evaluation in the software industry. Conventional methods, such as coding tests or interviews, often fail to capture the full spectrum of cognitive abilities underlying programming expertise. This study explores using electroencephalography (EEG) and machine learning to investigate neural correlates of programming skill. We analyzed an existing EEG dataset recorded during code comprehension from 37 programmers with 1 to 30 years of experience (8.1 +/- 6.3 years) to examine relationships between neural activity and expertise. Additionally, we conducted classification experiments using Random Forest classifiers with diverse features for binary (experts vs. novices) and multi-class (experts, intermediates, novices) setups. We identified EEG features and brain regions associated with programming expertise. Specifically, EEG entropy showed the strongest correlation with skill level. Furthermore, experts' brains were characterized by highly localized centro-frontal activation, whereas frontal activation in other groups was part of a more distributed network. Regarding classification, our setup achieved an average accuracy of 91.83% (binary) and 78.15% (multi-class) in stratified 10-fold cross-validation, while leave-one-subject-out validation achieved 85.00% and 58.80%, respectively. Individual frequency bands outperformed full-spectrum analyses, and both program comprehension and resting-state data yielded strong results. These findings demonstrate that EEG features effectively capture neural correlates across different skill levels and highlight the potential of neural data to complement traditional methods of skill assessment.
Robust Operational Space Control with Conformal Disturbance Bounds for Safe Redundant Manipulation
arXiv:2607.00424v1 Announce Type: new Abstract: Redundant robotic manipulators operating in constrained and human-interactive environments require accurate task-space tracking together with rigorous safety guarantees under dynamic uncertainties. Classical operational space computed torque controller (OSCTC) relies on accurate dynamic models and degrades in the presence of disturbances. In contrast, the data-driven paradigm of residual learning approximates disturbances as functions learned from full-state measurements, which are often noisy in practice, lack rigorous theoretical guarantees, and introduce additional design complexity. This paper proposes a robust OSCTC framework that integrates an extended state observer (ESO) with conformal prediction to combine model-based robustness and data-driven adaptability. The ESO estimates lumped disturbances directly in operational space without requiring full-state measurements as in residual learning, and a robust control barrier function (CBF) is constructed to enforce safety under uncertainty. However, robust CBFs require a known disturbance-variation bound to guarantee absolute safety, which often leads to conservatism in practice. To address this limitation, we further employ a sliding-window conformal prediction mechanism to estimate the bound online in a distribution-free manner, thereby achieving practical probabilistic safety guarantees. Experiments on a 7-DoF Franka Research 3 manipulator demonstrate millimeter-level tracking accuracy and real-time safe control at 1~kHz under various disturbances.
HyFL-CLIP: Hyperbolic Fine-Tuning of CLIP for Robust Long-Context Understanding
arXiv:2607.00428v1 Announce Type: new Abstract: CLIP (Contrastive Language-Image Pre-training) has become a de facto paradigm for image-text alignment, but it struggles with long-context descriptions (>77 tokens) due to absolute positional encoding and pretraining on short captions. In long contexts, sentences are often reordered, summarized, or partially omitted. Although prior works extend CLIP with longer positional encodings, they often suffer from degraded image-text alignment under such text perturbations. We attribute this limitation to the Euclidean contrastive objective, which enforces strict one-to-one matching and lacks explicit mechanisms for modeling hierarchical relationships between global context and its constituent elements. To address this issue, we propose HyFL-CLIP, a hyperbolic fine-tuning framework that distills the well-established text-image alignment learned in Euclidean CLIP into hyperbolic space via cross-manifold similarity distillation, leveraging its geometry to capture hierarchical and entailment relations. Our method models hierarchical semantics by linking summarized token-wise features, long-context descriptions, constituent short textual components, and images, capturing part-whole relationships via hyperbolic entailment with Einstein midpoint aggregation. Experiments on diverse benchmarks, including long-context cross-modal retrieval, cross-modal retrieval with caption perturbations, intra-modality retrieval, and short-text cross-modal retrieval, show that HyFL-CLIP achieves more robust long-context understanding. In particular, it yields up to 19.5% improvement in long-text cross-modal retrieval under textual perturbations over the best prior method. We also show HyFL-CLIP can be seamlessly integrated into other model frameworks by applying it to Stable Diffusion XL (SDXL).
Caption Bottleneck Models
arXiv:2607.00578v1 Announce Type: new Abstract: Concept Bottleneck Models (CBMs) provide interpretability by routing predictions through a layer of human-understandable concepts. However, defining an optimal concept set for a specific dataset remains an open challenge. Existing approaches rely on expensive expert annotations or LLM-generated lists based solely on class names. Even "open-vocabulary" variants typically depend on static concept sets, which restrict discovery and introduce label bias. Furthermore, traditional CBMs often suffer from information leakage, where unmodeled visual features bypass the bottleneck and compromise the integrity of the explanations. To overcome these limitations, we propose Caption Bottleneck Models (CaBM), a framework that circumvents the need for predefined concept sets by replacing rigid concept layers with free-form natural language. By representing images via LMM-generated captions and training a classifier strictly on this text, CaBM ensures a leakage-free architecture by construction. Additionally, by analyzing the text classifier post-training, CaBM autonomously discovers high-quality, dataset-specific concepts. Our results across fine- and coarse-grained benchmarks demonstrate that CaBM achieves competitive accuracy while preserving interpretability without the constraints of external dictionaries or manual labeling.
CellPrior-Net: Prior-Guided Nuclei Detection and Classification for H&E Whole-Slide Images
arXiv:2607.00802v1 Announce Type: new Abstract: Accurate nuclei detection and classification in hematoxylin and eosin (H and E) whole-slide images (WSIs) is a key task in computational pathology, particularly for quantitative analysis of the tumor microenvironment. However, this task remains highly challenging due to variations in nuclei morphology, staining procedures, scanners, organs, magnifications, and WSI artifacts. In addition, many existing pipelines rely on computationally demanding architectures and post-processing procedures, making gigapixel WSI analysis time consuming. In this work, CellPriorNet (CP Net) is proposed, an efficient nuclei detection and classification pipeline that utilizes a lightweight convolutional neural network architecture and hematoxylin (H) channel as prior information to enhance nuclei-aware feature learning. Extensive benchmarking was conducted against state of the art pipelines on 8 public and private datasets (total:10.4M nuclei) obtained from different organs, scanners, magnifications, and clinical centers. Experimental results demonstrate that CP Net achieves comparable performance while significantly reducing inference time. Furthermore, CellQuant Net was introduced, an end to end nuclei quantification pipeline, that integrates a quality assessment (QA) model to exclude regions with artifacts, followed by CP-Net cell detection and classification. The pipeline is publicly available on GitHub, and provides a potentially efficient and scalable framework for downstream computational pathology applications.
Not All Prediction Targets Keep Training-Free Diffusion Guidance on the Manifold
arXiv:2607.00647v1 Announce Type: new Abstract: Training-free guidance (TFG) steers a pretrained diffusion model toward a desired attribute at inference. To be effective, this guidance must be applied from the earliest, high-noise steps of sampling. Because its objective (a classifier or energy) is defined on clean images, $\epsilon$- and $v$-prediction models must first estimate the clean image $\hat{x}$ from the noisy state at each step, and the accuracy of that estimate determines how easily guidance drifts off the data manifold. $x$-prediction, a recent alternative, outputs the clean image directly, removing this source of error even at high noise. This is our motivation. We provide a theoretical analysis of how each prediction target shapes this accuracy, and introduce guided-class FID (Child FID), a metric that exposes the manifold damage standard evaluation misses. Experiments on a new fine-grained bird benchmark and on style transfer confirm that $x$-prediction keeps guided samples on the manifold most reliably, making it the strongest foundation for training-free guidance. Code is available at https://github.com/ManLuML/on-manifold-tfg
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
arXiv:2605.31597v3 Announce Type: replace Abstract: Measuring structured object understanding in vision foundation models remains challenging due to inconsistent evaluation protocols and limited part-level supervision. Semantic correspondence (SC) evaluates this capability by testing whether object parts can be matched across instances and categories under large variations in appearance, viewpoint, and geometry. To enable a systematic SC evaluation, we introduce SOCO, a new benchmark for Semantic Object Correspondence that introduces a taxonomy of correspondence types and provides consistent, functionally meaningful keypoint annotations across 100 categories and over 1M correspondence pairs. In addition, SOCO includes keypoint language descriptions, enabling the evaluation of large vision-language models (LVLMs) and their fine-grained part-level understanding. Comprehensive experiments reveal that (i) vision foundation backbones encode strong semantic structure but transfer correspondences poorly across related categories and only partially capture object-part position, (ii) LVLMs are stronger at text-prompted part localization than at visual-reference cross-image matching, exposing a gap between language-grounded localization and fine-grained visual correspondence, and (iii) correspondence performance predicts performance on dense downstream tasks, including segmentation, tracking, 3D pose estimation, and 3D detection, more strongly than ImageNet classification. Together, these findings position SOCO as a benchmark for structured, part-level representation quality in vision and multimodal foundation models.
Learning-based control of a single-DOF Aero system
arXiv:2607.00640v1 Announce Type: new Abstract: This paper presents a learning-based control framework that integrates feedback linearization with reinforcement learning for the adaptive control of nonlinear mechatronic systems. The control law is derived using Lyapunov stability analysis, ensuring closed-loop stability in the presence of modeling uncertainties and external disturbances. Feedback linearization serves as the main control framework, while a reinforcement learning component estimates and compensates for unmodeled dynamics and disturbances online. The learning module is based on the REINFORCE-with-baseline algorithm, which improves learning efficiency by reducing the variance of policy-gradient estimates and enabling stable policy updates during adaptation. The proposed controller is evaluated on a single-degree-of-freedom rotor-based AERO system. Results from simulations demonstrate accurate trajectory tracking, fast adaptation, and strong robustness against parameter variations and external disturbances. Overall, the proposed approach combines the analytical guarantees of Lyapunov-based control with the adaptability of reinforcement learning, providing an effective solution for controlling nonlinear mechatronic systems.
Shapley in Context: Explaining Financial Language with Domain Expertise
arXiv:2607.00856v1 Announce Type: cross Abstract: In recent years, large language models have achieved remarkable success and have seen growing adoption in financial applications. At the same time, explainability remains critical in finance, a domain characterized by high stakes and strict regulatory requirements. Although numerous methods have been proposed to explain black box machine learning models, the majority of these approaches are designed for general purpose tasks and do not incorporate domain specific knowledge. In this work, we study the explainability of financial textual data modeled by large language models through the lens of the Shapley value. Specifically, we investigate whether Shapley based attributions align with established financial domain knowledge. Through rigorous theoretical analysis and extensive empirical evaluations, we demonstrate that Shapley values can yield explanations that are consistent with financial reasoning and can offer meaningful insights into the model's behavior in text based financial applications.
Would You Marry Superintelligence?
arXiv:2607.00120v1 Announce Type: new Abstract: Emotional bonds between humans and AI companions are growing, and the question of whether a person may marry an AI system will soon move from speculative fiction into law. This chapter examines whether the autonomy-centered logic that has expanded marital choice among human beings can justify extending marital status to superintelligent companions. Following a scenario-envisioning exercise informed by anticipatory ethics, I argue that granting such status leads to socially unjust outcomes, even under the generous assumption of reliable superintelligence. Marriage as a socio-legal institution does more than ratify private agreement; it creates networks of mutual obligation, joins families, and makes each partner vulnerable to the other. A relationship sustained by corporate policy and continued payments is a subscription rather than a bond tested by time. Discussing wholesale marital status is therefore the wrong frame. Law should carve out targeted rights and protections for pressing needs arising from intimate human-AI relationships.
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization
arXiv:2607.00726v1 Announce Type: new Abstract: Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for feature extraction models exhibit dimensional bias, typically focusing on either semantic matching or temporal offset detection. Moreover, their data construction remains coupled, preventing independent assessment of temporal and semantic consistency. We propose AV-SyncBench, the first benchmark to fully separate temporal and semantic evaluation for audio-visual synchronization. Built from in-the-wild videos, it spans Voice, Music, and Sound across 10 scenarios and 5 challenge tasks. Data are automatically filtered and manually verified to ensure on-screen sound sources. The benchmark contains 3,269 videos and 38,390 samples, and we evaluate five representative models to quantify feature quality for alignment and downstream tasks. The code and dataset are available at: https://fgt7t6g.github.io/AV-SyncBench.
Room temperature valley coherence in monolayer WSe2 mediated by chiral nematic liquid crystal
arXiv:2607.01098v1 Announce Type: new Abstract: Valley coherence refers to a phase-coherent superposition of inequivalent momentum valleys, in which quantum information can be encoded in the relative valley phase. Chiral nematic liquid crystals, by imposing a flip of the spin angular momentum upon light reflection, provide an effective photonic environment for optically coupling excitons of the K and K' valleys in monolayer semiconducting transition-metal dichalcogenides. We experimentally demonstrate that using such liquid crystal as a substrate, it is possible through nearfield interaction to engineer a room temperature mechanism for inducing the intervalley coupling. Our results show that this approach provides a simple and scalable route toward valleytronic functionalities based on controlled coherent emission from valleys with opposite Berry curvature.
Demodulating Digital Holograms with Unknown Uniform Phase-shifts by Spiral Phase Transform
arXiv:2607.00126v1 Announce Type: new Abstract: The Spiral Phase Transform (SPT) is a generalization of the Hilbert transform for 2D signals and, as such, can be used for AC signal demodulation. However, phase demodulation with the SPT is complicated by a multiplicative term that depends on the fringe directional map. We derived an analytical formula for the twofold directional map and applied the SPT for the blind reconstruction of phase-shifted digital holograms. Possible phase ambiguities in the unfolded directional map were resolved by satisfying the spatial uniformity condition of the phase shifts. The method was experimentally verified using on-axis and off-axis digital holograms of specularly reflecting subjects.
DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
arXiv:2607.00889v1 Announce Type: new Abstract: We present DeWorldSG, a novel framework that generates spatio-temporally robust 3D Semantic Scene Graphs from RGB-D sequences. Existing methods often struggle to construct reliable 3D scene graphs due to unstable 3D object representations and missing relations caused by frame-wise inference. DeWorldSG addresses these issues by estimating instance-level geometric 3D Gaussian distributions through depth-guided filtering and representing each object as a probabilistic 3D node rather than a single projected point. To mitigate relational sparsity from frame-wise inference, our framework further aggregates spatiotemporal evidence across object pairs and refines relations using contextual priors derived from a world model (V-JEPA 2). Experiments on the 3DSSG and ReplicaSSG datasets demonstrate state-of-the-art (SoTA) performance in both object and predicate prediction, while producing temporally consistent scene structures. In particular, our method improves triplet recall by 77.4% and predicate recall by 23.2% over prior SoTA approaches, making it suitable for robotic manipulation and AR applications. Our code and models are open-sourced.
The effect of 20th century industrialization: Power station, acid rains, over-pumping, on an erstwhile uniform freshwater dune aquifer in Haifa Bay, Israel
arXiv:2607.00893v1 Announce Type: new Abstract: A small phreatic sand dune aquifer lies along the shore of Haifa Bay. It has been exploited for its freshwater resources since the 1930s. During this time the salinity has increased continuously, partly by seawater intrusion due to overpumping. The chemistry of the young aquifer water is laterally variable and is characterized by excess SO$_4^{2-}$, high $Sr^{2+}$ concentrations above that of modern seawater, high alkalinity, and markedly enriched $\delta^{13}C_{DIC}$ values. Acidic winter rains, formed from $SO_x$ and $NO_x$ gaseous emissions from a nearby power station, leach the dry deposition that accumulated across the dune surface during the dry summers. The acidity also partially dissolves the aragonite sea shells in the dune sands, remnants of a previous marine transgression. As a consequence, this adds $Sr^{2+}$, $Ca^{2+}$ excess, and alkalinity, while leading to enriched $\delta^{13}C_{DIC}$ values, particularly during the winter, at which time the radiocarbon activity in the DIC is observed to decrease.
Guaranteed Escape for a Bouncing Robot in Pipe Chains
arXiv:2607.00221v1 Announce Type: new Abstract: We study the symmetric bouncing of a point robot within orthogonally-joined rectangles with equal width, which we refer to as pipes. We provide an exhaustive case analysis of every trajectory pattern inside a single rectangular pipe segment, identifying the conditions under which the robot exits. We then extend the analysis to L-shaped pipes and, more generally, to linear chains of $k$ orthogonally connected pipe segments. We prove exit guarantees for the special angle $\alpha = \pi/4$. Furthermore, these results extend to pipes with curved joints.