arXiv:2607.16379v1 Announce Type: new
Abstract: One of us is a realist and believes that reality fundamentally consists in an external physical world, governed by laws of physics. Observers are either emergent (physicalism) or exist in addition to physical reality. The other defends a version of idealism and believes that reality fundamentally consists in first-person states, unembedded into worlds, on which laws of nature act directly. Shared "physical worlds" are merely emergent. The disagreement ultimately centers on the overall coherence of a formal model, known as Algorithmic Idealism, and its ability to resolve observer paradoxes, such as the Boltzmann brain paradox. Our aim is to confront this issue in the form of an adversarial collaboration, where we avoid misunderstanding each other, so that readers can see the true source of the disagreement, and decide for themselves. Our work represents an attempt to establish a form of exchange that may help overcome the fragmentation of scholarly communities in philosophy, physics, and elsewhere. The collaboration was conducted using a discipline-neutral template for theoretical adversarial collaboration developed in a companion paper [1], which we hope will be useful in other theoretical disciplines.
Science Journals
arXiv:2607.18052v1 Announce Type: new
Abstract: Over-the-air computation (OAC) enables efficient function aggregation in wireless networks by exploiting the superposition property of the multiple-access channel. However, practical deployment of OAC is severely challenged by the reliance on accurate carrier synchronization and coherent reception, which are costly and fragile, especially in short-range and low-complexity systems. In this work, we propose a \emph{self-coherent, synthesizer-free over-the-air computation framework} based on \emph{Kramers--Kronig (KK) reception}. By transmitting a biased aggregate waveform and employing direct detection followed by KK phase reconstruction at the receiver, the proposed scheme eliminates the need for explicit carrier recovery while preserving coherent-like signal aggregation.
We develop a signal-domain system model for multi-user OAC under KK reception and provide a synchronization-relaxation analysis demonstrating that the proposed architecture fundamentally removes carrier-frequency offset (CFO) sensitivity between transmitters and receiver. By shifting synchronization complexity away from strict carrier-phase tracking and eliminating distributed phase alignment requirements, the framework reduces control overhead and improves scalability in multi-user aggregation. A detailed per-symbol mean-squared error (MSE) characterization isolates the impact of channel mismatch and KK reconstruction noise, showing that the proposed self-coherent architecture approaches the theoretical performance limits of baseband OAC under practical operating conditions. Finally, we demonstrate that the approach is particularly well suited for mmWave and sub-THz systems, where oscillator phase instability otherwise represents a fundamental bottleneck to scalable coherent OAC.
arXiv:2607.18056v1 Announce Type: new
Abstract: Frontier large language models (LLMs) are increasingly integrated into scientific workflows, yet their growing biological capabilities may outpace current safeguards. To assess the biological risks of frontier models, we develop Intern-BioBreaker, a specialized bio-red-teaming model, together with an integrated computational-to-physical framework that couples model-level stress testing with wet-lab validation. Within this framework, Intern-BioBreaker generates targeted jailbreak prompts to test whether aligned models can be induced to provide operational guidance for safety-sensitive biological tasks or produce sequence-level outputs with potentially harmful properties. Selected sequence outputs are then carried forward for DNA synthesis, host expression, and orthogonal protein verification to assess whether model-generated designs can yield the intended biological products. Our evaluation reveals a concerning gap between text-level safeguards and the risks posed by capable scientific models: (i) Intern-BioBreaker outperforms baseline attack models and reveals widespread bio-risk jailbreak vulnerabilities across both open-weight and proprietary frontier LLMs, with several targets reaching near-saturated or 100% task-level attack success rate (ASR); (ii) in sequence-level case studies, GPT-5.5 can be induced to generate modified viral candidate sequences with pathogenic potential; the corresponding translated proteins may exhibit even stronger receptor-binding affinity and thus enhanced infection potential; and (iii) end-to-end verification shows that selected model-generated biological designs are not merely textual artifacts, but can be physically realized under controlled experimental settings. These findings underscore the need for stronger biological red-teaming, nucleic acid synthesis screening, and safety mechanisms that keep pace with model capabilities.
arXiv:2607.16383v1 Announce Type: new
Abstract: Verifying academic credentials remains difficult: records are held by individual institutions in proprietary systems, verification is slow and manual, and counterfeit qualifications are widespread. Blockchain-based registries have been proposed as a remedy, but existing systems tend to anchor certificate hashes without binding them to a verifiable identity, without an explicit mechanism to accredit issuing institutions, and without support for correcting or revoking credentials once issued. This paper investigates whether an infrastructure designed for regulated financial instruments can be repurposed to close these gaps. We present the design of a registry for identity-bound academic credentials that composes OnchainID self-sovereign identities (ERC-734/ERC-735) with the T-REX suite (ERC-3643): its trusted-issuer registry becomes an on-chain issuer-accreditation whitelist, and each certificate is represented as a signed, updatable claim bound to a student's identity and verifiable by any third party without a wallet, while sensitive fields are kept off-chain. We make explicit the tension between a transferable-security-token standard and non-transferable credentials, clarifying which of its guarantees carry over. We validate the design with a reference implementation covering the full certificate life cycle and evaluate it in terms of gas cost, scalability, latency, and security, quantifying the overhead relative to a hash-anchoring baseline.
arXiv:2607.16540v1 Announce Type: new
Abstract: Understanding and characterizing human color perception is a longstanding research goal. One of the most traditional approaches is looking for the human color discrimination thresholds, the minimum chromatic differences perceptible to human observers. In recent years, deep neural networks have become the standard networks for computer vision tasks. In particular, deep vision encoders, foundation models trained on large-scale visual data, map images into latent feature representations. Despite the widespread use of deep vision encoders, few studies have investigated whether their internal representations exhibit human-like discrimination thresholds. In this work, we present a large-scale exploratory study probing the chromatic sensitivity of more than 50 pretrained vision encoders, including convolutional networks and vision transformers, against human discrimination thresholds. Using controlled chromatic stimuli at multiple chroma levels, we compare model-derived chromatic discrimination thresholds with human discrimination ellipses through a region-overlap metric (mIoU). Our analysis reveals generally weak alignment between model representations and human perceptual thresholds across all model families, with the best mIoU < 0.25. Moreover, we find that self-supervised encoders consistently outperform supervised ones, while language-supervised models show the most polarized behavior, occupying both the top and bottom of the ranking. These findings suggest that human-like chromatic sensitivity does not emerge naturally from current large-scale visual training objectives for any of the analyzed architectures.
arXiv:2607.18058v1 Announce Type: new
Abstract: Although a complete characterisation of the probability distribution in the phase space of turbulent flows remains elusive, accurately sampling this distribution is essential for both synthetic turbulence generation and turbulent flow reconstruction. Motivated by these applications, we examine to what extent a machine-learned distribution can approximate the physical invariant distribution of turbulent channel flow at $\mathrm{Re}_\tau=180$. We assess three important properties of the approximation: physical ensemble statistics, consistent conditional sampling, and dynamical invariance. To this end, a flow-based generative model is trained on a minimal conditional flow unit, which we define as the smallest domain outside which conditional fields, given a single observation at the domain centre, are indistinguishable from unconditional fields in terms of mean-square discrepancy to other conditional fields. We also introduce a consistent procedure for sampling from the conditional learned distribution. Comparisons with direct numerical simulation show that synthetic turbulent fields reproduce key statistical and dynamical features of turbulence, including intermittency and nonlinear energy transfer. The consistency of conditional sampling is demonstrated in a flow reconstruction problem, and subsequently used to generate synthetic turbulent velocity fields on a large domain. When adopted as initial conditions in direct numerical simulations, these fields yield physical and statistically stationary ensemble statistics, indicating that the learned distribution provides a good approximation to the natural distribution of the turbulent dynamical system.
arXiv:2607.18060v1 Announce Type: new
Abstract: Long-horizon robotic tasks require diverse capabilities that no single policy can reliably provide. Heterogeneous policies offer complementary strengths, but orchestrating them requires reasoning over uncertain capability boundaries and cross-policy distribution mismatch, which are largely overlooked by existing planning methods built on homogeneous, predefined skills with fixed applicability. We propose RoboHarness, a unified framework that encapsulates independently developed robot control systems as reusable agentic skills. Although instantiated in this work with VLAs, RL policies, and task-and-motion planning (TAMP) systems, RoboHarness is designed as a general framework compatible with a broader range of robot policies, such as navigation policies, model predictive controllers, and world-action models. RoboHarness uses multi-modal execution memory and online evidence to characterize policy capability boundaries for capability-aware decomposition and routing. To stabilize policy handoffs, its Memory Bridge retrieves execution trajectories associated with the next policy, estimates its in-distribution state region, and guides the robot toward that region without joint policy retraining. Extensive experiments on three public benchmarks, 500 customized tasks, and 135 real-robot experiments demonstrate effective capability-aware routing and stable policy orchestration, yielding substantial improvements in zero-shot long-horizon planning and out-of-distribution robustness.
arXiv:2607.16979v1 Announce Type: new
Abstract: Repository mining studies increasingly analyze AI-evidence projects, yet it remains unclear how to measure whether architectural changes create deferred stabilization obligations. A natural metric, post-anchor stabilization density, counts tests, CI gates, documentation, and fixes appearing after a durable boundary is introduced. We show that this metric fails. In a diff-level study of 338 high-visibility 2026 GitHub repositories, anchors are common (321 of 338 contain real changed-file anchor evidence), but a controlled 308-event contiguous-window experiment finds no post-anchor uplift: broad and strict stabilization signals both yield median post/pre density ratios near 1.0, and non-anchor controls are equally dense. We introduce stabilization regimes, six recurring patterns that explain why the density metric fails, and use human validation to calibrate them. Two independent coders label 100 stratified candidate-anchor events from blinded packets (kappa = 0.50 on debt attribution). The validation exposes a two-layer trap: many candidate anchors are not durable boundaries (45 of 100), and even among valid anchors in this calibration sample the no-uplift result holds: only 3 of 100 events survive as attributable delayed obligations; the remaining 97 are explained by classifier error, anchor-local hardening, background maintenance, or pre-anchor hardening. Post-anchor density conflates pervasive maintenance with genuine debt; controlled designs with regime-aware attribution are necessary before repository mining can reliably identify stabilization obligations.
arXiv:2607.17593v1 Announce Type: new
Abstract: Class Incremental Learning (CIL) aims to learn new concepts consistently from a data stream without forgetting. Unlike typical CIL methods which need to learn a model from scratch, pre-trained model (PTM) can easily adapt to a new task with fine-tuning. However, existing PTM-based CIL methods fail to achieve a trade-off between performance and computational expenditure, i.e., they either adopt the same parameter space so that leading catastrophic forgetting, or expand a new branch for each task but adding more computational cost. To this end, we propose MetrIc Learning with Expandable Subspace (Miles) to harness the prior information within pre-trained knowledge, thereby orchestrating an efficient expansion of the parameter space through guided optimization. Specifically, it decouples the learnable modules with the pre-trained model, exploiting prior information from intermediate features of the backbone network to enable more flexible parameter expansion. Then, a central loss is adopted to guide the new category to cluster towards the corresponding prototype in the new task subspace while incorporating an auxiliary distance regularization term to maintain metric equilibrium across tasks. Extensive experiments on six benchmark datasets demonstrate that Miles achieves state-of-the-art performance in various CIL settings.
arXiv:2607.17329v1 Announce Type: new
Abstract: Medical image segmentation models require both high accuracy and lightweight design to accommodate real-world medical applications. The deployment of these models on resource-limited medical platforms remains a significant challenge due to their high computational and parameter requirements. Existing pruning methods for model compression mostly overlook the intrinsic connections and similarity between the internal structures of complex deep neural networks. As a result, compressed models may not effectively retain the basic features of the pretrained network. To solve this problem, we propose a hierarchical clustering compression method for medical image segmentation models (MIS-HCC). This approach employs hierarchical clustering to partition channels and fuse their parameters efficiently. Specifically, it leverages the Wasserstein distance to represent similarity of channels within layers of pre-trained network, forming a similarity matrix that guides the clustering process. Channels within each cluster are then fused to produce a compressed network. Experimental results on three medical image datasets application demonstrate that MIS-HCC outperforms the state-of-the-art methods in both accuracy and compression efficiency, offering an effective solution for deploying medical image segmentation models on resource-limited medical platforms.
arXiv:2607.16543v1 Announce Type: new
Abstract: As enterprises increasingly adopt Software-as-a-Service (SaaS) platforms for mission-critical functions, onboarding these services has emerged as a complex challenge extending well beyond procurement and basic security review. In regulated environments, SaaS onboarding must address multiple interdependent control domains, including Third-Party Risk Management (TPRM), cybersecurity assessment, Identity and Access Management (IAM), and disaster recovery (DR), which are often executed in isolation, resulting in delayed go-lives, duplicated assessments, unclear ownership, and residual operational risk. This paper proposes a control-driven, end-to-end SaaS onboarding framework that integrates TPRM, cybersecurity, IAM, and DR into a unified lifecycle model. The framework introduces a stage-based approach spanning intake and risk scoping, architecture validation, identity design, resilience assessment, and post-production governance. Key contributions include: (1) a structured SaaS onboarding lifecycle emphasizing sequencing and dependency management across control domains; (2) a cross-domain control mapping highlighting failure modes caused by siloed reviews; and (3) practical design patterns for secure connectivity, federated identity, least-privilege access, and shared-responsibility disaster recovery. Supported by operational lessons from enterprise-scale implementations and a reference governance checklist, the framework helps organizations reduce onboarding friction, improve auditability, and strengthen the security and resilience posture of SaaS-enabled platforms.
arXiv:2607.16820v1 Announce Type: cross
Abstract: We investigate the existence and dynamics of two-dimensional solitary waves in a quantum droplet environment described by the extended Gross-Pitaevskii equation featuring logarithmic mean-field and Lee-Huang-Yang interactions. In the modulationally stable regime of the background, we employ suitable multiscale asymptotic methods to derive effective nonlinear integrable models corresponding to the Kadomtsev-Petviashvili and Davey-Stewartson equations. Based on these reduced models, we construct approximate analytical solutions describing line solitons, algebraically localized lump solitons, ring solitons, and exponentially localized dromions embedded on the droplet background. The dynamical robustness of these solutions is monitored through numerical simulations. Line, lump and ring solitons stay closest to the theoretical predictions, although progressively deviate due to the emergence of small-amplitude radiation, while dromions depart from their analytical waveform the most, although they roughly maintain their shape. Our results unveil unprecedented multidimensional soliton solutions in models featuring the competition of mean-field and quantum fluctuations and as such are amenable to current ultracold atom experiments.
arXiv:2607.17117v1 Announce Type: new
Abstract: Sparse autoencoders (SAEs) decompose language model activations into sparse features, but standard SAEs encode each token independently and do not expose information that persists across a sequence. We introduce Persistent Sparse Autoencoders (Persistent SAEs), which extend standard SAEs by learning a persistence coefficient for each feature, allowing the model to learn which features should persist and for how long. Our experiments show that they retain competitive reconstruction quality while learning a spectrum of feature timescales: fast features behave as locally interpretable detectors, whereas slow features concentrate topic-level information in a persistent state. Moreover, as shown in a prompt-injection monitoring case study, slow features preserve detection signals and remain causally effective over long contexts. These results suggest that Persistent SAEs open up new opportunities for interpreting and monitoring language models through persistent semantic representations.
arXiv:2607.18144v1 Announce Type: new
Abstract: Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molecular design and have shown competitive performance in pocket-conditioned molecular generation. However, their ability to reason about physics and 3D spatial environments is largely underexplored. In this work, we systematically analyze whether current general-purpose LLMs are capable of navigating complex 3D constraints compared to established baselines such as specialized diffusion models. We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. To enable this evaluation, we introduce 3D-Fit - a token-efficient benchmarking strategy for assessing LLM performance on multi-conditioned spatial molecule generation. Our findings reveal a clear pattern in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.
arXiv:2607.16812v1 Announce Type: new
Abstract: The rapid emergence of sixth-generation (6G) networks and the low-altitude economy has accelerated the evolution of wireless infrastructures toward air-ground integrated coverage networks (AGICNs), which seamlessly fuse terrestrial and aerial communication resources. However, existing AGICN studies primarily focus on coverage enhancement, while ignoring sustainability. Pursuing sustainable AGICNs introduces new challenges due to the multidimensional resource coupling across heterogeneous air-ground segments. In view of this, this paper presents a comprehensive survey and tutorial on sustainable AGICNs, aiming to balance coverage capacity with carbon efficiency in low-altitude economies. An integrated sensing, communication, and computation (ISCC)-driven architecture, which enables dynamic resource orchestration through closed-loop control, is proposed. We thus introduce a multi-dimensional sustainability metric system, which covers operational efficiency, task-oriented performance, and full lifecycle carbon emissions, to quantify energy and carbon footprints. We review enabling technologies, including artificial intelligence, hybrid precoding, integrated sensing and communication, and simultaneous wireless information and power transfer, and discuss their integration into the ISCC framework to minimize energy consumption while maintaining robust coverage. Experimental results on a real-world testbed demonstrate a 20% reduction in power consumption while achieving over 90% coverage probability, highlighting the feasibility of sustainable AGICNs for future green networks.
arXiv:2607.17336v1 Announce Type: new
Abstract: Drift detection is a core component of production machine learning monitoring systems, where detectors are used to compare incoming data with a reference distribution and trigger alerts when changes occur. However, these detectors are often evaluated in research settings that emphasize detection accuracy under synthetic shifts, while overlooking false alarms under continuous monitoring. In production environments, models are monitored repeatedly over time and across many features, and even small false positive rates can accumulate into frequent alerts, leading to alarm fatigue. We empirically analyze false positive behavior across five commonly used drift detectors: PSI, KS, MMD, LSDD, and adversarial validation. Consistent with existing literature, PSI exhibits strong sensitivity to batch size, producing frequent false alarms at small sample sizes; however, we further observe that its behavior stabilizes and improves substantially once batch sizes exceed approximately 200 samples. In contrast, KS, MMD, and LSDD display persistent fluctuations across batch sizes, while remaining comparatively more reliable than PSI in low-data regimes. Applying a Bonferroni correction reduces false positive rates, but often at the cost of reduced true positive sensitivity, reinforcing the well-known stability - sensitivity trade-off in drift detection. This work provides a systematic comparison of false positive behavior across multiple drift detectors under continuous monitoring conditions. We identify tradeoffs across detector families and provide practical guidelines for selecting and calibrating drift detectors in production ML systems.
arXiv:2607.18033v1 Announce Type: new
Abstract: As with most ``in the wild'' collections of the natural world, the North America Camera Trap Images (NACTI) dataset exhibits long-tailed class imbalance, with the largest class covering over 50% of its 3.7M images. Building on the PyTorch Wildlife model, we systematically evaluate Long-Tail Recognition (LTR) methodologies to benchmark species recognition performance, including specialised loss functions and LTR-sensitive regularisation. Our optimised configuration achieves state-of-the-art 99.40% Top-1 accuracy on the NACTI test split, significantly outperforming standard baselines and previously reported top performances. To assess robustness under domain shifts (e.g., night-time captures, occlusion, motion-blur), we extend our evaluation across three independent reduced-bias test sets (including ENA-Detection, Caltech Camera Traps and Missouri Camera Traps). Across these out-of-distribution (OOD) evaluations, our LTR-enhanced model consistently demonstrates substantially stronger generalisation capabilities compared to standard cross-entropy approaches. However, qualitative and quantitative analyses underline that current LTR optimisations cannot fully overcome representational bottlenecks, resulting in catastrophic predictive breakdown for rare `Tail' classes under severe domain shift. For maximum reproducibility, all dataset splits, key code, and network weights are published with this paper at https://github.com/ZehuaLiuY/Species-Classification.
arXiv:2607.16854v1 Announce Type: cross
Abstract: A hypergraph is \emph{linear} if every pair of vertices is contained in at most one hyperedge. For a family $\mathcal{F}$ of $r$-uniform hypergraphs, the linear Tur'an number $ex^{\mathrm{lin}}_r(n,\mathcal{F})$ is the maximum number of hyperedges in an $n$-vertex $\mathcal{F}$-free linear $r$-uniform hypergraph. Extending the work of Gy'arf'as, Ruszink'o, and S'ark"ozy on $3$-uniform linear hypertrees, we study linear Tur'an numbers for higher uniformity.
We determine the linear Tur'an number of the $r$-uniform linear star $S_k^r$, proving [ex^{\mathrm{lin}}_r(n,S_k^r)\le \frac{n(k-1)}{r},] with equality exactly for $(k-1)$-regular linear $r$-uniform hypergraphs, whenever they exist. We also construct dense $T_k^r$-free hypergraphs showing that, under suitable divisibility and design-existence assumptions, [ex^{\mathrm{lin}}_r(n,T_k^r)\ge \frac{n(k-1)}{r}] for every linear $r$-uniform hypertree $T_k^r$ with $k$ hyperedges.
We then study all linear hypertrees with four hyperedges. For the broom $B_4^r$, we prove [ex^{\mathrm{lin}}_r(n,B_4^r)\le \frac{(r+1)n}{r},] and characterize the extremal hypergraphs as disjoint unions of Steiner systems $S(2,r,r^2)$, whenever such systems exist. For the crown $E_4^r$, we establish [ex^{\mathrm{lin}}_r(n,E_4^r)\le \frac{(2r-1)n}{r},] together with a lower-bound construction leaving only a constant-factor gap.
For the linear path $P_4^r$, we construct $P_4^r$-free hypergraphs with $(r+1)n/r$ hyperedges and conjecture this is the optimal general bound. We verify the conjecture for connected hypergraphs under suitable degree conditions. Finally, for $r=4$, we identify counterexamples to a key structural claim in a previous proof of Zhang and Wang, give a new proof that [ex^{\mathrm{lin}}_4(n,P_4^4)\le \frac{5n}{4},] and show that equality holds precisely for disjoint unions of Steiner systems $S(2,4,16)$.
arXiv:2607.17121v1 Announce Type: new
Abstract: Skeleton-based emotion recognition from body motion remains challenging because emotional expressions are often characterized by subtle dynamic and relational motion cues, and hard labels may not fully capture ambiguity among related emotion categories. For the DIEM-A task in the MMAC ACII 2026 Challenge, we propose a multi-branch skeleton-based emotion recognition framework that combines a 6D rotation-based branch, a part-aware kinetic multi-stream branch, and a metadata-conditioned weak label distribution learning (LDL) branch. The branches are trained independently and fused by a probability-level ensemble at inference time. In 10-fold leave-performer-out cross-validation, the proposed framework improves Accuracy from 0.271 to 0.366 and Macro-F1 from 0.252 to 0.353 over the rotation-based baseline. Explainability ablations show that velocity and bone streams, as well as arm and leg regions, provide important cues for recognizing emotional body motion.
arXiv:2607.17888v1 Announce Type: new
Abstract: In this work, we investigate the nonlinear dynamics of isolated current-carrying edge-localized mode (ELM) filaments using a reduced electromagnetic fluid model in slab geometry. Numerical simulations show that unidirectional parallel current significantly suppresses radial filament velocity and reduces the outward propagation velocity by weakening the curvature-driven interchange force. The reduction in radial velocity is found to follow a modified scaling relation, demonstrating that increasing current progressively weakens outward filament propagation. Analysis of the vorticity equation shows that the electromagnetic current source changes from a dipolar structure to a remarkable spiral pattern, and overcomes the conventional curvature drive in the nonlinear phase. This current-driven source directly imprints its topology on the vorticity field, resulting in spiral vorticity, enhanced angular momentum, increased rotational energy, and localized shear layers. The filament therefore undergoes a transition from a conventional propagating state to a rotationally self-organized electromagnetic structure. These findings demonstrate that parallel current acts as an effective electromagnetic vorticity source and provides new insight into the nonlinear dynamics of ELM filaments in tokamak edge plasmas.
arXiv:2607.17122v1 Announce Type: new
Abstract: Scope 3 greenhouse gas (GHG) emissions account for the majority of corporate carbon footprints, yet remain difficult to analyze at scale due to sparse disclosures, heterogeneous report document formats, and limited evidence traceability. Existing approaches typically rely on large language models to extract emissions information from ESG reports, but often lack explicit evidence grounding or depend on costly manual annotation and verification to ensure extraction reliability. To address these challenges, we propose Scope3Trace, an evidence-grounded information extraction framework designed to extract interpretable and traceable Scope 3 emissions information from real-world ESG and sustainability reports. The framework integrates a document information extraction pipeline that performs PDF collection and OCR parsing, LLM-assisted page localization and table reconstruction, and hybrid rule-LLM extraction of organization- and building-level emissions disclosures with evidence-grounded verification. Building upon this framework, we further contribute a dual-level, evidence-grounded, multimodal dataset comprising organization-level Scope 3 disclosures extracted from heterogeneous sustainability reports. Scope3Trace enables reliable extraction and transparent integration of heterogeneous sustainability disclosures, achieving high accuracy in extracting Scope 1-3 totals and category-level disclosures from sustainability reports.
arXiv:2607.16877v1 Announce Type: cross
Abstract: The increasing complexity of next-generation wireless networks has driven the integration of artificial intelligence (AI) into wireless communications. However, most existing studies focus on developing task-specific deep learning techniques for single scenarios, which limits their ability to generalize across diverse tasks, channel conditions, and system configurations. To address this generalization bottleneck, we propose a hierarchical wireless foundation model (WFM) for multi-task optimization. The proposed WFM couples an upstream foundation channel encoder (FCE) with a downstream foundation optimization decoder (FOD) via geometry-aware cross-attention. Specifically, the FCE extracts task-agnostic channel representations via self-supervised masked reconstruction while the FOD generates multi-task optimization decisions through differentiable output heads. Moreover, a hybrid supervised-to-unsupervised training strategy is employed to overcome the performance ceiling of purely supervised learning, and the modular architecture of the WFM enables efficient adaptation to unseen communication tasks with minimal parameter overhead. Simulation results show that the proposed WFM learns high-fidelity channel representations and achieves competitive multi-task optimization performance while substantially reducing optimization inference latency relative to numerical baselines. Furthermore, it exhibits robust generalization to unseen propagation environments, varying constraint parameters, and heterogeneous system configurations.
arXiv:2607.18062v1 Announce Type: new
Abstract: This paper focuses on the problem of Embodied Task Planning, where an agent is required to execute a sequence of atomic actions within an interactive environment to complete a user-specified task. Though a variety of simulators and datasets have previously been built for this task, these efforts are largely isolated, with each using its own observation format, action type, and task domain. This fragmentation complicates comprehensive model evaluation and hinders the scalability of training data. As an effort towards generalizable embodied planning, we propose UniETP, a unified interface integrating four commonly-used simulators (AI2-THOR, VirtualHome, Habitat, BEHAVIOR). UniETP is characterized by both standardization and diversity. On one hand, it formalizes all the simulators into a consistent observation and action space, and builds an evaluation system to support complicated task goal. On the other hand, it enhances task diversity and complexity across dimensions like task logic, instance grounding, and instruction understanding, constructing a new dataset with varied levels of difficulty in an automatic manner. Extensive experiments on the proposed benchmark are conducted to evaluate the embodied planning capabilities of recent models and analyze the performance bottlenecks. Codes and data will be available at https://github.com/woyut/UniETP .
arXiv:2607.17684v1 Announce Type: cross
Abstract: We study universal monotonicity and Frank--Wolfe stability properties for atomic splittable congestion games. Specifically, we characterize the largest resource cost class for which the associated variational inequality operator is monotone for every game. This characterization is given by a curvature inequality involving the first two derivatives of the allowable cost functions and the number of players. Our framework yields exact characterizations for universal monotonicity, strict monotonicity, and strong monotonicity; the strict and strong variants require corresponding stricter curvature conditions. We then draw a perhaps surprising connection to learning dynamics in atomic splittable congestion games. We show that the very same curvature condition also characterizes universal \emph{local and global stability} of the Euclidean-regularized Frank--Wolfe dynamics on arbitrary convex strategy spaces, provided the cost class is closed under positive affine transformations. Finally, we study games on simplices and show that an interior equilibrium of the regularized Frank--Wolfe dynamic is locally exponentially stable, even without the curvature condition.
arXiv:2607.16981v1 Announce Type: new
Abstract: An agent acting under partial observability must decide when to gather information and which observations are worth their cost. Standard POMDPs value information only through its eventual effect on reward. The $\rho$-POMDP framework instead rewards uncertainty reduction directly, through a belief-dependent utility $\rho$, but in practice both the choice of $\rho$ and the weight placed on it are tuned by hand for every task. We show that active inference removes this tuning entirely. Minimizing Expected Free Energy (EFE) is exactly equivalent to solving a $\rho$-POMDP whose utility is expected information gain, and the exploration weight is fixed at $w=1$ because the variational bound expresses pragmatic and epistemic value in the same units (nats). We prove this equivalence for observe-then-commit POMDPs and extend it to factored observation POMDPs, a broader class that covers interleaved observe-act problems such as non-destructive testing and mobile sensing, where gathering information leaves the hidden state unchanged. Experiments support the theory. Across environments ranging from the classic Tiger problem to RockSample and a new Structural Inspection benchmark with over 65,000 states, the untuned weight matches or outperforms reward-only planning at the same horizon, avoids the over-exploration of bonuses tuned per task, and sits near the reward-maximizing knee of the success-reward Pareto frontier. The practical payoff is an exploration objective that works out of the box. In applications such as fault detection and medical screening, where every test has a price and every missed fault has a cost, EFE supplies a belief-dependent utility that is derived rather than tuned.