arXiv:2606.11632v1 Announce Type: new Abstract: Agentic infrastructure introduces a critical control-plane authorization problem: non-deterministic reasoning systems can propose high-stakes mutations to production resources, yet existing security mechanisms -- such as identity and access management (IAM), policy engines, consensus protocols, and audit logs -- either enforce static, context-unaware permissions or merely record actions post-execution. This paper introduces the Sovereign Assurance Boundary (SAB), a certificate-bound runtime admission layer for autonomous execution authority. SAB intercepts agent proposals at an assurance airlock, compiles them into typed execution contracts $C$, and binds these contracts to cryptographic evidence digests $H(E)$ and policy versions. The contracts are then routed through consequence-aware certification paths. Upon successful admission, the system emits a signed Sovereign Assurance Certificate ($\Omega$) that is strictly scoped to a specific execution identity, revocation epoch, and validity window. Finally, a sovereign execution broker verifies $\Omega$ and performs fresh pre-execution revocation and drift checks before invoking infrastructure APIs. We detail the airlock-broker architecture, formalize its admission and revocation invariants, and report preliminary feasibility measurements from a Go prototype evaluated over 2,500 admission attempts. Ultimately, this broker-enforced model prevents autonomous reasoning from directly mutating state, transforming delegated execution authority into a cryptographically verifiable, evidence-bound, revocable, and replayable runtime artifact.
Science Journals
arXiv:2606.11640v1 Announce Type: new Abstract: Few-shot tabular learning provides a cost-effective approach for real-world applications where annotation is costly and collecting sufficient samples for new tasks is difficult. Existing Traditional and LLM-based methods have demonstrated effectiveness in few-shot scenarios. However, traditional methods need additional training on unlabeled or generated data, which incur significant computational overhead. In addition, LLM-based methods that directly feed raw tabular data into LLMs raise privacy and compliance concerns. More importantly, both paradigms largely overlook the semantic relationships between features, which provide structural and semantic prior for constructing a semantic graph. Semantic graph is essential for modeling meaningful feature interactions in few-shot scenarios. In this paper, we propose TAROT, a GNN-based framework that encodes the structural and semantic prior by constructing and refining a task-adaptive semantic graph from this prior, thereby improving predictive performance in few-shot tabular learning. TAROT first encodes heterogeneous tabular data into unified node semantic representations via a Unified Semantic Tabular Node Encoder (USTNE). Then, it prompts LLMs to infer the semantic relationship between features based on the task description and feature names to construct a semantic graph. To mitigate structural noise introduced by the hallucination of LLMs, TAROT introduces Task-adaptive Semantic Graph Refinement that prunes spurious or task-unrelated edges and adds missing task-related ones, aligning the graph structure with the downstream objective. Finally, a GNN performs message passing over the refined graph to capture task-related semantic dependencies for prediction. Extensive experiments on various few-shot tabular learning benchmarks demonstrate the superior performance of TAROT, establishing it as a state-of-the-art approach in this domain.
arXiv:2606.11642v1 Announce Type: new Abstract: How far can we reduce the number of physical keys if we endow an ambiguous keyboard with modern language models? Fewer keys increase hardware design freedom in constrained settings such as assistive devices and mobile form factors. This paper systematically evaluates text entry systems using 2-5 physical keys combined with language-model-based disambiguation. On a 300-sentence English corpus (100 sentences each for Business / Conversational / Technical), we compare key counts (2-5), letter-to-key mappings (layout-based / frequency-based / intentionally worst-case), and decoders (Trie-only, GPT-2 beam search, GPT-4o selection). We find that 3 keys + GPT-4o achieves character error rate (CER) 9.46% and word error rate (WER) 12.20%, reducing CER by 59% relative to 2 keys (CER 23.3%). At 3 keys, the key-stream entropy is 1.54 bits/char; while increasing to 5 keys improves accuracy (CER 5.4%), the marginal gains diminish. Mapping choice has a small impact under standard designs ({\Delta}CER < 0.5 pp), and even an intentionally worst mapping degrades CER by only +0.5 pp, whereas Technical sentences yield roughly twice the error rate of Business. These results suggest that, in our evaluated offline setting under a strong LM prior, 3 keys are a practical minimum for general English.
arXiv:2606.11645v1 Announce Type: new Abstract: Micro-gesture analysis attracts increasing attention for inferring spontaneous emotion from subtle body movements. Micro-gesture online recognition, which localizes and classifies each gesture instance in untrimmed videos, is a core task in the 4th EI-MiGA-IJCAI Challenge. Compared with typical temporal action detection, MGR emphasizes the localization and classification of actions, requiring the model to output the start time, end time, and category of each micro-gesture. Moreover, since micro-gestures are highly spontaneous, relying solely on a single modality makes it difficult to capture the complete and accurate multi-modal cues. In this work, we propose DyFADet+, which extends DyFADet into a dual-stream RGB-skeleton framework. In our model, both modalities are projected into shared multi-scale temporal embeddings and fused through a gated residual module, which adaptively injects skeleton motion into the RGB representation rather than using naive concatenation. Finally, these fused features are decoded by a Dynamic TAD head for online classification and boundary regression. On the SMG dataset, our method achieves an F1 score of 40.88, ranking 2nd in the Micro-gesture Online Recognition track.
arXiv:2606.11647v1 Announce Type: new Abstract: On a charged interface, broken inversion symmetry permits a large field-linear Pockels response through $\chi^{(2)}(\omega;\omega,0)$; at the water/ITO interface $|r_{13}|$ has been reported to reach the $10^{2}\,\mathrm{pm/V}$ order. The coexisting third-order DC Kerr term $\chi^{(3)}(\omega;\omega,0,0)$ -- small in bulk water ($|\chi^{(3)}_{\mathrm{bulk}}|\sim5.5\times10^{-21}\,\mathrm{m^2/V^2}$) -- had not been jointly parameterized with the Pockels term along the fundamental-frequency ($\omega$) electro-optic path. Superimposing an AC modulation and a DC bias in 0.1\,M NaCl mixes the $\chi^{(3)}$ contribution with the 1f response through the cross-term $2 s_{1133}\,E_{\mathrm{DC}}$, so that the water refractive-index modulation $\Delta n_{\mathrm{water}}$ varies linearly with $V_{\mathrm{WE}}$; a model-assisted linear fit then determines both terms from a single AC$+$DC sweep. At $V_{\mathrm{WE}}=0\,\mathrm{V}$ vs Ag/AgCl, $|r_{13}|=(1.18 \pm 0.06_{\mathrm{PZC}})\times10^{2}\,\mathrm{pm/V}$ and, under $\eta_{\mathrm{DC}}=1$, the thickness-normalized DC Kerr coefficient $|s_{1133}/d_{\mathrm{EDL}}|=33.0 \pm 5.6\,\mathrm{pm/V^{2}}$. Across physically reasonable $d_{\mathrm{EDL}}$ ($0.6$--$1.6\,\mathrm{nm}$), the interfacial DC Kerr susceptibility reaches $|\chi^{(3),\mathrm{int}}_{1133}| \approx (2\text{--}5.5)\times10^{-20}\,\mathrm{m^2/V^2}$, several-fold above the visible-range bulk-water value. This response is a property of the specific interface, tunable through the choice of electrode, electrolyte, and solvent rather than intrinsic to bulk water. Amid renewed interest in the Kerr response of water (including recent THz-band optical Kerr studies), the method directly probes this DC Kerr term along the $\omega$ path and complements SHG/SFG (the $2\omega$ path).
arXiv:2606.11652v1 Announce Type: new Abstract: This paper investigates reinforcement learning (RL) methods for improving tool-calling capabilities in multimodal small language model (SLM) agents. While existing works have explored various reward designs to improve agentic tool-calling ability, these approaches face inherent limitations for SLM training, especially under multimodal scenarios. First, many existing methods evaluate tool use correctness through exact matching against certain ground-truth or predefined formats. However, this assumption is often unsuitable for multimodal tasks, where multiple tool use paths may be valid and annotated tool trajectories are typically unavailable. Second, such sparse and brittle binary rewards provide little guidance on how to improve the underlying decision process, making them particularly difficult for multimodal SLM to learn from. To address these issues, we propose Input Attribution-Aware Policy Optimization (IAPO), an RL algorithm for improving tool use in multimodal SLM by aligning the model's attribution across input components with that of a stronger teacher. Experiments on Qwen2.5-VL-3B show that the proposed method improves visual question answering accuracy by an average of 3% across six test sets compared with existing visual tool use work, by helping the model attend to the most relevant input evidence.
arXiv:2606.11666v1 Announce Type: new Abstract: Open-set source tracing is increasingly framed as a verification problem, motivating the use of pairwise metric-learning objectives from biometrics. We thus compare global anchoring and pairwise verification under matched backbones and a fixed data and epoch budget on MLAAD (in-domain) and STOPA (out-of-domain). In our runs, global anchoring yields lower in-domain error (8.61% EER) than pairwise variants (12-15% EER), even with rival mining and XLS-R finetuning. Because pairwise objectives optimize similarity directly, they concentrate variance into fewer embedding directions, reducing resolution among closely related generators. To test if this drives the drop, we impose a similar bottleneck to the globally supervised baseline, yet the baseline remains competitive. Together with an embedding-space analysis ($k_{99}$), these results suggest that the gap is not explained by dimensionality alone, but rather by the pairwise objective's shaping of the retained directions.
arXiv:2606.11674v1 Announce Type: new Abstract: We present SpAArSIST, a deployment-oriented refinement of the widely used AASIST graph pooling backend for self-supervised learning (SSL) based anti-spoofing. Motivated by redundant operations in public implementations, we replace learned pooling and stack-node attention with explicit, lightweight choices: separate train and inference graph pooling ratios $(k_{\mathrm{tr}},k_{\mathrm{inf}})$, magnitude-based node scoring, and mean aggregation of graph nodes. The best overall configuration (rank 1) cuts backend compute by 20.7% (195.045M $\rightarrow$ 154.706M MACs) and model size by 4.1% (611.8k $\rightarrow$ 586.4k params), while improving out-of-domain robustness on In-the-Wild to 2.82% EER and 0.078 minDCF (from 4.64% and 0.133) and remaining competitive on ASVspoof5. We further provide a composite selection score that summarizes accuracy, calibration, and compute to support balanced deployment-oriented model choice.
arXiv:2606.11698v1 Announce Type: new Abstract: Model watermarking safeguards AI model intellectual property by embedding distinctive knowledge that induces unique behavioral signatures. The primary technical challenge lies in ensuring watermark robustness against various post-processing attacks on the watermarked model. Model extraction attacks emerge as the most severe threat, where adversaries exploit prediction outputs to train surrogate models that illegally replicate the original model's functionality. In this work, we propose a rehearsal-based watermark embedding framework to enhance the robustness of model watermarks against model extraction attacks. By simulating the extraction process, our method leverages the loss of a \textit{simulated stolen model} on a trigger set as a training signal to fine-tune the watermark knowledge within the target model. This fine-tuning step encourages the watermark to be embedded in a way that boosts transferability, thereby increasing its chances of persisting and remaining detectable in stolen models. Comprehensive experiments conducted under diverse settings demonstrate that the proposed method significantly improves the robustness of model watermarks against both model extraction and subsequent watermark removal attacks.
arXiv:2606.11701v1 Announce Type: new Abstract: Multireference alignment (MRA) is the task of recovering a hidden "signal" vector, given many noisy copies that have been cyclically shifted by unknown offsets. This task belongs to the class of orbit recovery problems, in which the observed samples are affected by some group action. These problems have a variety of practical motivations, including the reconstruction of 3-dimensional molecular structure from cryogenic electron microscopy (cryo-EM) images. We consider two variants of MRA: dihedral MRA, where the cyclic group is replaced by the dihedral group, allowing for reversals of the vector in addition to shifts; and projected MRA, where the observations are passed through a projection operator akin to the tomographic projection present in cryo-EM. We apply the method of moments and aim to recover the signal from the third moment tensor of the samples. This inverse problem is well understood for basic MRA, but for the variants we consider there is no polynomial-time algorithm known to succeed for generic signals. We give the first such algorithm for both of these variants. Our method requires the signal length to be a power of two, and recursively subdivides the problem into smaller problems of half the size. The algorithm's success for generic signals is proven, conditional on a conjecture about the rank of a certain symbolic matrix of polynomials. For any given problem size, this conjecture can be verified on a computer.
arXiv:2606.11710v1 Announce Type: new Abstract: This paper presents ERN-Net, an Evolving Reason Node-Net for efficient document image binarization. ERN-Net enhances degradation-sensitive regions, such as faint strokes, broken characters, and noisy backgrounds, through evolving reason nodes and multi-scale reasoning. We further compare ResNet-101, ConvNeXt-Tiny, and ConvNeXt-Base, and find that ConvNeXt-Tiny provides the best practical trade-off between accuracy and memory usage. In addition, DIBCO-based pretraining improves binarization performance without increasing model memory consumption, requiring only about 1.5 additional training hours. Experiments on DIBCO-style benchmarks show that ERN-Net is effective under low-data and low-memory settings.
arXiv:2606.11729v1 Announce Type: new Abstract: Industry has embraced Zero Trust (ZT) architectural tenets and implementations for cloud-native environments, following stricter security requirements to both internal and external tenants. Among others, these approaches combine fine-grained identity management and monitoring for both inventorying and better analysing the devices' security posture for overall protection, along with strict separation of concerns and isolation to enforce minimal privilege. Networking-wise, ZT approaches rely as well on isolation and least privilege; enacted by separate, secure tunnels per tenant connecting to a given infrastructure. Such implementations can also be applied to the connectivity within and towards experimental infrastructures. In this sense, this work contributes the design and evaluation of a cloud-native VPN-as-a-Service (VPNaaS) that can be (i) easily orchestrated to deploy on-the-fly, separate tunnels per each tenant remotely connecting to the infrastructure; (ii) integrated with common Identity and Access Management (IAM) tools, key to ZT deployments; and (iii) adapt to computing- or entropy- constrained environments. This solution is customisable and allows, among others, to select from RSA or Elliptic Curves (EC) as key generation algorithm and their parameters to achieve more secure keys and adapt to resource-constrained environments.
arXiv:2606.11730v1 Announce Type: new Abstract: How should one design efficient chemically open optical cavities for molecular strong coupling? Addressing this question is important for the development of soft-cavity platforms for dynamically tunable light--matter interactions, where direct access to confined electromagnetic modes is essential. Conventional cavity figures of merit such as $Q/\sqrt{V}$ and cooperativity successfully describe spectral confinement and dissipation but do not fully capture the role of linewidth asymmetry between cavity and molecular degrees of freedom. Here, we systematically investigate strong coupling between TDBC dye molecules and whispering gallery modes of polystyrene microspheres by varying the microsphere radius over a broad range. To quantify the robustness of strong coupling, we define the parameter $\chi = \frac{g}{\max(\kappa,\gamma)}$, where $g$ is the coupling strength, while $\kappa$ and $\gamma$ denote the cavity and molecular linewidths, respectively. Although the coupling strength decreases monotonically with increasing cavity size due to mode-volume scaling, we find that $\chi$ exhibits a pronounced maximum near the condition $\kappa \approx \gamma$. This observation suggests that linewidth matching is not merely a criterion for improved spectral visibility, but reflects a dissipation-matching condition that optimizes the robustness of coherent light--matter exchange in soft-cavities. Our results provide an alternative framework for designing morphology-dependent cavities for molecular strong coupling.
arXiv:2606.11734v1 Announce Type: new Abstract: This paper presents a novel class of high-order mass-, energy- and momentum-preserving exponential integrators for solving the derivative nonlinear Schr\"{o}dinger equation. Firstly, we reformulate the original system into an exponential supplementary variable system based on the idea of the exponential supplementary variable approach, and then the reformulated system is discretized by using the standard Fourier pseudo-spectral method in space and the high-order prediction and correction Lawson Runge-Kutta method in time, respectively. The proposed method is highly efficient, temporally high-order accurate, and simultaneously preserves the mass, energy and momentum in the discrete setting. Finally, numerical experiments validate the accuracy and energy-preserving properties.
arXiv:2606.11740v1 Announce Type: new Abstract: We study whether grounded reasoning supervision from abundant 2D medical images can improve 3D medical VQA when both input types are aligned through a common reasoning interface. We introduce UniReason-Med, a single-checkpoint framework that processes either a 2D image or a slice-serialized 3D volume at inference time, generating interleaved textual reasoning and localized visual evidence through shared box syntax, region-token injection, and a common grounded reasoning policy. To train this interface, we construct UniMed-CoT, a 220K instruction-tuning dataset with interleaved textual reasoning and grounded visual evidence, including 170K 2D and 50K 3D samples. Through supervised fine-tuning followed by outcome-level reinforcement learning, UniReason-Med learns to generate grounded reasoning traces without IoU/Dice-based localization rewards during RL. Data-mixture and component ablations show that joint 2D+3D grounded supervision substantially improves 3D reasoning over 3D-only training, while grounding and region-token injection consistently benefit both 2D and 3D tasks. These results suggest that a shared grounded reasoning interface can transfer reasoning structure from 2D images to slice-serialized volumetric medical understanding. The code and data are publicly available at https://github.com/IQuestLab/unireason-med.
arXiv:2606.11745v1 Announce Type: new Abstract: Visual causal reasoning is essential for understanding and intervening in the physical world, requiring identification of causal variables from visual inputs and reasoning over intervention effects. Despite recent progress, large vision--language models (VLMs) remain brittle at such tasks, especially for interventional and counterfactual queries over multi-image inputs. Most existing explorations inject causal knowledge via textual prompts, leaving causal mechanisms external to model execution and limiting reliable control during inference. To address this problem, we propose BridgeVLM, which internalizes visual causal reasoning by inducing a causal graph from multi-image inputs and converting it into structured Causal Tokens executed by RAMP layers injected into the LLM decoder for causal message passing. We further introduce a unified training interface M3S for fine-grained causal supervision from different granularities (local/global level). BridgeVLM achieves 54.4% accuracy on intervention tasks on CausalVLBench (vs. 33.2% with prompt-level supervision), improves results on Causal3D from 43.6% to 49.0%, and substantially improves causal structure learning on CausalVLBench ($F_1$: 33.4% $\rightarrow$ 75.1%).
arXiv:2606.11764v1 Announce Type: new Abstract: Recently, constructions of linear complementary pairs (LCPs) of codes and linear complementary dual (LCD) codes on function fields have attracted considerable attention due to the wide range of applications of these codes. Such constructions rely on non-special divisors of degrees $g$ and $g-1$. In this work, we investigate Kummer extensions defined by $y^m = f(x)$ with $f(x)\in\mathbb{F}_q(x)$ and establish an arithmetic characterization of non-special divisors whose support can contain non-totally ramified places. Based on this characterization, we explicitly construct non-special divisors of degree $g-1$ on the GK curve. Moreover, utilizing pure gaps, we explicitly provide several families of effective non-special divisors of degree $g$ on Kummer extensions with the same multiplicities. We then develop a general framework for constructing LCPs of algebraic geometry (AG) codes on Kummer extensions. By virtue of canonical divisors, we show that the security parameters of LCPs of AG codes can be determined within this framework, which also enables the construction of LCD AG codes. Finally, we illustrate our results with representative examples, including LCPs of codes on the GK curve and LCD codes on quotients of the Hermitian curve.
arXiv:2510.22353v2 Announce Type: replace Abstract: We present evidence of the dynamical hysteretic nature of dissipation in unsteady turbulent flows. Wind tunnel experiments and direct numerical simulations in oscillating flows show that, at stationary mean Reynolds number, the dissipation constant is larger for decelerating flows. Consequently, a periodic behavior of the flow produces a hysteresis cycle, whose area scales with a parameter combining the Strouhal number and the relative amplitude of the forcing. This phenomenon can be explained and quantified through the influence of the unsteady term in the Karman-Howarth equation, with implications for a wide range of out-of-equilibrium systems.
arXiv:2606.11771v1 Announce Type: new Abstract: In this paper, we propose a segment-wise soft robotic antenna (SRA) system, where each soft robotic arm referred to as a tentacle, comprises multiple independently controllable segments with bending, elongation-retraction, and sweeping motions. By adjusting segment motion parameters, the positions of surface-mounted antennas are reconfigured, distinguishing it from conventional reconfigurable antenna (RA) systems. Based on this model, we propose two antenna deployment schemes: the segmented end-antenna configuration (SEAC), where fixed antennas are mounted at the segment ends and reconfigured via segment motions; and the hybrid end-and-intermediate antenna configuration (HEIAC), where RAs are further integrated as intra-segment antennas. In HEIAC, soft-robot segment deformation provides large-scale spatial reconfiguration, while RAs enable fine-grained adjustment. For SEAC, we formulate a sum-rate maximization problem accounting for inter-segment connectivity and the nonlinear mapping from segment deformation parameters to antenna coordinates, and develop a penalty dual decomposition-projected gradient ascent (PDD-PGA) algorithm. For HEIAC, we jointly optimize segment deformation, intra-segment antenna positions, and antenna activation using a block coordinate descent (BCD)-PDD-PGA algorithm with greedy backward antenna selection. Simulation results demonstrate that the proposed schemes substantially outperform fixed-position antenna arrays and conventional RA baselines. In particular, SEAC and HEIAC achieve 37.9% and 32.1% sum-rate gains over conventional 3D reconfigurable arrays, respectively, while SEAC provides up to a 49.3% gain in compact array deployments.
arXiv:2606.11772v1 Announce Type: new Abstract: Originally motivated by creating first-person computer visualizations within Riemannian manifolds -- the author was led to study deformable-body mechanics, as rigid-body mechanics is not available in a generic Riemannian manifold due to its lack of nontrivial isometry group. Hyperelasticity is a particularly nice sub-category of continuum mechanics in which a deformable, elastic body's behavior is determined by a stored energy density function. This allows problems to be posed variationally, and powerful tools brought to bear on studying and solving them. This article presents numerical simulations of static solutions to a particular class of problems in hyperelastic mechanics in 2-dimensional Riemannian manifolds in which a flat hyperelastic body $B$ is embedded into a region $\Omega$ in a nowhere-flat surface $S$ of revolution $z=z\left(r\right)$ such that $\left|K\left(r\right)\right|$ decreases as $r\to\infty$, where $K$ denotes the Gaussian curvature of $S$. For example, the funnel $z=-r^{-1}$ or the paraboloid $z=\frac{1}{2}r^{2}$. Because $B$ is flat, the body can't achieve a zero-stored-energy configuration, and restorative forces arise in the body to move it toward a region of lower stored energy -- meaning, toward a flatter configuration. With the addition of a gravitational potential $U\left(r\right)=z\left(r\right)$ on $S$, forces act on the body to pull it toward $r=0$. If the body has sufficient stiffness and remains within the region $\Omega$, then the body has an equilibrium configuration in which the body's deformation-response forces perfectly cancel the gravitational forces. Such a configuration represents a kind of "levitation" phenomenon within this surface. The numerical implementation of this problem will be detailed and the resulting numerical solutions and various consequences discussed.
arXiv:2606.11779v1 Announce Type: new Abstract: The need for detecting and sorting batteries is drastically increasing for many applications. This study proves the potential of transfer learning in predicting whether the image contains a battery or not, the location and identifying three types of batteries, namely: prismatic, pouch, and cylindrical Lithium-Ion Batteries (LIB). Particularly, it focuses on the transfer learning method in two applications: Training a large-scale dataset to detect electronic devices using a pre-trained YOLOv5m, then using these latter trained weights to detect and classify the batteries. The precision of battery detection achieves 94%, which outperforms the pretrained YOLOv5m weights with 5%, in 22 ms inference time.
arXiv:2606.11780v1 Announce Type: new Abstract: We establish conditions for embedding a corpus of $N$ documents as $d$-dimensional vectors such that every $k$-subset $S \subseteq [N]$ is realizable as a result of top-$k$ retrieval by some query vector. Recent work shows that $d = O(k)$ suffices for such embeddings to exist in $\mathbb{R}^d$, independently of $N$. We theoretically prove that this corpus-independent bound is specific to infinite precision. With $B$ bits per coordinate, perfect top-$k$ retrieval requires $Bd = \Omega(k \ln N)$; thus, at any fixed precision, the dimension must grow at least logarithmically with $N$. Specializing to a $\ell_2$-normalized $B$-bit uniform scalar quantization model, we also identify a threshold on the precision $B^{*} = O(\ln \ln N)$ below which no dimension suffices, together with two further regimes that bound the feasible $(B, d)$ pairs. Our result implies that in practical vector databases and dense retrieval systems where quantization is standard, the embedding dimension and possibly the precision must grow with the corpus size.
arXiv:2606.11797v1 Announce Type: new Abstract: Studies on rodents such as mice have shown the capabilities to adapt their behavior when dealing with changing parameters (``drift'') of the environment even if no information about change is provided (uncertainty) -- a behavior that can be modeled by forgetting mechanisms. Non-stationary Reinforcement Learning (NSRL) deals with adapting state-of-the-art RL methods to deal with changing environments: these however usually require (partially) perfect information about the drift such as ``task IDs'' or ``context''. To mitigate the effects of drift, this work develops \emph{Space-sampled Value Decay} as an explicit forgetting mechanism for value-based deep RL architectures as a simple yet effective approach. In particular we demonstrate and discuss positive effects but also limitations in achieved returns for modifications of Deep Q-networks (DQN) and Soft Actor-Critic (SAC) when evaluated on non-stationary environments.
arXiv:2606.11801v1 Announce Type: new Abstract: We propose novel semi-decoupled and fully-decoupled iterative algorithms for efficiently solving the fully-coupled nonlinear four-field thermo-poroelastic model discretized in space by discontinuous Galerkin method on polytopal grids. We present the model problem, its four-field formulation, and the arbitrary-order weighted symmetric interior penalty scheme exploited for its spatial discretization. Such a scheme is robust with respect to strong heterogeneities in the model coefficients. Then, we present the two solution strategies and prove that under suitable conditions both schemes are convergent. A wide set of numerical simulations is presented to assess the convergence and robustness properties of the proposed method. Moreover, we test the scheme with literature and physically sound test cases for proof-of-concept applications in the geophysical context.
arXiv:2606.11805v1 Announce Type: new Abstract: Text-conditioned 3D generation has progressed rapidly for images and isolated objects, but producing a hand-object mesh remains challenging: the output must preserve language semantics, cross-view consistency, object geometry, articulated hand shape, and physically plausible contact. We present TextHOI-3D, a staged framework that uses generated multi-view observations as an explicit interface between text-conditioned visual generation and geometry-aware hand-object recovery. TextHOI-3D learns a compact VQ token space for fixed-camera hand-object observations, predicts multi-view visual tokens from text with a CLIP-conditioned visual autoregressive model, and recovers a unified hand-object mesh through prior initialization, multi-view joint optimization, and anti-penetration refinement. The design separates semantic generation from geometric recovery while keeping both stages connected by a discrete multi-view representation. On HO3D-derived evaluations, the multi-view setting reduces object CD from 17.26 mm to 4.92 mm and penetration volume from 5.3721 cm^3 to 0.2193 cm^3 compared with a single-view counterpart, while improving hand errors and surface F-scores. These results support multi-view visual tokens as an effective intermediate representation for text-driven 3D hand-object mesh creation.