Forskningsradar

Science Journals

Peer-reviewade publikationer — 62341 artiklar

Masked Language Flow Models
arXiv:2606.27617v1 Announce Type: new Abstract: Masked Diffusion Models (MDMs) promise fast, parallel language generation, but their reverse transition factorises across token positions -- an approximation that breaks down in the few-step sampling regime where parallel generation ought to provide the greatest efficiency gains. Flow Language Models (FLMs) sidestep this limitation by learning a continuous flow that transports noise toward clean sequences represented in Euclidean space, inducing a flow map that can be distilled for single-step generation. However, this makes complex tasks requiring multi-step reasoning problematic for FLMs, as FLMs are forced to decode every token during generation. To address this, we introduce Masked Language Flow Models (MLFMs), which incorporate masking into FLMs using a continuous stochastic interpolant to bridge partially masked and clean sequences. This design enables conditional generation via continuous flows and allows pretrained MDMs to be converted into MLFMs through a simple, lightweight adaptation. Leveraging this flexibility, we propose a novel sampler that alternates continuous denoising with the discrete unmasking of confident tokens to better support multi-step reasoning. We evaluate our approach on GSM8K and MT-Bench and find, for the first time, that flow-based language models can be scaled to solve downstream reasoning and instruction-following tasks.
Optothermal Actuation of Unidirectional Thermo-osmotic Flows
arXiv:2606.27735v1 Announce Type: new Abstract: In this paper, we experimentally demonstrate the microscale direction control of thermoosmotic flows using a focused-laser heating. The key is the off-center laser irradiation on an immobilized light-absorbing microparticle, which generates a nonuniform, asymmetric heat source. The resulting thermo-osmotic flows are evaluated using the optically trapped particle tracking velocimetry (ot-PTV), presented in our preceding paper (T. Tsuji, et al., Physical Review Fluids 11, 034901 (2026)). It is shown that the flow characteristics can be modulated by the ionic strength of a sample solution and/or the surface molecular coating of the substrate. In particular, the significance of ionic strength on thermo-osmotic flows are discussed based on the surface potential of the substrate measured by frequency-modulated atomic force microscopy.
From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection
arXiv:2606.27751v1 Announce Type: new Abstract: This report investigates the extension of pretrained General-Purpose Audio Tagging (GP-AT) models toward spatially grounded Sound Event Localization and Detection (SELD). The proposed AT2SELD framework couples a pretrained AT backbone with compact First-Order Ambisonics (FOA) spatial processing, track-wise SED and Cartesian DOA estimation, permutation aware supervision, and calibration. It characterizes how semantic audio priors support localization-aware scene analysis under data, computation, and deployment constraints. The framework is developed through informed multi-stage Neural Architecture Search (NAS). Stage 1 shows that spectral FOA descriptors, based on magnitude, phase, and Intensity Vectors (IVs), provide the most reliable interface for semantic-to-spatial transfer. Stage 2 identifies early residual spatial encoding as the main capacity-sensitive component, while late track-wise abstraction and recurrent smoothing act mainly as refinement stages. Stage 3 shows that late cross-stitch coupling improves semantic-spatial interaction, whereas early fusion is costlier and less effective. Diagnostic evaluation analyzes the selected architecture under class balancing, focal loss, activity-conditioned DOA supervision, threshold calibration, and transfer across STARSS23, TAU2019, TAU-NIGENS2020, and TAU-NIGENS2021. Focal loss improves the activity point, active-only DOA supervision mitigates inactive target dominance, and validation-selected thresholds recover calibration without replacing spatial learning. Cross-dataset and oracle-activity analyses indicate strong fixed source localization on TAU2019, transferable representations from TAU NIGENS2021, and meaningful but uncertain behavior on STARSS23. Overall, GP-AT priors appear promising for SELD design when embedded in spatial-aware architectures and optimized through integrated calibration and deployment oriented strategies.
Statistical equilibria of two-dimensional turbulent flows for generic initial vorticity fields on a sphere, calculated on the basis of the original Miller-Robert-Sommeria theory
arXiv:2606.27778v1 Announce Type: new Abstract: Based on the original Miller-Robert-Sommeria theory, we explicitly compute a statistical equilibrium of two-dimensional turbulent flow on a sphere for a generic initial vorticity field introduced in a previous study. The macroscopic vorticity field corresponding to the obtained statistical equilibrium has a quadrupole structure. The resulting quadrupole structure is topologically consistent with the final state of the long-term time integration of the vorticity equation. However, the statistical equilibrium does not predict the formation of concentrated vortices as seen in the time integration. We also calculate statistical equilibria for the initial vorticity field with a planetary vorticity term, and find a change of statistical equilibria from quadrupole states to zonally symmetric states as the angular velocity of the sphere increases. The quadrupole statistical equilibria show nearly linear relations between the macroscopic vorticity and the macroscopic stream function, implying that higher-order Casimir invariants are virtually ineffective even when all Casimir invariants are considered. The discrepancy between the equilibria and the time integration results emphasizes the importance of mixing barriers, which prevent the relaxation of the evolving vorticity field to the statistical equilibria and allow the point-vortex-like dynamics of coherent vortices to persist.
Continual Learning for Sequential Personalization of Small Language Models: A Stability Monitoring Analysis
arXiv:2606.27634v1 Announce Type: new Abstract: Small Language Models (SLMs) are increasingly being considered for deployment on edge devices such as laptops, enabling private, low-latency, and locally personalized applications. However, personalization requires models to adapt over time to evolving user- or task-specific data, placing them in a continual learning setting. This creates the risk of catastrophic forgetting, where learning new information degrades performance on previously learned tasks or broader model capabilities. Recent benchmarks such as TRACE have shown that continual fine-tuning can significantly degrade the general abilities of aligned large language models. In this work, we present a study for sequential LoRA personalization of SLMs. We save model checkpoints after each adaptation stage and evaluate them on current tasks, previously seen tasks, and a fixed reference set. This checkpoint-level protocol enables us to monitor task performance, forgetting, and reference set drift over time. We show that lightweight reference set distributional diagnostics can reveal model-specific instability patterns during sequential LoRA personalization of SLMs, including cases where task-level metrics alone hide harmful adaptation. We hope this can highlight new research avenues for monitoring stability of SLMs in a continual learning setting.
SpikeVLA: Vision-Language-Action Models with Spiking Neural Networks
arXiv:2606.27807v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have become a dominant paradigm for embodied intelligence. However, most existing approaches are built on large-scale transformers, resulting in substantial inference latency and energy consumption that limit their practical deployment in low-power, real-time scenarios. We propose SpikeVLA, a spiking VLA architecture for embodied navigation with energy-efficient inference, consisting of three key components. (i) A spiking vision encoder, Spike-V, that replaces dense continuous layers with event-driven spiking layers to reduce the energy consumption of visual representation learning. (ii) A multi-modal spiking large language model, Spike-L, that reformulates cross-modal reasoning with spiking dynamics and token-level event-driven sparsity to further lower computational cost. (iii) A spiking action policy network, Spike-A employs Laplacian-kernel population coding with a multi-layer fully connected SNN, and decodes spiking activities into stable and robust continuous control with energy-efficient inference under low-power constraints. Experiments on navigation and robotic control tasks show that SpikeVLA significantly reduces energy consumption and computational cost while maintaining competitive performance, highlighting its potential for low-power, real-time embodied intelligence.
Temporal-Emerged Prompting for Segment Anything in Multiframe Infrared Small Target Detection
arXiv:2606.27655v1 Announce Type: new Abstract: Accurately localizing and segmenting small targets in low signal-to-noise ratio (SNR) infrared sequences remains a challenging task. Since targets are often indistinguishable from the background in individual frames, existing methods, even when equipped with advanced foundation model and powerful inter-frame association mechanisms, still fail to detect them. Motivated by the observation that targets tend to emerge gradually from the background over time and become distinguishable, we propose Temporal-Emerged Prompting for Segment Anything Model (TEP-SAM), a principled framework designed to explicitly exploit such temporal-emerged cues to modulate and prompt SAM. TEP-SAM operates by jointly modeling global motion patterns and local motion deviations to locate potential targets. It further enhances target region features by leveraging motion discrepancy, thereby generating temporal-emerged cues for SAM and enabling non-interactive segmentation. By bridging large-scale semantic pretraining with task-specific temporal modeling, TEP-SAM effectively adapts SAM to the challenging multiframe infrared small target detection task. Extensive experiments demonstrate the effectiveness of our approach, particularly under severely low-SNR conditions and in complex dynamic background.
LLM-Assisted Model-Based GUI Testing for Vue.js Web Applications
arXiv:2606.27665v1 Announce Type: new Abstract: Vue.js is a popular framework for building modern web applications. As Vue.js functionality and tooling support grow, ensuring its reliability (through automated testing) is becoming increasingly important. Although model-based testing has been successfully used to automate graphical user interface (GUI) testing on other platforms, its application to Vue.js remains challenging: Transition candidates, which are spread across router configurations and single-file components (SFCs), must be concretized and normalized into an executable page transition graph (PTG) for testing. To address this, we propose the LLMVue framework, which uses a large language model (LLM) to generate a PTG from Vue.js source code. LLMVue infers component hierarchies and route transitions, merging them into a unified PTG across multiple SFCs. We evaluated LLMVue on a collection of ten open-source Vue.js projects from GitHub, using GPT-4o as the LLM backbone. The constructed graphs demonstrate high precision and recall, with low graph edit distance. LLMVue -guided testing also significantly improves the coverage and exploration efficiency, compared to a random exploration baseline (with the same time constraints). To the best of our knowledge, this is the first use of LLMs for model-based GUI testing of Vue.js applications using source-level PTG extraction.
Geometry-Preserving Reduced-Order Modeling via Immersed Tensor Decomposition (ITD)
arXiv:2606.27674v1 Announce Type: new Abstract: Body-fitted finite-element methods deliver high-order accuracy but hinge on a clean, watertight, conforming mesh, a requirement that breaks down for the geometrically imperfect CAD assemblies, image-based volumetric data, and voxel-native designs that pervade biomedical engineering and additive manufacturing, where mesh generation has become the dominant cost of the analysis cycle. Immersed methods on regular background Cartesian grids sidestep body-fitted meshing, but classical implementations integrate over irregular cut subdomains, destroying the tensor-product structure that enables separable, reduced-order methods such as tensor decomposition. In this paper we propose the \emph{Immersed Tensor Decomposition} (ITD) framework, which couples a mesh-free geometric representation via body-fitted function with the separable C-HiDeNN-TD reduced-order solver to enable large-scale simulation directly on regular background voxel meshes. The geometry is encoded in three steps: a signed-distance function represents the boundary, a body-fitted function $\Phi$ approximates it with controllable error, and a low-rank Tucker decomposition provides model-order reduction; for a fixed grid spacing $h$, accuracy is improved by raising the approximation order of C-HiDeNN interpolation up to degree $p$ with a linear background mesh. The central contribution is an exact Dirichlet formulation that enforces the boundary condition strongly by multiplying the trial function with $\Phi$, so that $u=g$ holds by construction without any variational penalty or interface quadrature. We establish an a priori error estimate for the formulation and assess it on canonical 2D/3D domains, demonstrating optimal convergence and robustness on non-Cartesian geometries discretized by regular voxel meshes.
Two-Stage Cross-Domain Cervical Abnormality Screening with Cytopathological Image Synthesis and Knowledge Distillation
arXiv:2606.27678v1 Announce Type: new Abstract: Cross-domain diagnosis remains a major challenge in cervical cell pathology due to pronounced domain shifts across institutions and the subtle visual differences among disease stages, which jointly impair model generalization. To address these issues, this paper proposes a two-stage framework for cross-domain cervical cell detection. In the first stage, we propose the Spatially-Continuous Unpaired Neural Schr\"odinger Bridge (SC-UNSB), which constructs a synthetic intermediate domain to mitigate cross-domain distribution shifts by modeling image translation as an entropy-regularized optimal transport process. In the second stage, we propose a dual-level feature alignment strategy within a knowledge distillation, which progressively aligns shallow structural features and deep semantic representations to facilitate the transfer of domain-invariant knowledge from the source to the target model. Experimental results demonstrate that the proposed method effectively mitigates domain shift and category ambiguity, improving the cross-domain detection performance.
Intuition-Guided Latent Reasoning for LLM-Based Recommendation
arXiv:2606.27684v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated impressive reasoning capabilities in complex problem-solving tasks, motivating their use for preference reasoning in recommender systems. Latent reasoning, which operates in continuous hidden spaces rather than discrete tokens, has recently emerged as a promising paradigm for LLM-based recommendation. However, existing methods often start from unconstrained reasoning points, where hidden representations are misaligned with target item embeddings, leading to suboptimal reasoning trajectories. Inspired by cognitive neuroscience, which suggests that human multi-step reasoning is guided by intuition as a latent prior, we propose \emph{IntuRec}, a two-stage framework that anchors latent reasoning with \emph{recommendation intuition}. In the extraction stage, the LLM-based recommender generates a top-$K$ candidate set based on users' histories as the source of intuition. In the injection stage, the candidate set is transformed into a preference-aligned intuition embedding using self- and cross-attention mechanisms, which initializes the reasoning start point and guides subsequent latent reasoning. By providing a semantically grounded starting point, IntuRec efficiently explores the preference space along more accurate reasoning trajectories. Extensive experiments on multiple real-world datasets demonstrate that IntuRec consistently outperforms state-of-the-art baselines. We release our code at https://github.com/Ten-Mao/IntuRec.
Mitigating LLM-based p-Hacking by Preregistering for the Next LLM
arXiv:2606.27687v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate, classify, and annotate data whose outputs feed downstream hypothesis tests. However, LLM-based research is easy to p-hack: a researcher can tune the prompts, decoding parameters, or output format until a desired result is reached. We propose a protocol to mitigate p-hacking in LLM-based research: preregistering the experiment and eligible models, and then running it on the first eligible LLM that is released after the preregistration. The researcher finalizes the procedure on current models, preregisters the analysis plan together with a set of eligible future models, and runs the confirmatory analysis on the first eligible model released afterward. Because this model does not exist at commitment time, it cannot be hacked against; furthermore, configurations that hack one model frequently do not transfer to the next. We evaluate the protocol on two tasks whose true values are known. Across 20 models from four providers and 11 LLM-analysis configurations, the protocol would have blocked successful transfer of the p-hack in 73.9% and 72.7% of cases in the two tasks. Additional analyses reveal that mitigation remains substantial under several stress tests. Finally, putting money where our mouth is, we followed our own protocol and preregistered our experiment. The preregistered experiment confirmed the protocol's effectiveness: out of the 7 configurations that hacked the prior model, the hacking failed to carry over in 6 configurations on the first eligible model released afterward.
PerturbCellRL: Verifier-Guided Reinforcement Learning for Single-Cell Perturbation Prediction
arXiv:2606.27752v1 Announce Type: new Abstract: Single-cell perturbation models can reduce costly wet-lab screening by predicting how cells respond transcriptionally to interventions. While recent generative models improve population-level prediction, individual generated cells are not explicitly checked for biological consistency. We introduce PerturbCellRL, a reinforcement learning (RL) framework that post-trains a pretrained single-cell transcriptomic generator using a suite of cell-level verifiers as rewards. These verifiers define four rewards: Pearson top-k similarity, RMSE top-k proximity, DE Spearman, and Pathway activity. The Pathway activity verifier rewards cells whose pathway responses match known perturbation biology. We evaluate PerturbCellRL on multiple genetic and chemical perturbation benchmarks. Across these benchmarks, PerturbCellRL improves over the pretrained flow-matching generator on reward-aligned evaluation metrics and a held-out evaluation metric. Moreover, PerturbCellRL remains competitive with state-of-the-art methods on population-level metrics. Together, these results frame trustworthy single-cell prediction as verifier-guided generative alignment, moving beyond matching expression distributions toward predictions whose single-cell perturbation effects are explicitly checked for biological consistency.
Mixed-Precision For Energy Efficient Computations
arXiv:2606.27949v1 Announce Type: new Abstract: As simulations grow more realistic, the pursuit of higher accuracy results in extended computation times and substantial power consumption. This study explores mixed-precision computing as a promising strategy to address these challenges, leveraging computer arithmetic tools to optimize performance. Using Reactor Simulator and LULESH benchmarks as case studies, we evaluated the potential of mixed-precision strategies to reduce both time-to-solution and energy-to-solution. For Reactor Simulator, we achieved a 30% reduction in both metrics without compromising accuracy. Similarly, for LULESH, results demonstrated up to a 30% improvement in time-to-solution and a 25% reduction in energy-to-solution.
Listwise Explanation of Embedding-Based Rankings via Semantic Chunk Grouping
arXiv:2606.27980v1 Announce Type: new Abstract: Dense embedding rankers score documents through contextual sentence- and passage-level representations. Yet many listwise explanation methods still attribute rankings to isolated words. This feature-unit mismatch leaves word-level features too fragmented for dense semantic ranking. We introduce ChunkGroupSHAP, a listwise Shapley method that clusters semantically related chunks into shared cross-document features. Masking a group perturbs all documents with related evidence, attributing rankings at a granularity closer to dense representations while preserving the listwise setup. Our findings across MS MARCO, FinanceBench, AILACaseDocs, and FinQA with E5 rankers and BM25 show that the best explanation unit is setting-dependent: word features for lexical BM25, corpus-level groups for dense rankers, and query-local grouping for heterogeneous web retrieval. Feature units should thus follow both the ranker's representational granularity and the structure of the retrieved corpus.
Constructions and Characterizations of $s$-Plateaued Partitions
arXiv:2606.27776v1 Announce Type: new Abstract: Bent partitions play a significant role in constructing bent functions and have rich connections with coding theory and combinatorics. In this paper, we introduce $s$-plateaued partitions, which generalize the bent partitions. Let $\Gamma=\{A_{i}, 1 \leq i \leq K\}$ be a partition of $V_{n}^{(p)}$, where $V_{n}^{(p)}$ is an $n$-dimensional vector space over the prime field $\mathbb{F}_{p}$ and $p \mid K$. Then $\Gamma$ is called an $s$-plateaued partition of $V_{n}^{(p)}$ of depth $K$ if each $p$-ary function $f: V_{n}^{(p)} \rightarrow \mathbb{F}_{p}$ for which every $j \in \mathbb{F}_{p}$ has exactly $\frac{K}{p}$ of sets $A_{i}$ in $\Gamma$ in its preimage set, is a $p$-ary $s$-plateaued function. By using an $s$-plateaued partition, a large number of $p$-ary $s$-plateaued functions, vectorial $s$-plateaued functions and generalized $s$-plateaued functions can be constructed. In particular, $0$-plateaued partitions are just bent partitions. In general, $s$-plateaued partitions are much more complicated than bent partitions. We analyze the possible cardinality of $A_{i}$ of an $s$-plateaued partition. We give some explicit constructions of $s$-plateaued partitions for which any generated $p$-ary $s$-plateaued function has no nonzero linear structure. We give a characterization of an $s$-plateaued partition $\Gamma=\{A_{i}, 1 \leq i \leq K\}$, where $p$ is odd, $K \geq 5$ and $-A_{i}=A_{i}, 1 \leq i \leq K$. Based on which, we show that if $p \geq 5$, then the preimage set partition of a $p$-ary $s$-plateaued function $f: V_{n}^{(p)} \rightarrow \mathbb{F}_{p}$ with $f(x)=f(-x)$ is an $s$-plateaued partition if and only if $f$ is of $(p-1)$-form, where $n+s$ is even.When $s=0$, we partially address an open problem on whether a bent partition $\Gamma$ of $V_{n}^{(p)}$ of depth $p^{\frac{n}{2}}$ must be obtained from spreads.
Learning Complementary Action Modeling from Automotive Maintenance Instructions
arXiv:2606.27808v1 Announce Type: new Abstract: A minute lexical variation can reverse the procedural meaning of an instruction even when the rest of the sentence remains unchanged. In automotive maintenance instructions, this pattern often appears when an action phrase turns an instruction into its procedural counterpart. The entities, modifiers, and surrounding context remain largely invariant, while the action phrase determines the procedural relation. We define this task as Complementary Action Modeling (CAM). Given a maintenance instruction, the goal is to identify or generate its procedural counterpart by modifying the action phrase while preserving the remaining sentence context. This task focuses on three aspects: distinguishing complementarity from surface similarity, controlling generation at the action-phrase level, and evaluating relational correctness using retrieval, overlap-based, and human evaluation. Using a German automotive maintenance dataset, we examine these questions through candidate matching and controlled Seq2Seq generation. The results show that complementary maintenance instructions are best modeled as procedural associations grounded in subtle lexical cues. They should therefore not be treated as ordinary cases of sentence similarity or synonym-based paraphrasing.
Booster Lab: A Data-Centric Pipeline for Learning Deployable Humanoid Locomotion Policies
arXiv:2606.27813v1 Announce Type: new Abstract: Humanoid robot motion learning requires not only task-oriented control policies but also physically feasible and natural behaviors that can be transferred to real robots. However, robot-feasible motion data are often scarce: raw human demonstrations may be incompatible with the robot morphology, open-source clips vary in quality, and simulation-collected robot trajectories still require feasibility checking. To address these challenges, we propose a data-centric training and deployment pipeline that integrates motion data curation, real-to-sim model adaptation, AMP-based reinforcement learning, and sim-to-real deployment. We validate the framework on the Booster T1 robot and further provide preliminary cross-platform validation on Booster K1.
Scalable and Differentiable Point-Cloud Registration Using Maximum Mean Discrepancy
arXiv:2606.27818v1 Announce Type: new Abstract: We present MMD-Reg, a novel correspondence-free approach to point-cloud registration that is differentiable and has linear computational complexity in the number of points. We model registration as a nonlinear least-squares problem based on the Maximum Mean Discrepancy, approximated using random Fourier features. The resulting objective can be solved efficiently with standard methods such as Levenberg-Marquardt, and the solution is differentiable via the implicit function theorem. This allows MMD-Reg to be used as a differentiable optimization layer within end-to-end trainable models, supporting registration under challenging conditions such as poor initial alignment and partial overlap. We demonstrate this Neural MMD-Reg formulation by integrating the layer with a set transformer, training the resulting model in supervised and unsupervised settings, and comparing its performance against recent learning-based methods. We also evaluate standalone MMD-Reg, comparing its accuracy and scalability against widely used non-learning-based registration methods.
Exploring and Exploiting Synchrony Limitations of Time-Triggered Network-Agnostic Guardians
arXiv:2606.27819v1 Announce Type: new Abstract: Time-triggered communication protocols rely on trusted components known as guardians to enforce adherence to predetermined network schedules. Network-agnostic guardians offer an efficient and scalable distributed solution with reduced implementation cost and complexity compared to network-aware alternatives. However, this efficiency is based on the guardian's dependence on the controlled node for clock synchronization, which introduces a vulnerability: a malicious node can exploit this dependency to launch timing attacks against its guardian and eventually interfere with messages from other nodes on the network. In this paper, we establish a theoretical lower bound on the attainable clock synchronization precision between a node and its network-agnostic guardian. Building on this result, we introduce a timing attack that leverages the unavoidably imperfect clock synchrony to cause controlled and undetected de-synchronization of the guardian. The attack enables a malicious node to cause collisions with targeted critical network messages. We evaluate the effectiveness of the attack using a FlexRay field bus network model implemented in the OMNeT++ simulation framework. Our results show that the attack is able to remain undetected with 100% success and disrupts the transmission of the critical messages of the target node by causing collisions with them with 100% success.
Differential Privacy over Hamming Codes
arXiv:2606.27849v1 Announce Type: new Abstract: We consider the transmission of the outputs of counting queries over a binary symmetric channel (BSC), where Hamming codes are employed as the channel encoder. Since the channel is inherently noisy, this transmission already provides a degree of privacy protection ``for free'', albeit at the cost of reduced utility in the form of decoding errors. A natural question is whether this privacy can be further improved (i) without any additional real-time obfuscation of the data, such as injecting artificial noise prior to transmission, and (ii) without increasing the end-to-end error probability. In this work, we answer this question in the affirmative by deriving an optimal codeword arrangement that strictly improves differential privacy guarantees while incurring no real-time computational overhead and no degradation in utility.
ScaLe-INR: Scale and Learn Implicit Neural Representations
arXiv:2606.27862v1 Announce Type: new Abstract: Implicit Neural Representations (INRs) parameterized by multilayer perceptrons excel at modeling continuous signals. However, a key challenge persists as INRs fundamentally suffer from spectral bias and information cross-talk. When a single network attempts to capture multi-scale phenomena, high-frequency weight updates destructively interfere with the underlying low-frequency structural approximation. We introduce Scale and Learn INR (ScaLe-INR), a novel multi-branch architecture that resolves these limitations by explicitly matching the signal's frequency spectrum with the optimal operating region of the INR. Drawing upon the Fourier inverse scaling theorem we demonstrate that applying directional coordinate scaling expands a network's representational bandwidth along specific spatial axes. To mathematically enforce functional disentanglement and minimize task-specific information leakage between branches, we propose a Directional Edge Guidance Loss, a spatially-conditioned sparsity prior derived from ground-truth gradients. By constraining the high-frequency branches to act as strict, localized edge-filters, ScaLe-INR eliminates spectral cross-talk, accelerates convergence, and achieves high-fidelity signal reconstruction on complex multi-scale topologies. We evaluate ScaLe-INR across diverse reconstruction and inverse tasks, demonstrating substantial performance gains over existing state-of-the-art (SOTA) methods. The proposed architecture improves upon the nearest baselines by +5.16 dB in image reconstruction and +0.65 dB in image denoising. Furthermore, it achieve an impressive figure of 50.02 dB on audio reconstruction and 0.999 IOU(Intersection Over Union) on 3D reconstruction which beats the all SOTA models.
GNBAN: Graph Neural Basis Attention Networks for Long-Horizon Forecasting over Large Entity Sets
arXiv:2606.27863v1 Announce Type: new Abstract: Demand forecasting at the bottom of a retail hierarchy requires predicting tens of thousands of correlated long-horizon series across products, stores, and regions. Modern systems must scale across massive catalogs, capture shared demand dynamics, and remain interpretable enough to be trusted. Classical statistical methods need a separate model per series and are hard to manage at scale; deep autoregressive models struggle as the joint state grows to tens of thousands of dimensions; and recent graph-based forecasters, while capturing cross-entity dependencies, often produce opaque long-horizon forecasts. We propose GNBAN (Graph Neural Basis Attention Network), an end-to-end architecture combining heterogeneous graph representation learning with an interpretable basis-decomposition head. Retail data are represented directly as a heterogeneous graph derived from the relational schema, so a single model serves the entire catalog. Rather than predicting the horizon directly, GNBAN decomposes each forecast into trend, seasonal, and generic components. Its key innovation is a per-basis attention mechanism: each basis function keeps its own learnable query and retrieves information independently from the entity's historical neighborhood, letting different bases specialize to distinct temporal patterns while preserving interpretability. On two large-scale benchmarks, M5 Walmart and Favorita Grocery Sales, evaluated under matched protocols, GNBAN improves volume-weighted WRMSSE by roughly 4-5% over a matched graph baseline. Qualitative analysis shows the learned decomposition exposes trend, seasonal, and residual demand drivers without post-hoc explanation methods. These results demonstrate that scalable relational forecasting and interpretable forecast decomposition can be achieved together in a unified graph-based framework.
FMO-xTB: Fragment molecular orbital method with GFN1-xTB for large-scale quantum-mechanical simulations
arXiv:2606.28022v1 Announce Type: new Abstract: We present the fragment molecular orbital method (FMO) combined with the GFN1-xTB extended tight-binding approach (FMO-xTB) for efficient quantum-mechanical calculations of large molecular systems. Both the two-body (FMO2) and three-body (FMO3) expansions are formulated, and fully analytic energy gradients including the response contribution from the self-consistent embedding potential are derived and implemented. The FMO-xTB method inherits the broad element coverage of GFN1-xTB, which employs element-specific rather than atom-pair-specific parameters and is parameterized for all spd-block elements up to radon(Z = 86), representing a significant practical advantage over FMO- DFTB approaches. The accuracy of FMO-xTB is systematically benchmarked against non-fragmented xTB calculations for water clusters, anthracene aggregates, and pentacene supercells. FMO3-xTB reproduces the reference energies with deviations on the order of 10^-4 Hartree for organic semiconductor systems. The covalent bond fragmentation capability using the hybrid orbital projection (HOP) boundary treatment is also implemented with fully analytic gradients and validated for polyalanine alpha-helices and B-DNA double helices, yielding FMO3-xTB energy deviations on the order of 10^-6 Hartree for polyalanine and in the millihartree range for B-DNA. Near-linear scaling is achieved with effective scaling exponents between b= 1.06 and b= 1.28, compared to cubic scaling for non-fragmented xTB. Parallelization over multiple CPU cores yields significant speed ups, and a complete energy and gradient evaluation of a pentacene supercell containing 23760 atoms is feasible within minutes on a single computing node, enabling routine molecular dynamics simulations of systems with tens of thousands of atoms. The method is implemented in the DIALECT software package.
SEADA: An efficient methodology for optimizing mixed-precision DNNs on multi-precision spatial architectures
arXiv:2606.27884v1 Announce Type: new Abstract: Mixed-precision computation has been introduced in deep neural networks (DNNs) as an effective approach to reduce latency, energy consumption, and memory footprint. However, efficiently mapping mixed-precision networks onto multi-precision spatial architectures poses several challenges. These include determining the appropriate precision for each layer, balancing layer-wise accuracy sensitivity to quantization against architectural heterogeneity and system-level constraints, and accurately estimating the system-level cost of heterogeneous precision assignments. This work presents SEADA, an efficient methodology designed to address these challenges. SEADA comprises: (i) a configurable system-level analytical cost model of a multi-precision spatial accelerator architecture; (ii) a fast mapping tool that identifies near-optimal mappings of DNN workloads onto the target integer accelerator; (iii) analytical models for floating-point layers to estimate the overall benefits of mixed-precision execution; and (iv) a per-layer precision selection methodology based on bit-level entropy, enabling efficient assignment across multiple numerical precisions. SEADA's efficiency provides designers with a robust framework for the design-space exploration of multi-precision architectures.