Forskningsradar

Science Journals

Peer-reviewade publikationer — 56237 artiklar

Towards Objective Dysgraphia Detection: A Multi-Branch Deep Learning Approach for Online Handwriting Analysis
arXiv:2607.09826v1 Announce Type: new Abstract: Dysgraphia is a specific learning disability that is prevalent among school-age children. It affects handwriting coherence, quality, fluency, and legibility, often hindering academic achievement and early learning development. This motor coordination disorder is typically diagnosed through subjective assessments based on clinician observation, which can be timeconsuming and prone to variability. In this paper, we introduce a deep learning-based framework for objective dysgraphia detection using online handwriting data captured via digitizing tablets. The proposed framework relies on two complementary branches: the first pipeline extracts both handcrafted and embedding-based kinematic features directly from raw temporal signals, while the second leverages image-based representations of the temporal signals generated using continuous wavelet transforms (CWT) and Gramian Angular Fields (GAF). The resulting features are then fused to leverage the complementary strengths of both representations. The four representations were evaluated separately and jointly using the publicly available DiaGraMo dataset, showing that the fusion of GAF, MOMENT, and hand-crafted kinematic features outperforms each individual representation, as well as other fusion schemes. These findings highlight the potential of the complementarity of image and signal based representations for more objective dysgraphia detection.
Water Reflection Detection Using Symmetric Attention
arXiv:2607.10749v1 Announce Type: new Abstract: Reflections of water pose a significant challenge for computer vision systems, as standard deep learning models frequently confuse objects with their mirror images, producing spurious false positives and negatives in tasks such as object detection and semantic segmentation. As a result, detecting reflection axes in natural-water scenes is pivotal for reliable object detection and scene understanding. To mitigate this issue, we leverage the intrinsic imperfect reflective symmetry of water and introduce a Symmetry-Aware Water Reflection Detection Network, namely, SAWRD-Net, that couples dihedral group-equivariant convolutions with a matrix-decomposition decoder in an end-to-end framework. First, dihedral group convolutional layers extract geometry-consistent feature maps that explicitly encode both rotational and mirror symmetries. A Multi-scale Reflection Equivariant block then aggregates features across scales and employs a symmetric-attention mechanism to highlight reflection-relevant regions. The proposed matrix-decomposition decoder factorizes high-dimensional features into compact low-rank parameter and confidence spaces, after which the network directly regresses keypoints on the reflection axis. Then a robust principal component analysis fits the final axis. Evaluated on the largest available water reflection scene data set, SAWRD-Net achieves a true-positive rate of 0.890 against human annotations, outperforming all existing water reflection detectors.
APHABAMAS: An analytical phantom-based scheme for assessing the accuracy of high-resolution 3D MRI motion-artifact simulations
arXiv:2607.09945v1 Announce Type: new Abstract: Purpose: Motion compromises the utility of high-resolution 3D MRI, an established tool in quantitative neuroimaging research. Deep learning-based methods have shown promise for mitigating motion-induced artifacts, but their development typically requires simulated motion-corrupted data. Several open-source tools exist for this task, each implementing different algorithms. However, no scheme currently exists for evaluating the accuracy of these simulations, making it difficult for users to choose the most suitable tool. Developing such a scheme is the aim of this study. Methods: The essential ingredient of the desired scheme is a ground-truth reference simulation that does not suffer from sampling-induced error. To meet this requirement, the proposed scheme, APHABAMAS, leverages a digital phantom whose representations in both the image and Fourier domains can be expressed analytically under arbitrary rigid-body transformations. Results: APHABAMAS is used to quantify the sampling-induced errors of three existing simulation algorithms, establishing their first definitive accuracy-based ranking. Conclusions: APHABAMAS provides a rigorous tool for assessing the accuracy of high-resolution 3D MRI motion-artifact simulations. It allows the accuracy-based ranking of existing simulation algorithms to be established, thereby enabling informed selection of the most suitable algorithm for synthesizing motion-corrupted data.
UniPose9D: Universal Category-Agnostic Object Pose Estimation
arXiv:2607.09985v1 Announce Type: new Abstract: Object pose estimation is a fundamental problem in 3D vision. Although recent state-of-the-art approaches achieve strong performance, they often overfit to existing benchmarks and exhibit limited generalization to novel categories and unseen scenes. We propose UniPose9D, a category-agnostic foundation model for 9D object pose estimation: given an instance mask/ROI and either an RGB-D observation or an RGB image with predicted depth, the model estimates rotation, translation, and metric size without category labels, CAD models, mean-shape priors, or reference views. Specifically, UniPose9D samples point pairs from the observed object geometry and uses DINOv2 and PointNet features to predict NOCS coordinates for each pair. To improve accuracy, we introduce a point-pair-based RANSAC N-hop Kabsch--Umeyama algorithm with an adaptive threshold. We further employ flow matching to address symmetric ambiguities and construct a large-scale training set by curating and aligning pose annotations from existing public datasets. Experiments across six datasets show that a single unified model can match or surpass specialist methods while generalizing to unseen objects and in-the-wild scenarios. Our code and model are available on https://github.com/qq456cvb/UniPose9D.
An asymptotic-preserving reduced-order method for parametrised rarefied gas flow by proper generalised decomposition
arXiv:2607.10085v1 Announce Type: new Abstract: Modelling rarefied gas flow using the Boltzmann equation is vital in many areas. Due to the high dimensionality and coexistence of multiple characteristic scales, conventional solution strategies to this equation incur prohibitively high computational costs and are inadequate for rapid response in engineering design simulations. Based on proper generalised decomposition (PGD), we propose an \textit{a priori}, asymptotic-preserving reduced-order method to solve the high-dimensional, parametrised Shakhov kinetic model equation. The method reduces the original problem to a few low-dimensional problems by formulating separated representations for the low-rank solution, thereby mitigating the curse of dimensionality. To capture the hydrodynamic asymptotics, we incorporated solutions of some synthetic equations into the PGD algorithm. This treatment allows the PGD solver to automatically reduce to a macroscopic solver for the Navier-Stokes equations, whose solution naturally exhibits low-rank structure. By treating the rarefaction parameter as an additional coordinate, a parametrised solution can be computed once and for all over the entire range of rarefaction, enabling fast multiple queries to any points in the parameter space. Numerical examples are presented to demonstrate the capability of the method to simulate rarefied gas flow with certain accuracy and a significant reduction in computational costs.
Next-Dense-Stride Prediction for Multimodal Autoregressive Visual Modeling
arXiv:2607.09892v1 Announce Type: cross Abstract: We introduce DenseAR, a new generative paradigm that reformulates autoregressive image generation as coarse-to-fine next-dense-stride prediction using a compact single-scale tokenizer. Our key insight is that traversing a single-scale latent grid with progressively denser strides naturally captures the transition from global structure to fine detail. This addresses two limitations of existing autoregressive models at once: the slow inference of raster-order autoregression, which DenseAR avoids by predicting multiple tokens in parallel, and the heavy cost of multi-scale approaches, which need long, multi-resolution token sequences to achieve coarse-to-fine prediction. Building on our efficient framework and the flexibility of autoregressive modeling, we further extend DenseAR to a unified model that handles multiple modalities and imaging tasks within a single backbone. We validate DenseAR on both medical and natural images. On multi-contrast brain MRI, a single DenseAR model unifies cross-modal translation, modality-conditioned generation, and tumor segmentation, while remaining competitive with task-specific methods. On ImageNet, DenseAR improves class-conditional generation quality (FID and IS) over both a single-grid baseline without stride ordering and a multi-scale tokenizer-based baseline.
Understanding Chemical Short-Range Order in CoNiV via Mode Analysis
arXiv:2607.10775v1 Announce Type: new Abstract: We analyze chemical short-range order in equiatomic fcc NiCoV using molecular-dynamics snapshots generated with a machine-learned interatomic potential. Radial distribution functions identify stable coordination shells, while shell-resolved Warren-Cowley parameters and bond probabilities reveal continued chemical ordering after the radial structure has largely converged. The dominant signal is V-V avoidance in the first shell and V-V enrichment in the second shell, consistent with an L1$_2$-like local ordering tendency, while the third-shell response remains weak. Lagged Jensen-Shannon diagnostics show that bond statistics relax more slowly than the RDF. Principal component analysis of per-replica-centered bond probabilities resolves three collective modes: a V-sublattice ordering amplitude, a Ni-Co redistribution mode, and a Co-V exchange-like mode. These results show that scalar RDF convergence can miss slow chemical relaxation, and that shell-resolved bond statistics provide a compact route for tracking SRO development in multicomponent alloys.
A Generalized Deep Non-negative Matrix Factorization Approach for SAR Automatic Target Recognition
arXiv:2607.09779v1 Announce Type: new Abstract: The deep nonnegative matrix factorization (DNMF) technique is proposed to address the low interpretability of deep learning-based methods in extracting multilayer features from synthetic aperture radar (SAR) target samples. However, existing DNMF methods employ a layer-by-layer decomposition strategy, which is prone to causing error accumulation and local optimum, thereby hindering a consistent improvement in recognition accuracy as the number of layer increases. In this paper, a robust multilayer feature extraction method, termed generalized deep non-negative matrix factorization (G-DNMF), is proposed to address the above challenges in SAR automatic target recognition (ATR). The G-DNMF aims global optimality and derives the update rules for each parameter using lagrangian multiplier method. The new update formula indicates that both the DNMF method based on the encoding matrix and the mixing matrix are special cases of the proposed method, theoretically demonstrating the universality of proposed method. In general, the proposed method discards the layer-by-layer decomposition strategy, thereby effectively mitigating the risk of local optima and eliminating error accumulation, leading to a significant improvement in DNMF's multi-layer feature extraction capability. The experimental results, by presenting the feature images extracted from each layer by G-DNMF and the reconstructed original images, verified the proposed method's pure additive understanding of multi-layer features and demonstrated its interpretability. The experimental results based on MSTAR and OpenSARship datasets show that G-DNMF outperforms existing DNMF algorithms and their derivatives in terms of stability and recognition performance.
Fully Dynamic Edge Connectivity in $\tilde{O}(n^{12/13})$ Time
arXiv:2607.10689v1 Announce Type: new Abstract: In the (fully) dynamic edge connectivity problem, the goal is to maintain the edge connectivity $\lambda_G$ of an $n$-vertex graph $G$ that undergoes edge insertions and deletions. Our main result is a randomized algorithm for maintaining edge connectivity in dynamic simple graphs using worst-case update and query time $\tilde{O}(n^{12/13})$, for all values of $\lambda_G$. This is the first algorithm that has $o(n)$ update and query time, as all existing algorithms achieve this only when $\lambda_G$ is below $n^{1/11}$ or above $n^{1/2}$ (up to polylogarithmic factors). We then use the tools developed for this purpose to design two additional algorithms. The first one is a deterministic algorithm for the exact same task, that uses $n^{1+o(1)}$ worst-case update and query time or $\tilde{O}(n)$ amortized update and query time; this gives a polynomial improvement over existing deterministic algorithms. The second one is a deterministic algorithm for the same task but in dynamic unweighted multigraphs, that uses $\tilde{O}(n^{3/2})$ worst-case update and query time.
Benchmarking Dynamic Affective Reasoning: A Viewer-Centric Video Emotion Dataset
arXiv:2607.10238v1 Announce Type: new Abstract: Video emotion analysis is typically framed as a static classification problem, treating each clip as an independent labeled unit. However, such a formulation overlooks a key psychological fact: emotions change as a result of cumulative reactions to consecutive causal events. To bridge this gap, we introduce Dynamic Affective Reasoning, the first large-scale benchmark for viewer-centric affect transitions and causal reasoning over consecutive video events. DAR contains 15,087 videos and 36,908 event-aligned affective segments annotated with 27 emotion categories. Unlike existing video-based emotion datasets, DAR presents a new viewer-centric perspective on fine-grained emotional expressions and transitions, and provides dense, temporally grounded, and causally explicit reasoning chains. Based on DAR, we formally define three challenging tasks: affective segmentation, fine-grained emotion classification, and affective reasoning. Complementing this benchmark, we propose DAR-R1, a two-stage framework that combines supervised fine-tuning with Group Relative Policy Optimization. Experiments across 10+ MLLMs show that DAR-R1 sets a new state-of-the-art for dynamic affective reasoning, in terms of both emotional localization and affective reasoning. Project page: https://github.com/Zhang-Zhiyan/DAR.
What Does Your Short-Answer VQA Score Actually Measure? Evaluator-Dependent Instability in Multimodal Short-Answer Benchmarks
arXiv:2607.10240v1 Announce Type: new Abstract: Short-answer VQA benchmarks conflate two distinct quantities: whether a model's answer is semantically correct, and whether that answer matches the surface form expected by the automatic evaluator. We study this conflation across six vision--language models and six benchmarks, using a human-validated semantic judge (97.6% precision) to audit over 37k official errors. A second text-only judge reproduces the same benchmark-level false-negative pattern, showing that the effect is not an artifact of a single audit model. On text-rich benchmarks, up to half of these errors are semantically acceptable answers penalized purely for surface-form mismatch. This instability is structured by answer type: extractive and multi-span answers are far more evaluator-sensitive than scalar answers. Benign prompt and context rewrites further destabilize official outcomes, flipping item-level correctness at substantial rates without changing the underlying task. A deterministic CPU-only contract repair confirms that the undercount is partially recoverable. These findings imply that official short-answer VQA scores should be accompanied by semantic audits and answer-type diagnostics to remain interpretable.
Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety
arXiv:2607.10112v1 Announce Type: new Abstract: Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settings. We introduce \textsc{Minionese}, a multilingual jailbreak benchmark spanning 18 languages, 4 resource tiers, and 4 perturbation types (standard translation, code-switching, transliteration, and translationese), paired with a geometric mechanistic analysis of refusal failure across language tiers. We show that each attack type produces a distinct vulnerability profile: transliteration vulnerability is mediated by script identity, code-switching maintains effectiveness through the lowest-resource tier, and a sharp safety regime transition between Tiers 2 and 3 is consistent across all models. Mechanistically, low-resource jailbreaks succeed by routing harmful content through a geometrically misaligned subspace that projects insufficiently onto the refusal directions, leaving the refusal mechanism intact but untriggered. These findings show that English-only safety evaluations are insufficient; they require accounting for script family, perturbation type, and per-language alignment coverage. The benchmark and analysis code is at https://github.com/Brentkong/Minionese-Comprehensive-Benchmark-and-Mechanistic-Study-of-Multilingual-LLM-Safety.git.
Sharper Analysis of Single-Loop Methods for Bilevel Optimization
arXiv:2607.10263v1 Announce Type: new Abstract: Bilevel optimization underpins many machine learning applications, including hyperparameter optimization, meta-learning, neural architecture search, and reinforcement learning. While hypergradient-based methods have advanced significantly, a gap persists between theoretical guarantees and practical single-loop implementations required for efficiency. We bridge this gap by establishing sharper convergence results for single-loop approximate implicit differentiation (AID) and iterative differentiation (ITD) methods, leveraging our proposed analytical framework, decoupled norm analysis (DNA). For AID, we improve the convergence rate from $\mathcal{O}(\kappa^6/K)$ to $\mathcal{O}(\kappa^5/K)$, where $\kappa$ is the condition number of the inner-level problem. For ITD, we prove that the asymptotic error is $\mathcal{O}(\kappa^2)$, exactly matching the known lower bound and improving upon the previous $\mathcal{O}(\kappa^3)$ guarantee. Numerical experiments on synthetic and real tasks corroborate our theoretical findings.
An End-to-End Hybrid Quantum--Classical Sampling Workflow for Discrete Markov Random Fields: A Reproducible Case Study
arXiv:2607.09893v1 Announce Type: cross Abstract: Sampling from discrete Markov random fields (MRFs) is a hard problem. We study amplitude-encoded i.i.d. sampling for small MRFs where $2^n$ target probabilities are precomputed classically. This removes quantum exponential speedup but allows a clean comparison against classical MCMC based on independent circuit samples ($\tau \approx 1$). Across 60 instances spanning five graph families (1k-step burn-in, 3k retained samples), the mean ESS ratios of Quantum to Single-Site Gibbs, Block Gibbs, Tuned-Block, and Parallel Tempering are $16.35$, $7.29$, $1.82$, and $1.79$, showing modern classical samplers substantially close this gap. Amortizing $O(2^n)$ preprocessing into wall-clock time, exact inverse-CDF sampling yields $17.7\text{M}$ ESS/s versus $488\text{K}$ ESS/s for the quantum sampler ($36\times$ mean rate, $153\times$ per-instance), confirming no wall-clock advantage. We characterize MCMC autocorrelation costs and benchmark amplitude-encoded state preparation at $n \in \{8,10,12\}$. An MPS scaling study ($n \le 40$) shows bond dimension $\chi=32$ achieves $F=0.721\pm0.059$ at $n=40$. Finally, a matched-budget VQC vs. MPS comparison at $n \in \{8,10,12\}$ shows VQC fidelities fall far below MPS: $(F_{\mathrm{VQC}}, F_{\mathrm{MPS}}) = (0.31, 0.99), (0.21, 0.96), (0.17, 0.88)$ at compressions $10.7\times$, $34.1\times$, and $113.8\times$.
Policy-Driven CT-Agent: Modeling Phase-Aware Diagnostic Control for Clinically Consistent CT Reasoning
arXiv:2607.10748v1 Announce Type: new Abstract: Computed Tomography (CT) diagnosis often relies on dynamic selection of imaging phases, such as non-contrast, arterial, or venous phases, based on preliminary findings, clinical suspicion, and diagnostic guidelines. This phase-wise decision process is critical for reducing unnecessary radiation exposure while supporting timely staging and treatment planning. However, phase-selection protocols can vary across hospitals, regions, and guidelines, while most existing CT-based AI methods assume that all phases are available and focus on static tasks under a fixed imaging phase, failing to model whether additional phases are required. This limitation stems from heterogeneous multi-phase representations, the need for knowledge-guided phase control beyond visual cues, and the lack of supervision for phase-sufficiency decisions in existing datasets. To address these challenges, we propose Policy-Driven CT-Agent (PD-CTAgent) for clinically consistent CT phase selection and diagnostic reasoning. PD-CTAgent introduces a Clinical Structure Abstraction Module (CSAM) to harmonize heterogeneous CT phases into a unified, phase-aware evidence representation. Based on this representation, a Knowledge-Guided Diagnostic Control Model (KDCM) evaluates phase sufficiency and iteratively requests additional phases when necessary. The policy-driven agent design further allows PD-CTAgent to flexibly follow different institutional, regional, or guideline-specific diagnostic protocols. Together, PD-CTAgent bridges static CT analysis and real-world clinical workflows. Experiments on two public datasets, LIDC and MCT-LTDiag, and one private dataset demonstrate its effectiveness and clinical consistency. Code will be made public upon acceptance.
The Nuts and Bolts of Natural Language to SQL Translation: A Systematic Analysis of Model Pipeline Optimisation Approaches and their Interactions
arXiv:2607.10911v1 Announce Type: new Abstract: In the age of large language models, Natural Language to SQL (NL2SQL) translation remains an open problem with many useful applications. We explore interactions between several NL2SQL pipeline extensions to inspire development of more lightweight models. Specifically, we integrate the NatSQL intermediate representation, include a preprocessing step and a fine-tuning step based on synthetic data, and develop a novel reranker model to improve SQL selection in the final beam. We perform an ablation study supplemented by a Shapley analysis of these different components integrated with two backbone architectures, SmBoP and RASAT. We find that simply combining all of them does not lead to best results, but that their impact depends on their interactions with the baseline system, as well as each other.
Coupled Tensor-Matrix Recovery via Proximal Alternating Linearized Minimization, with an Application to Workforce Skill and Small-Business Health Estimation
arXiv:2607.10163v1 Announce Type: cross Abstract: We study recovery of a low-rank tensor $\mathcal{T}$ and a low-rank matrix $M$ from sparse, noisy observations. $\mathcal{T}$ and $M$ share one mode. We relax tensor rank using the nuclear norm of the mode-1 unfolding. This unfolding carries the coupling. It also has an exact proximal operator. We couple $\mathcal{T}$ and $M$ through a learned linear operator $G$. We prove a minimizer exists for the ridge-stabilized penalized objective. We prove that a proximal alternating linearized minimization (PALM) scheme converges to a critical point, for the algorithm as implemented, by verifying the hypotheses of a known nonconvex block-coordinate convergence theorem against our objective and identifying which conditions come from this problem's structure. For the matrix-only sub-problem, we state a proven sampling bound from matrix completion theory. For the coupled problem, we prove a sample-complexity result for a sequential sub-case: a separately-known coupling operator recovers $M$ from $\mathcal{T}$'s recovery accuracy alone, with no observations of $M$ needed. For the fully joint, alternately-estimated case, we state a conjecture and test it empirically, including a low-density regime where coupling does not help. We report multi-seed synthetic experiments with mean and standard deviation across sampling densities, an asymmetric-density experiment, and convergence curves, and we explain why recovery error stays high at low density. We apply the framework to workforce-skill and small-business-health estimation. Every application-specific choice is a proposed design, not a validated result; we have not run the framework on deployed data.
Homological invariants of edge ideals of the multiple extended complete split-like graphs
arXiv:2607.10300v1 Announce Type: cross Abstract: We study the graphs $MECS_{b,n}^a \cong \overline{K}_a \join \big(n(K_b+K_2)\big)$, obtained by attaching an independent set of size $a$ to $n$ disjoint copies of the block $K_b+K_2$. For $n=1$, we get $MECS_{b,1}^a$, and recover the results of one-block case studied in [Anand, Gupta, Rather and Singh, Homological invariants of some complete split-like graphs, Beitr. Algebra Geom. (2025)]. Using Hochster's formula, tensor products of minimal free resolutions over disjoint variable sets, and the Betti-number formula for graph joins, we derive explicit descriptions of the independence complex, independence polynomial and its analytic properties, Hilbert series, linear and quadratic Betti strands, regularity, projective dimension, and several structural invariants of $MECS_{b,n}^a$. We further classify the well-covered and unmixed members, compute induced matching numbers, show that the family is never Cohen--Macaulay, and record algorithmic procedures for evaluating the Betti data.
Demographic Prompting at Scale: When More Attributes Hurt LLM--Human Agreement
arXiv:2607.10590v1 Announce Type: new Abstract: We investigate how annotator demographic attributes, supplied as prompt cues, shape the alignment between large language model (LLM) predictions and human annotations across five tasks. Using five open-source LLMs, we systematically vary the number and composition of demographic components in the prompt, spanning every combination from single-attribute through full-attribute configurations. Our experiments reveal three principal findings. First, alignment consistently peaks with one to three high-signal attributes and degrades under the full attribute set, establishing a clear over-specification threshold. Second, the overall magnitude of demographic influence on human annotations does not predict which attributes improve LLM alignment; instead, both the learnability and the directional coherence of each attribute's annotation signal need to be considered jointly. Third, neuron probing reveals that specialized activation correlates with alignment gains only under coherent annotation signals, and that activation volume alone does not imply steerability. Together, these results demonstrate that demographic prompting is not a monolithic intervention: its utility is highly context-dependent, shaped by attribute signal quality, task characteristics, and model architecture.
Small but Tubby: A Magnetic Loop Antenna Made from 100 mm Copper Tubing
arXiv:2607.10828v1 Announce Type: new Abstract: This paper presents the electrical model, key equations, and practical construction of a small transmitting magnetic loop antenna built from unusually large 100 mm diameter copper tubing. The large conductor surface area and wide-area transitions to the vacuum capacitors were designed to minimize resistive losses. The frequency range from 1.8 MHz to 31 MHz is unusually wide. Frequency, impedance matching, and azimuth are all adjusted automatically by servo motors. A novel feature is the routing of the control wiring inside the loop conductor, allowing the motor to be mounted without electrical insulation from the loop conductor. Indoor losses originate predominantly from near-field coupling to the environment rather than from the antenna itself. Temperature-rise measurements confirm that the bulk of the dissipated power is absorbed by the environment, not by the antenna components. The conducted H-field measurements demonstrate good agreement between the measured H-field and the theoretical free-space H-field calculated from the antenna geometry and an estimated loop current. The loop current was estimated from the measured antenna bandwidth and the applied transmit power. The antenna was developed for indoor operation where outdoor installation is not possible.
SPORT: Structure-Aware Prototype Disentanglement for Incomplete Multi-View Clustering
arXiv:2607.10413v1 Announce Type: new Abstract: Prototype-based Incomplete Multi-view Clustering has recently attracted increasing attention by exploiting prototypes as semantic anchors for missing-view imputation. However, existing approaches are still limited in three aspects. First, they typically focus on enforcing cross-view prototype consistency, while ignoring view-specific information embedded in prototypes, thus limiting multi-view expressiveness. Second, most methods rely on instance-level contrastive learning that only aligns paired samples across views, failing to preserve cluster-level relational structures. Third, missing-view imputation is usually performed using global prototypes alone, without considering local geometric neighborhood structures, leading to inaccurate recovery of missing representations. To address these limitations, we propose a novel framework termed Structure-aware PrOtotype disentanglement foR incomplete multi-view clusTering (SPORT), which explicitly disentangles shared and view-specific components of prototypes while preserving cluster-level relational structures. Specifically, we decouple prototypes into orthogonal shared and view-specific components, aligning only shared components to capture consensus semantics while de-correlating view-specific components to preserve complementary information. Meanwhile, a structure-aware contrastive learning mechanism is incorporated to explicitly model cluster-level relationships during cross-view representation learning. Furthermore, a hybrid imputation strategy integrates global prototype matching with local neighborhood matching, enabling joint exploitation of semantic prototypes and manifold structures for missing-view recovery. Extensive experiments on six benchmark datasets show that SPORT achieves superior performance over state-of-the-art methods under various missing rates.
Sequential compliance decisions of firms on cross-border data flows: An institutionally anchored decision support system
arXiv:2607.10620v1 Announce Type: new Abstract: The economic value of data arises from its flow across organizations and national borders. Yet increasingly stringent data governance regimes are turning cross-border transfer into an institutionally constrained sequential decision, in which firms repeatedly weigh compliance costs against the value of data flows. From the perspective of a data-exporting firm, this paper develops an institutionally anchored decision support system. It converts regulatory rules into a computable minimal compliance mapping and models the firm's weekly decisions as a finite-horizon Markov decision process (MDP), with compliance represented as a hard constraint rather than a penalty term. The resulting problem is solved using masked deep reinforcement learning, while counterfactual path advantages provide interpretable signals to support the firm's cross-border data flow decisions. Experiments show that the policies learned within the system outperform the baselines considered and deliver interpretable, auditable decision support. Local processing concentrates in states where the business value of small lawful transfers does not cover their compliance costs, and the localization boundary shifts systematically as the regime tightens. Credential acquisition is front-loaded within the compliance year, and shallow decision trees reproduce the policy's decisions with high fidelity. Treating the persistent-friction weight as a continuous representation of regulatory strictness further reveals an absorb-then-adjust pattern, in which expected rewards decline before observable behavior changes, implying that assessments based only on behavioral indicators may understate the burden already borne by firms. Moreover, the system is not tied to any specific regulation and can be transferred to other jurisdictions and rule-based compliance problems.
Photonic-Crystal Microresonator Frequency Combs in the O-band
arXiv:2607.10025v1 Announce Type: new Abstract: Photonic-crystal microresonators (PhCRs) are a powerful platform for generating Kerr frequency combs. Because Kerr-soliton dynamics in PhCRs are largely decoupled from the operating wavelength, the comb output can be engineered through customization of the device layer. Here, we demonstrate a tantalum pentoxide (tantala) PhCR platform that supports 1310 nm and 1550 nm band operation, and we explore high-efficiency O-band soliton microcombs with all-semiconductor laser pumps. We engineer the PhCRs with silicon dioxide cladding and normal dispersion with intrinsic quality factors exceeding $7\times10^{6}$. By pumping bandgap modes, we obtain robust and efficient soliton comb formation at a 200 GHz mode spacing. Our PhCRs enable systematic tuning from narrowband to broadband comb states within a single device geometry. The combs exhibit low relative intensity noise approaching the shot-noise limit, indicating stable phase-matching in the PhCR. Using a second resonator coupler, we amplify the comb output off-chip, demonstrating a pathway to high-power O-band sources. These results establish PhCR engineering in the tantala platform as a scalable approach to wavelength-agile, low-noise microcombs for applications in communications, sensing, and signaling.
Direct Numerical Simulation of Fully Developed Turbulent Channel Flow Based on the Corrected Navier-Stokes Equations
arXiv:2607.10224v1 Announce Type: new Abstract: Direct numerical simulations (DNS) of fully developed turbulent channel flows at Re_tau = 550 were performed to investigate the corrected Navier-Stokes (CNS) equations. Grounded in the fluid kinematics of Rortex, the CNS abandons the Stokes' isotropic hypothesis and applies the shearing-only constitutive relation by explicitly eliminating the controversial stretching terms in the stress tensor. Comparisons with the DNS data from the traditional Navier-Stokes (TNS) suggest that the CNS inherently rectifies the near-wall momentum transport. The removal of stretching-induced dissipation shifts the inner- and buffer-layer boundaries towards the wall and effectively suppresses the overshoot in the mean velocity profile in TNS. The turbulence statistics demonstrate a multiscale kinetic energy redistribution that intensifies the near-wall production-dissipation cycle. Furthermore, the topological delineation of instantaneous coherent structures, namely the discovery and definition of rotational/non-rotational interface (RNRI) via the velocity gradient tensor (VGT) discriminant (Delta = 0), confirms that the CNS is capable of capturing the highly complex and interwoven vortical structures. Ultimately, the spectral proper orthogonal decomposition (SPOD) unveils and elucidates that the shearing-only mechanisms intrinsically modulate the spatiotemporal energy cascade, promoting denser and more inclined vortices while enhancing turbulence intermittency by fragmenting the coherent packets. Overall, by isolating and detecting the shearing-only mechanism with physics purity, the DNS based on the CNS provides a more refined perception of the intrinsic dynamics of wall-bounded turbulence, offering a more physical soundness model to capture the interactions among the multiscale coherent structures and the improved capability in predicting the wall-bounded turbulence.
GigaAM Multilingual: Foundation Model for Underrepresented Languages
arXiv:2607.10371v1 Announce Type: cross Abstract: Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challenge of building robust foundation models for underrepresented Central Asian languages (Kazakh, Kyrgyz, Uzbek). We present GigaAM Multilingual, a Conformer encoder pre-trained on 2M hours of audio using a HuBERT-style objective. Crucially, we introduce a cluster-level data balancing strategy during pre-training and a domain-aware sampling method during fine-tuning to mitigate head-language dominance. In controlled comparisons, our approach outperforms strong open pretrained encoders (Whisper Large v3, Omnilingual-1B) on target languages, achieving significant gains on spontaneous speech while maintaining efficiency. We release the foundation encoder and ASR model, offering a proven recipe for effective multilingual adaptation under realistic data imbalance.