arXiv:2604.24592v3 Announce Type: replace-cross Abstract: We present a technique to investigate the stationary states of a system of a collisionless system confined by an external potential and coupled to boundary reservoirs through prescribed reinjection rules. We consider a family of boundary conditions parametrized by an integer $n$, corresponding to different velocity distributions imposed at the boundaries, generalizing the standard flux-weighted Maxwellian scheme. By combining Liouville's theorem with the boundary injection rule, we derive an explicit analytical expression for the stationary distribution function. This framework provides a direct link between microscopic boundary dynamics and macroscopic stationary profiles. We show that thermal equilibrium is recovered only for the standard flux-weighted injection method, while for all other cases the system relaxes to manifestly non-thermal stationary states. The resulting density and temperature profiles exhibit non-trivial spatial structures, including non-monotonic behaviour and temperature gradients induced by the boundary conditions alone. Analytical predictions for stationary moments are obtained in closed form for representative cases and are nicely reproduced by particle-based numerical simulations.
Science Journals
arXiv:2312.13683v2 Announce Type: replace-cross Abstract: The next-generation wireless networks are envisioned to jointly support high-rate communications and ubiquitous sensing. Ultra-Massive Multiple-Input Multiple-Output (UM-MIMO) offers abundant spatial Degrees of Freedom (DoFs) for both functions, yet its large aperture shifts electromagnetic propagation into the near field, invalidating conventional far-field (plane-wave) assumptions. While near-field channel modeling has been studied, existing channel estimation methods are inadequate: on-grid designs suffer from non-orthogonal codebooks, and off-grid methods lack convergence guarantees, yielding unreliable estimates. Moreover, channel estimation and localization are typically designed in isolation, preventing the exchange of information that could otherwise enable mutual performance improvement. To address this difficulty, we propose a unified framework that exploits near-field characteristics to jointly design channel estimation and cooperative localization. Specifically, we develop a Variational Newtonized Near-field Channel Estimation (VNNCE) algorithm that extracts position-aware soft information from the channel, and a Gaussian Fusion Cooperative Localization (GFCL) method that leverages this information across multiple Base Stations (BSs) for enhanced accuracy.
arXiv:2606.00149v2 Announce Type: replace-cross Abstract: Two-state stochastic models, where motion alternates between distinct dynamical modes, are widely observed in complex systems. Here we study the Two-State Random Walk (TSRW), which switches between a continuous-time random walk (CTRW) rest state and a standard L'evy walk (LW) motion state, each with power-law distributed sojourn times. Using anomalous diffusion decomposition, we show that TSRWs exhibit a generic coexistence of Joseph (correlation), Noah (heavy-tailed increments), and Moses (aging) effects. Strikingly, although classical L'evy walks alone possess only the Joseph effect, both Noah and Moses effects emerge in TSRWs solely due to stochastic switching with the CTRW phase. Our results demonstrate that coupling between dynamical states can fundamentally reshape the mechanisms driving anomalous diffusion, offering a minimal yet powerful framework for transport in heterogeneous and intermittently switching environments.
arXiv:2604.25965v2 Announce Type: replace-cross Abstract: Deep learning models are widely deployed in safety-critical domains, but remain vulnerable to adversarial attacks. In this paper, we study the adversarial robustness of NTK neural networks in the context of nonparametric regression. We establish minimax optimal rates for adversarial regression in Sobolev spaces and then show that NTK neural networks, trained via gradient flow with early stopping, can achieve this optimal rate. However, in the overfitting regime, we prove that the minimum norm interpolant is vulnerable to adversarial perturbations.
arXiv:2606.02667v2 Announce Type: replace-cross Abstract: Let $f(k,s)$ denote the minimum integer $m$ such that any family $\mathcal{F}$ consisting of $k$-sized sets of cardinality at least $m$ always contain a sunflower of size $s$. The Erd\H{o}s-Rado Sunflower Conjecture states that for every $s >2$, there is an constant $C=C(s)$ such that $f(k,s) \leq C^k$. In this paper, we prove the conjecture for shifted families.
arXiv:2606.09360v1 Announce Type: new Abstract: Open-domain open-vocabulary detection (ODOVD) requires detectors to generalize to both novel categories and unseen domains, making it more challenging than open-vocabulary detection. Existing methods typically train open-vocabulary detectors together with domain generalization modules from scratch, leading to high training cost. we propose ExDet, a lightweight category-domain collaborative generalization framework for ODOVD that enhances the cross-category and cross-domain generalization of existing detectors. ExDet consists of Text-Guided Extrapolation (TGE), a lightweight Detector-Compatible Rectification (DCR) module, and ExRPN. Specifically, TGE exploits the DeltaSpace property of vision-language models (VLMs) to infer category- and domain-aware proxy visual prototypes from text. DCR is learned from the TGE-generated prototypes in a detector training-free and real-data-free manner, and is inserted after the classification head at inference to rectify representations toward a detector-compatible source-domain visual distribution, thereby enhancing classification for targets from novel categories and unseen domains. ExRPN recalibrates proposal scores by combining semantic similarity with RPN confidence, improving recall for novel and domain-shifted objects while providing better support for subsequent classification and DCR. ExDet achieves SOTA performance on OD-LVIS, OV-LVIS, Objects365, and MSOSB.
arXiv:2606.09380v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a leading paradigm for improving the reasoning ability of large language models through outcome-based supervision. However, verifiable rewards frequently become uninformative at the group level: when all sampled traces of a given prompt receive identical rewards, group-relative advantage estimation provides no gradient signal, even though the traces may differ substantially in reasoning quality. We propose Reasoning Arena, an adaptive training framework that routes such non-diverse reward groups to a judge system instead of discarding them. Beyond examining the final answer, Reasoning Arena constructs trace tournaments, where reasoning traces are compared head-to-head to expose finer-grained preferences within the group, converting reasoning quality into rich relative reward signals. To make reward estimation efficient, rather than exhaustively comparing every pair, each new trace is evaluated against a small, dynamically updated pool of previously generated traces as anchors to efficiently establish a relative ranking. We then fit a Bradley-Terry model on the incomplete comparison graph, enabling scalable RL integration without quadratic pairwise comparisons. Empirical results demonstrate that Reasoning Arena consistently outperforms the RLVR baseline by 7.6% on average in competition mathematics and coding benchmarks. By converting otherwise wasted zero-advantage samples into useful gradient updates, our method accelerates training by 27% to 41%, saving nearly 50% of generation compute, and substantially improves overall reasoning performance.
arXiv:2606.09383v1 Announce Type: new Abstract: Conventional dynamics analysis of the human body is often constrained by the need for contact force and torque sensors and controlled laboratory environments. To address this issue, this study proposes an opticalmechanics kinematic-dynamic integrated estimation framework for multibody systems. Specifically, a constrained multibody model is established to describe the system dynamics, while image-measured kinematic quantities are used as non contact inputs for dynamic estimation. The unknown joint torque is then identified through a genetic-algorithm based optimization by minimizing the discrepancy between model-predicted and image-measured kinematic quan tities. Experimental validation on an air-bearing platform showed that the wrist joint torque estimated from image data achieved a mean absolute error of 0.46 Nm compared with sensor measurements. In the forward prediction test, the model-predicted angular velocity achieved a mean absolute error of 0.006 rad/s relative to the image-measured results. This study demonstrates the potential of combining image measurement and mechanical modeling for non-contact dynamic estimation in scenarios where direct force and torque measurement is difficult.
arXiv:2606.03275v2 Announce Type: replace Abstract: A low-order Galerkin model is developed for a rotating fluid annulus driven by localized heating at the outer bottom periphery, with uniform cooling at the inner cylindrical wall. The model retains the full cylindrical geometry and employs Bessel-function radial eigenfunctions satisfying physically correct Dirichlet-Neumann boundary conditions. A dual-series least-squares procedure determines the conductive base state under the mixed thermal boundary condition. Galerkin projection onto the leading radial and vertical basis functions yields a 10-variable dynamical system governing the mean meridional overturning, thermal wind, baroclinic wave amplitudes, and their nonlinear interactions. Linear stability analysis yields explicit critical Rayleigh numbers for both mean and wave instabilities, showing that rotation raises Ra_c in proportion to T^2. The model reproduces the Nu ~ Ra^(1/4) scaling, rotational suppression at low Ra, and the boundary-layer-dominated flow structure observed in companion axisymmetric simulations.
arXiv:2501.11755v2 Announce Type: replace-cross Abstract: Current self-supervised learning methods for 3D medical imaging rely on simple pretext formulations and organ- or modality-specific datasets, limiting their generalizability and scalability. We present 3DINO, a cutting-edge SSL method adapted to 3D datasets, and use it to pretrain 3DINO-ViT: a general-purpose medical imaging model, on an exceptionally large, multimodal, and multi-organ dataset of ~100,000 3D medical imaging scans from over 10 organs. We validate 3DINO-ViT using extensive experiments on numerous medical imaging segmentation and classification tasks. Our results demonstrate that 3DINO-ViT generalizes across modalities and organs, including out-of-distribution tasks and datasets, outperforming state-of-the-art methods on the majority of evaluation metrics and labeled dataset sizes. Our 3DINO framework and 3DINO-ViT will be made available to enable research on 3D foundation models or further finetuning for a wide range of medical imaging applications.
arXiv:2502.01332v3 Announce Type: replace-cross Abstract: The coherent equalization problem consists in designing a quantum system acting as a mean-square near-optimal filter for a given quantum communication channel. The paper develops an improved method for the synthesis of transfer functions for such equalizing filters, based on a linear quantum system model of the channel. The method draws on a connection with the two-disk problem of ${H}_{\infty}$ control for classical (i.e., non-quantum) linear uncertain systems. Compared with the previous methods, the proposed method applies to a broader class of linear quantum communication channels.
arXiv:2507.10091v2 Announce Type: replace-cross Abstract: This paper focuses on extending the use of Minkowski Tensors to analyze anisotropic signals in cosmological data, focusing on those introduced by redshift space distortion. We derive the ensemble average of the two translation-invariant, rank-2 Minkowski Tensors ($W_1^{0,2}$ and $W_2^{0,2}$) for a matter density field that is perturbatively non-Gaussian in redshift space. This is achieved through the Edgeworth expansion of the joint probability density function of the field and its derivatives, expressing the ensemble averages in terms of cumulants up to cubic order. Our goal is to connect these theoretical predictions to the underlying cosmological parameters, allowing for parameter estimation by measuring them from galaxy surveys. The work builds on previous analyses of Minkowski Functionals in both real and redshift space and addresses the effects of Finger-of-God velocity dispersion and shot noise. We validate our predictions by matching them to measurements of the Minkowski Tensors from dark matter simulation data, finding that perturbation theory is a qualified success. Non-perturbative Finger-of-God effects remain significant at relatively large scales $R_G \lesssim 20 \, h^{-1} \, {\rm Mpc}$ and are particularly pronounced in the components parallel to the line of sight.
arXiv:2606.09392v1 Announce Type: new Abstract: Efficient acquisition, storage, and utilization of traffic data are critical challenges in spatio-temporal data management. Most traffic data systems collect and store observations at fixed, coarse-grained temporal intervals to reduce storage and computation costs. However, such coarse-grained data severely limits downstream applications that require predictions at a finer temporal granularity. Collecting and maintaining fine-grained traffic data across all locations and time periods would impose a substantial burden on database storage and preprocessing pipelines. To address this temporal granularity mismatch, we formulate a novel problem: predicting fine-grained future traffic using coarse-grained sampled data. We propose the Spatial-Temporal Refinement Predictor (STRP), a granularity-aware framework for spatio-temporal data systems. STRP integrates two components: Tree Convolution for efficient and interpretable spatial dependency modeling, and Inverse Dilated Convolution for progressive temporal extrapolation. STRP supports two practical prediction settings: window-based and duration-based, to handle different forms of granularity mismatch. Experiments on six benchmark datasets show that STRP significantly outperforms state-of-the-art baselines in both accuracy and efficiency. Our work offers a practical and interpretable approach to managing granularity mismatches in spatio-temporal traffic data systems.
arXiv:2507.13503v2 Announce Type: replace-cross Abstract: Monte Carlo algorithms, like the Swendsen-Wang and invaded-cluster, sample the Ising and Potts models asymptotically faster than single-spin Glauber dynamics do. Here, we generalize both algorithms to sample Potts lattice gauge theory by way of a $2$-dimensional cellular representation called the plaquette random-cluster model. The invaded-cluster algorithm targets Potts lattice gauge theory at criticality by implementing a stopping condition defined in terms of homological percolation, the emergence of spanning surfaces on the torus. Simulations for $\mathbb Z(2)$ and $\mathbb Z(3)$ lattice gauge theories on the cubical $4$-dimensional torus indicate that both generalized algorithms exhibit much faster autocorrelation decay than single-spin dynamics and allow for efficient sampling on $4$-dimensional tori of linear scale at least $40$.
arXiv:2606.09435v1 Announce Type: new Abstract: Multilingual dictionaries are among the most valuable documentary resources for low-resource and endangered languages, yet many remain available only as scans. For many decades, their digitization and conversion into a machine-readable format was nearly impossible due to language-specific scripts, complex multi-column layouts full of entries with abbreviations and cross-references. Recent vision-language models offer a promising solution, but it is unclear how well they preserve characters, markup, and process lexicographic structure. We introduce MUDIDI, a two-stage framework for multi-lingual dictionary digitization. Stage One evaluates the quality of character recognition and markup preservation; Stage Two focuses on dictionary entry segmentation with subsequent mapping into a machine-readable lexicographic schema, SIL's Multi-Dictionary Formatter. We also release a dataset that consists of human-annotated lexicographic entries collected from 30 public-domain dictionaries featuring diverse writing systems, language families, and formats. We benchmark OCR systems, general-purpose Large Language Models (LLMs), and Vision Language Models (VLMs) on the dataset, demonstrating superior performance of LLMs across most writing systems and languages in both stages, and provide practical guidelines on improving the results for more challenging scenarios. Finally, we show that supplementing additional information, such as dictionary introduction, to the LLMs can improve the quality of the digitized dictionary. Github: https://github.com/DavidSamuell/MUDIDI-Pipeline-for-Digitization-of-Multilingual-Dictionary/
arXiv:2606.09436v1 Announce Type: new Abstract: The emerging AC/multi-terminal DC grids are regarded as a promising solution for accommodating the increasing integration of renewable energy sources. This work proposes an optimization framework to address transmission switching (TS) problems arising in practical operational scenarios, such as maintenance scheduling, contingency management, and fault restoration. Unlike most existing studies, the proposed framework considers the role of communication networks in TS operations and develops an optimal information-power flow (OIPF) model. The OIPF model captures the impact of information flows on circuit breaker actions while incorporating communication-related costs, thereby better reflecting practical operational decision-making processes. To ensure computational tractability, the resulting optimization problem is formulated as a mixed-integer second-order cone programming (MISOCP) model through convex relaxations, polygonal approximations, and Big-M reformulations. Numerical case studies illustrate the applicability of the proposed OIPF model and indicate its potential in supporting transmission switching decisions.
arXiv:2606.09447v1 Announce Type: new Abstract: We present AliyunConsoleAgent, a web agent framework for automated documentation verification in real-world cloud consoles. Major cloud platforms encompass hundreds of products with rapid feature iteration, causing console UIs to frequently diverge from their corresponding documentation. Verifying that documented procedures accurately reflect the current console and can be executed end-to-end demands an estimated 4 million recurring inspections annually, yet manual coverage remains below 1%. While agent systems built on frontier proprietary models achieve high success rates, their prohibitive cost and data privacy constraints preclude large-scale deployment. We propose a two-stage training paradigm: supervised fine-tuning (SFT) on distilled frontier-model trajectories, followed by reinforcement learning using Group Relative Policy Optimization (GRPO) and a dual-channel outcome reward model in real cloud environments. To support large-scale RL training, we construct a high-determinism rollout system featuring Terraform-based resource pre-provisioning and LLM-driven on-demand provisioning, which effectively isolates environment noise from the training signal. We further introduce a rule-based reward evaluation protocol grounded in backend audit logs, providing objective, reward-hacking-resistant outcome judgment. Our model evolves from mechanical instruction following to autonomous decision-making with cloud console and product-specific understanding. Experiments on a challenging 278-task benchmark where the best frontier model achieves only 65.34% demonstrate that AliyunConsoleAgent-32B achieves a 63.52% mean success rate -- a 20.24 percentage-point improvement over the base model, narrowing the gap to the best frontier proprietary model to 1.82 pp (bootstrap 95% CI [-1.27, 7.39]) -- at 92% lower inference cost.
arXiv:2508.12913v2 Announce Type: replace-cross Abstract: We investigate spectral fluctuations in multilayer networks within the random matrix theory (RMT) framework to characterize universal and non-universal features. The adjacency matrix of a multilayer network exhibits a block structure, with diagonal blocks representing intra-layer connections and off-diagonal blocks encoding inter-layer connections. Applying appropriate scaling factors for these blocks, we equalize variances across inter- and intra-layers, enabling direct comparison of spectral statistics. We analyze eigenvalue spectra across multilayer network configurations with varying inter- and intra-layer connectivities. Introducing a crossover model for bilayer networks, we capture the smooth transition of spectral properties from block-diagonal (two independent GOEs) to single-layer (one GOE) statistics as the relative strength of inter-layer to intra-layer connection varies. Furthermore, we analyze interatomic distance networks derived from protein crystal structures, including 1EWT, 1EWK, and 1UW6, to demonstrate applicability. Our findings reveal that the universality of spectral fluctuations persists across multilayer network architectures and highlight RMT as a robust tool for probing topological and dynamical complexities of real-world networks.
arXiv:2509.02916v4 Announce Type: replace-cross Abstract: We report the first results on ultracold neutron production from a new spallation-driven superfluid $^4$He (He-II) source at TRIUMF, which is being prepared for a new, precise measurement of the neutron electric dipole moment. A total of $(9.3 \pm 0.8)\times 10^{5}$ ultracold neutrons were observed at a proton beam current of \SI{37}{\uA}, when the target was irradiated for a period of \SI{60}{\s}. The results are in fair agreement with expectations based on a detailed simulation of neutron transport and ultracold neutron source cryogenics. There is some indication that the new source might not be as limited by the conduction of heat through the He-II as originally expected. The results indicate that the source is likely to make its ultimate production goals, once the liquid deuterium cold moderator system is completed, with the expectation that $5.7\times 10^7$~UCNs would be detected in the same experiment with full liquid levels. This would, for example, correspond to delivery of $1.4\times 10^6$~UCNs delivered to each of two nEDM measurement cells, and a statistical uncertainty of $1\times 10^{-27}~e$cm on the neutron EDM in 280 days of running.
No need to stay positive: a practical approach to direct numerical simulations of elastic turbulence
arXiv:2606.09468v1 Announce Type: new Abstract: Successfully performing direct numerical simulations of polymeric flows remains a major challenge in computational fluid mechanics. In addition to the velocity field, such simulations must resolve polymeric degrees of freedom, often expressed via the conformation tensor, $\mathbf{c}$, which captures the local stretch of polymer molecules. A key difficulty here lies in maintaining the physical requirement $\mathrm{Tr}\, \mathbf{c}>3$, which is not explicitly enforced by the governing equations. Consequently, simulations initiated from physical conditions may silently drift into unphysical states with $\mathrm{Tr}\, \mathbf{c}<0$, indicating a loss of positive-definiteness of the conformation tensor. Existing numerical methods to prevent this are costly, making direct numerical simulations of chaotic polymer flows, such as elastic turbulence, heavily reliant on high-performance computing. Here, we ask whether simulations that violate $\mathrm{Tr}\, \mathbf{c}>3$ can still yield meaningful physical insight into the underlying dynamics. We simulate a model dilute polymer solution driven through a plane channel at low Reynolds number and observe the transition to elastic turbulence. Our simulations exhibit two threshold resolutions: below the first, they become numerically unstable and exhibit a finite-time blow-up; above the second, they maintain positive-definiteness. In between, simulations remain stable and chaotic despite local violations of $\mathrm{Tr}\, \mathbf{c}>3$. Surprisingly, these violations do not affect mid-plane statistics of velocity, its gradients, or polymer stretch, which match results from fully positive-definite simulations. This suggests that resolving flow structures or key flow statistics may not require the extreme resolutions needed to preserve positive-definiteness, potentially lowering computational barriers for studying elastic turbulence.
arXiv:2606.09471v1 Announce Type: new Abstract: On-policy distillation (OPD) provides dense token-level supervision by asking a teacher to score student-generated rollouts. However, when the student drifts into an unrecoverable prefix, the teacher may locally agree with the degraded state, producing low reverse KL but little corrective training signal. We identify this persistent regime as a low-KL agreement trap. Further analyses show that tokens during and after such traps produce less useful supervision signals. We propose KAT (KL Agreement Trap Termination), an online OPD termination rule that detects persistent low-KL agreement with a dynamic training-adaptive threshold. By filtering weak supervision from degenerate agreement, KAT improves avg@k accuracy by 2.66% and pass@k by 3.43% across four mathematical benchmarks, while reducing average rollout length by 59.73%.
arXiv:2606.09484v1 Announce Type: new Abstract: Large language models (LLMs) have shown impressive performance on diverse reasoning tasks, yet their capacity for structural reasoning in graphs remains unclear. We investigate whether LLMs can genuinely understand graph isomorphism -a fundamental problem in graph theory. While LLMs achieve near-perfect accuracy on isomorphism detection, we show this performance is illusory. When identical graphs are presented with permuted node labels, LLMs fail to identify their isomorphism. This finding suggests that LLMs exploit patterns rather than reasoning about abstract graph structure. Since permutation invariance is a fundamental requirement for valid structural reasoning, these results indicate that success on graph reasoning benchmarks should not be interpreted as evidence of genuine topological understanding.
arXiv:2606.09489v1 Announce Type: new Abstract: Objective: Conformance checking in healthcare seeks to assess whether patient care pathways adhere to clinical guidelines. However, its practical application often depends on the availability of formal, machine-interpretable representations of guidelines, such as Computer-Interpretable Guidelines (CIGs), which are seldom available in real-world clinical settings. Methods: This work introduces a modular framework based on the orchestration of Large Language Models (LLMs) to support medical conformance checking directly from unstructured clinical and guideline texts, without requiring predefined CIGs. The proposed architecture integrates multiple LLMs and supporting components to extract patient traces from clinical discharge letters, identify normative rules from textual clinical guidelines, translate these rules into executable scripts, and compute a Trace Conformance Indicator to quantify compliance within the event log. Results: The framework was implemented and evaluated in the stroke care domain at the neurological ward of Alessandria Hospital. Hundreds of patient traces were automatically extracted from hospital data and assessed against 50 rules derived from the reference guideline. The analysis showed that more than 86\% of the available traces were conformant. Conclusion: The results demonstrate the feasibility of using orchestrated LLMs for practical healthcare conformance analysis. At the same time, the study provides evidence of a high level of adherence to stroke care guidelines at Alessandria Hospital.
arXiv:2606.09493v1 Announce Type: new Abstract: We investigate collisionless power absorption in resonant, low$-$pressure capacitively coupled plasmas (CCPs). In these radio-frequency (RF) discharges, the sheath capacitance almost exactly balances the plasma inductance, driving the total RF discharge voltage down to just a few volts. However, plasma persists not only in this ultra$-$low$-$voltage regime; it also generates ions that strike the electrodes with kinetic energies substantially exceeding the amplitude of the applied RF voltage. This counterintuitive behavior arises from the presence of a pronounced electrostatic potential well of approximately 40 V within the plasma bulk, which confines electrons while simultaneously accelerating ions toward the electrodes. We show that, under these resonant conditions, collisionless electron heating exhibits a fundamentally different behavior from the conventional paradigm of stochastic sheath heating mediated by electron$-$sheath interactions. Instead, the predominant energy transfer mechanism is bulk electron heating in RF electric fields via a primarily collisionless process that emerges from the synergistic action of: (i) a strongly amplified RF electric field within the plasma bulk, (ii) electron oscillatory motion (bouncing) within the plasma potential well, and (iii) electron scattering resulting from collisions with neutral atoms. Collectively, these phenomena give rise to a pronounced high$-$energy tail in the electron energy distribution function and thereby lead to a substantial enhancement of the ionization rates. As the gas pressure rises, the resonance is disrupted. At the same time, the region of maximum power absorption moves from the plasma core toward the edges and the sheath, which is accompanied by the disappearance of the high$-$energy electron population and a corresponding decrease in ionization rates
arXiv:2604.06278v4 Announce Type: replace-cross Abstract: Small regional datasets pose a dual statistical problem: correlated predictors inflate estimation variance, while flexible learners can become unstable because the available information per adaptive degree of freedom is limited. We examine this issue through predictive volatility, defined as the cross-sample dispersion and upper-tail behaviour of out-of-sample loss. Using simulation evidence reported for sparse linear, near-linear and heavy-tailed settings, we compare ordinary least squares, frequentist penalties, Bayesian shrinkage models, bounded-response and spatial specifications, and flexible machine-learning procedures. In the reported simulation results, regularised linear estimators generally dominate in the linear high-collinearity micro-sample settings and remain the most reliable overall, whereas tree-based methods become more competitive only when the signal is weakly nonlinear and the sample size is larger. In the empirical application to 34 Indonesian provinces, ridge yields the best leave-one-out performance, followed by elastic net and lasso. Across the Bayesian shrinkage specifications, ICT skills show the most consistent negative association with poverty, with the strongest support under horseshoe and spike-and-slab formulations. These results suggest that, in micro-sample regional modelling, the main constraint is limited information per effective degree of freedom rather than insufficient algorithmic flexibility.