arXiv:2508.00110v2 Announce Type: replace-cross
Abstract: Functional data present unique challenges for clustering due to their infinite-dimensional nature and potential sensitivity to outliers. An extension of the OCLUST algorithm to the functional setting is proposed to address these issues. The approach leverages the OCLUST framework, creating a robust method to cluster curves and trim outliers. The methodology is evaluated on both simulated and real-world functional datasets, demonstrating strong performance in clustering and outlier identification.
Science Journals
arXiv:2607.09115v2 Announce Type: replace
Abstract: Event-based vision and spiking neural networks (SNNs) are increasingly adopted for edge intelligence under strict latency and energy constraints. However, the vulnerability of event-based SNN object detection models to availability backdoor attacks remains insufficiently studied. This paper presents Event Burst Trigger (EBT), an availability backdoor attack targeting SNN-based object detection models. EBT injects carefully crafted event-based triggers into the training data, which induce temporally concentrated event streams during inference. These burst-like activations increase the number of phantom (i.e., spurious) object candidates, and consequently inflate the computational cost of the post-processing stage, particularly Non-Maximum Suppression (NMS). We evaluate EBT on SpikeYOLO, the state-of-the-art SNN-based object detector, under a poison-only threat model that does not require modifications to the model architecture, loss function, or inference pipeline. Experimental results show that while detection accuracy remains largely preserved, with mAP@0.5 decreasing by less than 0.099, the latency of the NMS stage increases by up to 38%. This indicates that NMS can become a dominant availability bottleneck in event-based SNN object detection. Experiments on an edge platform further show that the proposed attack elevates baseline resource utilization and reduces scheduling slack without inducing conspicuous peaks in resource usage. In addition, STRIP-based backdoor detection fails to reliably distinguish the proposed attack from benign inputs. These results characterize a previously underexplored availability backdoor threat in event-based SNN object detection systems.
arXiv:2605.07928v2 Announce Type: replace-cross
Abstract: Magnetohydrodynamic (MHD) simulations are indispensable research infrastructure in astrophysics today. In order to satisfy the solenoidal constraint of the MHD equations on discretized grids, modern simulation codes often employ either constrained transport (CT) with a staggered grid or divergence cleaning using an additional variable. We compare CT and Dedner's mixed divergence cleaning schemes systematically, and find that the divergence cleaning scheme can produce substantial artifacts in certain situations. Through numerical experiments including both idealized tests and practical applications, we show that the original implementation of Dedner's scheme becomes inaccurate when magnetic fields are strongly localized or when the timestep suddenly changes. We find that some previous results, such as the extremely rapid growth of magnetic fields during star formation in the early Universe, may be affected by the spurious behavior of the divergence cleaning scheme. We propose a few modifications to improve the robustness of the divergence cleaning method. Nevertheless, we find that the CT scheme is more accurate and reliable in many situations.
arXiv:2607.11873v1 Announce Type: new
Abstract: Institutions collect far more open-ended teaching-evaluation feedback than they read. A prior study introduced a validated protocol for classifying such comments by thematic category and sentiment, built from a documented annotation guide, an intra-annotator reliability measurement, stratified cross-validation, and a held-out evaluation on a Spanish institutional corpus with a frozen-encoder design. Two questions limit its reuse: whether a protocol fixed to 2019-era frozen embeddings stays competitive as representation methods advance, and whether it transfers to a second language. We re-run it on the original Spanish data across three representation generations, sparse lexical features, frozen transformer embeddings, and prompted large language models, and transfer its sentiment task to English with a balanced 45,000-comment corpus checked against an aspect-labeled education dataset. Treating paired comparisons as descriptive, we find the protocol durable: a 2026 frontier model posts the highest thematic F1 on the hardest Spanish task, yet shows no sentiment advantage over a cheap model and no descriptive separation from it on English, so model choice is a deployment decision, not a property of the method.
arXiv:2412.01283v3 Announce Type: replace-cross
Abstract: We investigate the structure of Kazhdan-Lusztig polynomials of the symmetric group by leveraging computational approaches from big data, including exploratory and topological data analysis, applied to the polynomials for symmetric groups of up to 11 strands.
arXiv:2206.07447v3 Announce Type: replace-cross
Abstract: In a magnetised plasma on scales well above ion kinetic scales, any constant-magnitude magnetic field, accompanied by parallel Alfv\'enic flows, forms a nonlinear solution in an isobaric, constant-density background. These structures, which are also known as spherically polarised Alfv\'en waves, are observed ubiquitously in the solar wind, presumably created by the growth of small-amplitude fluctuations as they propagate outwards in the corona. Here, we present a computational method to construct such solutions of arbitrary amplitude with general multi-dimensional structure, and explore some of their properties. The difficulty lies in computing a zero-divergence, constant-magnitude magnetic field, which leaves a single, quasi-free function to define the solution, while requiring strong constraints on any individual component of the field. Motivated by the physical process of wave growth in the solar wind, our method circumvents this issue by starting from low-amplitude Alfv\'enic fluctuations dominated by a strong mean field, then "growing" magnetic perturbations into the large-amplitude regime. We present example solutions with nontrivial structure in one, two, and three dimensions, demonstrating a clear tendency to form very sharp gradients or discontinuities, unless the solution is one dimensional. As well as being useful as an input for other calculations, particularly the study of parametric decay, our results provide a natural explanation for the extremely sharp field discontinuities observed across magnetic-field switchbacks in the low solar wind.
arXiv:2603.29637v2 Announce Type: replace-cross
Abstract: Using a field theory equivalent to a lattice version of the Poland-Scheraga model, the phase diagram for a long DNA molecule is derived in closed form. For the generalized model with excluded-volume interactions a one-loop renormalization group calculation shows that there are two stable fixed points. At both fixed points, the excluded-volume effect plays a role. At the fixed point reached when the original excluded-volume effect is weak, the phase transition is continuous. At the other fixed point, the phase transition is first order.
arXiv:2503.12692v2 Announce Type: replace-cross
Abstract: The self-organization of microbial ecosystems involves a large variety of mechanisms, ranging from biochemical signaling to population dynamics. Among these, the role of motility regulation has been little studied, despite the importance of active migration processes. Here we show how weak, random motility regulation suffices to induce complex forms of organization in bacterial mixtures comprising a large number of coexisting strains. First, we simulate microscopic models of run-and-tumble bacteria whose self-propulsion speeds are weakly regulated by the local density of each strain, mimicking the impact of weak, random metabolic interactions. Our simulations reveal that, as the heterogeneity of the interaction network increases, the system undergoes a phase transition leading to the emergence of distinct, spatially segregated communities. To account for these results and assess their robustness, we use random-matrix theory to analyze the hydrodynamic description of the bacterial mixture, obtaining a quantitative agreement with our microscopic simulations. Our results hold for a variety of motility-regulation mechanisms and highlight the need to characterize the role of motility regulation in experimentally relevant situations.
arXiv:2503.15854v3 Announce Type: replace-cross
Abstract: Stiefel-Whitney classes are topological invariants of vector bundles, and those of the tangent bundle capture essential features of a manifold, such as whether it is orientable and how it can be embedded in Euclidean space. We present an algorithm that computes these classes for the tangent bundle directly from a finite sample of points. Starting from the point cloud, we build a filtration of simplicial complexes and compute its persistent cohomology, and then apply the Wu formula, which recovers the Stiefel-Whitney classes from the cup product and the Steenrod squares alone, without estimating tangent spaces or a smooth structure. The key step, finding the Wu classes, reduces to solving a system of linear equations, so the computation runs in polynomial time in the number of simplices. We prove that whenever the sample recovers the shape of a closed manifold, the computed classes agree with the true Stiefel-Whitney classes of its tangent bundle, and that this remains true even when the data carry spurious topological features on which the Steenrod squares vanish, so the classes can be identified over a wide range of scales rather than only where the sample matches the manifold exactly. We illustrate the method on triangulated four-dimensional manifolds and on point clouds coming from image patches and from a molecular conformation space.
arXiv:2607.09680v1 Announce Type: cross
Abstract: Continuous cardiac monitoring in wearable devices demands classifiers that are simultaneously accurate, energy-efficient, and deployable on resource-constrained hardware. While deep neural network approaches have demonstrated high classification accuracy for electrocardiogram (ECG) arrhythmia detection, their substantial parameter counts and reliance on multiply-accumulate-intensive operations make them impractical for low-cost edge platforms. In this work, we propose ECG-LDC, a hardware-software co-design framework that adapts Low-Dimensional Computing (LDC) for real-time ECG arrhythmia classification. ECG-LDC employs a dual-encoder architecture with dedicated value and feature codebooks to independently encode morphological waveform features and RR-interval temporal features, enabling effective capture of both intra-beat and inter-beat cardiac dynamics. The framework encompasses data preprocessing, model training, and a hardware accelerator architecture prototyped on the Pynq-Z2 platform. Implemented using binary representations and XOR/XNOR-based operations, ECG-LDC achieves $97.18\%$ accuracy with a memory footprint of only $3.86\ \text{kB}$. ECG-LDC sacrifices approximately $1.8\%$ accuracy versus SOTA TinyML classifiers but achieves $11$~$ 570\times$ reduction in memory usage; among FPGA-based five-class arrhythmia classifiers, it delivers the highest accuracy with up to $2.4\times$ fewer LUTs and zero DSP block utilization, affirming its suitability for real-time arrhythmia detection on resource-constrained wearable platforms.
arXiv:2506.21426v2 Announce Type: replace
Abstract: Recent crises like the Covid-19 pandemic and geopolitical tensions have exposed vulnerabilities and caused disruptions of supply chains, leading to product shortages, increased costs, and economic instability. This has prompted growing efforts to assess systemic risk, namely the effects of firm disruptions on entire economies. However, the ability of firms to react to crises by rewiring their supply links has been largely overlooked, limiting our understanding of production networks resilience. Here, we study dynamics and determinants of firm-level systemic risk in the Hungarian economy from 2015 to 2022. We benchmark our results to a heuristic maximum entropy null model that generates randomized production networks while preserving the total input (demand) and output (supply) of each firm at the sector level. We show that the fairly stable set of firms with highest systemic risk undergoes a structural change during Covid-19, as those enabling economic exchanges become key players in the economy -- a pattern not reproduced by the null model. Although empirical systemic risk closely matches the null value prior to the pandemic, it becomes significantly lower afterwards, reflecting the emergence of a more resilient economy driven by firms' adaptive behavior. Furthermore, firms' international trade volume (being itself a channel of potential disruption) becomes a significant predictor of their systemic risk. However, international linkages alone cannot fully explain the observed trends, as imports and exports exert opposing effects on local systemic risk through the supply and demand channels.
arXiv:2607.11621v1 Announce Type: new
Abstract: Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language models not designed for clinical simulation can reproduce these patterns remains untested. We investigated (1) whether lesions or controlled perturbations to a multimodal language model can reproduce different types of errors in picture naming, and (2) whether the framework can reproduce the complete error profile of individual persons with aphasia (PWAs). Using LLaVA 1.6, we evaluated perturbation configurations that varied the layer, proportion, and amount of noise applied to model units. We examined 278 PWAs on the Philadelphia Naming Test, classifying responses into seven categories using a validated neural classifier. Six of seven response categories (correct, semantic, mixed, unrelated, neologism, no response errors) emerged at clinically-comparable proportions across distinct parameter space regions, with formal paraphasia being the exception. Searching the perturbation space revealed configurations that reproduced the individual error profile in at least six of seven categories for 97.8% of PWAs and in all seven categories for 79.5% of PWAs. Monte Carlo baselines confirmed that this matching reflects joint inter-category structure rather than marginal overlap. These results establish a quantitative framework for reproducing individual aphasic error patterns in picture naming. They suggest the potential for language models to serve as digital twins of individuals with post-stroke aphasia.
arXiv:2607.10024v1 Announce Type: cross
Abstract: We investigate the magnetic-field response of interstitial anionic electrons (IAEs) in two-dimensional electrides, using monolayer Ca$_2$N as a prototypical system. By computing the Landau-level (LL) spectrum of the electride bands forming the Fermi surface, we find a linear LL evolution with magnetic field that closely resembles the behavior of a nearly-free 2D electron gas (2DEG). The extracted cyclotron effective mass and Land\'e g-factor deviate moderately from their free-electron values, indicating that the IAEs retain a remarkably free-electron-like character. Furthermore, the energy dispersion of the electride bands remains insensitive to the choice of exchange-correlation functional (LDA vs.~PBEsol), indicating that local exchange and correlation effects have minimal influence on the IAEs. Overall, our findings provide fundamental insight into the quantum nature of electrides and open new avenues for exploring magnetic confinement, correlation effects, and emergent quantum phenomena in low-dimensional interstitial electronic systems.
arXiv:2510.01830v2 Announce Type: replace
Abstract: Object-Goal Navigation (ObjectNav) is a key capability for deploying mobile robots in everyday environments such as homes, schools, and workplaces. In this task, an agent must locate an instance of a target object category in previously unseen environments using only onboard perception, requiring the integration of semantic understanding, spatial reasoning, and long-horizon planning. Reinforcement learning (RL) has become a dominant paradigm for ObjectNav, yet modern systems involve numerous design choices across perception modules, policy architectures, and inference-time strategies. The relative impact of these components, however, remains poorly understood. In this work, we present a large-scale empirical study of modular RL-based ObjectNav systems. We decompose the navigation pipeline into three key components: perception, policy, and test-time enhancement, and conduct extensive controlled experiments to analyze their individual contributions. Our results suggest that improvements in perception quality and test-time strategies often yield larger performance gains than policy improvements alone, highlighting the importance of understanding how different components interact within modular navigation systems. Motivated by these findings, we introduce a unified framework for systematically studying modular ObjectNav systems. Guided by our analysis, we build an enhanced system that achieves state-of-the-art performance on the Gibson benchmark, improving SPL by 6.6% and success rate by 2.7% over prior methods. We also introduce a human expert baseline, achieving 98% success, highlighting the significant gap between current RL agents and human-level navigation. Finally, we provide practical insights and design recommendations for each module to help guide future research. Project page: https://honwang0054.github.io/What-matters-in-RL-ObjNav-web/.
arXiv:2607.10162v1 Announce Type: cross
Abstract: Sound symbolism, the human tendency to map speech sounds to perceptual qualities such as roundness or sharpness, arises primarily from the acoustics of speech rather than spelling. Whether Speech Language Models (SLMs) share this tendency remains open, as prior evaluations rely on text or images rather than real speech. We study it using genuine human speech recordings, comparing model judgments against human data across the auditory, crossmodal, and visual components of the effect. We find that SLMs' auditory judgments align poorly with human perception and miss the acoustic cues, such as spectral tilt, that drive human intuitions, and open-weight models cannot reliably link a heard sound to its corresponding shape. With a visual-only control ruling out shape perception, the weakness localizes to how speech is represented, suggesting that perceptual alignment depends not on stronger vision but on speech representations that capture the cues humans hear.
arXiv:2601.17090v2 Announce Type: replace
Abstract: Partial differential equations (PDEs) govern complex systems, yet neural operators often struggle to efficiently capture the long-range, nonlocal interactions inherent in their solution maps. We introduce Spectral Filtering Operator (SFO), a neural operator that parameterizes integral kernels using the Universal Spectral Basis (USB), a fixed, global orthonormal basis derived from the eigenmodes of the Hilbert matrix in spectral filtering theory. Motivated by our theoretical finding that the discrete Green's functions of shift-invariant PDE discretizations exhibit spatial Linear Dynamical System (LDS) structure, we prove that these kernels admit compact approximations in the USB. By learning only the spectral coefficients of rapidly decaying eigenvalues, SFO achieves a highly efficient representation. Across six benchmarks, including reaction-diffusion, fluid dynamics, and 3D electromagnetics, SFO achieves state-of-the-art accuracy, reducing error by up to 40% relative to strong baselines while using substantially fewer parameters.
arXiv:2204.09333v3 Announce Type: replace
Abstract: In recent years, with the rapid growth of Internet data, the number and types of scientific and technological resources are also rapidly expanding. However, the increase in the number and category of information data will also increase the cost of information acquisition. For technology-based enterprises or users, in addition to general papers, patents, and other resources, policies related to technology or the development of their industries should also belong to a type of scientific and technological resource. Extracting valuable science and technology policy resources from a huge amount of mixed-content data and providing accurate and fast retrieval will help break down information barriers and reduce information-acquisition costs, which has profound social significance and utility. This article focuses on the difficulties and problems in the field of science and technology policy and introduces related technologies and developments.
arXiv:2507.15148v2 Announce Type: replace-cross
Abstract: Accurately solving the Schr\"odinger equation remains a central challenge in computational physics, chemistry, and materials science. Here, we propose an alternative eigenvalue problem based on a system's autocorrelation function, avoiding direct reference to a wave function. In particular, we develop a rigorous approximation framework that enables precise frequency estimation from a finite number of signal samples. Our analysis builds on new results involving prolate spheroidal wave functions and yields error bounds that reveal a sharp accuracy transition governed by the observation time and spectral density of the signal. These results are very general and thus carry far. As one important example application we consider the quantum computation for molecular systems. By combining our spectral method with a quantum subroutine for signal generation, we define quantum prolate diagonalization (QPD) - a hybrid classical-quantum algorithm. QPD simultaneously estimates ground and excited state energies within chemical accuracy at the Heisenberg limit. An analysis of different input states demonstrates the robustness of the method, showing that high precision can be retained even under imperfect state preparation.
arXiv:2606.20859v2 Announce Type: replace-cross
Abstract: A fundamental assumption in statistics and machine learning is that ``the future looks like the past,'' formalized as exchangeability: the joint data distribution is order-invariant. In practice, this assumption is often violated due to distribution shifts over time. Early detection of exchangeability violations is crucial to prevent performance degradation and enable timely interventions like model retraining. Conformal test martingales offer a flexible, distribution-free framework for sequential exchangeability testing with guaranteed false-alarm rate control by betting against the uniformity of conformal p-values. While alternatives such as plug-in martingales and mixture-based strategies exist, computationally efficient baselines like the Simple Jumper are limited to detecting mean location shifts. We propose a family of conformal test martingales based on shifted Legendre polynomials that extend the Simple Jumper to higher-order moments. The Simple Legendre Jumper replaces linear betting functions with polynomials of arbitrary degree, enabling rapid detection of variance, skewness, and other higher-order deviations. The Product Legendre Jumper combines multiple polynomial degrees into a single betting function but suffers from exponential state-space growth, termed the jumping tax. To resolve this, we introduce the Variational Legendre Jumper, which employs a mean-field approximation to reduce complexity to constant time per step with minimal power loss, providing an expressive, scalable framework for real-time distribution shift monitoring.
arXiv:2606.22636v2 Announce Type: replace-cross
Abstract: We prove an explicit spectral-gap lower bound for the lazy swap chain on binary matrices with prescribed row and column sums. This chain is a standard sampler for fixed-margin null models in ecology, statistics, and network analysis. Kannan, Tetali, and Vempala (KTV) conjectured that it mixes rapidly for all feasible margins \citep{kannan1997simple}. We show that for every feasible set of margins on an $m\times n$ binary matrix, the lazy swap chain has spectral gap at least $$\binom{m}{2}^{-1}\binom{n}{2}^{-1}.$$ The bound is tight in the worst case. Thus, our result proves this KTV conjecture in a stronger quantitative form. The same spectral-gap bound also verifies the Mihail--Vazirani conjecture for fixed-margin 0/1-matrix polytopes.
The proof gives a new route to fixed-margin sampling that avoids stability assumptions and canonical-path constructions. We compare the swap chain with a two-row heat-bath chain and use a local-to-global spectral reduction to reduce the analysis from arbitrary $m\times n$ matrices to a three-row problem. The remaining three-row inequality is then proved by separating the scalar column-count sector from the non-scalar Johnson harmonic sectors.
The proof itself was generated by ChatGPT 5.5 Pro. The author's role was to pose the problem, guide the search direction, evaluate the AI-generated arguments, rewrite the proof, and take responsibility for the final form and validity of the result. The full proof of the main theorem has been formalized in Lean, and the accompanying formalization is available at the anonymous repository https://github.com/guanyangwang/ktv-swap-lean.
arXiv:2607.09322v2 Announce Type: replace
Abstract: In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluations of LLM-based medical agents have largely emphasized short-context knowledge QA and tool use. However, real-world medical care is inherently longitudinal, and clinicians must aggregate evidence across repeated visits, tests, and evolving treatments. Therefore, long-horizon interaction is essential for realistic assessment. LongMedBench is constructed via a reproducible pipeline that integrates MIMIC-IV admission records and clinical notes into time-series event streams and long-context memory datasets, enabling long-horizon, multi-session interactions between agents and a clinical environment. It comprises 335 patients, with 19.72 inpatient visits per patient on average and 44.91 medical events per visit. Guided by the long-horizon decision process, we propose an evaluation taxonomy with three suites: fact-based QA, temporal reasoning, and long-horizon decision-making. This taxonomy measures how agents understand and leverage historical patient information over extended horizons. Our experiments show that while recent LLMs can make good use of explicit timestamps, they have challenges in implicit time inference; The RAG and agent memory system can improve the performance of information retrieval tasks, but the performance of decision-making tasks is highly dependent on the model's immediate context.
LLM-Driven Collaborative Model for Untangling Commits via Explicit and Implicit Dependency Reasoning
arXiv:2507.16395v3 Announce Type: replace
Abstract: Atomic commits, which address a single development concern, are a best practice in software development. In practice, however, developers often produce tangled commits that mix unrelated changes, complicating code review and maintenance. Prior untangling approaches (rule-based, feature-based, or graph-based) have made progress but typically rely on shallow signals and struggle to distinguish explicit dependencies (e.g., control/data flow) from implicit ones (e.g., semantic or conceptual relationships). In this paper, we propose ColaUntangle, a new collaborative consultation framework for commit untangling that models both explicit and implicit dependencies among code changes. ColaUntangle integrates Large Language Model (LLM)-driven agents in a multi-agent architecture: one agent specializes in explicit dependencies, another in implicit ones, and a reviewer agent synthesizes their perspectives through iterative consultation. To capture structural and contextual information, we construct Explicit and Implicit Contexts, enabling agents to reason over code relationships with both symbolic and semantic depth. We evaluate ColaUntangle on two widely-used datasets (1,612 C# and 14k Java tangled commits). Experimental results show that ColaUntangle outperforms the best-performing baseline, achieving an improvement of 44% on the C# dataset and 82% on the Java dataset. These findings highlight the potential of LLM-based collaborative frameworks for advancing automated commit untangling tasks.
arXiv:2507.21759v3 Announce Type: replace
Abstract: With the large-scale integration of electric vehicles (EVs) in the distribution grid, the unpredictable nature of EV charging introduces considerable uncertainties to the grid's real-time operations. This can exacerbate load fluctuations, compromise power quality, and pose risks to the grid's stability and security. However, due to their dual role as controllable loads and energy storage devices, EVs have the potential to mitigate these fluctuations, balance the variability of renewable energy sources, and provide ancillary services that support grid stability. By leveraging the bidirectional flow of information and energy in smart grids, the adverse effects of EV charging can be minimized and even converted into beneficial outcomes through effective real-time management strategies. This paper explores the negative impacts of EV charging on the distribution system's real-time operations and outlines methods to transform these challenges into positive contributions. Additionally, it provides an in-depth analysis of the real-time management system for EV charging, focusing on state estimation and management strategies.
arXiv:2508.02091v3 Announce Type: replace
Abstract: Approximate nearest-neighbor search (ANNS) algorithms have become increasingly critical for recent AI applications, particularly in retrieval-augmented generation (RAG) and agent-based LLM applications. In this paper, we present CRINN, a new paradigm for ANNS algorithms. CRINN treats ANNS optimization as a reinforcement learning problem where execution speed serves as the reward signal. This approach enables the automatic generation of progressively faster ANNS implementations while maintaining accuracy constraints. Our experimental evaluation demonstrates CRINN's effectiveness across six widely-used NNS benchmark datasets. When compared against state-of-the-art open-source ANNS algorithms, CRINN achieves best performance on three of them (GIST-960-Euclidean, MNIST-784-Euclidean, and GloVe-25-angular), and tied for first place on two of them (SIFT-128-Euclidean and GloVe-25-angular). The implications of CRINN's success reach well beyond ANNS optimization: It validates that LLMs augmented with reinforcement learning can function as an effective tool for automating sophisticated algorithmic optimizations that demand specialized knowledge and labor-intensive manual refinement. Code can be found at https://github.com/deepreinforce-ai/CRINN
arXiv:2508.09904v3 Announce Type: replace
Abstract: Real-world forecasting requires models to integrate not only historical data but also relevant contextual information provided in textual form. While large language models (LLMs) show promise for context-aided forecasting, critical challenges remain: we lack diagnostic tools to understand failure modes, performance remains far below their potential, and high computational costs limit practical deployment. We introduce a unified framework of four strategies that address these limitations along three orthogonal dimensions: model diagnostics, accuracy, and efficiency. Through extensive evaluation across model families from small open-source models to frontier models including Gemini, GPT, and Claude, we uncover both fundamental insights and practical solutions. Our findings span three key dimensions: diagnostic strategies reveal the "Execution Gap" where models correctly explain how context affects forecasts but fail to apply this reasoning; accuracy-focused strategies achieve substantial performance improvements of 25-50%; and efficiency-oriented approaches show that adaptive routing between small and large models can approach large model accuracy on average while significantly reducing inference costs. These orthogonal strategies can be flexibly integrated based on deployment constraints, providing practitioners with a comprehensive toolkit for practical LLM-based context-aided forecasting. Code is made available at https://github.com/ashok-arjun/beyond-naive-prompting.