AI Dramatically Speeds Up Psychiatric Research Reviews
Researchers have developed an AI system that can screen thousands of research abstracts for systematic reviews in psychiatry, cutting manual work by 38% while maintaining 97% accuracy. The breakthrough could accelerate evidence synthesis across healthcare, reducing the time and cost of literature reviews that currently consume months of expert labor.
Originaltitel: Novel Abstract Screening Algorithm Using Delphi-Inspired Large Language Model Consensus for Systematic Reviews in Psychiatry: Nouvel algorithme de sélection des résumés utilisant un consensus issu d’un grand modèle de langage inspiré de la méthode Delphi pour les revues systématiques en psychiatrie
<p>Background: Large language models (LLMs) may reduce the burden associated with performing systematic reviews by prescreening abstracts from a literature search for eligibility for inclusion in full-text review. Methods: We developed an iterative, LLM-based workflow for screening abstracts: after manual specification of eligibility criteria and seed examples, an ensemble of five LLMs deliberates through a Delphi process to classify a batch of abstracts; these labels are used to train a logistic regression model that ranks the remaining abstracts and identifies a new batch of abstracts for LLM escalation until all abstracts are labelled by the LLM or probability thresholds. We tested our workflow on abstracts screened in three published systematic reviews in psychiatry. Our primary endpoint was the recall metric, and secondary endpoint was the work saved over sampling at 95% recall metric (WSS@95%). Results: In a dataset on autism biomarkers, 1,655 (35%) of 4,745 retrieved abstracts were judged to be relevant by the original authors. The Delphi–LLM workflow correctly identified 1,605 (97.0%) of these 1,655 abstracts (precision = 54.2%, WSS@95% = 38.1%). The performance metrics were better than non-LLM approaches (recall ≤ 91%, WSS@95 ≤ 26%), and, overall, balanced these metrics optimally compared to single-LLM agents (recall = 84.9–99.9%, WSS@95% = 16.7–39.8%). The recall and work saved metrics were similarly reliable and among the top in two low-prevalence datasets on an attention-deficit hyperactivity disorder treatment review (10% of 2,891 relevant) and a posttraumatic stress disorder trajectory review (7% of 4,453 relevant). For these two datasets, recall was 100.0% and 96.4%, and the WSS@95% was 17.3% and 18.5%, respectively. Conclusions: We presented the design and validation of a novel abstract screening workflow that centres around a Delphi-style aggregation process to harness the strengths of five open-source LLMs that can be run on consumer-level workstations. This multi-LLM workflow showed acceptable and reliable performance for use as an automated prescreening method to facilitate systematic reviews. © The Author(s) 2026. This article is distributed under the terms of the Creative Commons Attribution-NonCommercial 4.0 License (https://creativecommons.org/licenses/by-nc/4.0/) which permits non-commercial use, reproduction and distribution of the work without further permission provided the original work is attributed as specified on the SAGE and Open Access page (https://us.sagepub.com/en-us/nam/open-access-at-sage).</p>