Forskningsradar
← Hälsa & medicin
Hälsa & medicin 4.4 🇸🇪

AI models disagree on medical transcripts—and that's useful

When multiple AI speech-recognition systems transcribe doctor-patient conversations differently, those disagreements pinpoint likely errors, according to new research. For healthcare companies deploying AI scribes at scale, this finding offers a cheap way to prioritize which transcripts need human review—without needing expensive hand-verified reference data.

Originaltitel: Cross-model disagreement as a reference-free signal for prioritizing human review in medical speech transcription.

TL;DR — på svenska

# Modellskillnader kan effektivisera granskning av medicinsk tal-till-text Sjukhusskript som genererats av AI saknar ofta referensstandard för kvalitetskontroll i praktiken. Svenska forskare från Karolinska Institutet testade om oenighet mellan åtta tal-igenkänningssystem kan peka ut osäkra avsnitt i medicinsk transkription utan manuell referenstext. Åtta kommersiella och öppna ASR-system analyserade 50 medicinska utbildningsljudfiler. Resultatet: 72 procent av positionen visade starkt överensstämmelse, medan endast 2,5 procent var högriskområden. Vid lågt systemöverenskommelse ökade transkriptionsfel markant. Genom att flagga positioner med lågt överenskommande kunde forskarna identifiera 94 procent av verifierade fel samtidigt som endast 29 procent av texten behövde granskas manuellt. Metoden kan minska granskningsbelastning i klinisk miljö betydligt. Framtida validering på riktiga patientmöten krävs före implementering i verksamhet.

Abstrakt

INTRODUCTION: Ambient AI scribes generate transcripts at scale, but routine quality assurance is constrained by the absence of human-verified reference transcripts in most deployment settings. We evaluated whether disagreement among heterogeneous automatic speech recognition (ASR) systems can serve as an informative signal for localizing transcription uncertainty, using a public English-language medical-speech corpus rather than clinical encounter recordings. METHODS: Eight commercial and open-source ASR systems were applied to 50 medical-education audio clips (8 h 14 min). Multi-model outputs were aligned, and a leave-one-out consensus procedure was used to score per-model agreement while reducing circularity. RESULTS: Disagreement across models was sparse and localized: 72.1% of positions showed strong agreement (7-8 systems concordant), whereas only 2.5% were high-risk positions with minimal agreement (0-3 systems). Low-agreement regions were systematically enriched for meaning-bearing lexical differences, defined as lexical mismatches after excluding punctuation, contraction, numeric, and filler variation. A single-annotator human-corrected (HC) validation layer showed that transcription errors increased monotonically with decreasing agreement. At an illustrative post hoc threshold, flagging positions where six or fewer systems agreed selected 28.6% of tokens while recovering 93.7% of single-annotator HC-verified errors on this proxy corpus. DISCUSSION: These findings suggest that cross-model disagreement may help focus human review on a small number of likely error-prone transcript regions. However, agreement among all systems does not guarantee correctness, because shared errors may remain undetected by this approach. Validation on real clinical encounter data is required before operational deployment.

Generera ett redaktionellt utkast på svenska