“In our evaluation of more than 400 radiology reports, GPT-4 achieved precision of 90% and recall of 87% for true critical findings using a few-shot static prompting approach,” said first author Ish A. Talati, MD, from the department of radiology at Stanford University. Reports were manually reviewed to identify critical findings and categorized as true, known/expected, or equivocal. For true critical findings in the holdout test set, GPT-4 achieved 90.1% precision and 86.9% recall, compared to 75.6% and 77.4% for Mistral-7B. On the external test set, GPT-4 reached 82.6% precision and 98.3% recall, while Mistral-7B achieved 75.0% and 93.1%, respectively. “Effective identification of critical findings is essential for patient safety,” Talati and colleagues concluded.