Key Highlights

  • A Harvard study finds OpenAI’s o1-preview outperforming ER physicians on diagnostic accuracy — 67.1% triage accuracy versus 55.3% and 50.0% for two attendings, across 76 real cases. Notably, the model flagged a rare flesh-eating infection 12–24 hours before the human team did. The finding is on a 2024-era model — worth tracking what happens when frontier models get the same evaluation.

Research

AI shows its skills in the emergency room — The Rundown

Harvard researchers ran OpenAI’s o1-preview against two attending physicians on 76 emergency-department cases, scoring at the triage stage where information is sparsest. The model came in at 67.1% diagnostic accuracy versus 55.3% and 50.0% for the human doctors, and independent reviewers couldn’t tell AI-generated diagnoses from human ones. The standout case: o1-preview surfaced a rare flesh-eating infection 12–24 hours before the attending caught it.


References

  1. The Rundown, “AI shows its skills in the emergency room,” 2026-05-04