In a recent series of experiments led by researchers at the Beth Israel Deaconess Medical Center and Harvard Medical School, a preview of the OpenAI large language model (LLM) known as o1 bested human physicians in multiple tests of clinical and diagnostic reasoning. "We tested the AI model against virtually every benchmark, and it eclipsed both prior models and our physician baselines," said Harvard Medical School professor Arjun K. Manrai, one of the study's senior authors. The average score for o1 was 89 percent, compared to 34 percent for human doctors given the same test. The LLM especially outperformed human doctors when it came to diagnosis at the initial, triage stage, identifying "the exact or very close diagnosis….in 67.1% of cases," while the two human doctors did so in only 55.3 percent and 50 percent of cases, respectively. Measures like these could hamper AI's potential to diagnose diseases with more accuracy than human doctors alone can.