AI & Models
OpenAI model matches human doctors in emergency triage study
A Harvard study found OpenAI’s o1 model performed nominally better than or on par with internal medicine physicians in emergency triage, though experts caution against overhyping results.
A new study published this week in Science evaluated OpenAI’s o1 and 4o models against human physicians using 76 patients from the Beth Israel Deaconess Medical Center emergency room. The research was led by Harvard Medical School and conducted at Beth Israel Deaconess Medical Center, comparing the performance of the artificial intelligence models against two attending physicians—senior doctors who have completed residency training. During the initial triage, which is the process of determining the priority of patients’ treatments based on the severity of their condition, the o1 model achieved an exact or very close diagnosis in 67% of cases. In comparison, the first attending physician achieved an exact or close diagnosis 55% of the time, while the second attending physician hit the mark 50% of the time.
The study noted that at each diagnostic touchpoint, o1 performed nominally better than or on par with the two attending physicians and the 4o model. Arjun Manrai, who heads an AI lab at Harvard Medical School and is one of the study’s lead authors, stated that the researchers tested the AI model against virtually every benchmark, and it eclipsed both prior models and the physician baselines.
Despite these results, the study’s methodology has faced criticism from medical experts. Kristen Panthagani, an emergency physician who critiqued the study, pointed out that the research compared the AI’s diagnoses to those of internal medicine physicians rather than emergency room specialists. “If we’re going to compare AI tools to physicians’ clinical ability, we should start by comparing to physicians who actually practice that specialty,” said Kristen Panthagani, an emergency physician. She noted that as an emergency room doctor seeing a patient for the first time, the primary goal is not to guess an ultimate diagnosis, but rather to determine if the patient has a condition that could kill you.
Additionally, the study does not claim that AI is ready to make life-or-death decisions in the emergency room. The researchers emphasized that the models were only tested on text-based information, noting that existing studies suggest current foundation models are more limited when reasoning over non-text inputs. Adam Rodman, a Beth Israel doctor and one of the study’s lead authors, also warned that there is currently no formal framework for accountability regarding AI-generated diagnoses, and that patients still want humans to guide them through challenging treatment decisions.
Why it matters
The study highlights the potential for AI in clinical settings while underscoring the critical need for rigorous, specialty-specific validation before these tools can be trusted with high-stakes medical decisions.