I say this with all the love in my heart: Your ER doctor probably doesn't care about your exact diagnosis — and this is one of the many ways that media is failing in its reporting of the new paper in Science this week.
It's a good paper and extremely info-dense, but it's given me a lot of thought because the portion that seems to be getting all the headlines is its tiny section on my specialty, emergency medicine. So as your resident AI-advocate and-also-AI-skeptic ER doctor, here’s my take for just the ER portion of the study (which, again, is like 20% of the full paper):
The Science paper's ER evaluation
The "real-world ER" experiment compared two LLMs against two internal medicine attendings (not ER attendings; more on that below) on 79 cases. All four arms received identical EHR text packets at three touchpoints — triage, ED evaluation, admission — and then all four arms were asked to produce a five-item ranked differential diagnosis. The metric: did your differential include the eventual confirmed diagnosis? Seems good, right? As you get more information throughout the ER visit, how do all four arms do with capturing the exactly correct diagnosis?

And damn, AI (especially o1, the first reasoning/"thinking" model from OpenAI) — really appears to best humans here.
On my first reading of this paper, this metric made a ton of sense, but the more I thought about it, the less it sat right with me. It is a metric, but I'm not sure it's the one I want in the ED.
These weren't run of the mill ER patients
Just to be explicit here, they did not take 80 patients from the ER for this study. They took 80 patients from the ER who then got admitted and looked at those patients (technically 79, because one patient's data couldn't get collected).
But that’s a tiny portion of our ED patients! We discharge 85% of patients from the ED on average. The hardest part of emergency medicine is figuring out who to send home. Those patients were entirely excluded from this study. Might o1 have done just as well on non-admitted patients? Totally. Maybe even better. (But probably so would the humans, too.)
We don’t care about your diagnosis.
Okay look I’m being a little facetious, but only a little; while other doctors need to arrive at the exact right diagnosis to assign the exact right treatment and prognosis, that is not an ER doctor’s prime directive. My goals are very, very different:
- Stabilize patients who are dying.
- Admit patients who need admission.
- Don't send the wrong ones home.
Now to do 2 and 3, yes, of course, I generate a differential diagnosis. But once I know you have something that warrants admission (let's say your sodium is critically low or that you've had a stroke) — and I know enough to know what services you’ll need, I hand you off to one of my esteemed colleagues in those fields. In the stroke example: I need to know enough about the stroke — whether it's a hemorrhagic stroke or an ischemic one — to know whether to talk with the neurologists or the neurosurgeons, and perhaps if I need to transfer you out to another hospital. Or maybe I need to consider other possibilities like diseases mimicking a stroke. But once those answers are clarified, your final diagnosis doesn't really impact me and what I do as your ER doctor.
I care way more about crossing off dangerous diagnoses from my list than I do being exactly right about what’s going on with you. Whether your chest pain is acid reflux or costochondritis or a pulled muscle or anxiety doesn't matter — they all go home. But I do care about being confident that you're not having a heart attack that's just acting like heartburn. Whether your low sodium is due to SIADH or polydipsia doesn't really matter to me, either — the inpatient team will figure it out over 48 hours.
Just to defend myself here a bit: This is not because I’m lazy. This is not because ER doctors are stupid, contrary to popular social media belief. It’s because I am a sorting hat who has a continuous stream of patients coming down the conveyor belt that I cannot turn off for any reason, so I don’t have the time nor the resources to arrive at the exact right answer, nor do I need the exact right answer to help get the patient the resources they need for their problem.
So when the study asked, "Did the AI get the answer right from triage?" it's kind of a weird question to me. The paper highlights the largest gap at "initial ER triage, where there is the most urgency to make the correct decision." But correct decision about what, exactly? There is zero urgency for anyone in the ER — nurse, doctor, pharmacist, housekeeper — to generate a comprehensive differential at triage. Triage is another round of sorting. Its job is acuity assignment and recognizing time-critical patterns. If we just saw every patient in the ER in the order they arrived, then sure, there would be urgency to make “the correct decision.” But we do triage instead, because we all agree that a stroke should be seen before an ankle sprain.
Case in point: Sometimes I'll mention an interesting diagnosis I picked up to a triage nurse who saw the patient hours earlier. They mostly shrug. Not because they don't care, but because the diagnosis isn't their job. They weren’t even paying attention in that way, because it would distract them from their goal: assign the right triage category and get the patient into the right bed in the right order. That shrug is professional discipline. It's the system working correctly.
Emergency medicine is a very negative specialty.
Most of medicine is positive cognition: diagnose, treat, manage, refine. EM is largely negative cognition: rule out, exclude, mostly discharge, demonstrate that the dangerous thing isn't there. The 200 chest pains discharged appropriately leave no trace. The 50 headaches that weren't SAH don't generate a paper. Our primary work product, on most patients, is a series of negative findings supporting a defensible disposition. None of that is measured by scoring against a final diagnosis.
Discharging a missed emergency is catastrophic. Admitting a patient who didn't need admission is, at most, regrettable. Failing to nail the diagnosis during the ED stay often isn't even an error — the handoff to inpatient is designed to absorb that. The optimization isn't "get the diagnosis right." It's "minimize the rate at which dangerous things go home, subject to operational constraints." That's a fundamentally different problem.
❤️ Love you, hospitalists
I love my internal medicine colleagues — I couldn’t practice good emergency medicine without good hospitalists to hand off to. And I know many of these authors personally — they’re all great clinicians.
No offense, dear friends, but I would not want an internist working in ED triage, like the comparison found in this paper. (You would probably shutter to have an ER doc running the medical floors, too.) We are trained extremely differently; things that are obvious to me (like fix the nursemaids elbow in 5 seconds) might be management dilemmas to them, and often their case conferences debating primary vs second hyperparathyroidism will put any ER doctor to sleep in 5 seconds, too.
I'll take any help I can get, AI or otherwise
With open arms, I welcome AI — or any voluntary trained human — to help me take better care of patients in the ED. I truly do. But in emergency medicine, we have fundamentally different goals than much of non-emergency medicine.
- If you want to build AI that helps hospital administrators believe me that we need more resources because 10 patients just showed up who all might have heart attack as their final diagnosis, I’m all for it — but know that we will have probably sent 8 or 9 of them home.
- If you want to build AI that helps me not make a mistake like sending home a patient who’s got subtle signs of an emergency brewing: yes please.
- If you want to build AI that helps my patients understand their bodies, their disease, and the limitations of modern medicine, I’m all for it.
I absolutely think there's a world where data from this paper is valuable — could AI help sort patients at triage? Yes — there's already been studies supporting this.
And if you follow me, you know that I fundamentally believe medicine fundamentally must adopt new technology if we have any hope of keeping up with demand in healthcare. But when you read headlines that “AI bests doctors in the ER,” because the ER is viewed as a dramatic high stakes place like The Pitt, know that we’re still far, far, off from that reality.
(I also want to call out the authors here because every quote I’ve seen from them is crystal clear that they don’t want this study to be mis-interpreted by Big Tech or Big Media, even though the headlines you read imply the opposite. We need more of these studies and more debate and discussion, and they're the ones doing the difficult work to make this all happen.)
Always down to chat about AI and emergency medicine research! Help me... help me!
Peter Brodeur, MD, MA Thomas Buckley Zahir Kanjee Ethan Goh, MD Evelyn Ling Priyank Jain Raja-Elie Abdulnour Adrian Haimovich Daniel Morgan Jason Hom Robert Gallo Liam McCoy Haadi Mombini Daniele Restifo Eric Horvitz Jonathan H. Chen Arjun Manrai Adam Rodman