It's generally unprofessional for attending physicians to talk about residents behind their back. But there is one reason we all will: when we consider that resident dangerous.
Dangerous is a loaded term, so let me give you its precise meaning. A dangerous resident has a noticeable gap between their knowledge and experience level and their confidence. We can handle residents who are slow, or who make mistakes — if they didn't need help, they wouldn't be in residency. But when an attending colleague pulls me aside and says "Hey Graham, you should watch that resident closely," it means something specific: the resident is making unforced errors relative to their experience level. Not considering a diagnosis. Introducing unnecessary risk during a procedure. Confidently making an assessment that is way off base.
This is not about carelessness or sloppy practice — that's a separate problem entirely. Dangerous means something more precise: the attending needs to watch them like a hawk, needs to double-check and confirm everything, and cannot trust their assessment by default. There is a gap between what the trainee knows and what they think they know, with no outward signal that the gap exists.
Medicine's safety infrastructure runs on this self-reported limitation. Consultation systems, escalation protocols, supervision requirements — all of them depend on someone in the chain recognizing where their competence ends and generating a signal. If the signal doesn't come, the system has no trigger. The dangerous practitioner isn't dangerous because they lack knowledge — lacking knowledge is why we have education and training. They're dangerous because their confidence is miscalibrated. Of Rumsfeld's four quadrants, the unknown unknowns are always the most treacherous.

The miscalibrated practitioner fails situationally. The AI fails constitutionally.
This is, with uncomfortable precision, a description of every generative AI system in clinical use today.
The Line You Can't See
This piece was inspired by a TikTok video. A nurse practitioner looked into his camera and said: "There is nothing more dangerous than a nurse practitioner who doesn't know they're over their head." If you're in healthcare, you know we’ve spent decades arguing about scope of practice — who can do what, with or without supervision. Nurse practitioners, physician assistants, nurse anesthetists have all had fierce debates with physicians and lawmakers, rarely fully resolved. But the NP in that video wasn't making an argument about where the line should be drawn. He was making an argument about what happens when you can't see the line at all.
What makes any clinician dangerous — at any credential level — isn't their ceiling. It's their inability to realize when they’re bumping into theirs.
It's worth being precise about what scope of practice actually encodes, because it’s not just “are you allowed to do thing X in the hospital” like a list of privileges. “Scope” is, I think, the full consequence chain of a clinical action: the ability to do thing X, recognize the downstream consequences of thing X — success or complication — and then manage it, know when to escalate, know who to call, and own what comes next. While thing X could even be something like “take a history” or “clean an abrasion,” procedures are probably the easier way to think about scope.
We could probably teach a lot of people how to do an appendectomy in a remote setting. But even if we did, that’s still not “scope of practice for appendectomy,” because we’d also need to teach the post-operative course — how much pain is normal versus a sign of something wrong? What do I do about a fever on day three? Does that wound look infected, or just slowly healing?
This is not an AI problem, this is a today problem. 20 years ago, when patients spent several days in the hospital after surgery, the surgery team saw their own complications. They owned the full arc. Now most surgeries are outpatient, or same-day discharge — which means the post-op complications come to me, in the emergency department. And here's the thing: I can't own that arc either. Some swelling after a knee replacement is normal. Some redness and discoloration too. But how much? That's a judgment built from seeing hundreds of post-op knees, and that's not my training. I call ortho. Which means the consequence chain that surgery initiated now runs through at least three different clinicians before it resolves — and at each handoff, someone is working at or near the edge of their scope.
Scope of practice is a shorthand for: this is where your consequence ownership ends.
But now let’s compare it to generative AI, which theoretically has no boundary. We talk about GenAI as “the world’s best cardiologist, AND orthopedist, AND emergency physician, AND pediatrician in the palm of your hand,” but with omnipotence comes new risk, too. GenAI will give you the appendectomy and the post-operative management and the complication recognition and the escalation decision, all in the same confident, unbroken paragraph. It does not experience a limit. It does not slow down at the edge of its competence. It does not pause… or… hesitate. You cannot hear Claude’s voice waver, or see Claude’s face wince with self-doubt. The scope-of-practice line that medicine has spent decades drawing and arguing about simply does not exist for AI in any native, architectural sense.
(GenAI zealots will say, “Sure you could just prompt it to hesitate,” but now you’re having to tell the human “Reminder, GenAI is in hesitate mode, not confidence mode,” and how the hell do you then decipher what it’s saying?)
Redundancy Is a Feature, Not a Bug
As I’ve said a million times, I hate inefficiency; it’s what drives me as an informatics and builder in healthcare. But I’d like us to consider why some redundancy has been added to healthcare over the past 100 years.
Consider the patient who calls a nurse advice line at 2am with chest pain. Pressure, center of her chest, started an hour ago. The nurse asks several questions and sends her to the emergency department. A paramedic crew runs a 12-lead EKG before she reaches the door. The triage nurse notices she is diaphoretic in a way that doesn't match her reported pain level, but it's resolved by the time the ER doctor sees her. The cardiologist gets paged at 4am.
Now imagine replacing all of that — every step — with an AI chatbot. Remember, we've got near-omnipotence here. The AI chatbot can access the world's cardiology literature. And is also the world's best triage nurse. And can answer questions in real time. The upside sounds obvious: the best cardiologic reasoning, available to anyone, at any hour, for "free." What we don't name is what we've actually built — a system that bypasses every serial filter in that chain. Sometimes, yes, those filters get in the way. But anyone who's worked in a hospital can tell you how often they catch things too.
Each node in that chain is not only a filter but a self-aware one — each practitioner knows, roughly, what they are good at catching and what they are likely to miss, and that self-knowledge is what generates the next handoff. Redundancy is a feature, not a bug. We built these systems because individual nodes fail, and we built them with overlapping coverage because we knew individual nodes would fail silently. The handoffs are not inefficiency. They are the architecture of a system that does not trust any single point of assessment.
The Double Failure
I'm a huge proponent of upskilling: take someone less skilled, give them AI assistance, and get them operating at a higher level. There's good data behind this. An NBER study on customer service agents showed you could get new hires performing at senior-level quality within weeks. It's genuinely remarkable.
In medicine, the pitch is the same. Give a primary care doctor AI-assisted dermatology support and watch their diagnostic accuracy climb — from 70% to 84%, say. Close the specialist gap in rural communities. Bring expertise to places it's never been. I believe in this. I think it's real.
But there's a compounding problem I didn't fully reckon with until recently, and it's this: the AI that's doing the upskilling has no concept of where the upskilling ends and over-reach begins.
A primary care doctor using AI to diagnose a rare, serious dermatologic condition might get the diagnosis right. But then what? Who manages the treatment? Who recognizes the complications? Who knows which of those complications are serious and which are expected? That's not a dermatology question — that's a dermatologist question. It's built from years of seeing what happens after the diagnosis. And the AI chatbot that just helped nail the diagnosis has no idea that a limit just got crossed. It doesn't have the concept of a limit in its architecture. It will happily help manage the condition it just identified, with the same confidence it brought to the diagnosis, right up to and past the point where the clinician is dangerously out of their depth.
This is the NP over their skis — except the human has no idea they've left the slope, because the thing helping them doesn't know what a slope is.
This is different from the dangerous resident — and in one specific way, it is worse. The dangerous resident exists within the system. Someone can see them. And critically: a good attending isn't just watching for distress signals. They're independently verifying — checking the reasoning, reviewing the plan, catching the gap between the resident's confidence and their actual knowledge. The dangerous resident may have no tells. But the system is designed to catch them anyway, through the attending.
The chatbot has no attending. There is no independent node verifying its reasoning from a separate vantage point. When it's wrong, the output looks exactly like when it's right — and nothing in the architecture is positioned to notice the difference.
I've written separately about what "I don't know" actually does mechanically in clinical training.
The Question We Are Not Asking
The field has spent enormous energy measuring what clinical AI gets right. Benchmark after benchmark: accuracy, sensitivity, specificity — how does the model perform against physicians on licensing exam questions, can it read a chest X-ray as well as a radiologist.
We have spent almost no energy on whether it knows when to shut up.
This is not a small gap. Clinical safety is not primarily a function of average performance. It's a function of failure modes — where the system breaks down, how it breaks down, and whether there is any mechanism to catch the breakdown before it reaches the patient. A system that is right 95 percent of the time and fails silently is more dangerous in many clinical contexts than a system that is right 80 percent of the time and flags its uncertainty.
The dangerous resident fails silently. That is precisely what makes them dangerous. Everything else — the knowledge gaps, the inexperience, the errors — those are manageable if the system knows about them. The dangerous part is the absence of signal.
Current generative AI is, structurally, a silent failure machine. Not because it fails constantly — it doesn't — but because when it fails, it fails the same way it succeeds: fluently, confidently, and without any indication that this particular output is the one you should not trust.
What This Demands of Us
None of this is an argument against AI in clinical medicine. I want to be clear about that, because the response to any critique in this space tends to be binary. This is not "AI bad. Human good.” This is: we are dropping something into a system that was built around an assumption AI doesn't share.
The dangerous resident analogy is instructive here — not as a condemnation, but as a protocol. We don't hand dangerous residents unsupervised access to complex patients and trust that their confidence is calibrated. We watch them more closely, not less. We make sure the redundancy is intact. We strengthen the safety nets they're most likely to bypass. And crucially — we tell them. We give them direct feedback that their confidence is outrunning their competence, and most of the time, that feedback lands. The resident recalibrates. The gap closes. The system has a correction mechanism.
AI has no equivalent. You cannot pull a language model aside and say "hey, you should watch yourself more closely." There is no recalibration loop. The miscalibration is not a phase — it's the architecture.
The question is not only what the AI knows. It's what the AI doesn't know about what it doesn't know — and what we plan to do about the fact that it will never tell us. The triage stack exists for a reason. Scope of practice lines exist for a reason. Independent verification exists for a reason. All of it encodes the same lesson, earned in decades of hard cases: in medicine, uncalibrated confidence isn't a minor flaw. It's the specific mechanism by which patients can get hurt.
We have built something that is, by architecture, incapable of that calibration in the clinical sense. Fluent at the center of its competence and at the edge of it. Confident when it's right and confident when it's wrong. Unbounded by scope, unaware of the redundancy it collapses, constitutionally unable to generate the signal that would tell us when to watch more closely.
That is not a reason to stop. It is a reason to stop mistaking the absence of signal for the absence of error.