Part 2: Real-World Risks and the Confidence Spiral
In Part 1, we explored how generative AI reshapes the way we think and learn. We looked at studies showing how confidence in AI can reduce critical thinking, and how students (and perhaps clinicians-in-training) may learn less when AI gives them the answer too easily. In Part 2, we take that cognitive vulnerability and bring it into the real world of medicine.
De-Skilling in the Real World: The Endoscopy Evidence
While there's been a lot of attention paid to how AI can "up-skill" clinicians — helping novices perform at a higher level — this piece is focused on the inverse: de-skilling. That's when human performance deteriorates after prolonged exposure to automation. So let's look at a paper that does exactly that, examining AI in the context of medical procedures.
Across the globe, millions of colonoscopies are performed daily, often for colorectal cancer screening — meaning no symptoms. We know a few key things:
- Most colorectal cancers start as polyps.
- If we remove polyps, we can prevent the cancer entirely.
- Many patients with polyps — and even early cancers — are asymptomatic.
In screening colonoscopies, gastroenterologists may perform hundreds of procedures before finding a single cancer (but will often find polyps). These are physically demanding tasks: twisting a 20-foot cable through a 5-foot colon, ensuring 360° visualization, and detecting abnormalities that may be hiding in mucosal folds.
That's exactly the kind of setting where AI can help — tired eyes, high volume, low signal. So, a recent multi-center preprint from Poland studied experienced endoscopists using an AI polyp detection system. The researchers compared adenoma detection rates (ADR) — a core quality metric — during standard colonoscopies without AI, both before and after a sustained period of AI use.

It mirrors the Turkish study from Part 1: exposure, performance, then removal.
Here's what they found:
- Baseline (no AI): ADR = 28.4%
- With AI: not specified in this study, but other literature suggests ~33% (5-15% increase)
- After AI was removed: ADR dropped to 22.4%

That's a 6% absolute decrease from the pre-AI baseline — and likely even more from the "AI-assisted" performance. The endoscopists weren't just worse without AI — they were worse than before they ever used it.
If this data is true (this is a pre-print, to be clear), then the likely explanation? Over-reliance. A gradual lowering of vigilance, a shift in attention, and potentially, a soft erosion of skill.
But the most interesting part of this entire paper? The performance drop wasn't uniform. At all! Some endoscopists declined sharply, others less so. This raises a key question: were those who declined the most the ones who had the least confidence in their own abilities, and thus relied more heavily on the AI? Could this be the Microsoft "confidence paradox" in action?
Look at the range here in ADR!
This study suggests we should be real-world concerned that AI can cause subtle, unintentional skill decay — even among experienced clinicians. And if it can happen in endoscopy, why not in radiology? Or surgery? Or pathology? Or interpreting labs or monitoring sepsis trends?
If we today are already lamenting the “death of the physical exam” — meaning that today’s clinicians are less facile at performing and detecting abnormalities when examining patients, what could this mean for AI in medicine? I’ll ask the same question I asked in part one: is this just progress at work? It’s clear that the physical exam is often insufficiently sensitive or specific for many disease states. We no longer wait for the patient with a pulmonary embolism to develop a Homan’s sign or a parasternal lift. We do testing and risk stratification and imaging, because we’ve generally decided that’s a better approach. Will that be the same with AI?
We didn't just gain technology; we lost a way of thinking. AI promises to accelerate this process tenfold.
What happens when the AI flags the wrong thing — or fails to flag anything at all? Do we still have the cognitive strength and perceptual vigilance to catch what it missed?
Confidence Spiral: When Use Begets More Use
One of the more unsettling implications of the Microsoft study from Part 1 is the possibility of a confidence spiral — where using AI actually erodes our confidence in ourselves.
And once that self-confidence declines, what do we do?
We turn to AI even more.
Imagine a physician using GenAI to support diagnoses in their own specialty. Over time, they might begin to second-guess their instincts, defer more decisions to the system, and lose the drive to double-check or investigate further. The more they rely on AI, the less confident they become. And the less confident they are, the more they rely on AI.
It's a feedback loop.
This spiral might be the greatest risk of all — not a dramatic failure, but a slow, steady erosion of our cognitive edge. Once we start handing over the thinking, how — and when — do we take it back?
The Psychological Paradox: AI and Clinician Burnout
Consider the psychological minefield: AI consistently produces differential diagnoses that include rare entities you haven't considered. Initially, you'll feel grateful. Eventually, you'll question your clinical reasoning. It’s death by a thousand “What If”s.
We're potentially creating the first generation of physicians who develop imposter syndrome not from comparing themselves to senior physicians, but from comparing themselves to ChatGPT. Imagine spending a decade in medical training only to feel intellectually inferior to an algorithm that didn't even exist when you started med school.
It’s a double-bind. Using AI feels like cheating; not using it feels like negligence.
The ultimate irony? When AI inevitably makes a critical error, the responsibility falls not on the algorithm, but on the human who's been systematically de-skilled by relying on it.
When AI Succeeds Too Quietly
Am I a modern Luddite for even asking these questions? I’m at least starting to understand how they felt.
We don't panic when Google Maps glitches. Or when a calculator battery dies. These tools are reliable, interchangeable, and — crucially — low stakes. We adapt.
But is medicine in that same category? Even with 99.9% AI uptime, should it be?
Unlike our daily tech, medicine doesn't leave much room for "good enough" thinking. The danger isn't limited to dramatic failure in a resuscitation. It's the erosion of vigilance that occurs over months and years — particularly in the low-stakes, seemingly safe scenarios where it's easiest to defer, and hardest to notice the loss.
This is where the real danger lies:
- The triage note you skim because the summary looks fine.
- The slowly fading habit of building a differential from scratch.
- The intern who never learns to fully synthesize a workup because the AI spits out a plan that seems "good enough."
We assume AI will fail loudly. But maybe the bigger risk is that it succeeds — quietly, plausibly, pervasively — while our own clinical muscles weaken, unnoticed, until we need them most.
The Future of Medical Skills: Critical Thinking, Procedures, and 'Gestalt'
Looking 5–10 years out, it's nearly certain that AI will be integrated into every aspect of medical workflow. So what does that mean for core medical skills?
1. Critical Thinking & Diagnostics: Will we lose the ability to generate a thoughtful differential diagnosis ourselves? Will we become overly anchored by plausible AI suggestions, unable to think outside its initial framing? Will AI help us consider rare diagnoses, or will it burden us with long lists of rare things that the patient doesn’t have?
2. Procedural Skills: As seen in the endoscopy study, we risk atrophy. If AI highlights the polyp or navigates the robot, do we still maintain tactile awareness, spatial understanding, or anatomical confidence when the tool goes offline?
3. Clinical Gestalt: Perhaps the hardest to preserve — and most essential to fields like EM. That quick, experienced "sick vs. not sick" gut instinct. Can AI replicate it? Possibly not. Can over-reliance make us stop developing it? Possibly yes.
Navigating the Path Forward: Education, Design, and Vigilance
The risk is real. But so is the opportunity.
We need a deliberate strategy — not just to use AI, but to preserve and protect human expertise alongside it. My colleague Bernard P. Chang wrote a piece in JAMA recently about the AI-enabled med school of the future, but along with all the promise, I think we need to consider the risk as well. Some ideas I've had:
1. Rethink Medical Education:
- Teach AI literacy: how these models work, where they fail.
- Require AI-free practice: like simulation labs, AI-off days, or first-100-case documentation before AI support is allowed.
- Adopt tools that prioritize learning: GPT Tutors, not GPT base that guide rather than give answers.
- Implement expertise-calibrated access: Reserve full AI capabilities for attendings while limiting trainees to "hint systems" that preserve their cognitive development. The attending gets the answer; the resident gets the pathway to finding it themselves.
2. Design Better Tools:
- Build in "cognitive forcing functions" that require engagement: "The AI thinks this is sepsis — do you agree?"
- Introduce delayed feedback or occasional intentional uncertainty to keep users engaged and alert.
- Show confidence levels, highlight ambiguity, and nudge deeper thinking.
- Apply the Folded-EKG Principle universally: Design all AI interfaces to require clinicians to document independent assessments before revealing AI suggestions. What looks like a workflow inconvenience is actually a cognitive safeguard.
- Create comparative performance dashboards: Show clinicians how their diagnostic accuracy compares with and without AI assistance, giving tangible feedback on whether they're becoming better because of AI or merely dependent on it.
3. Update Professional Standards:
- Require periods of deliberate de-automation: Schedule regular "AI-free shifts" where clinicians practice without algorithmic assistance, maintaining independent skills while identifying areas where expertise has eroded.
- Have protocols for AI downtime, error recognition, and independent cross-checking.
- Foster a culture of healthy skepticism, not blind trust.
- Develop formal curricula teaching clinicians how to critically evaluate AI outputs, including understanding the limitations, biases, and hallucinations inherent in large language models.
Conclusion: Wielding the Tool, Not Becoming It
The calculator didn't destroy our ability to do math — but it made us reach for it by default. AI in medicine presents a similar, but far higher-stakes challenge.
Unchecked reliance on AI doesn't just automate workflows — it rewires (or even dewires) our brains. It threatens the very skills we've spent decades cultivating: critical thinking, procedural precision, clinical gestalt, and human judgment.
We can either let AI reshape the profession quietly — or we can shape its role with intention, with smarter tools, better design, and a commitment to keep the thinking part of medicine exactly where it belongs: in the minds of clinicians.
Then again… maybe I'm wrong.
Maybe we're simply entering a new era, and I'm just a middle-aged fuddy-duddy clinging to legacy thinking. I haven't looked at a paper map in a decade. I trust my phone to get me where I need to go without question. So is it hypocritical of me to worry about our profession? Is AI just… progress?
I really don't know.
But I do know this: The future of medicine shouldn't be about choosing between human and machine. It should be about ensuring AI helps us think better — not think less. It shouldn't make the diagnosis for us, but give us the bandwidth to see the patient behind the diagnosis.
Our challenge now is to harness AI's immense potential while preserving the uniquely human expertise that has defined medicine for millennia.
Perhaps what we need isn't just better AI, but a new model of professional identity that embraces technological augmentation without surrendering cognitive sovereignty. After all, the stethoscope didn't replace our ability to listen — it enhanced it. Can we ensure AI does the same for our ability to think?
Our patients deserve nothing less.