Part 1: Thinking Less, Learning Less
Quick: what's 17 x 19?
While I'd imagine most of you could work this out, you probably reached for your phone. It's less brain work, it's faster, and you're likely more confident in its answer than your own mental math under pressure. I’d also posit that before calculators were commonplace, many could probably do 17 x 19 in their heads rather quickly; they had to.
This isn't a lament about lost arithmetic skills. But I’m not exactly sure it’s a parallel, either. Generative AI is transforming the world, but it isn’t just a calculator handling arithmetic. It’s also algebra and calculus. And it’s meatball recipes, and it’s travel tips for Cairo in April, and it’s Harry Potter fan fiction, and it’s reasoning and thinking itself.
But most critically, GenAI can weaken or replace our clinical thinking pathways — the very mental frameworks we use to diagnose patients, interpret findings, and make treatment decisions. When these skills atrophy, the consequences aren't just inefficiency or inconvenience — they're potentially life-altering for our patients. My worry is that AI adoption in medicine risks degrading our core clinical competencies (diagnosis, management, prognosis, critical thinking) over time, creating a dangerous dependency that masks itself as progress.
That being said, I’m firmly in the “AI has tremendous opportunity for me and my patients” camp. Predictive AI may improve diagnostic accuracy and prognosis. And Generative AI may reduce cognitive load and enhance efficiency, too. But I’m still apprehensive; if medicine teaches you anything, it’s that there’s no free lunch in life. You fix one thing and risk another.
- Make the kidneys happy? The heart or lungs may complain.
- Take opiates for your broken ankle? Shut down your bowels with constipation or make you nauseous, dizzy, or delirious.
- Start to rely on AI? Well, that’s what we’re here to discuss.
Because the risk isn’t that AI makes us lazy — it’s that it quietly rewires our willingness and ability to think critically, to learn deeply, and to maintain skills that medicine demands we keep sharp — even when AI is present.
This is Part 1 of a two-part deep dive on the risks of over-reliance on AI in medicine. In this first piece, we’ll explore how AI affects the way we think and learn — with insights from Generative AI users, high school students, and a few uncomfortable parallels to medical education.
The Confidence Paradox and the Shift in Thinking

Lee et al. from Carnegie Mellon and Microsoft Research
Many of you reading this probably use Generative AI like me, accessing it at least a few times a week if not every single day. But what’s it doing to our brains, our thought patterns, and our own cognitive abilities? Researchers from Microsoft surveyed hundreds of knowledge workers who used GenAI so that the researchers could understand how the knowledge workers were using it, why they were using it, and then tried to determine the factors that impacted its usage. Their findings revealed an interesting "confidence paradox":

Confidence in one's ability impacts AI usage
- Higher confidence in the AI's ability to do a task correlated with less perceived effort and less critical thinking by the human user. Essentially: "I know AI is good at this — perhaps better than me — so I’m going to trust it over my own ability, I’m going to put in less effort, and I won’t have to think as hard about it."
If you work in finance, you may not be a particular expert at marketing, so if you’re asked to come up with a marketing plan you may have several related thoughts:
- *I’m not skilled at marketing… *
- *And even if I put in effort, I’m not sure I’ll be any good at it… *
- And GenAI is probably going to be better at coming up with a marketing plan, so I won’t expend as much effort and will just let GenAI handle this task for me.
But it turns out the inverse is true as well.
- Higher self-confidence in one's own ability to do the task correlated with more perceived critical thinking effort, especially in evaluating the AI's output. Translation: "I know this area so well that I’ll do a better job than AI could. Even if I use AI, I'd scrutinize it closely, which is probably the same amount of effort of just doing it myself."
Our finance expert, asked to generate a simple accounting spreadsheet, may just build the spreadsheet themselves, because they’ve done this hundreds of times.
This insight from the Microsoft researchers completely resonates with me personally; I almost never use Generative AI tools when it comes to emergency medicine (arguably my most seasoned and advanced skillset), but I use it all the time for numerous other areas because
- I think it’s probably better at all those other areas than I am
- I’d rather not expend the cognitive effort to, say, generate a business plan, or plan a trip to Italy, or edit this article you’re reading.
And while this is probably fairly intuitive (if you can have a “free” “expert” help you with a task, of course you’re going to say yes), it’s an insightful point particularly because it depends on one’s own self-assessment and self-confidence.
What this means for medicine
This insight alone has tremendous — and I mean tremendous – implications for medicine. Would a physician or nurse trust a generative AI outside of their specialty over their own knowledge? Should they? While those of us in emergency medicine often know a little bit about a lot of specialties, what about the pediatric endocrinologist with questions about his adult prostate? They likely have little confidence in their knowledge in urology.
Even more importantly, I worry that they will commit themselves to simply less cognitive effort as well. Get an answer about prostatitis from ChatGPT and instead of investigating the condition further or taking the steps to find a review article, download it, read it, and analyze it, does the pediatrician just accept ChatGPT at face value, regardless of its accuracy?
My other major concern here — and a recurring theme in this article — is around medical education and training. How can a medical student, pharmacy resident, or student in any discipline or training program actually become an expert without going through the work of becoming an expert? Invariably this requires work and sacrifice and failure. And if they never develop that expertise, then they never develop their own self-confidence, which we just established above is key to trusting themselves instead of relying on GenAI.
If they simply trust the AI implicitly, they might review the AI's output less rigorously. This saves time, but potentially embeds errors or misses nuances.
It’s important to emphasize: this effect isn’t specific to AI. It’s a reflection of how human cognition works. We naturally offload effort when we believe a tool is better or more efficient than we are. That’s not new — we’ve done this with GPS, calculators, even with EHRs.
What is new is the breadth and depth of offloading that GenAI enables. We’re not just delegating one task like multiplication or driving directions — we’re delegating many. Writing notes. Interpreting results. Summarizing studies. Generating differentials. Making treatment plans. That’s what makes this shift so consequential. GenAI isn’t just a tool. It’s a cognitive collaborator — and maybe, over time, a cognitive replacement if we’re not careful.
The AI 'Crutch' and Harm to Learning

Bastani et. al from Penn
Speaking of learning, now let’s move to Turkey, where researchers conducted a large randomized trial in high schools, evaluating GPT-4 based math tutors to help students with learning. In the trial they offered two AI versions and a control group:
- "GPT Base": The standard ChatGPT interface and interactions. Big text input box and a submit button.
- "GPT Tutor": This was the base layer but with added safeguards, providing incremental hints and avoiding direct answers, based on teacher input and known common mistakes. “GPT Tutor” offered guidance, hints, and help, but will avoid answering the question directly for the student.
- Control: No AI help for math education, just standard teaching.
Now here’s the wrinkle: they didn’t just look at the students’ performance before and after these GPT models were introduced. In phase 2 of the trial, they took the models away. Want to guess the results?
- With the AI, students performed significantly better than the control group (GPT Base: 48% improvement; GPT Tutor: 127% improvement). AI clearly helped during the task.
- However, when access was removed and students later took an unassisted exam, those who had used the standard "GPT Base" performed significantly worse (a 17% reduction) than the control group who never had AI access at all.
- And the GPT tutor group? The negative learning effect was largely mitigated by the safeguards in the "GPT Tutor" group – they performed similarly to the control group on the unassisted exam.

Students with GPT Base did worse on the exam compared to practice.
The conclusion? The standard AI interface was actually worse than a crutch or training wheels, since people did worse once they were taken away! As we saw in the Microsoft study, students using the “GPT Base” option probably relied on it to get answers without engaging deeply with the underlying concepts. When the AI was removed, their independent ability had diminished. (Because either they never learned the concepts well in the first place, relying on ChatGPT entirely, or learned the concepts on a more superficial level, without applying the concepts themselves.)
What this means for medicine
This is perhaps the most concerning finding for medical education. Imagine medical students or residents using AI tools to help interpret ECGs, radiology reports, or lab results. Or using it to generate differential diagnoses. If the AI simply provides the "answer" (or a highly plausible one), are trainees truly learning the foundational principles? Are they developing the pattern recognition and reasoning skills needed to function independently? I cannot even begin to tell you how much of medicine is based on understanding the fundamental principles of how the body works and how and why it fails — and then offering tests or treatments that "make sense" even if there’s not evidence for them.
Are we potentially creating the first generation of physicians who develop imposter syndrome not from comparing themselves to senior physicians, but from comparing themselves to ChatGPT? Imagine spending a decade in medical training only to feel intellectually inferior to an algorithm that didn't even exist when you started med school. And calling back to the Microsoft study, if you continue to feel inferior, you’ll never achieve mastery if AI is available.
This study from Turkey suggests that without careful design and integration – without safeguards – AI tools could inadvertently harm the development of core clinical competencies, even while appearing to boost performance in the short term.
But fear not.
This automation bias has been known in medicine for a long time, and I think we can continue to find ways to fight against it — like using systems like the GPT Tutor scenario.
One example: for decades now, EKG machines have printed out their own computerized interpretation of the EKG at the top of the paper. It includes heart rate, rhythm abnormalities, and signs of heart ischemia or heart attack. Knowing that it’s human nature to look at the computer interpretation at the top, what would my attendings always do before handing me an EKG? Fold the top down.
This forces the learner to first do their own interpretation of the EKG before peaking at what the computer thinks. (And at least for now, it’s a well-known fact that the computer analysis is often incorrect, which gives even more motivation for the trainee to learn to read the EKG independently as well.)
Image from Xitter
Wrapping Up Part 1: The Thinking Trap
So far, we’ve seen that AI doesn’t just change what we do — it changes how we think and how we learn. It offloads cognitive effort, may reshape our self-confidence, and, in many cases, prevents deep (human) learning altogether. Not because it’s broken or bad — but because it’s convenient and helpful. Because it works just well enough to stop us from doing the hard thinking ourselves.
Short-term gain. Long-term brain drain.
But what happens when we take these effects out of the classroom, and into the real world of clinical care? When the crutch becomes part of daily practice, what happens when it's suddenly taken away?
That’s what we’ll explore in Part 2, where we dive into the evidence of real-world de-skilling in clinical medicine, the risks of AI eroding clinician confidence, and what we can do to design tools and training that enhance human expertise instead of replacing it.
Look. This stuff is here — and it's only going to be coming on louder and stronger over the next few years. Maybe "it's fine" and this is "progress." It's not as if I feel guilt or remorse when I reach for a calculator. But I think medicine should actively choose this path, giving it careful thought — not just slowly allowing it to creep into our lives like some passive diffusion gradient. This is not a path that we can easily retrace our steps on if we're wrong.
Thank you to the authors from the Microsoft study Hank Lee Advait Sarkar Lev Tankelevitch Ian Drosos Sean Rintel Richard Banks Nicholas Wilson and the Turkey study Hamsa Bastani Osbert Bastani Alp Süngü Haosen Ge, Ph.D. Özge Kabakcı and Rei Mariman!
Update: Part 2 is here.