Google's AMIE and the World That Guidelines Imagine

There is quiet, unglamorous work in medicine that will never make it on The Pitt. It saves way more lives. It is the work of often primary care but specialists do it too: keeping patients from falling through the cracks. Diabetes. Heart failure. COPD. Hypertension. Chronic kidney disease. Depression. Anticoagulation. Cancer surveillance. Diet screening. End of life counseling. Dose adjustments. Motivational interviewing. Monitoring labs. Disability paperwork. Medication refills. Specialist follow-up. All of these things are constantly, actively, trying to kill people.

When people say "My doctor doesn't do anything," I usually stare at them blankly, because that list and "doesn't do anything" does not comport in my brain.

Multiply that list by thousands of patients per doctor, and yes — it's really hard to score 100% on that checklist. This is where healthcare leaks value, and where the value has already been agreed upon. Not "I wonder if this new drug would help" but "We know this medication works, we just need to make sure the patient gets the prescription, can fill it, and remembers to take it." A never-ending spinning plate.

Google DeepMind's AMIE paper takes aim at exactly this. Unlike MIRA (Part 1), which tested whether an agent could navigate a simulated ER, AMIE asks whether conversational AI can help manage disease over time — retrieving guidelines, reasoning across visits, producing structured management plans. Where MIRA was ordering CT scans and admitting patients, AMIE is building care plans it can manage longitudinally.

If Part 1 was about what AI can't perceive — the messiness of real patients, the chaos of real EDs — Part 2 is about what AI can't do: deliver the care it recommends.

What AMIE Did in this Paper

AMIE was tested in a randomized, blinded virtual OSCE — a structured clinical exam where both AMIE and primary care physicians managed simulated patients with chronic conditions across multiple visits. The cases covered common longitudinal management scenarios: diabetes, hypertension, heart failure, and similar bread-and-butter primary care.

The headline result: AMIE was non-inferior to PCPs on overall management reasoning, and outperformed them on precision of investigations and alignment with clinical guidelines. Blinded physician raters evaluated both sets of management plans without knowing which came from a human and which from an AI.

Article content

The red and blue hexagons are familiar in AMIE papers: AI does a pretty damn good job

Two things worth noting for context. First, the simulated patients were cooperative, articulate, and consistent — the same clean-text problem we saw in MIRA. Second, and more importantly for this essay: AMIE operated without the friction of reality. No formulary restrictions. No prior authorization. No patient who stopped the medication six months ago because of side effects and never told anyone. No exhausted caregiver. No social barriers. No competing priorities.

AMIE made its care plans in the world guidelines imagine. The results were impressive. That world just doesn't exist.

(The other thing I’ve said is that Google’s AMIE needs to get deployed in the real world, because the model is in for a rude awakening.)

The Promise Is Real

Guideline-directed care is hard — not because doctors don't know their guidelines, but because real care is fragmented, time-limited, and constrained by patients' ability to pay and life circumstances.

A patient with diabetes not on a medication they probably should be on. A heart failure patient missing one pillar of evidence-based therapy. A high-risk medication with overdue monitoring labs. These are real failures — not always malpractice, but real. And we hold doctors responsible for them as quality indicators while giving them almost no power to actually fix them.

The best version of AMIE isn't a robot doctor issuing orders from the cloud. It's a clinical safety layer — a do-not-miss system, an exception-checker, a second set of eyes that never gets tired of asking: did the medication get started? Did the dose get titrated? Did kidney function get rechecked? Did the patient fall through the cracks?

Medicine needs that.

Guidelines Don't Apply a Lot of the Time

Here's something that doesn't get said enough: guidelines are broad strokes, not rules. They're population-level tools derived from trials that enrolled specific patients under controlled conditions. The person in front of you is almost never that patient.

Just look at a guideline summary like the one from the ACG on the management of acute pancreatitis: any doctor in their sleep could make up a patient scenario that would invalidate any of these guidelines. It's not hard for us, because that's what we do all day long: decide if this guideline should apply or not apply based on the real human being we're dealing with.

Real patients are multimorbid, preference-sensitive, financially constrained, and socially complicated in ways no guideline anticipates. You can't write a guideline that perfectly applies to every patient — it would either be so broad as to be useless, or so exhaustive it would be a thousand pages long and still miss someone. So guidelines deliberately leave room for judgment. That's not a bug. That's the point.

The problem is that AI doesn't always know that. A system reasoning from guidelines may treat them as rules — and flag every deviation as a failure, when in reality the deviation 𝘪𝘴 the clinical judgment.

The Care Gap Trap

If doctors don't roll their eyes at guidelines, they still feel frustration — because there are two different problems that look identical from the outside.

  • The first: a ball got dropped. A patient should be on a medication, it got missed, nobody followed up. That's a real care gap.
  • The second: we already tried the guideline-based option, and it didn't work out.

The patient had a side effect. The patient refused. The patient couldn't afford it. Insurance wouldn't cover it. The patient doesn't believe it would help. The records proving it happened are somewhere else. Something about this patient makes the guideline inapplicable — they're on hospice, they're too frail, they have a competing diagnosis.

Some of those are true missed opportunities. Some are wise exceptions. Some are health-system failures.

To a guideline machine, they look identical.

The AI Guideline Machine

And this is where I worry about tools like AMIE — but not AMIE itself, just any generative tool like this: their ability to be abused by the powerful stakeholders in medicine. This concept actually comes from another article this week, from Dr. Isaac Kohane, Editor-in-Chief of NEJM AI. He writes this week:

The danger is not that AI will fail to produce guideline-like prose. The danger is that it will produce too much of it or plausible 'guidance' for situations in which no trustworthy guidance exists. And because the output will look polished, be mapped to citations, and be expressed in the familiar syntax of evidence-based medicine, it will be easy for health systems, payers, regulators, vendors, and professional societies to mistake linguistic competence for clinical legitimacy.

I’ve written this before but I’ll say it again: I don’t know that we have a less-biased, less-conflicted arbiter of “right” and “wrong” or “good” and “bad” than a good physician who’s following their oath to their patient. Any other stakeholder here has an incentive to say that the care is too much or not enough. And I don’t care how wicked smart and well-trained your Generative AI tool is, it can and will still be influenced by context, bias, and potentially — secretly — who’s paying for it.

Kohane goes on to suggest that the “AI Guideline Machine” should be used for failures of omission. “Hey, you might have forgotten to restart the patient’s diuretic” for example might save billions of dollars in hospital admissions. Or it might be completely reasonable to not restart, either. And the best way to determine that is to trust and support the patient’s doctor. They're the just-right answer. Not the algorithm, not the payer. The doctor who knows the patient.

The Right Role

I'm more optimistic about AMIE than MIRA — not despite these concerns, but because of them.

MIRA is optimized for the EHR version of medicine: complete the workflow, enter the order, match the record. Its reality gap is perceptual — it can't see what the simulation doesn't show it.

AMIE points toward where healthcare actually fails patients: missed follow-up, incomplete medication optimization, unrecognized safety issues, failure to close the loop. That's the right target. And AI could be extraordinary at it — not because doctors are careless, but because the system is too complex for unaided human memory, built on brittle workflows that fail silently.

But only if we build it as decision support, not surveillance. A system that says have you considered this — not one that says provider failed guideline.

MIRA tests medicine as the EHR imagines it. AMIE tests medicine as guidelines imagine it. Neither tests medicine as patients live it.

That's the reality gap.

Machines are very good at finding gaps.

Wisdom is knowing which ones are real, and which ones actually matter.

Tagging paper authors — great work, as always: Valentin Liévin Anil Palepu Wei-Hung Weng Khaled Saab David Stutz Yong Cheng Kavita Kulkarni S. Sara Mahdavi Joëlle Barral Dale Webster Katherine Chou Avinatan Hassidim Yossi Matias James Manyika Ryutaro Tanno Vivek Natarajan Adam Rodman Tao Tu Alan Karthikesalingam Mike Schäkermann