AI in Medicine 12 min read

Medical Education Does Not Need an AI Course. It Needs a New Clinical Curriculum.

Preparing physicians for agentic AI requires more than an elective. Medical education must redesign how students reason independently, delegate work, verify machine output, and retain clinical judgment.

Listen to this post

Medical Education Does Not Need an AI Course. It Needs a New Clinical Curriculum.

0:00 / 0:00
A physician teaches medical students about clinical AI while standing beside a healthcare robot
On this page18 sections

Michael Turken, MD, MPH, left a question beneath my essay, The Physician Above the Loop.

“Medical education is nowhere near ready for this transition. What do you think medical education will look like in ten years?”

My first answer was simple.

I may be in the minority, but I do not think this problem is solved by adding an artificial intelligence course to the existing curriculum.

Medical education itself will have to change.

Students must learn how to decompose clinical problems for agents, verify machine-generated work, preserve independent reasoning, and contextualize automated recommendations within the life of a patient.

The danger is not only that a future physician will use AI incorrectly.

The greater danger is a physician who remains technically “in the loop” but lacks the bedrock knowledge and system literacy required to supervise what the machine is doing.

AI competency cannot remain an elective.

It must become part of how physicians learn to think.

I. The AI Course Problem

Medical schools know that artificial intelligence matters.

The predictable institutional response will be to create a course.

Students will learn the definitions of machine learning, large language models, neural networks, bias, hallucination, and prompt engineering. They may discuss privacy, regulation, and ethics. They may complete a project or evaluate a chatbot.

That course may be useful.

It will not be sufficient.

Agentic AI will not enter medicine as an isolated subject. It will enter history-taking, documentation, diagnostic reasoning, imaging, medication reconciliation, evidence retrieval, patient communication, prior authorization, and longitudinal care.

A technology that enters every clinical workflow cannot be contained inside one course.

This is the AI course problem.

The institution acknowledges the technology without redesigning the educational system that the technology is changing.

Students learn about AI on Tuesday.

On Wednesday, they return to a curriculum built around the assumption that the physician personally performs every meaningful cognitive task.

The contradiction remains intact.

II. The Unit of Clinical Work Is Changing

Medical education has traditionally trained physicians to produce answers.

What is the diagnosis?

What is the differential?

What test should be ordered next?

What treatment is indicated?

What does the evidence show?

These questions remain essential.

But the future physician will increasingly encounter a different set of problems.

Which system should be assigned this task?

What information should the system receive?

What constraints should govern its work?

Which parts of its output require independent verification?

Why do two clinical agents disagree?

What evidence did the system retrieve?

What did it omit?

When should the entire machine-generated frame be rejected?

The unit of work is changing from answer production to supervised intelligence.

That does not reduce the need for medical knowledge.

It changes what medical knowledge must accomplish.

Knowledge will no longer be needed only to generate an answer.

It will also be needed to recognize when a generated answer is wrong.

III. Bedrock Knowledge Becomes More Important

There is an understandable argument that medical students should memorize less because machines can retrieve more.

That argument is only partially correct.

Students do not need to memorize every provisional guideline, drug interaction, or classification system. Information with a short half-life belongs in a reliable retrieval system.

But medicine still has bedrock knowledge.

Anatomy.

Physiology.

Pathophysiology.

Pharmacology.

Probability.

The natural history of disease.

The relationship between a clinical finding and the mechanisms that can produce it.

This knowledge cannot be outsourced completely because it forms the internal error-detection system of the physician.

A clinician who understands fetal physiology can recognize when a recommendation about surveillance or delivery does not fit the clinical picture.

A clinician who understands renal physiology can recognize when a medication plan ignores the consequences of declining filtration.

A clinician who understands pretest probability can recognize when a confident diagnostic recommendation rests on weak evidence.

The machine can retrieve a fact.

Bedrock knowledge allows the physician to determine whether the fact belongs in this decision.

Medical education should therefore distinguish between two types of knowledge:

Knowledge that must remain inside the physician.

Knowledge that can be retrieved safely at the point of care.

We have not defined that boundary carefully enough.

Agentic medicine will force us to.

IV. Medical Education Needs Two Reasoning Tracks

Future physicians should learn in two distinct modes.

I call this the dual-track curriculum.

Closed-Agent Reasoning

The student reasons without machine assistance.

She takes the history.

She performs the examination.

She constructs the problem representation.

She creates the differential diagnosis.

She decides what information is missing.

She explains the mechanism behind her conclusions.

The purpose is not nostalgia.

It is preservation.

Closed-agent reasoning protects the clinical cognition required to recognize automation failure.

Open-Agent Reasoning

The student receives access to one or more clinical agents.

She must decide which tasks to delegate, how to structure the request, and how to evaluate the result.

The agent may summarize the chart, generate a differential, retrieve evidence, identify drug interactions, draft documentation, or challenge the student’s plan.

The student is then evaluated on what she does with the output.

Did she accept it too quickly?

Did she identify unsupported claims?

Did she inspect the source?

Did she notice missing information?

Did she recognize that the recommendation was generally correct but wrong for this patient?

Both tracks are necessary.

Closed-agent reasoning builds the physician.

Open-agent reasoning teaches the physician how to supervise the machine.

V. The Clinical Vignette Must Become Messier

Traditional medical questions are unusually cooperative.

The relevant history is present.

The laboratory values are already available.

The important findings have been separated from the noise.

The question has a defined endpoint.

Real patients do not arrive this way.

Agentic medical education should expose students to incomplete, fragmented, and contradictory information.

The chart should contain duplicated medications.

The outside hospitalization should be poorly summarized.

The family history should conflict with an earlier note.

The patient should introduce a new concern halfway through the encounter.

The AI-generated summary should omit one clinically important detail.

The evidence agent should retrieve a guideline that is authoritative but outdated.

The documentation agent should produce a polished note that subtly changes the physician’s reasoning.

These are not artificial complications.

They are the clinical environment.

The student should be evaluated on whether she can discover the problem before attempting to solve it.

Information acquisition is clinical reasoning.

Medical education should treat it that way.

VI. Students Must Learn the Five Duties Above the Loop

In the original essay, I described five duties for physicians supervising agentic systems.

Those duties can become the longitudinal structure of an AI-integrated medical curriculum.

1. Delegate

Students should learn how to decompose a clinical problem into bounded tasks.

They should know which tasks can be delegated safely, which require escalation, and which should remain human.

“Review this patient” is not an adequate assignment.

“Reconcile these medication lists, identify discrepancies, show the source of each conclusion, and escalate any medication with uncertain status” is a supervised clinical task.

Delegation requires precision.

2. Verify

Students should demand provenance.

Where did the recommendation come from?

How current is the evidence?

Which patient facts were used?

Which facts were ignored?

What part of the answer is retrieved evidence, and what part is model inference?

Verification should not be an afterthought.

It should be part of the workflow.

3. Preserve

Students need protected opportunities to reason without AI assistance.

This should not be framed as punishment or resistance to technology.

It is cognitive maintenance.

A physician who has never developed independent clinical reasoning cannot supervise machine reasoning.

Preservation is therefore a safety requirement.

4. Contextualize

The technically correct recommendation is not always the clinically correct recommendation.

Students must learn to reconcile machine output with patient preferences, comorbidities, affordability, family circumstances, culture, risk tolerance, and uncertainty.

An agent can identify what is usually recommended.

The physician must decide what should be done for this patient.

5. Judge

The student must make the final decision.

Accept the output.

Reject it.

Modify it.

Escalate it.

Then explain why.

Judgment is not the final click after the machine has completed the work.

Judgment is the clinical act that determines whether the work should enter the patient’s care.

VII. Assessment Must Change

A student who uses AI to generate a correct answer has not necessarily demonstrated competence.

A student who rejects an incorrect AI recommendation may have demonstrated something more important.

Medical assessment must learn to distinguish the two.

Future examinations should include agent-assisted clinical encounters in which the system is intentionally imperfect.

The agent should occasionally:

  • retrieve the wrong guideline,
  • omit a contraindication,
  • anchor on an early diagnosis,
  • misattribute a statement,
  • overstate the certainty of weak evidence,
  • or recommend a plan that ignores the patient’s stated goals.

The student should not be told where the error is.

Finding it is the examination.

We should assess more than whether the final answer is correct.

We should assess:

  • how the student framed the problem,
  • what work was delegated,
  • which sources were inspected,
  • how uncertainty was handled,
  • whether independent reasoning was preserved,
  • and why the final recommendation was accepted or rejected.

The audit trail should become part of the grade.

In agentic medicine, the path to the answer is part of the answer.

VIII. Faculty Development Is the Immediate Bottleneck

Students cannot be trained to supervise systems their instructors do not understand.

This is the faculty problem.

A medical school can purchase an AI platform quickly.

It cannot create experienced faculty supervision as quickly.

Clinical educators will need practical familiarity with agent behavior, retrieval systems, model limitations, automation bias, data provenance, workflow design, and human checkpoints.

They do not all need to become software engineers.

They do need to understand enough of the system to teach its safe use.

Otherwise, faculty will be placed in one of two weak positions.

Some will prohibit tools they cannot evaluate.

Others will permit tools they cannot supervise.

Neither position creates competence.

Medical schools should therefore build faculty AI laboratories before they build student AI requirements.

Faculty need protected environments where they can test systems against real clinical tasks, inspect failures, compare workflows, and establish appropriate boundaries.

The curriculum should be designed from observed failure modes.

Not vendor demonstrations.

IX. Professionalism Must Include Machine Supervision

Medical professionalism has traditionally addressed honesty, confidentiality, informed consent, conflicts of interest, accountability, and the physician-patient relationship.

Agentic care adds another layer.

What did the physician delegate?

Was the patient informed?

Which system handled the information?

Could the physician inspect what the system did?

Who was responsible for monitoring performance?

What happened when the system encountered uncertainty?

Who owned the final decision?

These are not only technical questions.

They are professional questions.

A physician should not be able to avoid responsibility by saying that the recommendation came from the system.

The use of an agent changes how the work was performed.

It does not remove accountability for the work.

Students should learn this before the clinical workflow makes the distinction feel inconvenient.

X. What the Curriculum Should Look Like

AI should not be confined to a semester.

It should appear longitudinally.

During the preclinical years, students should learn bedrock mechanisms and identify where model-generated explanations become oversimplified or incorrect.

During clinical-skills training, they should compare unaided histories with agent-generated summaries and identify lost context.

During clerkships, they should use agents for bounded tasks while documenting delegation, provenance, uncertainty, and verification.

During evidence-based medicine, they should distinguish retrieved evidence from generated synthesis.

During ethics and professionalism, they should address disclosure, accountability, privacy, bias, and the limits of automated authority.

During residency, they should supervise multi-agent workflows, manage exceptions, evaluate performance over time, and participate in system design.

The goal is not to produce physicians who know more terminology about artificial intelligence.

The goal is to produce physicians who can remain clinically authoritative while working inside systems that generate more intelligence than any individual can produce alone.

XI. The Physician We Are Training

Medical education has always been built around an image of the physician.

The physician remembers.

The physician retrieves.

The physician synthesizes.

The physician documents.

The physician decides.

That image is changing.

The future physician will still examine, reason, communicate, and decide. But she will also delegate, supervise, audit, and manage machine-generated work.

She will need enough medicine to recognize when the system is wrong.

Enough technical literacy to understand why it may be wrong.

Enough humility to accept useful machine input.

Enough independence to reject it when necessary.

And enough judgment to know the difference.

That physician will not be produced by adding one AI course to a curriculum designed for another era.

Medical education must train the physician above the loop.

Preserve the knowledge. Teach the supervision. Keep responsibility attached to the decision.


This essay extends the argument developed in The Physician Above the Loop: Preparing Medicine for Agentic AI.

Share this article

Share X / Twitter Bluesky LinkedIn

Related articles

A physician supervises a clinical AI robot during a consultation with a pregnant patient
AI in Medicine Featured

The Physician Above the Loop: Preparing Medicine for Agentic AI

Clinical intelligence is becoming abundant. Physicians must learn to delegate, verify, contextualize, and judge the work of increasingly autonomous clinical AI systems.

· 14 min read
agentic-aiclinical-aiautonomous-ai
Layered, glowing knowledge graph representing a physician's longitudinal memory of a patient across years of clinical data
AI in Medicine Featured

AI's Next Breakthrough Is Not a Bigger Brain. It Is a Memory You Can Trust.

The AI race has been about model size for years. The real bottleneck is memory: trustworthy, auditable knowledge that persists across years, not conversations. Medicine will feel this shift first.

· 10 min read
ai-agentsagentic-aiknowledge-systems
Clinical AI interface contrasting a polished answer with evidence layers, uncertainty warnings, and a physician review gate.
AI in Medicine

Fluent Answers Are Not Clinical Judgment

Language models can make uncertain medical information sound finished. The problem is not fluency. The problem is mistaking fluency for accountable clinical reasoning.

· 8 min read
clinical-ailarge-language-modelsclinical-judgment
Chukwuma Onyeije, MD, FACOG

Chukwuma Onyeije, MD, FACOG

Maternal-Fetal Medicine Specialist

MFM specialist at Atlanta Perinatal Associates. Founder of CodeCraftMD and OpenMFM.org. I write about building physician-owned AI tools, clinical software, and the case for doctors who code.