Benchmark analysis / 4 min read

Why is the simulated patient part of an AgentClinic result?

Separate the tested doctor model, simulated patient, measurement agent and diagnosis moderator when interpreting AgentClinic.

The short answer

AgentClinic tests a doctor agent through conversation with a simulated patient and measurement agent. This introduces a valuable evaluation question: can the model obtain the information needed to reach a diagnosis? It also introduces a measurement question: what information does the simulation actually provide? The final journal publication reports historical model comparisons. Our interpretation keeps the candidate model and its surrounding simulation together, because changing either can change the meaning of the observed diagnostic accuracy.

Name all roles in the experiment

The doctor is the candidate being evaluated. The patient and measurement agents expose information about the case, while the moderator evaluates the final diagnosis. The benchmark distributes information among those roles instead of placing an entire case in one prompt. This makes information gathering observable, but it also means that the candidate’s input is partly produced during the experiment.

A reproducibility description should identify every participating model, its role and the relevant instructions. A label such as “GPT-4 accuracy” is incomplete if the patient, measurement provider or moderator differs between runs. These are components of the test instrument, not incidental implementation details.

Distinguish a source case from its simulated conversation

AgentClinic uses several case sources, including transformed exam questions, clinical records and multimodal case challenges. A source record can ground a scenario without making the generated dialogue an observation of an actual encounter. The language model still decides how to express information within its assigned role.

Our deduction is that realism has several layers: the underlying facts, the distribution of those facts, the dialogue behavior and the consequences of an action. Evidence about one layer should not be silently transferred to the others. A convincing patient response can still be more explicit, more cooperative or differently phrased than a real patient’s response.

Control what the patient reveals

The paper investigates different patient models and reports changes in diagnostic performance. That finding makes the patient configuration important when comparing candidate doctors. An apparent improvement can come from better questioning, from a more informative simulated patient, or from their interaction. The reported endpoint alone cannot resolve that attribution.

For a new experiment, our proposed analysis artifact is a reveal log. Record which relevant facts were available initially, which were requested and which appeared in responses. Reviewers can then inspect whether the candidate failed to ask, whether the simulator failed to answer, or whether an available fact was interpreted incorrectly. This is a proposed diagnostic aid, not a published AgentClinic metric.

An interaction budget defines the task

The main comparison limits the doctor’s patient and measurement interactions. A shorter budget changes the cost of an unnecessary question and the opportunity to recover from a misunderstanding. More turns also change the context that the candidate must manage. It is not enough to say that two experiments both used dialogue.

Keep budget, stop rules and tool availability next to accuracy. When comparing a tool-augmented agent with a plain agent, describe what additional evidence or memory was available. Otherwise the result can be misread as a pure comparison of model backbones when it is actually a comparison of differently equipped systems.

Use diagnosis as one endpoint with a visible boundary

A correct final diagnosis does not by itself show that the information-gathering path was efficient, that every intermediate statement was supported or that a proposed treatment was appropriate. Those questions can matter even when the moderator accepts the final answer. Conversely, an incorrect diagnosis deserves investigation before attributing the error to a single capability.

The official code exposes configuration choices that make this boundary inspectable. Use the paper for the reported experiment and the repository for the implementation being run; do not assume the current checkout exactly reproduces a historical row. A useful result label identifies the source collection, doctor model, simulator models, interaction cap and judgment method together.

References & further reading

These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.

  1. AgentClinic: a multimodal benchmark for tool-using clinical AI agents ↗Schmidgall et al.; npj Digital Medicine. Final journal article, publishing historical model experiments.
  2. AgentClinic expanded preprint and appendices ↗Schmidgall et al.. Versioned source for dataset inventory and statistical appendix.
  3. AgentClinic official code and release notes ↗Samuel Schmidgall and collaborators. MIT-licensed code; MIMIC-IV cases require source-data approval.

Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.

Continue reading.

All guides →