The environment changes the score
The study varies the patient model and observes different diagnostic accuracy. [4]
An agent result is a property of both the tested policy and its simulation. Record the patient model alongside the doctor model.
Independent benchmark analysis / 2026 journal article; expanded inventory from v5 appendix
AgentClinic evaluates diagnostic dialogue through interacting doctor, patient, measurement and moderator agents. Its datasets mix transformed exam cases, clinical records and image-based case challenges. We examine what changes when a model must request information rather than receive a complete vignette. The simulation makes information gathering inspectable, but it also introduces another source of variation: the model playing the patient. Our analysis separates the candidate doctor model, the simulated environment and the diagnosis judge. The historical results below describe this particular combination, rather than a general claim about autonomous clinical competence.
01 / What is being tested?
Data origin. MedQA and MedMCQA exam-derived scenarios, MIMIC-IV record-derived scenarios, NEJM case challenges and translated scenarios; conversations themselves are simulated. [4][5][6]
Selected single-diagnosis cases in the paper.
Appendix H.1 [5]Published principal doctor-model comparison.
Comparison of models [4]Doctor, patient, measurement and moderator.
Figure 1 [4]Different case fields are exposed to different simulated roles.
[4]The doctor asks the patient questions or requests measurements.
[4]The doctor concludes before the interaction budget is exhausted.
[4]The moderator compares the conclusion with the reference diagnosis.
[4][6]Dataset anatomy
Exam-derived text scenarios.
Selected clinical record scenarios.
Multimodal case challenges.
Selected core collections, not the entire expanded inventory; specialty and language derivatives are omitted to avoid double-counting. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Higher is better under a fixed simulation.
03 / Measured evidence
Paper-reported results / selected rows
2026 journal article, Figure 2; GPT-4 simulated patient and measurement agents, 20-interaction cap.
Published measurements only. The ± quantities are reproduced as reported, not relabeled as confidence intervals.
Source: Comparison of models; Figure 2 [4]
04 / Our original analysis
The study varies the patient model and observes different diagnostic accuracy. [4]
An agent result is a property of both the tested policy and its simulation. Record the patient model alongside the doctor model.
AgentClinic-Lang translates MedQA-derived scenarios. [5]
Language coverage expands evaluation conditions without automatically adding independent clinical problems.
05 / Scope of the evidence
Generated patient responses and moderator judgments can affect measured performance. [4][6]
Exam transformations and selected single-diagnosis records limit coverage of ambiguity and multimorbidity. [5]
The publication date does not make these older model configurations a current frontier leaderboard. [4]
The website says 24 biases while the final paper says 23; this site avoids a bias-count headline. [4]
Evidence trail
Schmidgall et al.; npj Digital Medicine. Final journal article, publishing historical model experiments.
Schmidgall et al.. Versioned source for dataset inventory and statistical appendix.
Samuel Schmidgall and collaborators. MIT-licensed code; MIMIC-IV cases require source-data approval.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.