Independent benchmark analysis / 2026 journal article; expanded inventory from v5 appendix

AgentClinic

The simulated patient is part of the test instrument.

AgentClinic evaluates diagnostic dialogue through interacting doctor, patient, measurement and moderator agents. Its datasets mix transformed exam cases, clinical records and image-based case challenges. We examine what changes when a model must request information rather than receive a complete vignette. The simulation makes information gathering inspectable, but it also introduces another source of variation: the model playing the patient. Our analysis separates the candidate doctor model, the simulated environment and the diagnosis judge. The historical results below describe this particular combination, rather than a general claim about autonomous clinical competence.

01 / What is being tested?

The task, before the score.

input
An initial encounter objective; further history and findings obtained through simulated dialogue.
output
A final diagnosis judged against the case diagnosis.
unit
One simulated diagnostic encounter.
setting
Published comparison: GPT-4 patient/measurement agents and at most 20 interactions.

Data origin. MedQA and MedMCQA exam-derived scenarios, MIMIC-IV record-derived scenarios, NEJM case challenges and translated scenarios; conversations themselves are simulated. [4][5][6]

MedQA cases
215

Expanded scenario inventory.

Dataset summary; release notes [5][6]
MIMIC-IV cases
200

Selected single-diagnosis cases in the paper.

Appendix H.1 [5]
NEJM cases
120

Image and dialogue case challenges.

Results; dataset summary [4][5]
Interaction cap
20

Published principal doctor-model comparison.

Comparison of models [4]
Agent roles
4

Doctor, patient, measurement and moderator.

Figure 1 [4]
  1. 01

    Distribute case information

    Different case fields are exposed to different simulated roles.

    [4]
  2. 02

    Gather evidence

    The doctor asks the patient questions or requests measurements.

    [4]
  3. 03

    Commit to a diagnosis

    The doctor concludes before the interaction budget is exhausted.

    [4]
  4. 04

    Judge the conclusion

    The moderator compares the conclusion with the reference diagnosis.

    [4][6]

Dataset anatomy

Three core scenario collections

MedQA-derived

Exam-derived text scenarios.

215 scenarios[5]
MIMIC-IV-derived

Selected clinical record scenarios.

200 scenarios[5]
NEJM-derived

Multimodal case challenges.

120 scenarios[5]

Selected core collections, not the entire expanded inventory; specialty and language derivatives are omitted to avoid double-counting. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Diagnostic accuracy

Higher is better under a fixed simulation.

The moderator checks the final diagnosis against the scenario diagnosis.

Scoring definition

Accuracy = encounters judged correct / evaluated encounters × 100

Patient and measurement model, moderator, tool access, case subset and interaction budget must match. [4][6]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Historical AgentClinic-MedQA comparison

2026 journal article, Figure 2; GPT-4 simulated patient and measurement agents, 20-interaction cap.

Diagnostic accuracy · %
050100
Reported
Claude-3.5Paper reports ±3.3 alongside the mean.
62.1%
GPT-4Paper reports ±3.3 alongside the mean.
51.6%
GPT-4oPaper reports ±3.4 alongside the mean.
34.2%

Published measurements only. The ± quantities are reproduced as reported, not relabeled as confidence intervals.

Source: Comparison of models; Figure 2 [4]

04 / Our original analysis

What follows from the design?

01

The environment changes the score

Published evidence

The study varies the patient model and observes different diagnostic accuracy. [4]

Our interpretation

An agent result is a property of both the tested policy and its simulation. Record the patient model alongside the doctor model.

02

Real source records do not make dialogue real

Published evidence

MIMIC-IV records seed cases while language models generate interactions. [4][5]

Our interpretation

This supports studying simulated information gathering, not claims that actual patients would provide equivalent answers.

03

Translations are related cases

Published evidence

AgentClinic-Lang translates MedQA-derived scenarios. [5]

Our interpretation

Language coverage expands evaluation conditions without automatically adding independent clinical problems.

04

Action and conversation are complementary

Published evidence

AgentClinic judges diagnosis after dialogue; MedAgentBench grades record tasks and proposed payloads. [4][1]

Our interpretation

A combined evaluation should preserve both endpoints rather than average their percentages.

05 / Scope of the evidence

Where this benchmark stops.

Simulator and judge dependence

Generated patient responses and moderator judgments can affect measured performance. [4][6]

Source and case selection

Exam transformations and selected single-diagnosis records limit coverage of ambiguity and multimorbidity. [5]

Historical models

The publication date does not make these older model configurations a current frontier leaderboard. [4]

Known source inconsistencies

The website says 24 biases while the final paper says 23; this site avoids a bias-count headline. [4]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Public implementation; source-specific data access
License
Code MIT; underlying record and case-source terms remain separate.
Conditions
MIMIC-IV-derived data requires approval. Third-party case/image reuse rights are not established by the code license.
[6]

Evidence trail

Read the originals.

  1. AgentClinic: a multimodal benchmark for tool-using clinical AI agents ↗

    Schmidgall et al.; npj Digital Medicine. Final journal article, publishing historical model experiments.

  2. AgentClinic expanded preprint and appendices ↗

    Schmidgall et al.. Versioned source for dataset inventory and statistical appendix.

  3. AgentClinic official code and release notes ↗

    Samuel Schmidgall and collaborators. MIT-licensed code; MIMIC-IV cases require source-data approval.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗