Benchmark analysis / 4 min read

How should MedAgentBench and AgentClinic be compared?

Build a capability comparison between record actions and diagnostic dialogue without averaging non-equivalent benchmark scores.

The short answer

MedAgentBench and AgentClinic answer complementary questions about medical agents. One centers on interacting with a virtual record system; the other centers on gathering diagnostic information through simulated dialogue. A useful comparison starts with their input and output contracts, rather than a combined percentage. This guide proposes a capability map that preserves the evidence each benchmark provides. The mapping is our analysis of the published protocols, not an experiment showing that one benchmark predicts performance on the other.

Map the request to its completion event

In a record task, completion may mean returning the right value or constructing an acceptable payload. In a diagnostic conversation, completion means reaching a diagnosis that the moderator accepts. The fact that both systems use tools does not make their endpoints interchangeable. One result can improve while the other remains unchanged.

Our proposed first row in a comparison is therefore a sentence: given this information, the agent must produce this output, and this grader decides whether it passed. Writing that sentence exposes hidden differences in tool permissions, available context and answer format before anyone sees a model ranking.

Separate access to evidence from use of evidence

A MedAgentBench agent retrieves records through a fixed interface. An AgentClinic doctor elicits information from simulated roles. Both require acquiring information, but the failure opportunities differ. An API request can be malformed; a conversational question can be ambiguous; either can produce an answer that the model then misinterprets.

Use these distinctions to organize error review. Label a failure by where the evidence path broke: request formation, evidence delivery, interpretation or final output. This taxonomy is an analytical proposal. It should not be presented as though the original papers measured identical categories or provided directly comparable error rates.

Keep the grader in the comparison

MedAgentBench uses reference solutions and authored checks. AgentClinic uses a moderator to evaluate a diagnosis. Both implement a definition of success, but their judgments have different sensitivities. Strict formatting can dominate one result; diagnostic equivalence and moderator behavior can influence the other.

A comparison table should therefore show the evaluator rather than hide it in a methods footnote. If an engineering team replaces a parser or a research team changes a moderator prompt, preserve the earlier configuration as a separate experiment. The benchmark title alone cannot tell a reader whether a score difference reflects changed competence or changed measurement.

Choose a local question before combining evidence

Consider two illustrative projects: a record-navigation assistant and a diagnostic interview assistant. The first has a direct reason to inspect query and proposed-action tasks. The second has a direct reason to inspect information elicitation and diagnosis. Both might benefit from additional tests, but their most relevant evidence differs.

Our recommendation is a task-level evidence matrix, with one column for what a benchmark measures and another for what the project still needs to establish. Avoid adding the benchmark percentages together. A numerical combination would require justified weights and comparable endpoints, neither of which follows simply from both benchmarks being healthcare-related.

Pin versions before discussing progress

MedAgentBench’s early 100-task preprint and later 300-task paper are distinct snapshots. AgentClinic’s expanded datasets and later journal publication also require careful labeling. A newer publication date does not imply that every displayed model was evaluated recently, and a current repository can evolve after the reported experiments.

Preserve paper date, dataset inventory, code revision and complete model configuration. Then describe progress in the narrowest supportable terms: improvement on the same endpoint under matched conditions. The resulting comparison is more useful than a universal winner because it tells the next evaluator exactly which capability has evidence and which question remains open.

References & further reading

These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.

  1. MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents ↗Jiang, Black, Geng et al.; Stanford University. Primary 300-task paper; historical results, not a current leaderboard.
  2. MedAgentBench initial 100-task preprint ↗Jiang, Black, Geng et al.; Stanford University. Version history only; its scores must not be mixed with the 300-task evaluation.
  3. MedAgentBench official implementation ↗Stanford Machine Learning Group. Repository MIT license applies to code; repository access is not a separate patient-data license.
  4. AgentClinic: a multimodal benchmark for tool-using clinical AI agents ↗Schmidgall et al.; npj Digital Medicine. Final journal article, publishing historical model experiments.
  5. AgentClinic official code and release notes ↗Samuel Schmidgall and collaborators. MIT-licensed code; MIMIC-IV cases require source-data approval.

Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.

Continue reading.

All guides →