Independent benchmark analysis / 300-task paper; arXiv v2

MedAgentBench

A successful read and a successful action are different claims.

MedAgentBench turns an instruction into a sequence of requests against a virtual FHIR record environment. Its useful distinction is between returning information and preparing an action. We analyze the 300-task paper snapshot, including the separate query and action results. Our central interpretation is that an aggregate score hides a decision about workload composition. The baseline also checks proposed write payloads without executing them against the server, which places a precise boundary around the word success. This is an independent reading of the authors’ evaluation, with no new model runs.

01 / What is being tested?

The task, before the score.

input
Clinician instruction, hospital-specific context and access to nine FHIR functions.
output
A final answer or a proposed POST payload.
unit
One patient-specific task attempt.
setting
Virtual EHR; pass@1; maximum eight interaction rounds.

Data origin. Physician-authored tasks grounded in deidentified Stanford STARR records; synthetic identifying details and patient-level date jittering. [1][2][3]

Tasks
300

150 query tasks and 150 action tasks.

Section 2.5.1 [1]
Patient profiles
100

Profiles are shared across tasks; tasks are not 300 independent patients.

Table 2 [1]
Record elements
785,207

Four reported FHIR resource groups.

Table 2 [1]
Task categories
10

Grouped into seven broad categories in Table 1.

Table 1 [1]
Interaction budget
8 rounds

Invalid actions or exceeding the budget count as failures.

Sections 2.4.1–2.4.3 [1]
  1. 01

    Resolve the task

    Read instruction and hospital context before choosing a FHIR function.

    [1]
  2. 02

    Retrieve

    GET requests return raw server responses to the agent.

    [1]
  3. 03

    Propose an action

    POST payloads receive local checks and a simulated success response; baseline writes are not executed.

    [1]
  4. 04

    Finish and grade

    Compare query answers with reference solutions; apply authored checks to action payloads.

    [1]

Dataset anatomy

The paper balances reads and proposed actions

Queries

GET-only task subgroup.

150 tasks[1]
Actions

Requires a proposed POST, sometimes after retrieval.

150 tasks[1]

Counts describe the 300-task paper; they are not deployment traffic weights. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.

02 / Measurement

Task success rate

Higher is better within the same task/configuration.

Rule-based success for one attempt; query answers and proposed action payloads have different checks.

Scoring definition

SR = successful task attempts / all task attempts × 100

Pin task version, orchestrator, parser, reference solution and interaction cap. This is not a patient-outcome rate. [1]

03 / Measured evidence

Results, with their conditions attached.

Paper-reported results / selected rows

Paper-reported query performance

arXiv v2, February 2025; 150 queries, pass@1, eight-round baseline; temperature zero for these selected models.

Query task success rate · %
050100
Reported
Claude 3.5 Sonnet v2Query subgroup; not action success.
85.33%
GPT-4oQuery subgroup; not action success.
72.00%
DeepSeek-V3Query subgroup; not action success.
70.67%
Gemini-1.5 ProQuery subgroup; not action success.
52.67%

Historical selected rows. Table 3 does not provide uncertainty intervals.

Source: Table 3 [1]

Paper-reported results / selected rows

Paper-reported proposed-action performance

Same paper snapshot and orchestrator; 150 action tasks, with proposed POST payload validation.

Action task success rate · %
050100
Reported
Claude 3.5 Sonnet v2Action subgroup; payload checks, not executed writes.
54.00%
GPT-4oAction subgroup; payload checks, not executed writes.
56.00%
DeepSeek-V3Action subgroup; payload checks, not executed writes.
54.67%
Gemini-1.5 ProAction subgroup; payload checks, not executed writes.
71.33%

Keep this separate from query scores. No deployment reliability or current model ranking is implied.

Source: Table 3 [1]

04 / Our original analysis

What follows from the design?

01

The workload determines the headline

Published evidence

Claude’s query/action scores are 85.33/54.00; Gemini-1.5 Pro’s are 52.67/71.33. [1]

Our interpretation

A read-heavy workflow and a write-heavy workflow can favor different systems. Equal weighting is a benchmark design choice, not a universal utility function.

02

Action grading stops before persistence

Published evidence

The baseline sends GETs to the server but validates POST payloads locally. [1]

Our interpretation

A passing payload leaves server validation, persisted state and downstream execution untested. Treat these as additional integration questions.

03

Count tasks and people separately

Published evidence

300 tasks use 100 patient profiles. [1]

Our interpretation

A task-level denominator describes attempts; it does not establish evidence from 300 distinct patients or institutions.

04

Formatting belongs to the system

Published evidence

Invalid actions are failures under the specified parser. [1]

Our interpretation

Model knowledge and orchestrator compatibility jointly affect this result. Changing a parser changes the evaluated system.

05 / Scope of the evidence

Where this benchmark stops.

One source institution

STARR profiles do not establish transportability to other EHR systems or populations. [1]

Bounded workflow scope

The task inventory does not cover all specialties, teamwork or the consequences of real actions. [1]

Version sensitivity

The January 2025 v1 evaluated 100 tasks; do not combine its scores with these 300-task rows. [2][1]

Data rights require separate checking

MIT repository licensing does not by itself establish unrestricted reuse of patient-derived records. [3]

06 / Working with the benchmark

Access & reuse.

Open the author’s resource ↗
Availability
Official code and virtual environment available
License
Repository code: MIT. Separate patient-data reuse terms not verified.
Conditions
Consult the official repository and data distribution terms before obtaining or reusing records. This site reproduces aggregate metadata only.
[3]

Evidence trail

Read the originals.

  1. MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents ↗

    Jiang, Black, Geng et al.; Stanford University. Primary 300-task paper; historical results, not a current leaderboard.

  2. MedAgentBench initial 100-task preprint ↗

    Jiang, Black, Geng et al.; Stanford University. Version history only; its scores must not be mixed with the 300-task evaluation.

  3. MedAgentBench official implementation ↗

    Stanford Machine Learning Group. Repository MIT license applies to code; repository access is not a separate patient-data license.

Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.

Explore the assumptions ↗