A successful read and a successful action are different claims.
MedAgentBench turns an instruction into a sequence of requests against a virtual FHIR record environment. Its useful distinction is between returning information and preparing an action. We analyze the 300-task paper snapshot, including the separate query and action results. Our central interpretation is that an aggregate score hides a decision about workload composition. The baseline also checks proposed write payloads without executing them against the server, which places a precise boundary around the word success. This is an independent reading of the authors’ evaluation, with no new model runs.
Clinician instruction, hospital-specific context and access to nine FHIR functions.
output
A final answer or a proposed POST payload.
unit
One patient-specific task attempt.
setting
Virtual EHR; pass@1; maximum eight interaction rounds.
Data origin. Physician-authored tasks grounded in deidentified Stanford STARR records; synthetic identifying details and patient-level date jittering. [1][2][3]
Counts describe the 300-task paper; they are not deployment traffic weights. Bar lengths use the largest listed count as their reference; they are not percentages of a shared population.
02 / Measurement
Task success rate
Higher is better within the same task/configuration.
Rule-based success for one attempt; query answers and proposed action payloads have different checks.
Scoring definition
SR = successful task attempts / all task attempts × 100
Pin task version, orchestrator, parser, reference solution and interaction cap. This is not a patient-outcome rate. [1]
03 / Measured evidence
Results, with their conditions attached.
Paper-reported results / selected rows
Paper-reported query performance
arXiv v2, February 2025; 150 queries, pass@1, eight-round baseline; temperature zero for these selected models.
Query task success rate · %
050100
Reported
Claude 3.5 Sonnet v2Query subgroup; not action success.
85.33%
GPT-4oQuery subgroup; not action success.
72.00%
DeepSeek-V3Query subgroup; not action success.
70.67%
Gemini-1.5 ProQuery subgroup; not action success.
52.67%
Historical selected rows. Table 3 does not provide uncertainty intervals.
Claude’s query/action scores are 85.33/54.00; Gemini-1.5 Pro’s are 52.67/71.33. [1]
Our interpretation
A read-heavy workflow and a write-heavy workflow can favor different systems. Equal weighting is a benchmark design choice, not a universal utility function.
02
Action grading stops before persistence
Published evidence
The baseline sends GETs to the server but validates POST payloads locally. [1]
Our interpretation
A passing payload leaves server validation, persisted state and downstream execution untested. Treat these as additional integration questions.
Stanford Machine Learning Group. Repository MIT license applies to code; repository access is not a separate patient-data license.
Benchmark authors retain authorship of their work. This publication provides independent analysis; published rows are not Arcophos evaluation runs. Editorial method.