The short answer
MedAgentBench measures whether an agent completes a specified record task under a particular interaction protocol. The important first step is to identify the kind of completion being tested. Returning a requested value and preparing a valid order payload are different endpoints, even when both appear in one overall success rate. This analysis uses the February 2025, 300-task paper. Its results are useful evidence about that virtual environment and orchestrator, with a precise boundary around what the action grader verifies.
Start with the endpoint, then read the percentage
The paper divides its tasks evenly between queries and actions. Query grading compares an answer with a reference solution. Action grading checks a proposed POST payload against manually written rules. Our interpretation is that these should remain visible as separate capabilities when selecting an evaluation. A record assistant used to look up results has a different relevant workload from an assistant preparing orders, even if the same language model sits behind both interfaces.
An overall benchmark result therefore encodes a workload assumption. It is helpful for comparing systems under the paper’s fixed mix, but it does not automatically express the value of either system in a different workflow. Write the intended task mix beside the result before using the headline.
The published action stops at a payload boundary
The baseline performs GET requests against the FHIR environment but does not execute its proposed POST requests against the server. It checks their form, signals success to the agent and grades the saved payload. That design makes repeated evaluation easier, but the meaning of action success must follow the actual procedure.
Our deduction is that a valid proposed write and a successfully persisted write require different evidence. Server-side validation, conflicting updates, authorization, persistence and later downstream behavior would need separate tests in a real integration. This observation does not diminish the value of payload evaluation; it identifies the next boundary an engineering team must investigate.
The ordering changes when the endpoint changes
The selected published rows show Claude 3.5 Sonnet v2 leading the query subgroup while Gemini-1.5 Pro leads the action subgroup among those rows. The comparison is a concrete example of why a single leaderboard position can conceal an operational choice. The task categories, rather than a vendor’s general reputation, determine which result is relevant.
We do not convert those subgroup scores into an estimated deployment success rate. That would require the local distribution of tasks and evidence that subgroup performance transfers to it. A sensible use is to identify a candidate configuration and the kind of failure analysis needed next.
Treat parsing and budgets as evaluation components
An invalid action or an exhausted interaction budget is a failure under the paper’s procedure. A model may know the requested clinical fact and still fail to produce the exact request or final-answer form expected by the orchestrator. The published score is therefore evidence about the combined model and protocol.
If a team introduces a new tool schema, structured-output layer or recovery policy, it has changed the system. Keep those modifications in the experiment manifest. Compare the revised system on the same task version before attributing a gain to better reasoning. Otherwise the explanation may confuse a communication repair with an improvement in knowledge.
Carry a reproducibility note alongside the result
The official repository provides the implementation and setup route. Record the task inventory, code revision, reference solution, model snapshot and inference parameters used for a run. The January 2025 preprint evaluated a smaller inventory, so its percentages should not be appended to the later table as if the cases were identical.
A useful reading note ends with two bounded statements: what passed under the specified grader, and what remains outside that grader. For MedAgentBench, preserve the distinction between query correctness, valid proposed action and successful real-system execution. That makes the result easier to inspect and prevents an efficient benchmark from carrying an unsupported claim about an entire care workflow.
References & further reading
These original sources support the methods discussed here. Our suggested planning steps are editorial guidance, not an endorsement by the source authors.
- MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents ↗Jiang, Black, Geng et al.; Stanford University. Primary 300-task paper; historical results, not a current leaderboard.
- MedAgentBench initial 100-task preprint ↗Jiang, Black, Geng et al.; Stanford University. Version history only; its scores must not be mixed with the 300-task evaluation.
- MedAgentBench official implementation ↗Stanford Machine Learning Group. Repository MIT license applies to code; repository access is not a separate patient-data license.
Published by Arcophos. Educational material, not clinical advice or a claim of regulatory compliance. Read our editorial method.