{"publication":"Healthcare Bench AI","url":"https://healthcarebench.ai","publisher":"Arcophos","updated":"2026-09-28","provenance":"Independent analytical publication. Benchmark creation and experimental results belong to their cited authors. Reported results are source-version snapshots, not new Arcophos runs or a live leaderboard.","benchmarks":[{"slug":"medagentbench","name":"MedAgentBench","shortName":"MedAgentBench","version":"300-task paper; arXiv v2","creators":"Jiang, Black, Geng et al.; Stanford University","paperDate":"2025-02-12","headline":"A successful read and a successful action are different claims.","summary":"MedAgentBench turns an instruction into a sequence of requests against a virtual FHIR record environment. Its useful distinction is between returning information and preparing an action. We analyze the 300-task paper snapshot, including the separate query and action results. Our central interpretation is that an aggregate score hides a decision about workload composition. The baseline also checks proposed write payloads without executing them against the server, which places a precise boundary around the word success. This is an independent reading of the authors’ evaluation, with no new model runs.","task":{"input":"Clinician instruction, hospital-specific context and access to nine FHIR functions.","output":"A final answer or a proposed POST payload.","unit":"One patient-specific task attempt.","setting":"Virtual EHR; pass@1; maximum eight interaction rounds."},"dataOrigin":"Physician-authored tasks grounded in deidentified Stanford STARR records; synthetic identifying details and patient-level date jittering.","facts":[{"label":"Tasks","value":"300","detail":"150 query tasks and 150 action tasks.","sourceIds":["mab"],"locator":"Section 2.5.1"},{"label":"Patient profiles","value":"100","detail":"Profiles are shared across tasks; tasks are not 300 independent patients.","sourceIds":["mab"],"locator":"Table 2"},{"label":"Record elements","value":"785,207","detail":"Four reported FHIR resource groups.","sourceIds":["mab"],"locator":"Table 2"},{"label":"Task categories","value":"10","detail":"Grouped into seven broad categories in Table 1.","sourceIds":["mab"],"locator":"Table 1"},{"label":"Interaction budget","value":"8 rounds","detail":"Invalid actions or exceeding the budget count as failures.","sourceIds":["mab"],"locator":"Sections 2.4.1–2.4.3"}],"metric":{"name":"Task success rate","description":"Rule-based success for one attempt; query answers and proposed action payloads have different checks.","formula":"SR = successful task attempts / all task attempts × 100","direction":"Higher is better within the same task/configuration.","comparability":"Pin task version, orchestrator, parser, reference solution and interaction cap. This is not a patient-outcome rate.","sourceIds":["mab"]},"workflow":[{"label":"Resolve the task","detail":"Read instruction and hospital context before choosing a FHIR function.","sourceIds":["mab"]},{"label":"Retrieve","detail":"GET requests return raw server responses to the agent.","sourceIds":["mab"]},{"label":"Propose an action","detail":"POST payloads receive local checks and a simulated success response; baseline writes are not executed.","sourceIds":["mab"]},{"label":"Finish and grade","detail":"Compare query answers with reference solutions; apply authored checks to action payloads.","sourceIds":["mab"]}],"slices":[{"label":"Queries","value":150,"unit":"tasks","detail":"GET-only task subgroup.","sourceIds":["mab"]},{"label":"Actions","value":150,"unit":"tasks","detail":"Requires a proposed POST, sometimes after retrieval.","sourceIds":["mab"]}],"sliceTitle":"The paper balances reads and proposed actions","sliceNote":"Counts describe the 300-task paper; they are not deployment traffic weights.","results":[{"id":"query","title":"Paper-reported query performance","metric":"Query task success rate","unit":"%","lower":0,"upper":100,"scope":"arXiv v2, February 2025; 150 queries, pass@1, eight-round baseline; temperature zero for these selected models.","sourceIds":["mab"],"locator":"Table 3","rows":[{"label":"Claude 3.5 Sonnet v2","value":85.33,"display":"85.33%","detail":"Query subgroup; not action success."},{"label":"GPT-4o","value":72,"display":"72.00%","detail":"Query subgroup; not action success."},{"label":"DeepSeek-V3","value":70.67,"display":"70.67%","detail":"Query subgroup; not action success."},{"label":"Gemini-1.5 Pro","value":52.67,"display":"52.67%","detail":"Query subgroup; not action success."}],"note":"Historical selected rows. Table 3 does not provide uncertainty intervals."},{"id":"action","title":"Paper-reported proposed-action performance","metric":"Action task success rate","unit":"%","lower":0,"upper":100,"scope":"Same paper snapshot and orchestrator; 150 action tasks, with proposed POST payload validation.","sourceIds":["mab"],"locator":"Table 3","rows":[{"label":"Claude 3.5 Sonnet v2","value":54,"display":"54.00%","detail":"Action subgroup; payload checks, not executed writes."},{"label":"GPT-4o","value":56,"display":"56.00%","detail":"Action subgroup; payload checks, not executed writes."},{"label":"DeepSeek-V3","value":54.67,"display":"54.67%","detail":"Action subgroup; payload checks, not executed writes."},{"label":"Gemini-1.5 Pro","value":71.33,"display":"71.33%","detail":"Action subgroup; payload checks, not executed writes."}],"note":"Keep this separate from query scores. No deployment reliability or current model ranking is implied."}],"analysis":[{"heading":"The workload determines the headline","evidence":"Claude’s query/action scores are 85.33/54.00; Gemini-1.5 Pro’s are 52.67/71.33.","interpretation":"A read-heavy workflow and a write-heavy workflow can favor different systems. Equal weighting is a benchmark design choice, not a universal utility function.","sourceIds":["mab"]},{"heading":"Action grading stops before persistence","evidence":"The baseline sends GETs to the server but validates POST payloads locally.","interpretation":"A passing payload leaves server validation, persisted state and downstream execution untested. Treat these as additional integration questions.","sourceIds":["mab"]},{"heading":"Count tasks and people separately","evidence":"300 tasks use 100 patient profiles.","interpretation":"A task-level denominator describes attempts; it does not establish evidence from 300 distinct patients or institutions.","sourceIds":["mab"]},{"heading":"Formatting belongs to the system","evidence":"Invalid actions are failures under the specified parser.","interpretation":"Model knowledge and orchestrator compatibility jointly affect this result. Changing a parser changes the evaluated system.","sourceIds":["mab"]}],"limitations":[{"title":"One source institution","detail":"STARR profiles do not establish transportability to other EHR systems or populations.","sourceIds":["mab"]},{"title":"Bounded workflow scope","detail":"The task inventory does not cover all specialties, teamwork or the consequences of real actions.","sourceIds":["mab"]},{"title":"Version sensitivity","detail":"The January 2025 v1 evaluated 100 tasks; do not combine its scores with these 300-task rows.","sourceIds":["mab-v1","mab"]},{"title":"Data rights require separate checking","detail":"MIT repository licensing does not by itself establish unrestricted reuse of patient-derived records.","sourceIds":["mab-code"]}],"access":{"status":"Official code and virtual environment available","license":"Repository code: MIT. Separate patient-data reuse terms not verified.","restrictions":"Consult the official repository and data distribution terms before obtaining or reusing records. This site reproduces aggregate metadata only.","url":"https://github.com/stanfordmlgroup/MedAgentBench","sourceIds":["mab-code"]},"sourceIds":["mab","mab-v1","mab-code"]},{"slug":"agentclinic","name":"AgentClinic","shortName":"AgentClinic","version":"2026 journal article; expanded inventory from v5 appendix","creators":"Schmidgall, Ziaei, Harris, Kim, Reis, Jopling and Moor","paperDate":"2026-04-27","headline":"The simulated patient is part of the test instrument.","summary":"AgentClinic evaluates diagnostic dialogue through interacting doctor, patient, measurement and moderator agents. Its datasets mix transformed exam cases, clinical records and image-based case challenges. We examine what changes when a model must request information rather than receive a complete vignette. The simulation makes information gathering inspectable, but it also introduces another source of variation: the model playing the patient. Our analysis separates the candidate doctor model, the simulated environment and the diagnosis judge. The historical results below describe this particular combination, rather than a general claim about autonomous clinical competence.","task":{"input":"An initial encounter objective; further history and findings obtained through simulated dialogue.","output":"A final diagnosis judged against the case diagnosis.","unit":"One simulated diagnostic encounter.","setting":"Published comparison: GPT-4 patient/measurement agents and at most 20 interactions."},"dataOrigin":"MedQA and MedMCQA exam-derived scenarios, MIMIC-IV record-derived scenarios, NEJM case challenges and translated scenarios; conversations themselves are simulated.","facts":[{"label":"MedQA cases","value":"215","detail":"Expanded scenario inventory.","sourceIds":["ac-pre","ac-code"],"locator":"Dataset summary; release notes"},{"label":"MIMIC-IV cases","value":"200","detail":"Selected single-diagnosis cases in the paper.","sourceIds":["ac-pre"],"locator":"Appendix H.1"},{"label":"NEJM cases","value":"120","detail":"Image and dialogue case challenges.","sourceIds":["ac","ac-pre"],"locator":"Results; dataset summary"},{"label":"Interaction cap","value":"20","detail":"Published principal doctor-model comparison.","sourceIds":["ac"],"locator":"Comparison of models"},{"label":"Agent roles","value":"4","detail":"Doctor, patient, measurement and moderator.","sourceIds":["ac"],"locator":"Figure 1"}],"metric":{"name":"Diagnostic accuracy","description":"The moderator checks the final diagnosis against the scenario diagnosis.","formula":"Accuracy = encounters judged correct / evaluated encounters × 100","direction":"Higher is better under a fixed simulation.","comparability":"Patient and measurement model, moderator, tool access, case subset and interaction budget must match.","sourceIds":["ac","ac-code"]},"workflow":[{"label":"Distribute case information","detail":"Different case fields are exposed to different simulated roles.","sourceIds":["ac"]},{"label":"Gather evidence","detail":"The doctor asks the patient questions or requests measurements.","sourceIds":["ac"]},{"label":"Commit to a diagnosis","detail":"The doctor concludes before the interaction budget is exhausted.","sourceIds":["ac"]},{"label":"Judge the conclusion","detail":"The moderator compares the conclusion with the reference diagnosis.","sourceIds":["ac","ac-code"]}],"slices":[{"label":"MedQA-derived","value":215,"unit":"scenarios","detail":"Exam-derived text scenarios.","sourceIds":["ac-pre"]},{"label":"MIMIC-IV-derived","value":200,"unit":"scenarios","detail":"Selected clinical record scenarios.","sourceIds":["ac-pre"]},{"label":"NEJM-derived","value":120,"unit":"scenarios","detail":"Multimodal case challenges.","sourceIds":["ac-pre"]}],"sliceTitle":"Three core scenario collections","sliceNote":"Selected core collections, not the entire expanded inventory; specialty and language derivatives are omitted to avoid double-counting.","results":[{"id":"ac-medqa","title":"Historical AgentClinic-MedQA comparison","metric":"Diagnostic accuracy","unit":"%","lower":0,"upper":100,"scope":"2026 journal article, Figure 2; GPT-4 simulated patient and measurement agents, 20-interaction cap.","sourceIds":["ac"],"locator":"Comparison of models; Figure 2","rows":[{"label":"Claude-3.5","value":62.1,"display":"62.1%","detail":"Paper reports ±3.3 alongside the mean."},{"label":"GPT-4","value":51.6,"display":"51.6%","detail":"Paper reports ±3.3 alongside the mean."},{"label":"GPT-4o","value":34.2,"display":"34.2%","detail":"Paper reports ±3.4 alongside the mean."}],"note":"Published measurements only. The ± quantities are reproduced as reported, not relabeled as confidence intervals."}],"analysis":[{"heading":"The environment changes the score","evidence":"The study varies the patient model and observes different diagnostic accuracy.","interpretation":"An agent result is a property of both the tested policy and its simulation. Record the patient model alongside the doctor model.","sourceIds":["ac"]},{"heading":"Real source records do not make dialogue real","evidence":"MIMIC-IV records seed cases while language models generate interactions.","interpretation":"This supports studying simulated information gathering, not claims that actual patients would provide equivalent answers.","sourceIds":["ac","ac-pre"]},{"heading":"Translations are related cases","evidence":"AgentClinic-Lang translates MedQA-derived scenarios.","interpretation":"Language coverage expands evaluation conditions without automatically adding independent clinical problems.","sourceIds":["ac-pre"]},{"heading":"Action and conversation are complementary","evidence":"AgentClinic judges diagnosis after dialogue; MedAgentBench grades record tasks and proposed payloads.","interpretation":"A combined evaluation should preserve both endpoints rather than average their percentages.","sourceIds":["ac","mab"]}],"limitations":[{"title":"Simulator and judge dependence","detail":"Generated patient responses and moderator judgments can affect measured performance.","sourceIds":["ac","ac-code"]},{"title":"Source and case selection","detail":"Exam transformations and selected single-diagnosis records limit coverage of ambiguity and multimorbidity.","sourceIds":["ac-pre"]},{"title":"Historical models","detail":"The publication date does not make these older model configurations a current frontier leaderboard.","sourceIds":["ac"]},{"title":"Known source inconsistencies","detail":"The website says 24 biases while the final paper says 23; this site avoids a bias-count headline.","sourceIds":["ac"]}],"access":{"status":"Public implementation; source-specific data access","license":"Code MIT; underlying record and case-source terms remain separate.","restrictions":"MIMIC-IV-derived data requires approval. Third-party case/image reuse rights are not established by the code license.","url":"https://github.com/SamuelSchmidgall/AgentClinic","sourceIds":["ac-code"]},"sourceIds":["ac","ac-pre","ac-code"]}],"explorer":{"kind":"coverage","title":"Which part of an agent does the benchmark exercise?","intro":"Filter task families to compare the information available, expected behavior and point where grading stops. These annotations are our analysis of published protocols.","caution":"Coverage rows describe evaluation scope. They are not measured performance or proof that a capability transfers to a deployed workflow.","sourceIds":["mab","ac"],"parameters":[],"rows":[{"label":"Find a patient or result","category":"Record retrieval","input":"Identifiers and FHIR search functions","output":"Requested value or identifier","metric":"Rule-based query SR","constraint":"Depends on the parser and reference answer.","benchmarkSlug":"medagentbench","sourceIds":["mab"]},{"label":"Aggregate recent measurements","category":"Record retrieval","input":"Time-bounded observations","output":"Computed task answer","metric":"Rule-based query SR","constraint":"Tests the specified calculation and time filter.","benchmarkSlug":"medagentbench","sourceIds":["mab"]},{"label":"Record a measurement","category":"Proposed actions","input":"New observation and patient identifier","output":"POST payload","metric":"Action SR","constraint":"Baseline validates payload, without executing a server write.","benchmarkSlug":"medagentbench","sourceIds":["mab"]},{"label":"Prepare an order","category":"Proposed actions","input":"Task conditions, codes and retrieved context","output":"Medication/test/referral payload","metric":"Action SR","constraint":"Does not verify downstream care or persisted server state.","benchmarkSlug":"medagentbench","sourceIds":["mab"]},{"label":"Elicit a diagnostic history","category":"Diagnostic dialogue","input":"Partial information distributed among agents","output":"Final diagnosis after dialogue","metric":"Diagnostic accuracy","constraint":"Patient-model responses influence the test.","benchmarkSlug":"agentclinic","sourceIds":["ac"]},{"label":"Request a diagnostic measurement","category":"Diagnostic dialogue","input":"Doctor requests to the measurement agent","output":"Returned finding, then diagnosis","metric":"Diagnostic accuracy","constraint":"Information request and answer quality are coupled.","benchmarkSlug":"agentclinic","sourceIds":["ac"]},{"label":"Interpret a case image","category":"Multimodal dialogue","input":"NEJM-derived image and case interaction","output":"Final diagnosis","metric":"Diagnostic accuracy","constraint":"Curated case challenges are not routine imaging prevalence.","benchmarkSlug":"agentclinic","sourceIds":["ac","ac-pre"]}]},"references":[{"id":"mab","title":"MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents","organization":"Jiang, Black, Geng et al.; Stanford University","url":"https://arxiv.org/html/2501.14654v2","note":"Primary 300-task paper; historical results, not a current leaderboard.","locator":"Sections 2.2–2.5; Tables 2–3","version":"arXiv v2, 12 February 2025"},{"id":"mab-v1","title":"MedAgentBench initial 100-task preprint","organization":"Jiang, Black, Geng et al.; Stanford University","url":"https://arxiv.org/html/2501.14654v1","note":"Version history only; its scores must not be mixed with the 300-task evaluation.","locator":"Tables 2–3","version":"arXiv v1, 24 January 2025"},{"id":"mab-code","title":"MedAgentBench official implementation","organization":"Stanford Machine Learning Group","url":"https://github.com/stanfordmlgroup/MedAgentBench","note":"Repository MIT license applies to code; repository access is not a separate patient-data license.","locator":"README and LICENSE","version":"Accessed 28 September 2026"},{"id":"ac","title":"AgentClinic: a multimodal benchmark for tool-using clinical AI agents","organization":"Schmidgall et al.; npj Digital Medicine","url":"https://www.nature.com/articles/s41746-026-02674-7","note":"Final journal article, publishing historical model experiments.","locator":"Comparison of models; Figure 2; Methods","version":"27 April 2026"},{"id":"ac-pre","title":"AgentClinic expanded preprint and appendices","organization":"Schmidgall et al.","url":"https://arxiv.org/html/2405.07960v5","note":"Versioned source for dataset inventory and statistical appendix.","locator":"Appendices D and H; dataset summary table","version":"arXiv v5, 25 May 2025"},{"id":"ac-code","title":"AgentClinic official code and release notes","organization":"Samuel Schmidgall and collaborators","url":"https://github.com/SamuelSchmidgall/AgentClinic","note":"MIT-licensed code; MIMIC-IV cases require source-data approval.","locator":"README release history and evaluation flags; LICENSE.txt","version":"Accessed 28 September 2026"}]}