# Healthcare Bench AI > Independent analysis of MedAgentBench and AgentClinic: task success, simulated dialogue, query/action differences and the limits of medical agent scores. Explore two different forms of agency: FHIR record tasks in MedAgentBench and diagnostic dialogue in AgentClinic. Our original analysis tracks what an agent receives, what it must do and what the grader actually checks. The published results retain their historical model configurations and benchmark versions. Use the coverage explorer to locate a capability, then inspect why successful retrieval, valid proposed actions and correct simulated diagnoses support different conclusions. ## Provenance Independent analysis published by Arcophos. Benchmark creation belongs to the credited authors. Result rows are selected paper-reported measurements with their source versions and evaluation conditions, not new Arcophos runs or a live leaderboard. ## Benchmark dossiers - [MedAgentBench](https://healthcarebench.ai/benchmarks/medagentbench/): A successful read and a successful action are different claims. Source version: 300-task paper; arXiv v2. - [AgentClinic](https://healthcarebench.ai/benchmarks/agentclinic/): The simulated patient is part of the test instrument. Source version: 2026 journal article; expanded inventory from v5 appendix. ## Original analyses - [What does MedAgentBench success actually measure?](https://healthcarebench.ai/guides/medagentbench-query-action-success/): Read MedAgentBench’s query and action results through its 300-task protocol, parser and proposed-write grading boundary. - [Why is the simulated patient part of an AgentClinic result?](https://healthcarebench.ai/guides/agentclinic-simulator-and-patient-model/): Separate the tested doctor model, simulated patient, measurement agent and diagnosis moderator when interpreting AgentClinic. - [How should MedAgentBench and AgentClinic be compared?](https://healthcarebench.ai/guides/compare-medagentbench-agentclinic/): Build a capability comparison between record actions and diagnostic dialogue without averaging non-equivalent benchmark scores. ## Inspect the evidence - [Evidence JSON](https://healthcarebench.ai/evidence.json): Task definitions, dataset facts, scoring rules, source-version results, our interpretations, and reference IDs. - [Sources](https://healthcarebench.ai/sources/): Original papers and repositories with evidence locators. - [Editorial method](https://healthcarebench.ai/methodology/): Source reconciliation and interpretation boundaries. - [About](https://healthcarebench.ai/about/): Ownership and corrections. Analysis updated: 2026-09-28