What does MedAgentBench success actually measure?
Read MedAgentBench’s query and action results through its 300-task protocol, parser and proposed-write grading boundary.
Original analysis / Healthcare Bench AI
Explore two different forms of agency: FHIR record tasks in MedAgentBench and diagnostic dialogue in AgentClinic. Our original analysis tracks what an agent receives, what it must do and what the grader actually checks. The published results retain their historical model configurations and benchmark versions. Use the coverage explorer to locate a capability, then inspect why successful retrieval, valid proposed actions and correct simulated diagnoses support different conclusions.
Read MedAgentBench’s query and action results through its 300-task protocol, parser and proposed-write grading boundary.
Separate the tested doctor model, simulated patient, measurement agent and diagnosis moderator when interpreting AgentClinic.
Build a capability comparison between record actions and diagnostic dialogue without averaging non-equivalent benchmark scores.