Benchmark dossiers / independent analysis

The benchmark behind the number.

Inspect what each benchmark asks, how its score is produced, and which conclusions its design can support. These are original analyses by Arcophos of the benchmark authors’ work. Any reproduced measurements are labeled with their published source and version.

Dossier01

300-task paper; arXiv v2

MedAgentBench ↗

A successful read and a successful action are different claims.

UnitOne patient-specific task attempt.MeasureTask success rate
Dossier02

2026 journal article; expanded inventory from v5 appendix

AgentClinic ↗

The simulated patient is part of the test instrument.

UnitOne simulated diagnostic encounter.MeasureDiagnostic accuracy
What we contribute. Source reconciliation, task and metric interpretation, and tools that make assumptions inspectable. We do not claim to have created these benchmarks or run the reported models.