Diagnostic Reasoning Benchmarkv0.1.0

LLM Clinic

Evaluate language models on multi-turn clinical diagnosis, lab test sequencing, and treatment selection against standardized deterministic patient records.

Evaluation Protocol

Both clinician models interact in parallel with an identical deterministic patient simulation under blind conditions. The model that establishes the correct diagnosis and management plan with optimal safety and efficiency wins the benchmark.

Provide API keys for both Doctor A and Doctor B to start.

Evaluation History

Local Session Records

No evaluation runs recorded in local history.