Benchmark / challenge
Open download
MIT
MedQA (USMLE)
The exam-style benchmark on which Med-PaLM 2, GPT-4 and MedGemma report headline accuracy; useful for comparing models, not for judging clinical safety.
At a glance
ProviderJin et al. (Columbia University)
AccessOpen download
LicenceMIT
questions (English)12 723
Size notesmultiple-choice medical licensing exam questions; also Mandarin and Traditional Chinese subsets
FormatsJSON
Labels and annotations
Correct answer per question; no oncology sub-labels, though many items are oncology cases.
Details
Not yet documented on this card.
Models trained or evaluated on it
- benchmark MedGemma Google (Health AI Developer Foundations)
- benchmark Med-PaLM 2 Google Research
Sources
This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.