Med-PaLM 2
Google's medical large language model, the first to reach expert-level scores on USMLE-style questions (86.5% on MedQA); available only through Google Cloud to selected partners, and largely succeeded by Gemini-based medical models.
At a glance
What it does
Med-PaLM 2 is a large language model tuned for the medical domain, designed to provide high-quality answers to medical questions. It scored 86.5% on USMLE-style questions from the MedQA benchmark — a level described as comparable to human experts — and underpinned the MedLM family of models that Google offered to healthcare organisations until MedLM's retirement in 2025.
Notable uses in oncology
In pairwise physician evaluations, its long-form answers to health questions were preferred on eight of nine criteria, including scientific factuality and medical reasoning. In oncology and other specialties it has been studied as an information-support tool; like all medical language models, it does not replace clinical judgement.
Tasks, data types and cancers
Architecture
Training data
Not disclosed. The publication states that the PaLM 2 base model underwent medical-domain finetuning, but it does not name the tuning corpus; only the evaluation benchmarks are named (MedQA, MedMCQA, PubMedQA, MMLU clinical topics).
Linked datasets
- benchmark MedQA (USMLE) Open download
Evaluation
How to run
Not yet documented on this card.
Regulatory status and intended use
Regulatory status is quoted from the source linked above and can change. Research-use-only models must not be used for clinical decisions.
Limitations and bias
- Closed and access-restricted; results cannot be independently reproduced.
- Benchmark scores on exam questions do not translate to clinical safety.
Sources
This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.