Cancer3.AIAI in OncologyAI Models › Med-PaLM 2
Clinical product Closed / commercial Research use only clinical-LLM

Med-PaLM 2

Google's medical large language model, the first to reach expert-level scores on USMLE-style questions (86.5% on MedQA); available only through Google Cloud to selected partners, and largely succeeded by Gemini-based medical models.

At a glance

DeveloperGoogle Research
Released2023-05-16
Licenceproprietary
AvailabilityClosed / commercial
KindClinical product
Regulatory statusResearch use only

What it does

Med-PaLM 2 is a large language model tuned for the medical domain, designed to provide high-quality answers to medical questions. It scored 86.5% on USMLE-style questions from the MedQA benchmark — a level described as comparable to human experts — and underpinned the MedLM family of models that Google offered to healthcare organisations until MedLM's retirement in 2025.

Notable uses in oncology

In pairwise physician evaluations, its long-form answers to health questions were preferred on eight of nine criteria, including scientific factuality and medical reasoning. In oncology and other specialties it has been studied as an information-support tool; like all medical language models, it does not replace clinical judgement.

Tasks, data types and cancers

Data typeClinical text
Clinical settingEducationBasic research
CancerPan-cancer
Inputmedical questions and clinical text prompts (multiple-choice and long-form)
Outputfree-text answers; the paper adds an 'ensemble refinement' prompting strategy that conditions the model on its own sampled reasoning chains

Architecture

FamilyDecoder-only transformer large language model, adapted to the medical domain
BackbonePaLM 2 (parameter count not disclosed by the developer)
Pre-trainingPaLM 2 base pretraining followed by medical-domain finetuning (Singhal K et al., arXiv:2305.09617 / Nat Med 2025)

Training data

Not disclosed. The publication states that the PaLM 2 base model underwent medical-domain finetuning, but it does not name the tuning corpus; only the evaluation benchmarks are named (MedQA, MedMCQA, PubMedQA, MMLU clinical topics).

Linked datasets

Evaluation

Benchmark / datasetMetricValueExternal validationSource
MedQA (USMLE) accuracy 86.5% yes Source

How to run

Not yet documented on this card.

Regulatory status and intended use

Regulatory statusResearch use only
Intended useResearch preview for partners; not a medical device.

Source →

Regulatory status is quoted from the source linked above and can change. Research-use-only models must not be used for clinical decisions.

Limitations and bias

  • Closed and access-restricted; results cannot be independently reproduced.
  • Benchmark scores on exam questions do not translate to clinical safety.

Sources

  1. Singhal K et al. Toward expert-level medical question answering with large language models. Nat Med 2025
  2. Singhal K et al. Towards Expert-Level Medical Question Answering with Large Language Models (preprint)
  3. Google Research — Med-PaLM

This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.