{"architecture":{"backbone":"PaLM 2 (parameter count not disclosed by the developer)","family":"Decoder-only transformer large language model, adapted to the medical domain","input":"medical questions and clinical text prompts (multiple-choice and long-form)","output":"free-text answers; the paper adds an 'ensemble refinement' prompting strategy that conditions the model on its own sampled reasoning chains","pretraining":"PaLM 2 base pretraining followed by medical-domain finetuning (Singhal K et al., arXiv:2305.09617 / Nat Med 2025)"},"article":null,"cancer_slugs":["pan-cancer"],"category":"clinical-LLM","confidence":"medium","datasets":[{"name":"MedQA (USMLE)","note":"","role":"benchmark","slug":"medqa"}],"developer":"Google Research","evaluation":[{"benchmark":"MedQA (USMLE)","dataset_slug":"medqa","external":true,"metric":"accuracy","source":"https://www.nature.com/articles/s41591-024-03423-7","value":"86.5%"}],"hf":null,"kind":"product","license":"proprietary","limitations":"- Closed and access-restricted; results cannot be independently reproduced.\n- Benchmark scores on exam questions do not translate to clinical safety.","links":{"demo":null,"docs":"https://sites.research.google/med-palm/","doi":"10.1038/s41591-024-03423-7","github":null,"huggingface":null,"paper":"https://www.nature.com/articles/s41591-024-03423-7","pmid":null},"modalities":["clinical-text"],"name":"Med-PaLM 2","notable_uses":"In pairwise physician evaluations, its long-form answers to health questions were preferred on eight of nine criteria, including scientific factuality and medical reasoning. In oncology and other specialties it has been studied as an information-support tool; like all medical language models, it does not replace clinical judgement.","openness":"closed","regulatory":{"intended_use_en":"Research preview for partners; not a medical device.","intended_use_pl":"Podgl\u0105d badawczy dla partner\u00f3w; nie jest wyrobem medycznym.","source_url":"https://sites.research.google/med-palm/","status":"research-only"},"regulatory_status":"research-only","release_date":"2023-05-16","run_snippet":"","settings":["education","basic-research"],"slug":"med-palm-2","sources":[{"label":"Singhal K et al. Toward expert-level medical question answering with large language models. Nat Med 2025","url":"https://www.nature.com/articles/s41591-024-03423-7"},{"label":"Singhal K et al. Towards Expert-Level Medical Question Answering with Large Language Models (preprint)","url":"https://arxiv.org/abs/2305.09617"},{"label":"Google Research \u2014 Med-PaLM","url":"https://sites.research.google/med-palm/"}],"summary":"Google's medical large language model, the first to reach expert-level scores on USMLE-style questions (86.5% on MedQA); available only through Google Cloud to selected partners, and largely succeeded by Gemini-based medical models.","tasks":["question-answering"],"training":{"summary_en":"Not disclosed. The publication states that the PaLM 2 base model underwent medical-domain finetuning, but it does not name the tuning corpus; only the evaluation benchmarks are named (MedQA, MedMCQA, PubMedQA, MMLU clinical topics).","summary_pl":"Nieujawnione. Publikacja podaje, \u017ce model bazowy PaLM 2 przeszed\u0142 dostrajanie w domenie medycznej, ale nie nazywa korpusu dostrajaj\u0105cego; nazwane s\u0105 wy\u0142\u0105cznie zbiory ewaluacyjne (MedQA, MedMCQA, PubMedQA, MMLU clinical topics)."},"updated_at":"2026-09-06T15:15:13.658172","url":"/ai-oncology/models/med-palm-2","usage":{},"verified_at":"2026-09-05T22:26:00.605039","verified_by":"editorial","version":null,"what_it_does":"Med-PaLM 2 is a large language model tuned for the medical domain, designed to provide high-quality answers to medical questions. It scored 86.5% on USMLE-style questions from the MedQA benchmark \u2014 a level described as comparable to human experts \u2014 and underpinned the MedLM family of models that Google offered to healthcare organisations until MedLM's retirement in 2025."}
