Cancer3.AIAI in OncologyAI Models › Datasets

AI in Oncology · AI Models

A researcher-grade catalog of AI models, datasets and open data needs in oncology — every card structured, sourced and dated.

Datasets, registries and benchmarks you can train or evaluate oncology models on — with access conditions, licences, sizes, annotations and the models already using them. { } export JSON

× Clear filters
16 / 26 datasets

AACR Project GENIE

American Association for Cancer Research; hosted on Synapse / cBioPortal
Registry / portal Free registration AACR GENIE data-use terms (Synapse registration)

The largest real-world clinical sequencing registry in oncology — the place to test whether a variant-effect or biomarker model generalises beyond TCGA.

200 000 patients clinical-grade tumour sequencing from 19+ institutions with limited clinical data; releases twice a year
Benchmark / challenge Open download public

The reference affinity dataset behind fifteen years of binding predictors; its size and composition are why dataset composition, not architecture, drives reported performance.

176 161 affinity measurements 114 alleles across six species — the historical training core for binding predictors, and a reminder of how small the measured world is

CPTAC — Clinical Proteomic Tumor Analysis Consortium

NCI Office of Cancer Clinical Proteomics Research
Registry / portal Free registration open (proteomics via PDC) / controlled (raw sequencing via dbGaP)

Proteogenomic companion to TCGA: the largest public set where protein-level measurements, genomics and images exist for the same tumours.

1 000 patients 10+ tumour types with proteogenomic profiling (proteome, phosphoproteome) on genomically characterised cases; images in TCIA
Dataset site used by models: 2 Open card →

IEDB automated benchmark (MHC class I)

La Jolla Institute for Immunology
Benchmark / challenge Open download public

The only prospective, third-party benchmark in the field. Its eight-year summary is sobering: leading methods are statistically indistinguishable, and a new method needs about four years before enough data accumulate to judge it.

runs continuously on newly deposited data, before it can leak into anyone's training set
Dataset site used by models: 3 Open card →

IEDB — Immune Epitope Database

La Jolla Institute for Immunology, funded by NIAID
Registry / portal Open download free to use; NIAID-funded public resource

The field's central repository of epitope data and the source of almost every training set for peptide–MHC models — and of their allele skew.

1 600 000 epitope-related records curated from published literature and direct submissions; includes MHC binding assays, MS-eluted ligands and T-cell assays
Dataset site used by models: 4 Open card →

IPD-IMGT/HLA Database

EMBL-EBI / Anthony Nolan Research Institute
Registry / portal Open download CC BY-ND 4.0 (see site terms)

The naming authority for HLA alleles: the catalogue whose size — thirty thousand names against roughly a hundred well-measured alleles — defines the central problem of this field.

30 894 named class I alleles as of June 2026: 9,279 HLA-A, 11,258 HLA-B, 9,416 HLA-C — while one patient carries at most six
Dataset site used by models: 1 Open card →

MedQA (USMLE)

Jin et al. (Columbia University)
Benchmark / challenge Open download MIT

The exam-style benchmark on which Med-PaLM 2, GPT-4 and MedGemma report headline accuracy; useful for comparing models, not for judging clinical safety.

12 723 questions (English) multiple-choice medical licensing exam questions; also Mandarin and Traditional Chinese subsets
Dataset site used by models: 2 Open card →

MHC Motif Atlas

Gfeller lab, University of Lausanne
Dataset Open download free for academic use

The reference collection of HLA binding motifs — and the clearest picture of the field's long tail: a million measured ligands still describe barely 135 of thirty thousand alleles.

1 000 000 ligands over a million ligands — but spread across only about 135 class I molecules, against 30,894 named class I alleles in IPD-IMGT/HLA (June 2026)
Dataset site used by models: 1 Open card →

Mono-allelic HLA class I peptidome (Sarkizova / Abelin)

Broad Institute / Dana-Farber Cancer Institute
Dataset Open download published supplementary data (see papers)

The engineered-cell peptidome that gave the field clean allele labels; fifteen of its alleles had no described motif before, and the panel covers at least one allele in 95% of people worldwide.

186 464 peptides 95 mono-allelic cell lines (31 HLA-A, 40 HLA-B, 21 HLA-C, 3 HLA-G), median 1,860 peptides per allele; the earlier Abelin 2017 set covered 16 alleles and >24,000 peptides
Dataset site used by models: 4 Open card →

Protein Data Bank (wwPDB / RCSB)

Worldwide Protein Data Bank
Registry / portal Open download CC0

The open archive of experimentally solved macromolecular structures, including oncology targets (kinases, KRAS, p53) and their drug complexes.

220 000 experimental structures X-ray, cryo-EM and NMR structures of proteins, nucleic acids and complexes; the training ground of AlphaFold
Dataset site used by models: 2 Open card →

PubMed / PMC Open Access Subset

US National Library of Medicine
Text corpus Open download PubMed metadata public domain; PMC OA articles under CC licences per article

The literature corpus behind BiomedBERT, BiomedCLIP and CONCH's caption data; the entry point for any oncology NLP or vision-language pretraining.

36 000 000 citations (PubMed); millions of full texts in PMC OA abstracts for all of PubMed; full text and figures for the open-access subset
Dataset site used by models: 5 Open card →

TCGA — The Cancer Genome Atlas (via NCI Genomic Data Commons)

NCI / NHGRI; hosted by the Genomic Data Commons
Registry / portal Free registration NIH GDS Policy — open tier; controlled tier via dbGaP

The reference multi-omics cancer cohort: molecular profiles, clinical outcomes and whole-slide images for 33 cancer types, downloadable through the GDC portal and API.

11 000 patients 30 000 diagnostic + tissue WSIs 33 cancer types; ~11,000 patients with molecular data; diagnostic slides for most cases; matched clinical follow-up
Dataset site used by models: 7 Open card →

TCIA — The Cancer Imaging Archive

NCI Cancer Imaging Program; hosted by the University of Arkansas for Medical Sciences
Registry / portal Open download mostly CC BY 3.0/4.0 per collection; some restricted collections

The main public archive of de-identified cancer imaging, organised into collections by disease and modality; the source of most public radiology training data in oncology.

200 collections hundreds of collections, tens of thousands of patients; DICOM with linked clinical and sometimes genomic data
Dataset site used by models: 1 Open card →

Cards follow the cancer3.ai model-card standard: AI_MODEL_CARD_STANDARD.md. Corrections and new entries: contact the editorial team; every fact needs a public source.

This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.