Cancer3.AIAI in OncologyAI Models › Datasets

AI in Oncology · AI Models

A researcher-grade catalog of AI models, datasets and open data needs in oncology — every card structured, sourced and dated.

Datasets, registries and benchmarks you can train or evaluate oncology models on — with access conditions, licences, sizes, annotations and the models already using them. { } export JSON

26 / 26 datasets

AACR Project GENIE

American Association for Cancer Research; hosted on Synapse / cBioPortal
Registry / portal Free registration AACR GENIE data-use terms (Synapse registration)

The largest real-world clinical sequencing registry in oncology — the place to test whether a variant-effect or biomarker model generalises beyond TCGA.

200 000 patients clinical-grade tumour sequencing from 19+ institutions with limited clinical data; releases twice a year
Benchmark / challenge Open download public

The reference affinity dataset behind fifteen years of binding predictors; its size and composition are why dataset composition, not architecture, drives reported performance.

176 161 affinity measurements 114 alleles across six species — the historical training core for binding predictors, and a reminder of how small the measured world is

BraTS — Brain Tumor Segmentation Challenge

BraTS organisers (University of Pennsylvania, Indiana University, MICCAI)
Benchmark / challenge Free registration challenge terms (registration on Synapse)

The long-running benchmark for brain tumour segmentation from MRI; nnU-Net-based methods have dominated its leaderboards.

1 250 patients ≥1,250 glioma cases since 2021; later editions add paediatric, metastasis, meningioma and Sub-Saharan Africa tracks
Dataset site used by models: 1 Open card →

CAMELYON16 / CAMELYON17

Radboud University Medical Center and partners (grand-challenge.org)
Benchmark / challenge Open download CC0

The canonical whole-slide benchmark for breast cancer lymph-node metastasis detection; still the first sanity check for any new pathology encoder.

1 399 WSIs CAMELYON16: 399 slides (270 train / 129 test) with pixel-level metastasis annotations; CAMELYON17: 1,000 slides from 5 centres with patient-level pN stage
Dataset site used by models: 2 Open card →

CPTAC — Clinical Proteomic Tumor Analysis Consortium

NCI Office of Cancer Clinical Proteomics Research
Registry / portal Free registration open (proteomics via PDC) / controlled (raw sequencing via dbGaP)

Proteogenomic companion to TCGA: the largest public set where protein-level measurements, genomics and images exist for the same tumours.

1 000 patients 10+ tumour types with proteogenomic profiling (proteome, phosphoproteome) on genomically characterised cases; images in TCIA
Dataset site used by models: 2 Open card →

Genecorpus-30M

Theodoris/Ellinor lab (Broad Institute, Massachusetts General Hospital)
Dataset Open download Apache-2.0

Genecorpus-30M is a pretraining corpus of about 30 million human single-cell transcriptomes assembled from 561 publicly available datasets, of which 27,406,217 cells passed quality filters. It was built to pretrain Geneformer (Theodoris et al., Nature 2023) and is released under Apache-2.0. Crucially for oncology use: cells with high mutational burden — malignant cells and immortalized cell lines — were deliberately excluded from the corpus.

Dataset site used by models: 1 Open card →

Genome in a Bottle (GIAB)

National Institute of Standards and Technology (NIST), USA
Open download Materiał referencyjny NIST; zgoda dawców z PGP na redystrybucję komercyjną

A public benchmark resource run by NIST within a consortium of public bodies, companies and academic centres. It contains seven thoroughly characterised human genomes with benchmark variant sets and high-confidence regions, and is used to measure how accurately a bioinformatics pipeline calls variants. It is not an oncology dataset and not an imaging dataset — it is a reference point for genome readout itself. The consortium works with the GA4GH Benchmarking Team on comparison standards. Distribution: the NIST reference material shop, the Coriell Institute and the Personal Genome Project.

7 patients 7 genomy referencyjne Seven characterised genomes: the pilot genome NA12878/HG001 from the HapMap collection and two Personal Genome Project trios — Ashkenazi Jewish (HG002, HG003, HG004) and Han Chinese (HG005, HG006, HG007).
Genomics (DNA) FASTQBAMVCF
Dataset site used by models: 1 Open card →

IEDB automated benchmark (MHC class I)

La Jolla Institute for Immunology
Benchmark / challenge Open download public

The only prospective, third-party benchmark in the field. Its eight-year summary is sobering: leading methods are statistically indistinguishable, and a new method needs about four years before enough data accumulate to judge it.

runs continuously on newly deposited data, before it can leak into anyone's training set
Dataset site used by models: 3 Open card →

IEDB — Immune Epitope Database

La Jolla Institute for Immunology, funded by NIAID
Registry / portal Open download free to use; NIAID-funded public resource

The field's central repository of epitope data and the source of almost every training set for peptide–MHC models — and of their allele skew.

1 600 000 epitope-related records curated from published literature and direct submissions; includes MHC binding assays, MS-eluted ligands and T-cell assays
Dataset site used by models: 4 Open card →

IPD-IMGT/HLA Database

EMBL-EBI / Anthony Nolan Research Institute
Registry / portal Open download CC BY-ND 4.0 (see site terms)

The naming authority for HLA alleles: the catalogue whose size — thirty thousand names against roughly a hundred well-measured alleles — defines the central problem of this field.

30 894 named class I alleles as of June 2026: 9,279 HLA-A, 11,258 HLA-B, 9,416 HLA-C — while one patient carries at most six
Dataset site used by models: 1 Open card →

ISIC Archive — International Skin Imaging Collaboration

ISIC (Memorial Sloan Kettering and partners)
Registry / portal Open download CC-0 / CC BY-NC per contributor

The public backbone of skin-cancer AI; strongly skewed towards light skin tones, which every model card built on it should say.

70 000 dermoscopic images tens of thousands of images with diagnosis; annual challenge subsets (e.g. 2020: 33,126 images)

MedQA (USMLE)

Jin et al. (Columbia University)
Benchmark / challenge Open download MIT

The exam-style benchmark on which Med-PaLM 2, GPT-4 and MedGemma report headline accuracy; useful for comparing models, not for judging clinical safety.

12 723 questions (English) multiple-choice medical licensing exam questions; also Mandarin and Traditional Chinese subsets
Dataset site used by models: 2 Open card →

MHC Motif Atlas

Gfeller lab, University of Lausanne
Dataset Open download free for academic use

The reference collection of HLA binding motifs — and the clearest picture of the field's long tail: a million measured ligands still describe barely 135 of thirty thousand alleles.

1 000 000 ligands over a million ligands — but spread across only about 135 class I molecules, against 30,894 named class I alleles in IPD-IMGT/HLA (June 2026)
Dataset site used by models: 1 Open card →

Mono-allelic HLA class I peptidome (Sarkizova / Abelin)

Broad Institute / Dana-Farber Cancer Institute
Dataset Open download published supplementary data (see papers)

The engineered-cell peptidome that gave the field clean allele labels; fifteen of its alleles had no described motif before, and the panel covers at least one allele in 95% of people worldwide.

186 464 peptides 95 mono-allelic cell lines (31 HLA-A, 40 HLA-B, 21 HLA-C, 3 HLA-G), median 1,860 peptides per allele; the earlier Abelin 2017 set covered 16 alleles and >24,000 peptides
Dataset site used by models: 4 Open card →

NLST — National Lung Screening Trial

NCI (Cancer Data Access System)
Dataset Controlled access NCI CDAS data-use agreement

The trial that established LDCT screening; its images plus outcomes trained Sybil and remain the only large public-by-application longitudinal LDCT cohort.

53 454 patients randomised trial of low-dose CT vs chest X-ray screening (2002–2009) with cancer and mortality follow-up; imaging available for a subset
Dataset site used by models: 1 Open card →

PANDA — Prostate cANcer graDe Assessment

Radboud UMC and Karolinska Institutet (Kaggle challenge)
Benchmark / challenge Free registration CC BY-NC-SA 4.0 (Kaggle competition rules)

Reference dataset and challenge for AI Gleason grading; the follow-up Nature Medicine paper validated algorithms on external cohorts.

10 616 biopsy WSIs the largest public prostate biopsy set; Gleason / ISUP grade per biopsy from two centres
Dataset site used by models: 1 Open card →

PatchCamelyon (PCam)

Veeling et al.; mirrored on Hugging Face
Benchmark / challenge Open download CC0

Small, fast, fully open patch-classification benchmark — the standard smoke test for image encoders in pathology and a common zero-shot evaluation set.

327 680 96×96 patches derived from CAMELYON16; binary label = tumour tissue in the central 32×32 region
3 508 downloads / month ♥ 12 updated 2024-05-25 Dataset site Hugging Face used by models: 1 Open card →

Protein Data Bank (wwPDB / RCSB)

Worldwide Protein Data Bank
Registry / portal Open download CC0

The open archive of experimentally solved macromolecular structures, including oncology targets (kinases, KRAS, p53) and their drug complexes.

220 000 experimental structures X-ray, cryo-EM and NMR structures of proteins, nucleic acids and complexes; the training ground of AlphaFold
Dataset site used by models: 2 Open card →

PubMed / PMC Open Access Subset

US National Library of Medicine
Text corpus Open download PubMed metadata public domain; PMC OA articles under CC licences per article

The literature corpus behind BiomedBERT, BiomedCLIP and CONCH's caption data; the entry point for any oncology NLP or vision-language pretraining.

36 000 000 citations (PubMed); millions of full texts in PMC OA abstracts for all of PubMed; full text and figures for the open-access subset
Dataset site used by models: 5 Open card →

TCGA — The Cancer Genome Atlas (via NCI Genomic Data Commons)

NCI / NHGRI; hosted by the Genomic Data Commons
Registry / portal Free registration NIH GDS Policy — open tier; controlled tier via dbGaP

The reference multi-omics cancer cohort: molecular profiles, clinical outcomes and whole-slide images for 33 cancer types, downloadable through the GDC portal and API.

11 000 patients 30 000 diagnostic + tissue WSIs 33 cancer types; ~11,000 patients with molecular data; diagnostic slides for most cases; matched clinical follow-up
Dataset site used by models: 7 Open card →

TCIA — The Cancer Imaging Archive

NCI Cancer Imaging Program; hosted by the University of Arkansas for Medical Sciences
Registry / portal Open download mostly CC BY 3.0/4.0 per collection; some restricted collections

The main public archive of de-identified cancer imaging, organised into collections by disease and modality; the source of most public radiology training data in oncology.

200 collections hundreds of collections, tens of thousands of patients; DICOM with linked clinical and sometimes genomic data
Dataset site used by models: 1 Open card →

Cards follow the cancer3.ai model-card standard: AI_MODEL_CARD_STANDARD.md. Corrections and new entries: contact the editorial team; every fact needs a public source.

This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.