Cancer3.AIAI in OncologyAI Models › Datasets

AI in Oncology · AI Models

A researcher-grade catalog of AI models, datasets and open data needs in oncology — every card structured, sourced and dated.

Datasets, registries and benchmarks you can train or evaluate oncology models on — with access conditions, licences, sizes, annotations and the models already using them. { } export JSON

× Clear filters
6 / 26 datasets

AACR Project GENIE

American Association for Cancer Research; hosted on Synapse / cBioPortal
Registry / portal Free registration AACR GENIE data-use terms (Synapse registration)

The largest real-world clinical sequencing registry in oncology — the place to test whether a variant-effect or biomarker model generalises beyond TCGA.

200 000 patients clinical-grade tumour sequencing from 19+ institutions with limited clinical data; releases twice a year

CPTAC — Clinical Proteomic Tumor Analysis Consortium

NCI Office of Cancer Clinical Proteomics Research
Registry / portal Free registration open (proteomics via PDC) / controlled (raw sequencing via dbGaP)

Proteogenomic companion to TCGA: the largest public set where protein-level measurements, genomics and images exist for the same tumours.

1 000 patients 10+ tumour types with proteogenomic profiling (proteome, phosphoproteome) on genomically characterised cases; images in TCIA
Dataset site used by models: 2 Open card →

Genome in a Bottle (GIAB)

National Institute of Standards and Technology (NIST), USA
Open download Materiał referencyjny NIST; zgoda dawców z PGP na redystrybucję komercyjną

A public benchmark resource run by NIST within a consortium of public bodies, companies and academic centres. It contains seven thoroughly characterised human genomes with benchmark variant sets and high-confidence regions, and is used to measure how accurately a bioinformatics pipeline calls variants. It is not an oncology dataset and not an imaging dataset — it is a reference point for genome readout itself. The consortium works with the GA4GH Benchmarking Team on comparison standards. Distribution: the NIST reference material shop, the Coriell Institute and the Personal Genome Project.

7 patients 7 genomy referencyjne Seven characterised genomes: the pilot genome NA12878/HG001 from the HapMap collection and two Personal Genome Project trios — Ashkenazi Jewish (HG002, HG003, HG004) and Han Chinese (HG005, HG006, HG007).
Genomics (DNA) FASTQBAMVCF
Dataset site used by models: 1 Open card →

IPD-IMGT/HLA Database

EMBL-EBI / Anthony Nolan Research Institute
Registry / portal Open download CC BY-ND 4.0 (see site terms)

The naming authority for HLA alleles: the catalogue whose size — thirty thousand names against roughly a hundred well-measured alleles — defines the central problem of this field.

30 894 named class I alleles as of June 2026: 9,279 HLA-A, 11,258 HLA-B, 9,416 HLA-C — while one patient carries at most six
Dataset site used by models: 1 Open card →

TCGA — The Cancer Genome Atlas (via NCI Genomic Data Commons)

NCI / NHGRI; hosted by the Genomic Data Commons
Registry / portal Free registration NIH GDS Policy — open tier; controlled tier via dbGaP

The reference multi-omics cancer cohort: molecular profiles, clinical outcomes and whole-slide images for 33 cancer types, downloadable through the GDC portal and API.

11 000 patients 30 000 diagnostic + tissue WSIs 33 cancer types; ~11,000 patients with molecular data; diagnostic slides for most cases; matched clinical follow-up
Dataset site used by models: 7 Open card →

Cards follow the cancer3.ai model-card standard: AI_MODEL_CARD_STANDARD.md. Corrections and new entries: contact the editorial team; every fact needs a public source.

This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.