Cancer3.AIAI in OncologyAI Models › Datasets

AI in Oncology · AI Models

A researcher-grade catalog of AI models, datasets and open data needs in oncology — every card structured, sourced and dated.

Datasets, registries and benchmarks you can train or evaluate oncology models on — with access conditions, licences, sizes, annotations and the models already using them. { } export JSON

× Clear filters
6 / 26 datasets

CAMELYON16 / CAMELYON17

Radboud University Medical Center and partners (grand-challenge.org)
Benchmark / challenge Open download CC0

The canonical whole-slide benchmark for breast cancer lymph-node metastasis detection; still the first sanity check for any new pathology encoder.

1 399 WSIs CAMELYON16: 399 slides (270 train / 129 test) with pixel-level metastasis annotations; CAMELYON17: 1,000 slides from 5 centres with patient-level pN stage
Dataset site used by models: 2 Open card →

CPTAC — Clinical Proteomic Tumor Analysis Consortium

NCI Office of Cancer Clinical Proteomics Research
Registry / portal Free registration open (proteomics via PDC) / controlled (raw sequencing via dbGaP)

Proteogenomic companion to TCGA: the largest public set where protein-level measurements, genomics and images exist for the same tumours.

1 000 patients 10+ tumour types with proteogenomic profiling (proteome, phosphoproteome) on genomically characterised cases; images in TCIA
Dataset site used by models: 2 Open card →

PANDA — Prostate cANcer graDe Assessment

Radboud UMC and Karolinska Institutet (Kaggle challenge)
Benchmark / challenge Free registration CC BY-NC-SA 4.0 (Kaggle competition rules)

Reference dataset and challenge for AI Gleason grading; the follow-up Nature Medicine paper validated algorithms on external cohorts.

10 616 biopsy WSIs the largest public prostate biopsy set; Gleason / ISUP grade per biopsy from two centres
Dataset site used by models: 1 Open card →

PatchCamelyon (PCam)

Veeling et al.; mirrored on Hugging Face
Benchmark / challenge Open download CC0

Small, fast, fully open patch-classification benchmark — the standard smoke test for image encoders in pathology and a common zero-shot evaluation set.

327 680 96×96 patches derived from CAMELYON16; binary label = tumour tissue in the central 32×32 region
3 508 downloads / month ♥ 12 updated 2024-05-25 Dataset site Hugging Face used by models: 1 Open card →

TCGA — The Cancer Genome Atlas (via NCI Genomic Data Commons)

NCI / NHGRI; hosted by the Genomic Data Commons
Registry / portal Free registration NIH GDS Policy — open tier; controlled tier via dbGaP

The reference multi-omics cancer cohort: molecular profiles, clinical outcomes and whole-slide images for 33 cancer types, downloadable through the GDC portal and API.

11 000 patients 30 000 diagnostic + tissue WSIs 33 cancer types; ~11,000 patients with molecular data; diagnostic slides for most cases; matched clinical follow-up
Dataset site used by models: 7 Open card →

TCIA — The Cancer Imaging Archive

NCI Cancer Imaging Program; hosted by the University of Arkansas for Medical Sciences
Registry / portal Open download mostly CC BY 3.0/4.0 per collection; some restricted collections

The main public archive of de-identified cancer imaging, organised into collections by disease and modality; the source of most public radiology training data in oncology.

200 collections hundreds of collections, tens of thousands of patients; DICOM with linked clinical and sometimes genomic data
Dataset site used by models: 1 Open card →

Cards follow the cancer3.ai model-card standard: AI_MODEL_CARD_STANDARD.md. Corrections and new entries: contact the editorial team; every fact needs a public source.

This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.