AI in Oncology · AI Models
A researcher-grade catalog of AI models, datasets and open data needs in oncology — every card structured, sourced and dated.
Datasets, registries and benchmarks you can train or evaluate oncology models on — with access conditions, licences, sizes, annotations and the models already using them. { } export JSON
AACR Project GENIE
The largest real-world clinical sequencing registry in oncology — the place to test whether a variant-effect or biomarker model generalises beyond TCGA.
BD2013 — MHC binding affinity benchmark
The reference affinity dataset behind fifteen years of binding predictors; its size and composition are why dataset composition, not architecture, drives reported performance.
BraTS — Brain Tumor Segmentation Challenge
The long-running benchmark for brain tumour segmentation from MRI; nnU-Net-based methods have dominated its leaderboards.
CAMELYON16 / CAMELYON17
The canonical whole-slide benchmark for breast cancer lymph-node metastasis detection; still the first sanity check for any new pathology encoder.
Most-used open mammography set with pathology-confirmed labels; film-based, so domain shift to modern digital mammography must be handled.
ClinicalTrials.gov
The registry that trial-matching systems such as TrialGPT retrieve from; oncology is its largest therapeutic area.
CPTAC — Clinical Proteomic Tumor Analysis Consortium
Proteogenomic companion to TCGA: the largest public set where protein-level measurements, genomics and images exist for the same tumours.
DepMap — Cancer Dependency Map
Public map of cancer vulnerabilities in cell lines — the training and validation ground for target-discovery and drug-response models.
Genecorpus-30M
Genecorpus-30M is a pretraining corpus of about 30 million human single-cell transcriptomes assembled from 561 publicly available datasets, of which 27,406,217 cells passed quality filters. It was built to pretrain Geneformer (Theodoris et al., Nature 2023) and is released under Apache-2.0. Crucially for oncology use: cells with high mutational burden — malignant cells and immortalized cell lines — were deliberately excluded from the corpus.
Genome in a Bottle (GIAB)
A public benchmark resource run by NIST within a consortium of public bodies, companies and academic centres. It contains seven thoroughly characterised human genomes with benchmark variant sets and high-confidence regions, and is used to measure how accurately a bioinformatics pipeline calls variants. It is not an oncology dataset and not an imaging dataset — it is a reference point for genome readout itself. The consortium works with the GA4GH Benchmarking Team on comparison standards. Distribution: the NIST reference material shop, the Coriell Institute and the Personal Genome Project.
HLA Ligand Atlas
A benign-tissue reference peptidome: the set you check a candidate neoantigen against to make sure healthy tissue does not present it too.
IEDB automated benchmark (MHC class I)
The only prospective, third-party benchmark in the field. Its eight-year summary is sobering: leading methods are statistically indistinguishable, and a new method needs about four years before enough data accumulate to judge it.
IEDB — Immune Epitope Database
The field's central repository of epitope data and the source of almost every training set for peptide–MHC models — and of their allele skew.
IPD-IMGT/HLA Database
The naming authority for HLA alleles: the catalogue whose size — thirty thousand names against roughly a hundred well-measured alleles — defines the central problem of this field.
ISIC Archive — International Skin Imaging Collaboration
The public backbone of skin-cancer AI; strongly skewed towards light skin tones, which every model card built on it should say.
LIDC-IDRI — Lung Image Database Consortium
The standard open CT dataset for lung nodule detection and characterisation; basis of the LUNA16 challenge.
MedQA (USMLE)
The exam-style benchmark on which Med-PaLM 2, GPT-4 and MedGemma report headline accuracy; useful for comparing models, not for judging clinical safety.
MHC Motif Atlas
The reference collection of HLA binding motifs — and the clearest picture of the field's long tail: a million measured ligands still describe barely 135 of thirty thousand alleles.
Mono-allelic HLA class I peptidome (Sarkizova / Abelin)
The engineered-cell peptidome that gave the field clean allele labels; fifteen of its alleles had no described motif before, and the panel covers at least one allele in 95% of people worldwide.
NLST — National Lung Screening Trial
The trial that established LDCT screening; its images plus outcomes trained Sybil and remain the only large public-by-application longitudinal LDCT cohort.
PANDA — Prostate cANcer graDe Assessment
Reference dataset and challenge for AI Gleason grading; the follow-up Nature Medicine paper validated algorithms on external cohorts.
PatchCamelyon (PCam)
Small, fast, fully open patch-classification benchmark — the standard smoke test for image encoders in pathology and a common zero-shot evaluation set.
Protein Data Bank (wwPDB / RCSB)
The open archive of experimentally solved macromolecular structures, including oncology targets (kinases, KRAS, p53) and their drug complexes.
PubMed / PMC Open Access Subset
The literature corpus behind BiomedBERT, BiomedCLIP and CONCH's caption data; the entry point for any oncology NLP or vision-language pretraining.
TCGA — The Cancer Genome Atlas (via NCI Genomic Data Commons)
The reference multi-omics cancer cohort: molecular profiles, clinical outcomes and whole-slide images for 33 cancer types, downloadable through the GDC portal and API.
TCIA — The Cancer Imaging Archive
The main public archive of de-identified cancer imaging, organised into collections by disease and modality; the source of most public radiology training data in oncology.
Cards follow the cancer3.ai model-card standard: AI_MODEL_CARD_STANDARD.md. Corrections and new entries: contact the editorial team; every fact needs a public source.
This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.