AI in Oncology · AI Models
A researcher-grade catalog of AI models, datasets and open data needs in oncology — every card structured, sourced and dated.
Datasets, registries and benchmarks you can train or evaluate oncology models on — with access conditions, licences, sizes, annotations and the models already using them. { } export JSON
AACR Project GENIE
The largest real-world clinical sequencing registry in oncology — the place to test whether a variant-effect or biomarker model generalises beyond TCGA.
BD2013 — MHC binding affinity benchmark
The reference affinity dataset behind fifteen years of binding predictors; its size and composition are why dataset composition, not architecture, drives reported performance.
ClinicalTrials.gov
The registry that trial-matching systems such as TrialGPT retrieve from; oncology is its largest therapeutic area.
CPTAC — Clinical Proteomic Tumor Analysis Consortium
Proteogenomic companion to TCGA: the largest public set where protein-level measurements, genomics and images exist for the same tumours.
DepMap — Cancer Dependency Map
Public map of cancer vulnerabilities in cell lines — the training and validation ground for target-discovery and drug-response models.
HLA Ligand Atlas
A benign-tissue reference peptidome: the set you check a candidate neoantigen against to make sure healthy tissue does not present it too.
IEDB automated benchmark (MHC class I)
The only prospective, third-party benchmark in the field. Its eight-year summary is sobering: leading methods are statistically indistinguishable, and a new method needs about four years before enough data accumulate to judge it.
IEDB — Immune Epitope Database
The field's central repository of epitope data and the source of almost every training set for peptide–MHC models — and of their allele skew.
IPD-IMGT/HLA Database
The naming authority for HLA alleles: the catalogue whose size — thirty thousand names against roughly a hundred well-measured alleles — defines the central problem of this field.
MedQA (USMLE)
The exam-style benchmark on which Med-PaLM 2, GPT-4 and MedGemma report headline accuracy; useful for comparing models, not for judging clinical safety.
MHC Motif Atlas
The reference collection of HLA binding motifs — and the clearest picture of the field's long tail: a million measured ligands still describe barely 135 of thirty thousand alleles.
Mono-allelic HLA class I peptidome (Sarkizova / Abelin)
The engineered-cell peptidome that gave the field clean allele labels; fifteen of its alleles had no described motif before, and the panel covers at least one allele in 95% of people worldwide.
Protein Data Bank (wwPDB / RCSB)
The open archive of experimentally solved macromolecular structures, including oncology targets (kinases, KRAS, p53) and their drug complexes.
PubMed / PMC Open Access Subset
The literature corpus behind BiomedBERT, BiomedCLIP and CONCH's caption data; the entry point for any oncology NLP or vision-language pretraining.
TCGA — The Cancer Genome Atlas (via NCI Genomic Data Commons)
The reference multi-omics cancer cohort: molecular profiles, clinical outcomes and whole-slide images for 33 cancer types, downloadable through the GDC portal and API.
TCIA — The Cancer Imaging Archive
The main public archive of de-identified cancer imaging, organised into collections by disease and modality; the source of most public radiology training data in oncology.
Cards follow the cancer3.ai model-card standard: AI_MODEL_CARD_STANDARD.md. Corrections and new entries: contact the editorial team; every fact needs a public source.
This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.