Cancer3.AIAI in OncologyAI Models › Datasets

AI in Oncology · AI Models

A researcher-grade catalog of AI models, datasets and open data needs in oncology — every card structured, sourced and dated.

Datasets, registries and benchmarks you can train or evaluate oncology models on — with access conditions, licences, sizes, annotations and the models already using them. { } export JSON

× Clear filters
6 / 26 datasets

IEDB automated benchmark (MHC class I)

La Jolla Institute for Immunology
Benchmark / challenge Open download public

The only prospective, third-party benchmark in the field. Its eight-year summary is sobering: leading methods are statistically indistinguishable, and a new method needs about four years before enough data accumulate to judge it.

runs continuously on newly deposited data, before it can leak into anyone's training set
Dataset site used by models: 3 Open card →

IEDB — Immune Epitope Database

La Jolla Institute for Immunology, funded by NIAID
Registry / portal Open download free to use; NIAID-funded public resource

The field's central repository of epitope data and the source of almost every training set for peptide–MHC models — and of their allele skew.

1 600 000 epitope-related records curated from published literature and direct submissions; includes MHC binding assays, MS-eluted ligands and T-cell assays
Dataset site used by models: 4 Open card →

IPD-IMGT/HLA Database

EMBL-EBI / Anthony Nolan Research Institute
Registry / portal Open download CC BY-ND 4.0 (see site terms)

The naming authority for HLA alleles: the catalogue whose size — thirty thousand names against roughly a hundred well-measured alleles — defines the central problem of this field.

30 894 named class I alleles as of June 2026: 9,279 HLA-A, 11,258 HLA-B, 9,416 HLA-C — while one patient carries at most six
Dataset site used by models: 1 Open card →

MHC Motif Atlas

Gfeller lab, University of Lausanne
Dataset Open download free for academic use

The reference collection of HLA binding motifs — and the clearest picture of the field's long tail: a million measured ligands still describe barely 135 of thirty thousand alleles.

1 000 000 ligands over a million ligands — but spread across only about 135 class I molecules, against 30,894 named class I alleles in IPD-IMGT/HLA (June 2026)
Dataset site used by models: 1 Open card →

Mono-allelic HLA class I peptidome (Sarkizova / Abelin)

Broad Institute / Dana-Farber Cancer Institute
Dataset Open download published supplementary data (see papers)

The engineered-cell peptidome that gave the field clean allele labels; fifteen of its alleles had no described motif before, and the panel covers at least one allele in 95% of people worldwide.

186 464 peptides 95 mono-allelic cell lines (31 HLA-A, 40 HLA-B, 21 HLA-C, 3 HLA-G), median 1,860 peptides per allele; the earlier Abelin 2017 set covered 16 alleles and >24,000 peptides
Dataset site used by models: 4 Open card →

Cards follow the cancer3.ai model-card standard: AI_MODEL_CARD_STANDARD.md. Corrections and new entries: contact the editorial team; every fact needs a public source.

This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.