AI in Oncology · AI Models
A researcher-grade catalog of AI models, datasets and open data needs in oncology — every card structured, sourced and dated.
Problems in oncology AI that are blocked on data. Each entry says why it matters and what exactly to collect — a starting point for anyone who wants to build a dataset. { } export JSON
Polish-language oncology clinical text corpus (de-identified pathology reports and discharge summaries)
Why it matters
Every clinical NLP model in this catalog was trained on English. Polish pathology reports, tumour-board notes and discharge summaries have their own vocabulary, inflection and templates; extracting stage, receptor status or treatment lines from them today requires hand-built rules. A public, consented corpus would let any team fine-tune BiomedBERT-class encoders or evaluate LLM extraction in Polish.
What to collect
What to collect
- 5,000–20,000 documents: histopathology reports (breast, colorectal, lung, prostate first), discharge summaries, tumour-board decisions.
- Document type, hospital type (academic / regional) and year as metadata; no free-text identifiers.
Labels
- Span-level entities: diagnosis (ICD-O-3), pTNM, grade, receptors (ER/PR/HER2), biomarkers (KRAS, EGFR, PD-L1), treatment line, response.
- Double annotation on ≥10% with inter-annotator agreement reported.
Format and release
- JSONL with BRAT/CoNLL-style offsets; annotation guideline published with the data.
- De-identification validated against a held-out audit; release under a data-use agreement through a Polish research institution.
Ethics
- Ethics-committee approval and a GDPR Art. 89 research basis documented up front; no dates finer than month.
Multi-centre H&E slides with treatment-response labels (neoadjuvant and immunotherapy)
Why it matters
Pathology foundation models are benchmarked on diagnosis and grading, where public labels exist. The clinically valuable question — will this patient respond to this treatment — has almost no public whole-slide data with outcomes, so every response model is single-centre and unverifiable.
What to collect
What to collect
- Pre-treatment diagnostic H&E WSIs (≥1,000 patients across ≥3 centres and ≥2 scanners) for one indication at a time: e.g. neoadjuvant chemotherapy in triple-negative breast cancer, anti-PD-1 in melanoma or NSCLC.
Labels
- Pathological complete response (RCB class) or RECIST best response and progression-free survival with follow-up ≥24 months.
- Treatment regimen, stage, receptor/biomarker status as covariates.
Format
- Pyramidal TIFF/SVS at 20× or 40× with scanner and stain batch recorded; clinical table in CSV with a data dictionary.
Release
- De-identified slides under a DUA; a held-out test centre kept private for a leaderboard.
European low-dose CT screening cohort with longitudinal outcomes for lung-cancer risk models
Why it matters
Sybil and similar risk models were trained on NLST (US, 2002–2009 scanners, heavy smokers). Europe is rolling out LDCT screening with different populations, protocols and never-smoker inclusion; without a shareable European cohort with outcomes, no model can be validated or recalibrated for these programmes.
What to collect
What to collect
- ≥10,000 participants with baseline and follow-up LDCT (DICOM, thin slice ≤1.5 mm), 3+ years of cancer-registry-linked outcomes.
- Acquisition parameters per scan; smoking history, age, sex.
Labels
- Lung cancer diagnosis (date, histology, stage) from the national cancer registry; nodule Lung-RADS category per screen.
Release
- Federated access or a controlled-access repository (EHDS-compatible) with a standard DUA; synthetic sample for tooling.
Paired radiology, pathology and genomics for rare cancers (sarcoma, neuroendocrine, paediatric CNS)
Why it matters
Foundation models degrade on rare entities because public cohorts contain a few dozen cases each. Rare cancers are exactly where a second opinion from a model would help most — and where paired modalities are needed to make any prediction trustworthy.
What to collect
What to collect
- Per patient: diagnostic imaging (CT/MRI), H&E slide, targeted or whole-exome sequencing, and treatment/outcome — for ≥200 patients per entity, pooled across reference centres and ERN networks.
Labels
- WHO 2020+ histological subtype with central review; molecular class; survival.
Release
- Controlled access with a common minimal data model; federated learning setup acceptable where transfer is impossible.
Immunotherapy response cohort with tumour sequencing, TCR repertoire and outcomes
Why it matters
Checkpoint inhibitors help a minority of patients, and biomarkers (PD-L1, TMB) are weak. Models that predict response from neoantigen load, HLA type and T-cell repertoire exist only on small private sets; a public cohort with standardised outcomes would make the field comparable.
What to collect
What to collect
- ≥500 patients treated with anti-PD-1/PD-L1 (melanoma, NSCLC, RCC): tumour WES/RNA-seq, HLA typing, blood TCR-seq at baseline, PD-L1 IHC.
Labels
- RECIST response, PFS/OS with ≥24 months follow-up, immune-related adverse events.
Format
- Processed matrices (variants MAF, expression TPM, TCR clonotypes) plus raw reads under controlled access.
Dermoscopy and clinical photographs across all skin tones with histology-confirmed diagnoses
Why it matters
The ISIC archive, the basis of nearly every skin-cancer classifier, is dominated by light skin. Melanoma in darker skin presents differently (acral, subungual) and is diagnosed later; models trained on ISIC have no evidence of working for those patients.
What to collect
What to collect
- ≥5,000 lesions with Fitzpatrick IV–VI representation ≥40%, dermoscopic + clinical photo pairs, standardised colour calibration.
Labels
- Histopathology-confirmed diagnosis for every excised lesion; Fitzpatrick type and Monk skin tone recorded; anatomical site.
Release
- CC BY-NC with a benchmark split stratified by skin tone.
Mono-allelic immunopeptidomes for the long tail of HLA alleles (non-European ancestry first)
Why it matters
Thirty thousand class I alleles are named; about 135 have measured ligands, and HLA-A*02:01 alone is a quarter of the benchmark data. Pan-allele models cover the rest by analogy, which is prediction, not knowledge — and the patients whose alleles were never measured are exactly the ones a personalised vaccine pipeline serves worst.
What to collect
What to collect
- Mono-allelic peptidomes (B721.221 or equivalent) for 50–100 alleles common in African, South and East Asian, Indigenous American and Middle Eastern populations and absent from current panels; prioritise by population frequency × current data gap.
Labels
- Peptide sequences with the introduced allele as the label, length distribution, and the derived motif; report the elution protocol (acid, solvent concentration, FDR) because those parameters alone shift the apparent repertoire.
Controls
- Include a cysteine-aware control (e.g. a genetic-screen comparison) so the known 5–10× under-detection of cysteine peptides can be corrected rather than learned as biology.
Format and release
- Peptide tables in CSV with raw spectra deposited in PRIDE/MassIVE; CC BY, so every model can train on them.
Experimentally tested negative peptides: presented but not immunogenic
Why it matters
Presentation models answer question two of three; nobody answers question three well. The reason is structural: mass spectrometry never reports true negatives, so 94% of a typical training set is assumed negatives, and there is almost no public data on peptides that were presented and still failed to trigger a T-cell response.
What to collect
What to collect
- For ≥2,000 peptides confirmed present on the cell surface (MS-verified), a matched T-cell assay result in donors of known HLA type: responder / non-responder with the assay and threshold stated.
- Include the patient's TCR repertoire where possible, and note prior therapy.
Labels
- Binary immunogenicity plus the raw readout (spot counts, tetramer frequency), never only the call.
Why this shape
- Negatives must be peptides that genuinely had the chance to be positive — random proteome fragments do not teach the difference between 'presented' and 'recognised'.
Release
- CC BY tables plus deposited raw data; a held-out split kept private for a public leaderboard.
Cards follow the cancer3.ai model-card standard: AI_MODEL_CARD_STANDARD.md. Corrections and new entries: contact the editorial team; every fact needs a public source.
This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.