AI in Oncology · AI Models
A researcher-grade catalog of AI models, datasets and open data needs in oncology — every card structured, sourced and dated.
Filter by task, data type, cancer, availability and regulatory status. Each card follows one model-card standard and links to Hugging Face, code, papers and the datasets it was trained or tested on. { } export JSON
BiomedBERT (PubMedBERT)
BERT pretrained from scratch on PubMed abstracts and PMC full text with a biomedical vocabulary; the workhorse encoder for named-entity recognition, relation extraction and classification over oncology literature and reports.
BiomedCLIP
CLIP-style vision-language model pretrained on PMC-15M — 15 million figure–caption pairs from biomedical papers — with a PubMedBERT text tower and a ViT-B image tower; supports zero-shot classification and retrieval across pathology, radiology and more.
CHIEF
Clinical Histopathology Imaging Evaluation Foundation model: trained on 60,530 whole-slide images across 19 anatomical sites and validated on 19,491 slides from 24 hospitals for cancer detection, tumour-origin prediction, genomic profiling and survival.
CONCH CONCH (v1)
Vision-language foundation model for pathology: an image encoder and a text encoder trained together on 1.17 million histopathology image–caption pairs, enabling zero-shot classification and image–text retrieval without labelled slides.
ESM-2 / ESMFold
Protein language models from 8 million to 15 billion parameters trained on UniRef sequences; embeddings power variant-effect and function prediction, and ESMFold predicts structure directly from a single sequence without MSAs.
Geneformer
Transformer pretrained on about 30 million single-cell transcriptomes (Genecorpus-30M) that learns gene-network context from ranked expression, enabling few-shot prediction of gene dosage effects, cell states and candidate therapeutic targets.
H-optimus-0
1.1-billion-parameter ViT-g/14 pathology encoder trained on more than 500,000 H&E slides (hundreds of millions of tiles), released under Apache-2.0 — one of the few large pathology foundation models with a permissive licence.
Phikon-v2
ViT-L pathology encoder trained with DINOv2 on PANCAN-XL — 456 million tiles from 58,359 whole-slide images that mix public cohorts (TCGA, CPTAC, GTEx and others) with private data — positioned for biomarker prediction.
Prov-GigaPath
Whole-slide foundation model with 1.3 billion parameters, pretrained on 1.3 billion tiles from 171,189 slides of real-world clinical data; pairs a DINOv2 tile encoder with a LongNet slide encoder that reasons over an entire slide.
scGPT
Generative pretrained transformer for single-cell multi-omics, trained on over 33 million cells, supporting cell-type annotation, batch integration, perturbation response prediction and gene-network inference.
Virchow2
Successor to Virchow: ViT-H/14 pretrained on 3.1 million whole-slide images from about 225,000 patients across 45 countries, at mixed magnifications (5×–40×), with pathology-specific augmentations.
UNI UNI (v1)
General-purpose self-supervised vision encoder for H&E histopathology tiles, pretrained on more than 100 million tiles from over 100,000 whole-slide images; the reference foundation model for pathology feature extraction.
Virchow Virchow (v1)
632-million-parameter vision transformer pretrained on 1.5 million whole-slide images from about 100,000 patients — the largest pathology pretraining set at its release — and used to build a pan-cancer detection model covering 17 cancer types, including rare ones.
Cards follow the cancer3.ai model-card standard: AI_MODEL_CARD_STANDARD.md. Corrections and new entries: contact the editorial team; every fact needs a public source.
This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.