{"architecture":{"family":"BERT-style transformer over rank-value gene encodings","input":"rank-encoded single-cell expression \u2014 2,048 genes in V1, 4,096 genes in V2","output":"cell and gene embeddings","params":"V1 ~10M; V2 checkpoints of 104M and 316M parameters","pretraining":"masked gene prediction on Genecorpus-30M (V1) or Genecorpus-104M (V2); a separate V2 variant is continually trained on about 14 million cancer transcriptomes"},"article":null,"cancer_slugs":["pan-cancer"],"category":"genomics","confidence":"high","datasets":[{"name":"Genecorpus-30M","note":"","role":"pretraining","slug":"genecorpus-30m"}],"developer":"Theodoris Lab (Gladstone Institutes / UCSF)","evaluation":[{"benchmark":"Corpus scaling and model quantization (Nature Computational Science, 27.03.2026) \u2014 resource use during fine-tuning","external":false,"metric":"fine-tuning time and GPU memory relative to the full-precision model","source":"https://www.nature.com/articles/s43588-026-00972-4","value":"quantization preserved the contextual gene and cell embedding space while requiring 15% of the time and 34% of the memory of the full model; the pretraining corpus was expanded to more than 100 million human single-cell transcriptomes"}],"hf":{"downloads":1698,"fetched_at":"2026-09-09T21:33:08Z","gated":false,"last_modified":"2026-05-26","library":"transformers","license":"apache-2.0","likes":311,"pipeline_tag":"fill-mask"},"kind":"foundation","license":"apache-2.0","limitations":"- Rank encoding discards absolute expression; batch effects still leak in.\n- Pretraining corpus is not cancer-specific; tumour microenvironment states may need extra fine-tuning.","links":{"demo":null,"docs":null,"doi":"10.1038/s41586-023-06139-9","github":null,"huggingface":"https://huggingface.co/ctheodoris/Geneformer","paper":"https://www.nature.com/articles/s41586-023-06139-9","pmid":null},"modalities":["single-cell","transcriptomics"],"name":"Geneformer","notable_uses":"","openness":"open-weights","regulatory":{"intended_use_en":"Research tool.","intended_use_pl":"Narz\u0119dzie badawcze.","source_url":"https://huggingface.co/ctheodoris/Geneformer","status":"not-applicable"},"regulatory_status":"not-applicable","release_date":"2023-05-31","run_snippet":"# generic transformers loader \u2014 see the model card for the task-specific head and preprocessing\nfrom transformers import AutoModel, AutoProcessor\nmodel = AutoModel.from_pretrained('ctheodoris/Geneformer')\nprocessor = AutoProcessor.from_pretrained('ctheodoris/Geneformer')\n","settings":["basic-research","drug-discovery"],"slug":"geneformer","sources":[{"label":"Theodoris CV et al. Transfer learning enables predictions in network biology. Nature 2023","url":"https://www.nature.com/articles/s41586-023-06139-9"},{"label":"Chen H, Venkatesh MS et al. Scaling and quantization of large-scale foundation model enables resource-efficient predictions in network biology. Nature Computational Science, 27.03.2026 (PMID 41896605)","url":"https://www.nature.com/articles/s43588-026-00972-4"},{"label":"Hugging Face \u2014 ctheodoris/Geneformer (karta modelu: warianty V1/V2, Genecorpus-104M, wariant nowotworowy CLcancer)","url":"https://huggingface.co/ctheodoris/Geneformer"}],"summary":"Transformer pretrained on about 30 million single-cell transcriptomes (Genecorpus-30M) that learns gene-network context from ranked expression, enabling few-shot prediction of gene dosage effects, cell states and candidate therapeutic targets.","tasks":["single-cell","gene-expression","feature-extraction","drug-discovery"],"training":{"size":"~30M cells (V1) / ~104M cells (V2) / ~14M cancer cells (continual-learning variant)","summary_en":"V1 (June 2021 corpus, 2023 release): Genecorpus-30M \u2014 about 29.9 million human single-cell transcriptomes from public datasets across tissues, not restricted to cancer. V2 (December 2024): Genecorpus-104M \u2014 about 104 million human single-cell transcriptomes, with 104M- and 316M-parameter checkpoints and an input window of 4,096 genes. A separate V2 variant (Geneformer-V2-104M_CLcancer) is continually trained on about 14 million cancer transcriptomes. Corpus composition per the developers' model card; the underlying datasets are not published as a single downloadable collection.","summary_pl":"V1 (korpus z czerwca 2021 r., publikacja 2023): Genecorpus-30M \u2014 ok. 29,9 mln ludzkich transkryptom\u00f3w pojedynczych kom\u00f3rek z publicznych zbior\u00f3w, z wielu tkanek, bez ograniczenia do nowotwor\u00f3w. V2 (grudzie\u0144 2024): Genecorpus-104M \u2014 ok. 104 mln transkryptom\u00f3w, warianty o 104 mln i 316 mln parametr\u00f3w, okno wej\u015bciowe 4096 gen\u00f3w. Osobny wariant V2 (Geneformer-V2-104M_CLcancer) jest douczany na ok. 14 mln transkryptom\u00f3w nowotworowych. Sk\u0142ad korpusu podajemy za kart\u0105 modelu autor\u00f3w; same zbiory \u017ar\u00f3d\u0142owe nie s\u0105 udost\u0119pnione jako jedna pobieralna kolekcja."},"updated_at":"2026-09-09T21:33:09.703720","url":"/ai-oncology/models/geneformer","usage":{"library":"transformers + geneformer package","notes_en":"Install from the Hugging Face repo; tokenize with the provided gene-median dictionary. Fine-tuning on a few hundred labelled cells is the intended workflow."},"verified_at":"2026-09-07T07:05:05.599962","verified_by":"Occe3C","version":null,"what_it_does":"Each cell becomes a sequence of genes ranked by expression; the model is fine-tuned for cell-type classification, disease-state prediction and in-silico perturbation (which gene knock-out shifts a malignant state)."}
