Text corpus
Open download
PubMed metadata public domain; PMC OA articles under CC licences per article
PubMed / PMC Open Access Subset
The literature corpus behind BiomedBERT, BiomedCLIP and CONCH's caption data; the entry point for any oncology NLP or vision-language pretraining.
At a glance
ProviderUS National Library of Medicine
AccessOpen download
LicencePubMed metadata public domain; PMC OA articles under CC licences per article
citations (PubMed); millions of full texts in PMC OA36 000 000
Size notesabstracts for all of PubMed; full text and figures for the open-access subset
Formatstext, JSON, PNG
Data typeBiomedical literature
CancerPan-cancer
Labels and annotations
MeSH indexing, article metadata, figure captions; no oncology-specific labels beyond MeSH.
Details
Not yet documented on this card.
Models trained or evaluated on it
- pretraining BiomedCLIP Microsoft Research
- pretraining BiomedBERT (PubMedBERT) Microsoft Research
- pretraining CONCH Mahmood Lab, Brigham and Women's Hospital / Harvard Medical School — figure–caption pairs from the PMC open-access subset
- training CONCH Mahmood Lab, Brigham and Women's Hospital / Harvard Medical School
- training BiomedCLIP Microsoft Research
Sources
This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.