{"count":8,"items":[{"cancer_slugs":["invasive-breast-carcinoma","colon-rectum-anal-canal","lung-bronchus","prostate"],"modalities":["clinical-text"],"slug":"polish-oncology-clinical-text","sources":[{"label":"BiomedBERT \u2014 English-only pretraining (this catalog)","url":"https://cancer3.ai/ai-oncology/models/biomedbert"}],"spec":"WHAT TO COLLECT.\n- 5,000\u201320,000 documents: histopathology reports (breast, colorectal, lung, prostate first), discharge summaries, tumour-board decisions.\n- Document type, hospital type (academic / regional) and year as metadata; no free-text identifiers.\nLABELS.\n- Span-level entities: diagnosis (ICD-O-3), pTNM, grade, receptors (ER/PR/HER2), biomarkers (KRAS, EGFR, PD-L1), treatment line, response.\n- Double annotation on \u226510% with inter-annotator agreement reported.\nFORMAT AND RELEASE.\n- JSONL with BRAT/CoNLL-style offsets; annotation guideline published with the data.\n- De-identification validated against a held-out audit; release under a data-use agreement through a Polish research institution.\nETHICS.\n- Ethics-committee approval and a GDPR Art. 89 research basis documented up front; no dates finer than month.","status":"open","title":"Polish-language oncology clinical text corpus (de-identified pathology reports and discharge summaries)","why":"Every clinical NLP model in this catalog was trained on English. Polish pathology reports, tumour-board notes and discharge summaries have their own vocabulary, inflection and templates; extracting stage, receptor status or treatment lines from them today requires hand-built rules. A public, consented corpus would let any team fine-tune BiomedBERT-class encoders or evaluate LLM extraction in Polish."},{"cancer_slugs":["invasive-breast-carcinoma","malignant-melanoma","lung-bronchus"],"modalities":["histopathology"],"slug":"wsi-treatment-response","sources":[{"label":"CHIEF \u2014 survival prediction relies on retrospective cohorts (this catalog)","url":"https://cancer3.ai/ai-oncology/models/chief"}],"spec":"WHAT TO COLLECT.\n- Pre-treatment diagnostic H&E WSIs (\u22651,000 patients across \u22653 centres and \u22652 scanners) for one indication at a time: e.g. neoadjuvant chemotherapy in triple-negative breast cancer, anti-PD-1 in melanoma or NSCLC.\nLABELS.\n- Pathological complete response (RCB class) or RECIST best response and progression-free survival with follow-up \u226524 months.\n- Treatment regimen, stage, receptor/biomarker status as covariates.\nFORMAT.\n- Pyramidal TIFF/SVS at 20\u00d7 or 40\u00d7 with scanner and stain batch recorded; clinical table in CSV with a data dictionary.\nRELEASE.\n- De-identified slides under a DUA; a held-out test centre kept private for a leaderboard.","status":"open","title":"Multi-centre H&E slides with treatment-response labels (neoadjuvant and immunotherapy)","why":"Pathology foundation models are benchmarked on diagnosis and grading, where public labels exist. The clinically valuable question \u2014 will this patient respond to this treatment \u2014 has almost no public whole-slide data with outcomes, so every response model is single-centre and unverifiable."},{"cancer_slugs":["lung-bronchus"],"modalities":["radiology-ct"],"slug":"european-ldct-screening-cohort","sources":[{"label":"Sybil \u2014 trained on NLST (this catalog)","url":"https://cancer3.ai/ai-oncology/models/sybil"},{"label":"NLST data card","url":"https://cancer3.ai/ai-oncology/datasets/nlst"}],"spec":"WHAT TO COLLECT.\n- \u226510,000 participants with baseline and follow-up LDCT (DICOM, thin slice \u22641.5 mm), 3+ years of cancer-registry-linked outcomes.\n- Acquisition parameters per scan; smoking history, age, sex.\nLABELS.\n- Lung cancer diagnosis (date, histology, stage) from the national cancer registry; nodule Lung-RADS category per screen.\nRELEASE.\n- Federated access or a controlled-access repository (EHDS-compatible) with a standard DUA; synthetic sample for tooling.","status":"open","title":"European low-dose CT screening cohort with longitudinal outcomes for lung-cancer risk models","why":"Sybil and similar risk models were trained on NLST (US, 2002\u20132009 scanners, heavy smokers). Europe is rolling out LDCT screening with different populations, protocols and never-smoker inclusion; without a shareable European cohort with outcomes, no model can be validated or recalibrated for these programmes."},{"cancer_slugs":["soft-tissue","net-nec-various-sites","cns-embryonal-tumours","bone-articular-cartilage"],"modalities":["radiology-mri","radiology-ct","histopathology","genomics"],"slug":"rare-cancer-paired-imaging-genomics","sources":[{"label":"Virchow \u2014 rare-cancer detection motivation (this catalog)","url":"https://cancer3.ai/ai-oncology/models/virchow"}],"spec":"WHAT TO COLLECT.\n- Per patient: diagnostic imaging (CT/MRI), H&E slide, targeted or whole-exome sequencing, and treatment/outcome \u2014 for \u2265200 patients per entity, pooled across reference centres and ERN networks.\nLABELS.\n- WHO 2020+ histological subtype with central review; molecular class; survival.\nRELEASE.\n- Controlled access with a common minimal data model; federated learning setup acceptable where transfer is impossible.","status":"open","title":"Paired radiology, pathology and genomics for rare cancers (sarcoma, neuroendocrine, paediatric CNS)","why":"Foundation models degrade on rare entities because public cohorts contain a few dozen cases each. Rare cancers are exactly where a second opinion from a model would help most \u2014 and where paired modalities are needed to make any prediction trustworthy."},{"cancer_slugs":["malignant-melanoma","lung-bronchus","kidney"],"modalities":["genomics","transcriptomics","proteomics"],"slug":"immunotherapy-response-multiomics","sources":[{"label":"AACR GENIE \u2014 real-world sequencing without immunotherapy outcomes at scale (this catalog)","url":"https://cancer3.ai/ai-oncology/datasets/aacr-genie"}],"spec":"WHAT TO COLLECT.\n- \u2265500 patients treated with anti-PD-1/PD-L1 (melanoma, NSCLC, RCC): tumour WES/RNA-seq, HLA typing, blood TCR-seq at baseline, PD-L1 IHC.\nLABELS.\n- RECIST response, PFS/OS with \u226524 months follow-up, immune-related adverse events.\nFORMAT.\n- Processed matrices (variants MAF, expression TPM, TCR clonotypes) plus raw reads under controlled access.","status":"open","title":"Immunotherapy response cohort with tumour sequencing, TCR repertoire and outcomes","why":"Checkpoint inhibitors help a minority of patients, and biomarkers (PD-L1, TMB) are weak. Models that predict response from neoantigen load, HLA type and T-cell repertoire exist only on small private sets; a public cohort with standardised outcomes would make the field comparable."},{"cancer_slugs":["malignant-melanoma","non-melanoma-skin-cancer-nmsc"],"modalities":["dermoscopy"],"slug":"diverse-skin-tone-dermoscopy","sources":[{"label":"ISIC Archive data card (this catalog)","url":"https://cancer3.ai/ai-oncology/datasets/isic-archive"}],"spec":"WHAT TO COLLECT.\n- \u22655,000 lesions with Fitzpatrick IV\u2013VI representation \u226540%, dermoscopic + clinical photo pairs, standardised colour calibration.\nLABELS.\n- Histopathology-confirmed diagnosis for every excised lesion; Fitzpatrick type and Monk skin tone recorded; anatomical site.\nRELEASE.\n- CC BY-NC with a benchmark split stratified by skin tone.","status":"open","title":"Dermoscopy and clinical photographs across all skin tones with histology-confirmed diagnoses","why":"The ISIC archive, the basis of nearly every skin-cancer classifier, is dominated by light skin. Melanoma in darker skin presents differently (acral, subungual) and is diagnosed later; models trained on ISIC have no evidence of working for those patients."},{"cancer_slugs":["pan-cancer"],"modalities":["immunopeptidomics"],"slug":"hla-tail-monoallelic-data","sources":[{"label":"cancer3.ai \u2014 Trzydzie\u015bci cztery litery zamka","url":"https://cancer3.ai/blog/modele-prezentacji-antygenu"},{"label":"Trevizani R et al. IEDB benchmark analysis. Brief Bioinform 2022","url":"https://academic.oup.com/bib/article/23/4/bbac259/6632617"}],"spec":"WHAT TO COLLECT.\n- Mono-allelic peptidomes (B721.221 or equivalent) for 50\u2013100 alleles common in African, South and East Asian, Indigenous American and Middle Eastern populations and absent from current panels; prioritise by population frequency \u00d7 current data gap.\nLABELS.\n- Peptide sequences with the introduced allele as the label, length distribution, and the derived motif; report the elution protocol (acid, solvent concentration, FDR) because those parameters alone shift the apparent repertoire.\nCONTROLS.\n- Include a cysteine-aware control (e.g. a genetic-screen comparison) so the known 5\u201310\u00d7 under-detection of cysteine peptides can be corrected rather than learned as biology.\nFORMAT AND RELEASE.\n- Peptide tables in CSV with raw spectra deposited in PRIDE/MassIVE; CC BY, so every model can train on them.","status":"open","title":"Mono-allelic immunopeptidomes for the long tail of HLA alleles (non-European ancestry first)","why":"Thirty thousand class I alleles are named; about 135 have measured ligands, and HLA-A*02:01 alone is a quarter of the benchmark data. Pan-allele models cover the rest by analogy, which is prediction, not knowledge \u2014 and the patients whose alleles were never measured are exactly the ones a personalised vaccine pipeline serves worst."},{"cancer_slugs":["pan-cancer"],"modalities":["immunopeptidomics","protein-sequence"],"slug":"immunogenicity-negatives","sources":[{"label":"cancer3.ai \u2014 Trzydzie\u015bci cztery litery zamka","url":"https://cancer3.ai/blog/modele-prezentacji-antygenu"},{"label":"Paul S et al. Benchmarking predictions of MHC class I restricted T cell epitopes. PLoS Comput Biol 2020","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC7274474/"}],"spec":"WHAT TO COLLECT.\n- For \u22652,000 peptides confirmed present on the cell surface (MS-verified), a matched T-cell assay result in donors of known HLA type: responder / non-responder with the assay and threshold stated.\n- Include the patient's TCR repertoire where possible, and note prior therapy.\nLABELS.\n- Binary immunogenicity plus the raw readout (spot counts, tetramer frequency), never only the call.\nWHY THIS SHAPE.\n- Negatives must be peptides that genuinely had the chance to be positive \u2014 random proteome fragments do not teach the difference between 'presented' and 'recognised'.\nRELEASE.\n- CC BY tables plus deposited raw data; a held-out split kept private for a public leaderboard.","status":"open","title":"Experimentally tested negative peptides: presented but not immunogenic","why":"Presentation models answer question two of three; nobody answers question three well. The reason is structural: mass spectrometry never reports true negatives, so 94% of a typical training set is assumed negatives, and there is almost no public data on peptides that were presented and still failed to trigger a T-cell response."}]}
