{"count":7,"items":[{"access":"open","article":null,"cancer_slugs":["pan-cancer"],"doi":null,"formats":["PDB","mmCIF"],"hf":null,"huggingface":null,"kind":"registry","license":"CC0","modalities":["protein-sequence","molecules"],"name":"Protein Data Bank (wwPDB / RCSB)","page":"/ai-oncology/datasets/pdb","provider":"Worldwide Protein Data Bank","size":{"items":220000,"notes_en":"X-ray, cryo-EM and NMR structures of proteins, nucleic acids and complexes; the training ground of AlphaFold","notes_pl":"struktury rentgenowskie, krio-EM i NMR bia\u0142ek, kwas\u00f3w nukleinowych i kompleks\u00f3w; grunt treningowy AlphaFold","unit":"experimental structures"},"slug":"pdb","summary":"The open archive of experimentally solved macromolecular structures, including oncology targets (kinases, KRAS, p53) and their drug complexes.","tasks":["protein-structure","drug-discovery"],"url":"https://www.rcsb.org/","verified_at":"2026-09-05T22:25:59.035915"},{"access":"open","article":"/blog/modele-prezentacji-antygenu","cancer_slugs":["pan-cancer"],"doi":null,"formats":["CSV","JSON"],"hf":null,"huggingface":null,"kind":"registry","license":"free to use; NIAID-funded public resource","modalities":["protein-sequence","immunopeptidomics"],"name":"IEDB \u2014 Immune Epitope Database","page":"/ai-oncology/datasets/iedb","provider":"La Jolla Institute for Immunology, funded by NIAID","size":{"items":1600000,"notes_en":"curated from published literature and direct submissions; includes MHC binding assays, MS-eluted ligands and T-cell assays","notes_pl":"kuratorowane z literatury i zg\u0142osze\u0144 bezpo\u015brednich; zawiera testy wi\u0105zania MHC, ligandy ze spektrometru i testy limfocyt\u00f3w T","unit":"epitope-related records"},"slug":"iedb","summary":"The field's central repository of epitope data and the source of almost every training set for peptide\u2013MHC models \u2014 and of their allele skew.","tasks":["peptide-mhc-binding","antigen-presentation"],"url":"https://www.iedb.org/","verified_at":"2026-09-05T22:25:59.114121"},{"access":"open","article":"/blog/modele-prezentacji-antygenu","cancer_slugs":["pan-cancer"],"doi":null,"formats":["CSV"],"hf":null,"huggingface":null,"kind":"benchmark","license":"public","modalities":["protein-sequence"],"name":"IEDB automated benchmark (MHC class I)","page":"/ai-oncology/datasets/iedb-benchmark","provider":"La Jolla Institute for Immunology","size":{"notes_en":"runs continuously on newly deposited data, before it can leak into anyone's training set","notes_pl":"dzia\u0142a na bie\u017c\u0105co na \u015bwie\u017co deponowanych danych, zanim mog\u0105 trafi\u0107 do czyjegokolwiek zbioru treningowego"},"slug":"iedb-benchmark","summary":"The only prospective, third-party benchmark in the field. Its eight-year summary is sobering: leading methods are statistically indistinguishable, and a new method needs about four years before enough data accumulate to judge it.","tasks":["peptide-mhc-binding","antigen-presentation"],"url":"http://tools.iedb.org/auto_bench/mhci/weekly/","verified_at":"2026-09-05T22:25:59.140589"},{"access":"open","article":"/blog/modele-prezentacji-antygenu","cancer_slugs":["pan-cancer"],"doi":"10.1038/s41587-019-0322-9","formats":["CSV","text"],"hf":null,"huggingface":null,"kind":"dataset","license":"published supplementary data (see papers)","modalities":["immunopeptidomics","protein-sequence"],"name":"Mono-allelic HLA class I peptidome (Sarkizova / Abelin)","page":"/ai-oncology/datasets/monoallelic-peptidome","provider":"Broad Institute / Dana-Farber Cancer Institute","size":{"items":186464,"notes_en":"95 mono-allelic cell lines (31 HLA-A, 40 HLA-B, 21 HLA-C, 3 HLA-G), median 1,860 peptides per allele; the earlier Abelin 2017 set covered 16 alleles and >24,000 peptides","notes_pl":"95 linii monoallelicznych (31 HLA-A, 40 HLA-B, 21 HLA-C, 3 HLA-G), mediana 1860 peptyd\u00f3w na allel; wcze\u015bniejszy zbi\u00f3r Abelin 2017 obj\u0105\u0142 16 alleli i ponad 24 000 peptyd\u00f3w","unit":"peptides"},"slug":"monoallelic-peptidome","summary":"The engineered-cell peptidome that gave the field clean allele labels; fifteen of its alleles had no described motif before, and the panel covers at least one allele in 95% of people worldwide.","tasks":["antigen-presentation","peptide-mhc-binding"],"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC7008090/","verified_at":"2026-09-05T22:25:59.152441"},{"access":"open","article":"/blog/modele-prezentacji-antygenu","cancer_slugs":["pan-cancer"],"doi":null,"formats":["CSV","PNG"],"hf":null,"huggingface":null,"kind":"dataset","license":"free for academic use","modalities":["immunopeptidomics","protein-sequence"],"name":"MHC Motif Atlas","page":"/ai-oncology/datasets/mhc-motif-atlas","provider":"Gfeller lab, University of Lausanne","size":{"items":1000000,"notes_en":"over a million ligands \u2014 but spread across only about 135 class I molecules, against 30,894 named class I alleles in IPD-IMGT/HLA (June 2026)","notes_pl":"ponad milion ligand\u00f3w \u2014 ale roz\u0142o\u017conych na zaledwie oko\u0142o 135 cz\u0105steczek klasy I, wobec 30 894 nazwanych alleli klasy I w IPD-IMGT/HLA (czerwiec 2026)","unit":"ligands"},"slug":"mhc-motif-atlas","summary":"The reference collection of HLA binding motifs \u2014 and the clearest picture of the field's long tail: a million measured ligands still describe barely 135 of thirty thousand alleles.","tasks":["antigen-presentation"],"url":"http://mhcmotifatlas.org/","verified_at":"2026-09-05T22:25:59.213233"},{"access":"open","article":"/blog/modele-prezentacji-antygenu","cancer_slugs":["pan-cancer"],"doi":null,"formats":["text","CSV"],"hf":null,"huggingface":null,"kind":"registry","license":"CC BY-ND 4.0 (see site terms)","modalities":["genomics","protein-sequence"],"name":"IPD-IMGT/HLA Database","page":"/ai-oncology/datasets/ipd-imgt-hla","provider":"EMBL-EBI / Anthony Nolan Research Institute","size":{"items":30894,"notes_en":"as of June 2026: 9,279 HLA-A, 11,258 HLA-B, 9,416 HLA-C \u2014 while one patient carries at most six","notes_pl":"stan na czerwiec 2026: 9279 HLA-A, 11 258 HLA-B, 9416 HLA-C \u2014 a jeden pacjent ma najwy\u017cej sze\u015b\u0107","unit":"named class I alleles"},"slug":"ipd-imgt-hla","summary":"The naming authority for HLA alleles: the catalogue whose size \u2014 thirty thousand names against roughly a hundred well-measured alleles \u2014 defines the central problem of this field.","tasks":["antigen-presentation"],"url":"https://hla.alleles.org/","verified_at":"2026-09-05T22:25:59.303091"},{"access":"open","article":"/blog/modele-prezentacji-antygenu","cancer_slugs":["pan-cancer"],"doi":null,"formats":["CSV"],"hf":null,"huggingface":null,"kind":"benchmark","license":"public","modalities":["protein-sequence"],"name":"BD2013 \u2014 MHC binding affinity benchmark","page":"/ai-oncology/datasets/bd2013","provider":"IEDB / Kim et al.","size":{"items":176161,"notes_en":"114 alleles across six species \u2014 the historical training core for binding predictors, and a reminder of how small the measured world is","notes_pl":"114 alleli w sze\u015bciu gatunkach \u2014 historyczny rdze\u0144 treningowy modeli wi\u0105zania i przypomnienie, jak ma\u0142y jest \u015bwiat zmierzony","unit":"affinity measurements"},"slug":"bd2013","summary":"The reference affinity dataset behind fifteen years of binding predictors; its size and composition are why dataset composition, not architecture, drives reported performance.","tasks":["peptide-mhc-binding"],"url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC4111843/","verified_at":"2026-09-05T22:25:59.496418"}]}
