{"architecture":{"backbone":"one hidden layer of 56 or 66 neurons; ensemble (4.0: 100 networks \u2014 2 architectures \u00d7 5 cross-validation splits \u00d7 10 seeds)","family":"feed-forward neural network ensemble (NNAlign_MA)","input":"9-aa binding core + 34-aa HLA pseudosequence, BLOSUM50-encoded (43 \u00d7 20 = 860 numbers) plus length/insertion/deletion features","notes_en":"PSEUDOSEQUENCE.\n- 34 positions chosen from crystal structures: within 4 \u00c5 of the peptide AND polymorphic across alleles. An allele stops being a name and becomes a 34-letter word, which is what makes prediction possible for alleles with zero measurements.\nVARIABLE LENGTH.\n- Everything is reduced to a 9-mer core by searching over insertions and deletions and keeping the best-scoring alignment; the model reports which core it chose.\nTWO HEADS, ONE SHARED LAYER.\n- Affinity measurements and eluted ligands each have their own output neuron but share the hidden layer, so the wide allele coverage of in-vitro data and the true presentation signal of mass spectrometry train one representation.\nMOTIF DECONVOLUTION (NNAlign_MA).\n- After 20 warm-up iterations on single-allele data, each peptide from a multi-allele sample is assigned to the best-scoring allele among the ones that sample actually carries, with score standardisation so a 'generous' allele cannot take everything.","notes_pl":"PSEUDOSEKWENCJA.\n- 34 pozycje wybrane ze struktur krystalicznych: bli\u017cej ni\u017c 4 \u00c5 od peptydu ORAZ polimorficzne mi\u0119dzy allelami. Allel przestaje by\u0107 nazw\u0105, a staje si\u0119 34-literowym s\u0142owem \u2014 i dlatego da si\u0119 przewidywa\u0107 dla alleli bez ani jednego pomiaru.\nZMIENNA D\u0141UGO\u015a\u0106.\n- Wszystko sprowadza si\u0119 do rdzenia dziewi\u0119cioliterowego przez przeszukiwanie wstawek i wyci\u0119\u0107 i wyb\u00f3r wariantu z najwy\u017csz\u0105 ocen\u0105; model podaje, kt\u00f3ry rdze\u0144 wybra\u0142.\nDWIE G\u0141OWY, JEDNA WSP\u00d3LNA WARSTWA.\n- Pomiary powinowactwa i peptydy ze spektrometru maj\u0105 osobne neurony wyj\u015bciowe, ale wsp\u00f3ln\u0105 warstw\u0119 ukryt\u0105, wi\u0119c szeroki przekr\u00f3j alleli z prob\u00f3wki i prawdziwy sygna\u0142 prezentacji ucz\u0105 jednej reprezentacji.\nDEKONWOLUCJA MOTYW\u00d3W (NNAlign_MA).\n- Po 20 iteracjach rozgrzewki na danych jednoallelicznych ka\u017cdy peptyd z pr\u00f3bki wieloallelicznej trafia do allelu z najwy\u017csz\u0105 ocen\u0105 spo\u015br\u00f3d tych, kt\u00f3re ta pr\u00f3bka faktycznie ma, z wyr\u00f3wnaniem skal, \u017ceby allel \u201ehojny\u201d nie zagarn\u0105\u0142 wszystkiego.","output":"two output neurons: predicted binding affinity and ligand probability; reported as %Rank","params":"small by modern standards \u2014 the point of this model is representation, not scale","pretraining":"supervised on binding affinity + MS-eluted ligands, with motif deconvolution of multi-allelic samples during training"},"article":"/blog/modele-prezentacji-antygenu","cancer_slugs":["pan-cancer"],"category":"immunopeptidomics","confidence":"high","datasets":[{"name":"IEDB \u2014 Immune Epitope Database","note":"binding-affinity measurements and eluted-ligand deposits","role":"training","slug":"iedb"},{"name":"Mono-allelic HLA class I peptidome (Sarkizova / Abelin)","note":"mono-allelic MS data give unambiguous allele labels","role":"training","slug":"monoallelic-peptidome"},{"name":"HLA Ligand Atlas","note":"multi-allelic tissue peptidomes, deconvolved during training","role":"training","slug":"hla-ligand-atlas"},{"name":"IPD-IMGT/HLA Database","note":"allele sequences from which the 34-residue pseudosequences are cut","role":"pretraining","slug":"ipd-imgt-hla"},{"name":"IEDB automated benchmark (MHC class I)","note":"","role":"benchmark","slug":"iedb-benchmark"}],"developer":"Health Tech, Technical University of Denmark (Nielsen lab)","evaluation":[{"benchmark":"MS-eluted ligand prediction vs NetMHCpan-4.0","external":true,"metric":"PPV / AUC","source":"https://academic.oup.com/nar/article/48/W1/W449/5837056","value":"clearly improved (paper)"},{"benchmark":"true T-cell epitopes vs NetMHCpan-4.0","external":true,"metric":"AUC","source":"https://academic.oup.com/nar/article/48/W1/W449/5837056","value":"comparable \u2014 consistent gain only for HLA-B and HLA-C"},{"benchmark":"vaccinia virus epitopes in mice (220 peptides tested experimentally)","external":true,"metric":"epitopes recovered in the top 0.04% of predictions","source":"https://pmc.ncbi.nlm.nih.gov/articles/PMC7274474/","value":"over half of the most immunogenic epitopes; 90% required going to 1.3\u20131.5% of the list"},{"benchmark":"IEDB automated benchmark, eight years","external":true,"metric":"ranking among top methods","source":"https://academic.oup.com/bib/article/23/4/bbac259/6632617","value":"leading methods statistically indistinguishable"}],"hf":null,"kind":"task-model","license":"free for academic use; commercial licence required (DTU)","limitations":"- Answers presentation, not immunogenicity: a well-presented peptide may still meet no matching T-cell receptor.\n- All negative training examples are assumed, never observed.\n- Inherits mass-spectrometry bias: cysteine-containing peptides are 5\u201310\u00d7 under-detected, and solvent conditions alone shifted detected ligands by more than twofold for HLA-A*02 while dropping HLA-A*30 by 25%.\n- Expression is not an input: the model does not know whether the source gene is transcribed in that tumour.\n- Rare alleles (especially HLA-C) are predicted by analogy, and the antibody used to collect data itself prefers some genes.","links":{"demo":null,"docs":"https://services.healthtech.dtu.dk/services/NetMHCpan-4.1/","doi":"10.1093/nar/gkaa379","github":null,"huggingface":null,"paper":"https://academic.oup.com/nar/article/48/W1/W449/5837056","pmid":null},"modalities":["protein-sequence","immunopeptidomics"],"name":"NetMHCpan-4.1","notable_uses":"The default first filter in neoantigen pipelines for individualised mRNA cancer vaccines: it narrows thousands of mutated peptides to the few dozen worth testing on the patient's own cells.","openness":"api-only","regulatory":{"intended_use_en":"Research software for narrowing candidate lists; not a medical device and not a basis for clinical decisions.","intended_use_pl":"Oprogramowanie badawcze do zaw\u0119\u017cania listy kandydat\u00f3w; nie jest wyrobem medycznym ani podstaw\u0105 decyzji klinicznych.","source_url":"https://services.healthtech.dtu.dk/services/NetMHCpan-4.1/","status":"research-only"},"regulatory_status":"research-only","release_date":"2020-05-22","run_snippet":"","settings":["basic-research","vaccine-design","immunotherapy"],"slug":"netmhcpan","sources":[{"label":"Reynisson B et al. NetMHCpan-4.1 and NetMHCIIpan-4.0. Nucleic Acids Res 2020","url":"https://academic.oup.com/nar/article/48/W1/W449/5837056"},{"label":"Jurtz V et al. NetMHCpan-4.0. J Immunol 2017","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC5679736/"},{"label":"Nielsen M et al. NetMHCpan (pseudosequence definition). PLoS ONE 2007","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC1949492/"},{"label":"Alvarez B et al. NNAlign_MA. Mol Cell Proteomics 2019","url":"https://pmc.ncbi.nlm.nih.gov/articles/PMC6885703/"},{"label":"cancer3.ai \u2014 Trzydzie\u015bci cztery litery zamka (explainer)","url":"https://cancer3.ai/blog/modele-prezentacji-antygenu"}],"summary":"The reference pan-allele predictor of MHC class I antigen presentation: one small neural network covers more than 11,000 MHC molecules because the groove itself is part of the input \u2014 a 34-residue pseudosequence next to the peptide.","tasks":["antigen-presentation","peptide-mhc-binding","neoantigen-prioritisation"],"training":{"consent_notes":"public deposited data; no patient-level identifiers","institutions":"IEDB deposits, published mono-allelic and multi-allelic immunopeptidomics studies","populations":"allele coverage is heavily skewed: HLA-A*02:01 alone is 24.38% of the IEDB benchmark datasets, and 66.24% of class I molecules have six datasets or fewer","size":"~850,000 measured peptides (+ assumed negatives to 13.2M points)","summary_en":"13,245,212 training data points \u2014 but only about 850,000 are laboratory measurements (binding affinities plus MS-eluted ligands). The rest are random natural peptides from UniProt assumed to be negatives, roughly 94% of the set: mass spectrometry never reports true negatives, so they have to be invented.","summary_pl":"13 245 212 punkt\u00f3w danych treningowych \u2014 ale tylko oko\u0142o 850 tysi\u0119cy to pomiary laboratoryjne (powinowactwa i peptydy ze spektrometru). Reszta to losowe peptydy naturalne z UniProt przyj\u0119te za negatywne, oko\u0142o 94% zbioru: spektrometria nigdy nie zwraca prawdziwych negatyw\u00f3w, wi\u0119c trzeba je za\u0142o\u017cy\u0107."},"updated_at":"2026-09-05T22:26:00.704889","url":"/ai-oncology/models/netmhcpan","usage":{"api":"https://services.healthtech.dtu.dk/services/NetMHCpan-4.1/","hardware":"CPU; a whole-proteome scan is minutes, not hours","library":"standalone binary / web server (DTU Health Tech)","notes_en":"Download requires an academic licence form; the binary is not on PyPI or Hugging Face. Read %Rank, not the raw score \u2014 and remember %Rank is a rank, not a probability: 0.1% is not ten times more likely to be presented than 1%.","notes_pl":"Pobranie wymaga formularza licencji akademickiej; binarki nie ma na PyPI ani Hugging Face. Czytaj %Rank, nie surow\u0105 ocen\u0119 \u2014 i pami\u0119taj, \u017ce %Rank to pozycja w rankingu, nie prawdopodobie\u0144stwo: 0,1% nie znaczy dziesi\u0119ciokrotnie wi\u0119kszej szansy ni\u017c 1%."},"verified_at":"2026-09-05T22:26:00.703103","verified_by":"editorial","version":"4.1","what_it_does":"Takes a peptide (8\u201314 aa) and an HLA allele and answers whether that pair will appear on the cell surface. Output is a %Rank \u2014 the peptide's position against a background of random natural peptides for that specific allele (0.5% strong binder, 2% weak) \u2014 plus the predicted 9-mer binding core and its offset."}
