Cancer3.AIAI in OncologyAI Models › HLAthena
Task-specific model API / hosted only Research use only immunopeptidomics

HLAthena

Presentation predictor trained on the mono-allelic peptidome of 95 cell lines — 186,464 peptides across 95 HLA alleles, fifteen of which had no described motif before — which is the dataset that changed this field more than any architectural idea.

At a glance

DeveloperBroad Institute / Dana-Farber Cancer Institute (Keskin, Wu, Carr labs)
Released2020-01-13
Licencefree web server; academic use
AvailabilityAPI / hosted only
KindTask-specific model
Regulatory statusResearch use only

What it does

Scores peptide–allele pairs and adds features beyond the peptide: gene expression of the source protein and cleavage context, which is why its predictions track real presentation more closely than binding-only tools.

Notable uses in oncology

The underlying allele panel covers at least one allele in 95% of people worldwide for each of HLA-A, -B and -C — the reason the 'long tail' of alleles shrank at all.

Tasks, data types and cancers

CancerPan-cancer
Inputpeptide + allele + source-gene expression + flanking cleavage context
Outputpresentation score / percentile

Architecture

Familyneural network with peptide, expression and cleavage features

Adding expression is what a peptide-only model structurally cannot do; the price is that you need RNA data for the sample.

Training data

95 mono-allelic B721.221 lines (31 HLA-A, 40 HLA-B, 21 HLA-C, 3 HLA-G), 186,464 unique peptides, median 1,860 per allele.

Training set size186,464 peptides / 95 alleles
InstitutionsBroad Institute, Dana-Farber

Linked datasets

Evaluation

Benchmark / datasetMetricValueExternal validationSource
top 0.1% of a 1:999 decoy list PPV reported by the authors at 1:999 — a stricter setup than several competitors use no Source

How to run

Libraryweb server
APIhttp://hlathena.tools/

No standalone weights; use the web tool or reuse the published peptidome (see the data card) to train your own model.

Regulatory status and intended use

Regulatory statusResearch use only
Intended useResearch use.

Source →

Regulatory status is quoted from the source linked above and can change. Research-use-only models must not be used for clinical decisions.

Limitations and bias

  • Web-only: not embeddable in an offline pipeline.
  • Expression features require matched RNA data, so it is not a universal 'peptide in, score out' tool.

Sources

  1. Sarkizova S et al. A large peptidome dataset improves HLA class I epitope prediction. Nat Biotechnol 2020
  2. cancer3.ai — Trzydzieści cztery litery zamka

This page is educational — it is not medical advice and does not replace consultation with an oncologist. Diagnostic and treatment decisions are made solely by specialist physicians.