Computational Biology · Haeryong High School
I study where a protein ends up inside the cell — and how much of that a language model already knows.
I'm a student researcher at Haeryong High School working in computational biology. My current project applies semi-supervised learning to protein subcellular localization prediction, built on embeddings from the ESM-2 protein language model — aiming to make accurate predictions possible even when labeled data is scarce.
MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQ…
a raw amino-acid sequence, no structure needed
Most localization datasets are small and expensive to label, while unlabeled protein sequences are abundant. My work explores whether ESM-2 embeddings already encode enough biological signal that a semi-supervised classifier can bridge that gap — turning sequence alone into a confident guess about where in the cell a protein belongs.
Model ESM-2 (Meta AI) · Approach semi-supervised pseudo-labeling · Task subcellular localization classification
Team Haeryong Cass
2026