PeptideBERT: A Language Model based on Transformers for Peptide Property Prediction

Guntuboina, Chakradhar; Das, Adrita; Mollaei, Parisa; Kim, Seongwon; Farimani, Amir Barati

Quantitative Biology > Biomolecules

arXiv:2309.03099 (q-bio)

[Submitted on 28 Aug 2023]

Title:PeptideBERT: A Language Model based on Transformers for Peptide Property Prediction

Authors:Chakradhar Guntuboina, Adrita Das, Parisa Mollaei, Seongwon Kim, Amir Barati Farimani

View PDF

Abstract:Recent advances in Language Models have enabled the protein modeling community with a powerful tool since protein sequences can be represented as text. Specifically, by taking advantage of Transformers, sequence-to-property prediction will be amenable without the need for explicit structural data. In this work, inspired by recent progress in Large Language Models (LLMs), we introduce PeptideBERT, a protein language model for predicting three key properties of peptides (hemolysis, solubility, and non-fouling). The PeptideBert utilizes the ProtBERT pretrained transformer model with 12 attention heads and 12 hidden layers. We then finetuned the pretrained model for the three downstream tasks. Our model has achieved state of the art (SOTA) for predicting Hemolysis, which is a task for determining peptide's potential to induce red blood cell lysis. Our PeptideBert non-fouling model also achieved remarkable accuracy in predicting peptide's capacity to resist non-specific interactions. This model, trained predominantly on shorter sequences, benefits from the dataset where negative examples are largely associated with insoluble peptides. Codes, models, and data used in this study are freely available at: this https URL

Comments:	24 pages
Subjects:	Biomolecules (q-bio.BM); Machine Learning (cs.LG)
Cite as:	arXiv:2309.03099 [q-bio.BM]
	(or arXiv:2309.03099v1 [q-bio.BM] for this version)
	https://doi.org/10.48550/arXiv.2309.03099

Submission history

From: Chakradhar Guntuboina [view email]
[v1] Mon, 28 Aug 2023 01:09:21 UTC (1,286 KB)

Quantitative Biology > Biomolecules

Title:PeptideBERT: A Language Model based on Transformers for Peptide Property Prediction

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Quantitative Biology > Biomolecules

Title:PeptideBERT: A Language Model based on Transformers for Peptide Property Prediction

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators