preprint Open access

TransMEP: Transfer learning on large protein language models to predict mutation effects of proteins from a small known dataset

  • bioRxiv (Cold Spring Harbor Laboratory)
  • Cold Spring Harbor Laboratory
Research footprint

At a glance

Citations
7
References
24
Comments
0
Paper overview

Abstract

Abstract Machine learning-guided optimization has become a driving force for recent improvements in protein engineering. In addition, new protein language models are learning the grammar of evolutionarily occurring sequences at large scales. This work combines both approaches to make predictions about mutational effects that support protein engineering. To this end, an easy-to-use software tool called TransMEP is developed using transfer learning by feature extraction with Gaussian process regression. A large collection of datasets is used to evaluate its quality, which scales with the size of the training set, and to show its improvements over previous fine-tuning approaches. Wet-lab studies are simulated to evaluate the use of mutation effect prediction models for protein engineering. This showed that TransMEP finds the best performing mutants with a limited study budget by considering the trade-off between exploration and exploitation. Graphical TOC Entry

Record transparency

Publication details

DOI
10.1101/2024.01.12.575432
OpenAlex
W4390870984
Document type
preprint
Language
EN
Source
bioRxiv (Cold Spring Harbor Laboratory)
Last metadata update
Community

Comments

Log in to join the discussion.

  1. No comments yet. Start the discussion.