Abstract
Protein language models (PLMs) trained solely on sequence data have significantly advanced our understanding of protein biology and achieved remarkable performance in protein prediction tasks. However, their lack of three-dimensional (3D) structural features limits their predictive power in applications that rely heavily on 3D conformation. To address this limitation, we developed two structure-aware PLMs, S-PLM1 and S-PLM2, that employ multi-view contrastive learning to align protein sequences with their 3D structures in a unified latent space. S-PLM1 represents structural information using contact maps encoded by a pretrained Swin-Transformer, while S-PLM2 directly encodes 3D backbone coordinates through a Geometric Vector Perceptron (GVP)-based model. The paired sequence-structure data were obtained from AlphaFoldDB. For both models, we designed efficient tuning strategies that enable optimal performance with minimal computational cost. Here, we present detailed protocols for adapting S-PLM1 and S-PLM2 for diverse protein applications. The protocols provide step-by-step guidance on generating structure-aware representations from S-PLMs, fine-tuning them for various protein prediction tasks, and using S-PLM2 to produce structure embeddings for structure-based downstream analyses. We also provide source code and Google Colab implementations for easy customization and deployment.
| Original language | English |
|---|---|
| Article number | e70377 |
| Number of pages | 24 |
| Journal | Current Protocols |
| Volume | 6 |
| Issue number | 5 |
| DOIs | |
| State | Published - May 2026 |
Bibliographical note
Publisher Copyright:© 2026 Wiley Periodicals LLC.
Funding
D.W. and D.X. are partially supported by the National Institutes of Health (grant R35GM126985). Q.S. and D.X. would like to acknowledge the National Institutes of Health (grant R01LM014510). This work used the GPU resources through allocations NAIRR240246 and NAIRR250134 from the National Artificial Intelligence Research Resource (NAIRR). This work also used Delta-GPU at NCSA through allocation CIS230053 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296. D.W. and D.X. are partially supported by the National Institutes of Health (grant R35GM126985). Q.S. and D.X. would like to acknowledge the National Institutes of Health (grant R01LM014510). This work used the GPU resources through allocations NAIRR240246 and NAIRR250134 from the National Artificial Intelligence Research Resource (NAIRR). This work also used Delta‐GPU at NCSA through allocation CIS230053 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603, and #2138296.
| Funders | Funder number |
|---|---|
| National Artificial Intelligence Research Resource | CIS230053 |
| National Institutes of Health (NIH) | NAIRR240246, NAIRR250134, R35GM126985, CIS230053, R01LM014510 |
| National Science Foundation Arctic Social Science Program | 2138286, 2138296, 2137603, 2138307, 2138259 |
Keywords
- protein analyses
- protein language model
- protein prediction
- protein sequence
- protein structure
ASJC Scopus subject areas
- General Neuroscience
- General Immunology and Microbiology
- General Biochemistry, Genetics and Molecular Biology
- General Pharmacology, Toxicology and Pharmaceutics
- Health Informatics
- Medical Laboratory Technology
Fingerprint
Dive into the research topics of 'Efficient Adaptation of Structure-Aware Protein Language Models for Diverse Protein Applications'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver