Reaching alignment-profile-based accuracy in predicting protein secondary and tertiary structural properties without alignment.


Journal

Scientific reports
ISSN: 2045-2322
Titre abrégé: Sci Rep
Pays: England
ID NLM: 101563288

Informations de publication

Date de publication:
09 05 2022
Historique:
received: 19 01 2022
accepted: 25 04 2022
entrez: 9 5 2022
pubmed: 10 5 2022
medline: 12 5 2022
Statut: epublish

Résumé

Protein language models have emerged as an alternative to multiple sequence alignment for enriching sequence information and improving downstream prediction tasks such as biophysical, structural, and functional properties. Here we show that a method called SPOT-1D-LM combines traditional one-hot encoding with the embeddings from two different language models (ProtTrans and ESM-1b) for the input and yields a leap in accuracy over single-sequence-based techniques in predicting protein 1D secondary and tertiary structural properties, including backbone torsion angles, solvent accessibility and contact numbers for all six test sets (TEST2018, TEST2020, Neff1-2020, CASP12-FM, CASP13-FM and CASP14-FM). More significantly, it has a performance comparable to profile-based methods for those proteins with homologous sequences. For example, the accuracy for three-state secondary structure (SS3) prediction for TEST2018 and TEST2020 proteins are 86.7% and 79.8% by SPOT-1D-LM, compared to 74.3% and 73.4% by the single-sequence-based method SPOT-1D-Single and 86.2% and 80.5% by the profile-based method SPOT-1D, respectively. For proteins without homologous sequences (Neff1-2020) SS3 is 80.41% by SPOT-1D-LM which is 3.8% and 8.3% higher than SPOT-1D-Single and SPOT-1D, respectively. SPOT-1D-LM is expected to be useful for genome-wide analysis given its fast performance. Moreover, high-accuracy prediction of both secondary and tertiary structural properties such as backbone angles and solvent accessibility without sequence alignment suggests that highly accurate prediction of protein structures may be made without homologous sequences, the remaining obstacle in the post AlphaFold2 era.

Identifiants

pubmed: 35534620
doi: 10.1038/s41598-022-11684-w
pii: 10.1038/s41598-022-11684-w
pmc: PMC9085874
doi:

Substances chimiques

Proteins 0
Solvents 0

Types de publication

Journal Article Research Support, Non-U.S. Gov't

Langues

eng

Sous-ensembles de citation

IM

Pagination

7607

Informations de copyright

© 2022. The Author(s).

Références

Proteins. 2018 May;86(5):592-598
pubmed: 29492997
Bioinformatics. 2007 May 15;23(10):1282-8
pubmed: 17379688
Adv Neural Inf Process Syst. 2019 Dec;32:9689-9701
pubmed: 33390682
BMC Bioinformatics. 2019 Dec 17;20(1):723
pubmed: 31847804
Sci Rep. 2016 Jan 11;6:18962
pubmed: 26752681
J Comput Chem. 2021 Jan 5;42(1):50-59
pubmed: 33058261
Genomics Proteomics Bioinformatics. 2019 Dec;17(6):645-656
pubmed: 32173600
Protein Sci. 1998 Feb;7(2):233-42
pubmed: 9521098
Biopolymers. 1983 Dec;22(12):2577-637
pubmed: 6667333
Proteins. 2019 Jun;87(6):520-527
pubmed: 30785653
IEEE Trans Pattern Anal Mach Intell. 2021 Jul 07;PP:
pubmed: 34232869
Comput Struct Biotechnol J. 2021 Feb 02;19:1145-1153
pubmed: 33680357
Bioinformatics. 2020 Dec 22;36(20):5021-5026
pubmed: 32678893
Nat Methods. 2019 Jul;16(7):603-606
pubmed: 31235882
Bioinformatics. 2000 Jul;16(7):613-8
pubmed: 11038331
Nature. 1974 Mar 22;248(446):338-9
pubmed: 4819639
Bioinformatics. 2016 Mar 15;32(6):843-9
pubmed: 26568622
Cell Syst. 2019 Apr 24;8(4):292-301.e3
pubmed: 31005579
Nucleic Acids Res. 2004 Jan 1;32(Database issue):D138-41
pubmed: 14681378
Curr Protoc Bioinformatics. 2013 Jun;Chapter 3:Unit3.1
pubmed: 23749753
Bioinformatics. 2019 Jul 15;35(14):2403-2410
pubmed: 30535134
Nat Biotechnol. 2017 Nov;35(11):1026-1028
pubmed: 29035372
Bioinformatics. 2021 May 13;:
pubmed: 33983382
Bioinformatics. 2000 Apr;16(4):404-5
pubmed: 10869041
Science. 2021 Aug 20;373(6557):871-876
pubmed: 34282049
J Comput Chem. 2018 Oct 5;39(26):2210-2216
pubmed: 30368831
Bioinformatics. 2018 Dec 1;34(23):4039-4045
pubmed: 29931279
Nature. 2021 Aug;596(7873):583-589
pubmed: 34265844
Nucleic Acids Res. 2017 Jan 4;45(D1):D170-D176
pubmed: 27899574
PLoS Comput Biol. 2017 Jan 5;13(1):e1005324
pubmed: 28056090
Bioinformatics. 2020 Jan 1;36(1):41-48
pubmed: 31173061
Proteins. 2019 Dec;87(12):1082-1091
pubmed: 31407406
Nat Commun. 2018 Jun 29;9(1):2542
pubmed: 29959318

Auteurs

Jaspreet Singh (J)

Signal Processing Laboratory, School of Engineering and Built Environment, Griffith University, Brisbane, QLD, 4111, Australia. jaspreetsingh2@griffithuni.edu.au.

Kuldip Paliwal (K)

Signal Processing Laboratory, School of Engineering and Built Environment, Griffith University, Brisbane, QLD, 4111, Australia. k.paliwal@griffith.edu.au.

Thomas Litfin (T)

Signal Processing Laboratory, School of Engineering and Built Environment, Griffith University, Brisbane, QLD, 4111, Australia.

Jaswinder Singh (J)

Signal Processing Laboratory, School of Engineering and Built Environment, Griffith University, Brisbane, QLD, 4111, Australia.

Yaoqi Zhou (Y)

Institute for Glycomics, Griffith University, Parklands Dr. Southport, Goldcoast, QLD, 4222, Australia. zhouyq@szbl.ac.cn.
Shenzhen Bay Laboratory, Institute for Systems and Physical Biology, Shenzhen, 518055, People's Republic of China. zhouyq@szbl.ac.cn.
Peking University Shenzhen Graduate School, Shenzhen, 518055, People's Republic of China. zhouyq@szbl.ac.cn.

Articles similaires

Selecting optimal software code descriptors-The case of Java.

Yegor Bugayenko, Zamira Kholmatova, Artem Kruglov et al.
1.00
Software Algorithms Programming Languages
Databases, Protein Protein Domains Protein Folding Proteins Deep Learning
1.00
Humans Magnetic Resonance Imaging Brain Infant, Newborn Infant, Premature
Humans Algorithms Software Artificial Intelligence Computer Simulation

Classifications MeSH