Lightweight ProteinUnet2 network for protein secondary structure prediction: a step towards proper evaluation.


Journal

BMC bioinformatics
ISSN: 1471-2105
Titre abrégé: BMC Bioinformatics
Pays: England
ID NLM: 100965194

Informations de publication

Date de publication:
22 Mar 2022
Historique:
received: 13 09 2021
accepted: 28 02 2022
entrez: 23 3 2022
pubmed: 24 3 2022
medline: 25 3 2022
Statut: epublish

Résumé

The prediction of protein secondary structures is a crucial and significant step for ab initio tertiary structure prediction which delivers the information about proteins activity and functions. As the experimental methods are expensive and sometimes impossible, many SS predictors, mainly based on different machine learning methods have been proposed for many years. Currently, most of the top methods use evolutionary-based input features produced by PSSM and HHblits software, although quite recently the embeddings-the new description of protein sequences generated by language models (LM) have appeared that could be leveraged as input features. Apart from input features calculation, the top models usually need extensive computational resources for training and prediction and are barely possible to run on a regular PC. SS prediction as the imbalanced classification problem should not be judged by the commonly used Q3/Q8 metrics. Moreover, as the benchmark datasets are not random samples, the classical statistical null hypothesis testing based on the Neyman-Pearson approach is not appropriate. We present a lightweight deep network ProteinUnet2 for SS prediction which is based on U-Net convolutional architecture and evolutionary-based input features (from PSSM and HHblits) as well as SPOT-Contact features. Through an extensive evaluation study, we report the performance of ProteinUnet2 in comparison with top SS prediction methods based on evolutionary information (SAINT and SPOT-1D). We also propose a new statistical methodology for prediction performance assessment based on the significance from Fisher-Pitman permutation tests accompanied by practical significance measured by Cohen's effect size. Our results suggest that ProteinUnet2 architecture has much shorter training and inference times while maintaining results similar to SAINT and SPOT-1D predictors. Taking into account the relatively long times of calculating evolutionary-based features (from PSSM in particular), it would be worth conducting the predictive ability tests on embeddings as input features in the future. We strongly believe that our proposed here statistical methodology for the evaluation of SS prediction results will be adopted and used (and even expanded) by the research community.

Sections du résumé

BACKGROUND BACKGROUND
The prediction of protein secondary structures is a crucial and significant step for ab initio tertiary structure prediction which delivers the information about proteins activity and functions. As the experimental methods are expensive and sometimes impossible, many SS predictors, mainly based on different machine learning methods have been proposed for many years. Currently, most of the top methods use evolutionary-based input features produced by PSSM and HHblits software, although quite recently the embeddings-the new description of protein sequences generated by language models (LM) have appeared that could be leveraged as input features. Apart from input features calculation, the top models usually need extensive computational resources for training and prediction and are barely possible to run on a regular PC. SS prediction as the imbalanced classification problem should not be judged by the commonly used Q3/Q8 metrics. Moreover, as the benchmark datasets are not random samples, the classical statistical null hypothesis testing based on the Neyman-Pearson approach is not appropriate.
RESULTS RESULTS
We present a lightweight deep network ProteinUnet2 for SS prediction which is based on U-Net convolutional architecture and evolutionary-based input features (from PSSM and HHblits) as well as SPOT-Contact features. Through an extensive evaluation study, we report the performance of ProteinUnet2 in comparison with top SS prediction methods based on evolutionary information (SAINT and SPOT-1D). We also propose a new statistical methodology for prediction performance assessment based on the significance from Fisher-Pitman permutation tests accompanied by practical significance measured by Cohen's effect size.
CONCLUSIONS CONCLUSIONS
Our results suggest that ProteinUnet2 architecture has much shorter training and inference times while maintaining results similar to SAINT and SPOT-1D predictors. Taking into account the relatively long times of calculating evolutionary-based features (from PSSM in particular), it would be worth conducting the predictive ability tests on embeddings as input features in the future. We strongly believe that our proposed here statistical methodology for the evaluation of SS prediction results will be adopted and used (and even expanded) by the research community.

Identifiants

pubmed: 35317722
doi: 10.1186/s12859-022-04623-z
pii: 10.1186/s12859-022-04623-z
pmc: PMC8939211
doi:

Substances chimiques

Proteins 0

Types de publication

Journal Article

Langues

eng

Sous-ensembles de citation

IM

Pagination

100

Subventions

Organisme : Silesian University of Technology
ID : 02/100/BK_21/0008

Informations de copyright

© 2022. The Author(s).

Références

J Mol Graph Model. 2017 Sep;76:379-402
pubmed: 28763690
Bioinformatics. 2019 Jul 15;35(14):2403-2410
pubmed: 30535134
Int J Mol Sci. 2021 Sep 30;22(19):
pubmed: 34638925
Nucleic Acids Res. 2000 Jan 1;28(1):235-42
pubmed: 10592235
BMC Bioinformatics. 2019 Dec 17;20(1):723
pubmed: 31847804
Curr Protoc. 2021 May;1(5):e113
pubmed: 33961736
Nat Methods. 2021 Feb;18(2):203-211
pubmed: 33288961
J Comput Chem. 2021 Jan 5;42(1):50-59
pubmed: 33058261
Proteins. 1999 Feb 1;34(2):220-3
pubmed: 10022357
IEEE Trans Pattern Anal Mach Intell. 2021 Jul 07;PP:
pubmed: 34232869
Proc Natl Acad Sci U S A. 2021 Apr 13;118(15):
pubmed: 33876751
Protein Sci. 2018 Jan;27(1):129-134
pubmed: 28875543
Bioinformatics. 2020 Nov 1;36(17):4599-4608
pubmed: 32437517
Biomolecules. 2020 May 14;10(5):
pubmed: 32423068
Source Code Biol Med. 2018 Apr 20;13:1
pubmed: 29713370
J Mol Biol. 1999 Sep 17;292(2):195-202
pubmed: 10493868
Biochemistry. 1974 Jan 15;13(2):222-45
pubmed: 4358940
Bioinformatics. 2018 Dec 1;34(23):4039-4045
pubmed: 29931279
Science. 1973 Jul 20;181(4096):223-30
pubmed: 4124164
Proteins. 2010 Apr;78(5):1137-52
pubmed: 19927325
J Mol Biol. 1978 Mar 25;120(1):97-120
pubmed: 642007
Nature. 2021 Aug;596(7873):583-589
pubmed: 34265844
Comput Struct Biotechnol J. 2021 Mar 25;19:1750-1758
pubmed: 33897979
Sci Rep. 2016 Jan 11;6:18962
pubmed: 26752681
Brief Bioinform. 2018 May 1;19(3):482-494
pubmed: 28040746
J Mol Biol. 1994 Jan 7;235(1):13-26
pubmed: 8289237
J Mol Biol. 1974 Oct 5;88(4):873-94
pubmed: 4427384
Biopolymers. 1983 Dec;22(12):2577-637
pubmed: 6667333
Nat Methods. 2011 Dec 25;9(2):173-5
pubmed: 22198341
Nucleic Acids Res. 2021 Jul 2;49(W1):W431-W437
pubmed: 33956157
Nature. 1958 Mar 8;181(4610):662-6
pubmed: 13517261
Proteins. 2019 Jun;87(6):520-527
pubmed: 30785653
Int J Pept Protein Res. 1988 Oct;32(4):269-78
pubmed: 3209351
J Bioinform Comput Biol. 2012 Aug;10(4):1250003
pubmed: 22809416
Proc Natl Acad Sci U S A. 1993 Aug 15;90(16):7558-62
pubmed: 8356056

Auteurs

Katarzyna Stapor (K)

Department of Applied Informatics, Silesian University of Technology, Akademicka 16, 44-100, Gliwice, Poland. Katarzyna.Stapor@polsl.pl.

Krzysztof Kotowski (K)

Department of Applied Informatics, Silesian University of Technology, Akademicka 16, 44-100, Gliwice, Poland.

Tomasz Smolarczyk (T)

Department of Applied Informatics, Silesian University of Technology, Akademicka 16, 44-100, Gliwice, Poland.

Irena Roterman (I)

Department of Bioinformatics and Telemedicine, Jagiellonian University Medical College, Medyczna 7, 30-688, Kraków, Poland.

Articles similaires

Databases, Protein Protein Domains Protein Folding Proteins Deep Learning
Animals Hemiptera Insect Proteins Phylogeny Insecticides
Humans Colorectal Neoplasms Biomarkers, Tumor Prognosis Gene Expression Regulation, Neoplastic

Classifications MeSH