Preprocessing Tandem Mass Spectra Using Genetic Programming for Peptide Identification.
Classification
Genetic programming
Mass spectrometry
Preprocessing
Tandem mass spectrometry
Journal
Journal of the American Society for Mass Spectrometry
ISSN: 1879-1123
Titre abrégé: J Am Soc Mass Spectrom
Pays: United States
ID NLM: 9010412
Informations de publication
Date de publication:
Jul 2019
Jul 2019
Historique:
received:
25
09
2018
accepted:
11
03
2019
revised:
15
01
2019
pubmed:
27
4
2019
medline:
18
12
2019
entrez:
27
4
2019
Statut:
ppublish
Résumé
One of the major challenges in proteomics is peptide identification from mass spectra containing high noise ratio and small number of signal (b-/y-ions) peaks. However, the accuracy and reliability of peptide identification in such highly imbalanced MS/MS data can be improved by applying a preprocessing step prior to peptide identification aiming at discriminating b-/y-ions from noise peaks in the spectra. In this study, we report a genetic programming (GP)-based preprocessing method for de-noising highly imbalanced and noisy CID MS/MS spectra. GP now becomes a popular machine learning method via automatic programming. GP preprocesses the highly noisy MS/MS spectra by classifying peaks as noise peaks or signal peaks in a binary classification manner. Meanwhile, a set of spectral fragment features based on the MS/MS fragmentation rules is extracted from the dataset to investigate their discriminating abilities by GP. A MS/MS spectral dataset containing thousands of spectra are used to train the GP model. As the GP tree-based representation has the capability for implicit feature selection during the evolutionary process, the evolved GP model with the selected features is compared with the best threshold-based method. The results show that the GP method improved the reliability of peptide identification and increased the identification rate of a de novo sequencing tool, PEAKS, to 99.4% from 80.1% achieved by the best threshold-based method. Moreover, the result of peptide identification by a database search tool, SEQUEST, using the data preprocessed by the GP method was statistically significant compared to the other methods.
Identifiants
pubmed: 31025295
doi: 10.1007/s13361-019-02196-5
pii: 10.1007/s13361-019-02196-5
doi:
Substances chimiques
Ions
0
Peptides
0
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Pagination
1294-1307Subventions
Organisme : Marsden Fund
ID : VUW1209, VUW1509 and VUW1615
Organisme : Huawei Industry Fund
ID : E2880/3663
Organisme : University Research Fund at Victoria University of Wellington
ID : 209862/3580, and 213150/3662
Références
Nucleic Acids Res. 2000 Jan 1;28(1):126-8
pubmed: 10592200
Electrophoresis. 1999 Dec;20(18):3551-67
pubmed: 10612281
Nature. 2000 Jun 15;405(6788):837-46
pubmed: 10866210
Nat Biotechnol. 2003 Mar;21(3):255-61
pubmed: 12610572
Rapid Commun Mass Spectrom. 2003;17(20):2337-42
pubmed: 14558135
J Proteome Res. 2004 Sep-Oct;3(5):958-64
pubmed: 15473683
J Bioinform Comput Biol. 2005 Jun;3(3):697-716
pubmed: 16108090
Anal Chem. 2007 Aug 1;79(15):5620-32
pubmed: 17580982
Proteomics. 2008 Aug;8(15):3019-29
pubmed: 18615428
BMC Bioinformatics. 2008 Jul 30;9:325
pubmed: 18664292
Proteomics. 2009 Nov;9(21):4978-84
pubmed: 19743429
Proteomics. 2012 Aug;12(14):2276-81
pubmed: 22887946
J Proteome Res. 2013 Jul 5;12(7):3223-32
pubmed: 23675732
J Am Soc Mass Spectrom. 1994 Nov;5(11):976-89
pubmed: 24226387
Proteome Sci. 2013 Nov 7;11(Suppl 1):S4
pubmed: 24565419
J Am Soc Mass Spectrom. 2015 Nov;26(11):1885-94
pubmed: 26122521
J Proteome Res. 2016 Dec 2;15(12):4423-4435
pubmed: 27748123
Science. 1997 Sep 5;277(5331):1453-62
pubmed: 9278503