Alternating EM algorithm for a bilinear model in isoform quantification from RNA-seq data.
Journal
Bioinformatics (Oxford, England)
ISSN: 1367-4811
Titre abrégé: Bioinformatics
Pays: England
ID NLM: 9808944
Informations de publication
Date de publication:
01 02 2020
01 02 2020
Historique:
received:
15
08
2018
revised:
13
06
2019
accepted:
09
08
2019
pubmed:
11
8
2019
medline:
18
9
2020
entrez:
11
8
2019
Statut:
ppublish
Résumé
Estimation of isoform-level gene expression from RNA-seq data depends on simplifying assumptions, such as uniform read distribution, that are easily violated in real data. Such violations typically lead to biased estimates. Most existing methods provide bias correction step(s), which is based on biological considerations-such as GC content-and applied in single samples separately. The main problem is that not all biases are known. We have developed a novel method called XAEM based on a more flexible and robust statistical model. Existing methods are essentially based on a linear model Xβ, where the design matrix X is known and is computed based on the simplifying assumptions. In contrast XAEM considers Xβ as a bilinear model with both X and β unknown. Joint estimation of X and β is made possible by a simultaneous analysis of multi-sample RNA-seq data. Compared to existing methods, XAEM automatically performs empirical correction of potentially unknown biases. We use an alternating expectation-maximization (AEM) algorithm, alternating between estimation of X and β. For speed XAEM utilizes quasi-mapping for read alignment, thus leading to a fast algorithm. Overall XAEM performs favorably compared to recent advanced methods. For simulated datasets, XAEM obtains higher accuracy for multiple-isoform genes. In a differential-expression analysis of a real single-cell RNA-seq dataset, XAEM achieves substantially better rediscovery rates in independent validation sets. The method and pipeline are implemented as a tool and freely available for use at http://fafner.meb.ki.se/biostatwiki/xaem/. Supplementary data are available at Bioinformatics online.
Identifiants
pubmed: 31400221
pii: 5545974
doi: 10.1093/bioinformatics/btz640
pmc: PMC9883676
doi:
Substances chimiques
Protein Isoforms
0
Types de publication
Journal Article
Research Support, Non-U.S. Gov't
Langues
eng
Sous-ensembles de citation
IM
Pagination
805-812Informations de copyright
© The Author(s) 2019. Published by Oxford University Press.
Références
Nat Methods. 2014 Jan;11(1):41-6
pubmed: 24141493
Bioinformatics. 2014 Feb 15;30(4):506-13
pubmed: 24307704
Nat Biotechnol. 2014 Sep;32(9):903-14
pubmed: 25150838
Bioinformatics. 2016 Jun 15;32(12):i192-i200
pubmed: 27307617
Nat Genet. 2013 Jun;45(6):580-5
pubmed: 23715323
Nat Biotechnol. 2014 May;32(5):462-4
pubmed: 24752080
Bioinformatics. 2009 May 1;25(9):1105-11
pubmed: 19289445
Bioinformatics. 2013 Jan 1;29(1):15-21
pubmed: 23104886
Nat Methods. 2008 Jul;5(7):621-8
pubmed: 18516045
Nat Genet. 2013 Oct;45(10):1113-20
pubmed: 24071849
Bioinformatics. 2013 Sep 15;29(18):2292-9
pubmed: 23821651
Physiol Rev. 2008 Oct;88(4):1341-78
pubmed: 18923184
BMC Genomics. 2017 Aug 7;18(1):583
pubmed: 28784092
Nat Biotechnol. 2010 May;28(5):511-5
pubmed: 20436464
Bioinformatics. 2009 Apr 15;25(8):1026-32
pubmed: 19244387
Nature. 2008 Mar 13;452(7184):230-3
pubmed: 18337823
Brief Bioinform. 2015 Jul;16(4):563-75
pubmed: 25256289
Nat Biotechnol. 2016 May;34(5):525-7
pubmed: 27043002
Nat Protoc. 2012 Mar 01;7(3):562-78
pubmed: 22383036
BMC Bioinformatics. 2010 Jan 19;11:35
pubmed: 20085625
Genome Biol. 2011;12(3):R22
pubmed: 21410973
J Comput Biol. 2000 Feb-Apr;7(1-2):203-14
pubmed: 10890397
Nat Methods. 2015 Apr;12(4):357-60
pubmed: 25751142
Bioinformatics. 2015 Sep 1;31(17):2778-84
pubmed: 25926345
Bioinformatics. 2016 Jul 15;32(14):2128-35
pubmed: 27153638
Nat Methods. 2017 Apr;14(4):417-419
pubmed: 28263959