Alternating EM algorithm for a bilinear model in isoform quantification from RNA-seq data.


Journal

Bioinformatics (Oxford, England)
ISSN: 1367-4811
Titre abrégé: Bioinformatics
Pays: England
ID NLM: 9808944

Informations de publication

Date de publication:
01 02 2020
Historique:
received: 15 08 2018
revised: 13 06 2019
accepted: 09 08 2019
pubmed: 11 8 2019
medline: 18 9 2020
entrez: 11 8 2019
Statut: ppublish

Résumé

Estimation of isoform-level gene expression from RNA-seq data depends on simplifying assumptions, such as uniform read distribution, that are easily violated in real data. Such violations typically lead to biased estimates. Most existing methods provide bias correction step(s), which is based on biological considerations-such as GC content-and applied in single samples separately. The main problem is that not all biases are known. We have developed a novel method called XAEM based on a more flexible and robust statistical model. Existing methods are essentially based on a linear model Xβ, where the design matrix X is known and is computed based on the simplifying assumptions. In contrast XAEM considers Xβ as a bilinear model with both X and β unknown. Joint estimation of X and β is made possible by a simultaneous analysis of multi-sample RNA-seq data. Compared to existing methods, XAEM automatically performs empirical correction of potentially unknown biases. We use an alternating expectation-maximization (AEM) algorithm, alternating between estimation of X and β. For speed XAEM utilizes quasi-mapping for read alignment, thus leading to a fast algorithm. Overall XAEM performs favorably compared to recent advanced methods. For simulated datasets, XAEM obtains higher accuracy for multiple-isoform genes. In a differential-expression analysis of a real single-cell RNA-seq dataset, XAEM achieves substantially better rediscovery rates in independent validation sets. The method and pipeline are implemented as a tool and freely available for use at http://fafner.meb.ki.se/biostatwiki/xaem/. Supplementary data are available at Bioinformatics online.

Identifiants

pubmed: 31400221
pii: 5545974
doi: 10.1093/bioinformatics/btz640
pmc: PMC9883676
doi:

Substances chimiques

Protein Isoforms 0

Types de publication

Journal Article Research Support, Non-U.S. Gov't

Langues

eng

Sous-ensembles de citation

IM

Pagination

805-812

Informations de copyright

© The Author(s) 2019. Published by Oxford University Press.

Références

Nat Methods. 2014 Jan;11(1):41-6
pubmed: 24141493
Bioinformatics. 2014 Feb 15;30(4):506-13
pubmed: 24307704
Nat Biotechnol. 2014 Sep;32(9):903-14
pubmed: 25150838
Bioinformatics. 2016 Jun 15;32(12):i192-i200
pubmed: 27307617
Nat Genet. 2013 Jun;45(6):580-5
pubmed: 23715323
Nat Biotechnol. 2014 May;32(5):462-4
pubmed: 24752080
Bioinformatics. 2009 May 1;25(9):1105-11
pubmed: 19289445
Bioinformatics. 2013 Jan 1;29(1):15-21
pubmed: 23104886
Nat Methods. 2008 Jul;5(7):621-8
pubmed: 18516045
Nat Genet. 2013 Oct;45(10):1113-20
pubmed: 24071849
Bioinformatics. 2013 Sep 15;29(18):2292-9
pubmed: 23821651
Physiol Rev. 2008 Oct;88(4):1341-78
pubmed: 18923184
BMC Genomics. 2017 Aug 7;18(1):583
pubmed: 28784092
Nat Biotechnol. 2010 May;28(5):511-5
pubmed: 20436464
Bioinformatics. 2009 Apr 15;25(8):1026-32
pubmed: 19244387
Nature. 2008 Mar 13;452(7184):230-3
pubmed: 18337823
Brief Bioinform. 2015 Jul;16(4):563-75
pubmed: 25256289
Nat Biotechnol. 2016 May;34(5):525-7
pubmed: 27043002
Nat Protoc. 2012 Mar 01;7(3):562-78
pubmed: 22383036
BMC Bioinformatics. 2010 Jan 19;11:35
pubmed: 20085625
Genome Biol. 2011;12(3):R22
pubmed: 21410973
J Comput Biol. 2000 Feb-Apr;7(1-2):203-14
pubmed: 10890397
Nat Methods. 2015 Apr;12(4):357-60
pubmed: 25751142
Bioinformatics. 2015 Sep 1;31(17):2778-84
pubmed: 25926345
Bioinformatics. 2016 Jul 15;32(14):2128-35
pubmed: 27153638
Nat Methods. 2017 Apr;14(4):417-419
pubmed: 28263959

Auteurs

Wenjiang Deng (W)

Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm 17177, Sweden.

Tian Mou (T)

Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm 17177, Sweden.

Krishna R Kalari (KR)

Department of Health Sciences Research, MN 55905, USA.

Nifang Niu (N)

Department of Molecular Pharmacology and Experimental Therapeutics, Mayo Clinic, Rochester, MN 55905, USA.

Liewei Wang (L)

Department of Molecular Pharmacology and Experimental Therapeutics, Mayo Clinic, Rochester, MN 55905, USA.

Yudi Pawitan (Y)

Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm 17177, Sweden.

Trung Nghia Vu (TN)

Department of Medical Epidemiology and Biostatistics, Karolinska Institutet, Stockholm 17177, Sweden.

Articles similaires

Selecting optimal software code descriptors-The case of Java.

Yegor Bugayenko, Zamira Kholmatova, Artem Kruglov et al.
1.00
Software Algorithms Programming Languages

Exploring blood-brain barrier passage using atomic weighted vector and machine learning.

Yoan Martínez-López, Paulina Phoobane, Yanaima Jauriga et al.
1.00
Blood-Brain Barrier Machine Learning Humans Support Vector Machine Software
Drought Resistance Gene Expression Profiling Gene Expression Regulation, Plant Gossypium Multigene Family
1.00
Humans Magnetic Resonance Imaging Brain Infant, Newborn Infant, Premature

Classifications MeSH