Comprehensive assessment of mRNA isoform detection methods for long-read sequencing data.
Journal
Nature communications
ISSN: 2041-1723
Titre abrégé: Nat Commun
Pays: England
ID NLM: 101528555
Informations de publication
Date de publication:
10 May 2024
10 May 2024
Historique:
received:
18
10
2022
accepted:
19
04
2024
medline:
11
5
2024
pubmed:
11
5
2024
entrez:
10
5
2024
Statut:
epublish
Résumé
The advancement of Long-Read Sequencing (LRS) techniques has significantly increased the length of sequencing to several kilobases, thereby facilitating the identification of alternative splicing events and isoform expressions. Recently, numerous computational tools for isoform detection using long-read sequencing data have been developed. Nevertheless, there remains a deficiency in comparative studies that systemically evaluate the performance of these tools, which are implemented with different algorithms, under various simulations that encompass potential influencing factors. In this study, we conducted a benchmark analysis of thirteen methods implemented in nine tools capable of identifying isoform structures from long-read RNA-seq data. We evaluated their performances using simulated data, which represented diverse sequencing platforms generated by an in-house simulator, RNA sequins (sequencing spike-ins) data, as well as experimental data. Our findings demonstrate IsoQuant as a highly effective tool for isoform detection with LRS, with Bambu and StringTie2 also exhibiting strong performance. These results offer valuable guidance for future research on alternative splicing analysis and the ongoing improvement of tools for isoform detection using LRS data.
Identifiants
pubmed: 38730241
doi: 10.1038/s41467-024-48117-3
pii: 10.1038/s41467-024-48117-3
doi:
Substances chimiques
RNA, Messenger
0
RNA Isoforms
0
Protein Isoforms
0
Types de publication
Journal Article
Research Support, Non-U.S. Gov't
Langues
eng
Sous-ensembles de citation
IM
Pagination
3972Subventions
Organisme : National Natural Science Foundation of China (National Science Foundation of China)
ID : 32170551
Organisme : National Natural Science Foundation of China (National Science Foundation of China)
ID : 32270835
Organisme : National Natural Science Foundation of China (National Science Foundation of China)
ID : 32250610202
Informations de copyright
© 2024. The Author(s).
Références
Wang, E. T. et al. Alternative isoform regulation in human tissue transcriptomes. Nature 456, 470–476 (2008).
pubmed: 18978772
pmcid: 2593745
doi: 10.1038/nature07509
Pan, Q., Shai, O., Lee, L. J., Frey, B. J. & Blencowe, B. J. Deep surveying of alternative splicing complexity in the human transcriptome by high-throughput sequencing. Nat. Genet. 40, 1413–1415 (2008).
pubmed: 18978789
doi: 10.1038/ng.259
McGuire, A. M., Pearson, M. D., Neafsey, D. E. & Galagan, J. E. Cross-kingdom patterns of alternative splicing and splice recognition. Genome Biol. 9, R50 (2008).
pubmed: 18321378
pmcid: 2397502
doi: 10.1186/gb-2008-9-3-r50
Baralle, F. E. & Giudice, J. Alternative splicing as a regulator of development and tissue identity. Nat. Rev. Mol. Cell Biol. 18, 437–451 (2017).
pubmed: 28488700
pmcid: 6839889
doi: 10.1038/nrm.2017.27
Marasco, L. E. & Kornblihtt, A. R. The physiology of alternative splicing. Nat. Rev. Mol. Cell Biol. 24, 242–254 (2023).
pubmed: 36229538
doi: 10.1038/s41580-022-00545-z
Branton, D. et al. The potential and challenges of nanopore sequencing. Nat. Biotechnol. 26, 1146–1153 (2008).
pubmed: 18846088
pmcid: 2683588
doi: 10.1038/nbt.1495
Rhoads, A. & Au, K. F. PacBio sequencing and its applications. Genom. Proteom. Bioinforma. 13, 278–289 (2015).
doi: 10.1016/j.gpb.2015.08.002
Garalde, D. R. et al. Highly parallel direct RNA sequencing on an array of nanopores. Nat. Methods 15, 201–206 (2018).
pubmed: 29334379
doi: 10.1038/nmeth.4577
Au, K. F. et al. Characterization of the human ESC transcriptome by hybrid sequencing. Proc. Natl Acad. Sci. USA 110, E4821–E4830 (2013).
pubmed: 24282307
pmcid: 3864310
doi: 10.1073/pnas.1320101110
Sharon, D., Tilgner, H., Grubert, F. & Snyder, M. A single-molecule long-read survey of the human transcriptome. Nat. Biotechnol. 31, 1009–1014 (2013).
pubmed: 24108091
pmcid: 4075632
doi: 10.1038/nbt.2705
Wang, Y., Zhao, Y., Bollas, A., Wang, Y. & Au, K. F. Nanopore sequencing technology, bioinformatics and applications. Nat. Biotechnol. 39, 1348–1365 (2021).
pubmed: 34750572
pmcid: 8988251
doi: 10.1038/s41587-021-01108-x
Logsdon, G. A., Vollger, M. R. & Eichler, E. E. Long-read human genome sequencing and its applications. Nat. Rev. Genet. 21, 597–614 (2020).
pubmed: 32504078
pmcid: 7877196
doi: 10.1038/s41576-020-0236-x
Kovaka, S. et al. Transcriptome assembly from long-read RNA-seq alignments with StringTie2. Genome Biol. 20, 278 (2019).
pubmed: 31842956
pmcid: 6912988
doi: 10.1186/s13059-019-1910-1
Tang, A. D. et al. Full-length transcript characterization of SF3B1 mutation in chronic lymphocytic leukemia reveals downregulation of retained introns. Nat. Commun. 11, 1438 (2020).
pubmed: 32188845
pmcid: 7080807
doi: 10.1038/s41467-020-15171-6
Tian, L. et al. Comprehensive characterization of single-cell full-length isoforms in human and mouse with long-read sequencing. Genome Biol. 22, 310 (2021).
pubmed: 34763716
pmcid: 8582192
doi: 10.1186/s13059-021-02525-6
Orabi, B. et al. Freddie: annotation-independent detection and discovery of transcriptomic alternative splicing isoforms using long-read sequencing. Nucleic Acids Res. 51, e11–e11 (2023).
pubmed: 36478271
doi: 10.1093/nar/gkac1112
Wyman, D. et al. A technology-agnostic long-read analysis pipeline for transcriptome discovery and quantification. Preprint at https://doi.org/10.1101/672931 (2019).
Al Kadi, M. et al. UNAGI: an automated pipeline for nanopore full-length cDNA sequencing uncovers novel transcripts and isoforms in yeast. Funct. Integr. Genomics 20, 523–536 (2020).
pubmed: 31955296
pmcid: 7283198
doi: 10.1007/s10142-020-00732-1
Kuo, R. I. et al. Illuminating the dark side of the human transcriptome with long read transcript sequencing. BMC Genomics 21, 751 (2020).
pubmed: 33126848
pmcid: 7596999
doi: 10.1186/s12864-020-07123-7
Chen, Y. et al. Context-aware transcript quantification from long-read RNA-seq data with Bambu. Nat. Methods 20, 1187–1195 (2023).
pubmed: 37308696
pmcid: 10448944
doi: 10.1038/s41592-023-01908-w
Prjibelski, A. D. et al. Accurate isoform discovery with IsoQuant using long reads. Nat. Biotechnol. 41, 915–918 (2023).
pubmed: 36593406
pmcid: 10344776
doi: 10.1038/s41587-022-01565-y
Hardwick, S. A. et al. Spliced synthetic genes as internal controls in RNA sequencing experiments. Nat. Methods 13, 792–798 (2016).
pubmed: 27502218
doi: 10.1038/nmeth.3958
Dong, X. et al. The long and the short of it: unlocking nanopore long-read RNA sequencing data with short-read differential expression analysis tools. NAR Genom. Bioinforma. 3, lqab028 (2021).
doi: 10.1093/nargab/lqab028
Dong, X. et al. Benchmarking long-read RNA-sequencing analysis tools using in silico mixtures. Nat. Methods 20, 1810–1821 (2023).
pubmed: 37783886
doi: 10.1038/s41592-023-02026-3
Pardo-Palacios, F. J. et al. Systematic assessment of long-read RNA-seq methods for transcript identification and quantification. Preprint at https://doi.org/10.1101/2023.07.25.550582 (2023).
Ono, Y., Asai, K. & Hamada, M. PBSIM: PacBio reads simulator-toward accurate genome assembly. Bioinformatics 29, 119–121 (2013).
pubmed: 23129296
doi: 10.1093/bioinformatics/bts649
Ono, Y., Asai, K. & Hamada, M. PBSIM2: a simulator for long-read sequencers with a novel generative model of quality scores. Bioinformatics 37, 589–595 (2021).
pubmed: 32976553
doi: 10.1093/bioinformatics/btaa835
Ono, Y., Hamada, M. & Asai, K. PBSIM3: a simulator for all types of PacBio and ONT long reads. NAR Genom. Bioinforma. 4, lqac092 (2022).
doi: 10.1093/nargab/lqac092
Wick, R. Badread: simulation of error-prone long reads. JOSS 4, 1316 (2019).
doi: 10.21105/joss.01316
Ai, Z. et al. Krüppel-like factor 5 rewires NANOG regulatory network to activate human naive pluripotency specific LTR7Ys and promote naive pluripotency. Cell Rep. 40, 111240 (2022).
pubmed: 36001968
doi: 10.1016/j.celrep.2022.111240
Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094–3100 (2018).
pubmed: 29750242
pmcid: 6137996
doi: 10.1093/bioinformatics/bty191
Pertea, G. & Pertea, M. GFF UTILITIEs: GffRead and GffCompare. F1000Res 9, 304 (2020).
doi: 10.12688/f1000research.23297.1
Pardo-Palacios, F. J. et al. SQANTI3: curation of long-read transcriptomes for accurate identification of known and novel isoforms. Nat. Methods https://doi.org/10.1038/s41592-024-02229-2 , 1–5 (2024).
Quinlan, A. R. & Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841–842 (2010).
pubmed: 20110278
pmcid: 2832824
doi: 10.1093/bioinformatics/btq033
Weber, L. M. et al. Essential guidelines for computational method benchmarking. Genome Biol. 20, 125 (2019).
pubmed: 31221194
pmcid: 6584985
doi: 10.1186/s13059-019-1738-8
Wright, D. J. et al. Long read sequencing reveals novel isoforms and insights into splicing regulation during cell state changes. BMC Genomics 23, 42 (2022).
pubmed: 35012468
pmcid: 8744310
doi: 10.1186/s12864-021-08261-2
Zhu, C. et al. Single-molecule, full-length transcript isoform sequencing reveals disease-associated RNA isoforms in cardiomyocytes. Nat. Commun. 12, 4203 (2021).
pubmed: 34244519
pmcid: 8270901
doi: 10.1038/s41467-021-24484-z
Legnini, I., Alles, J., Karaiskos, N., Ayoub, S. & Rajewsky, N. FLAM-seq: full-length mRNA sequencing reveals principles of poly(A) tail length control. Nat. Methods 16, 879–886 (2019).
pubmed: 31384046
doi: 10.1038/s41592-019-0503-y
Thomaidou, S. et al. Long RNA sequencing and ribosome profiling of inflamed β-cells reveal an extensive translatome landscape. Diabetes 70, 2299–2312 (2021).
pubmed: 34554924
doi: 10.2337/db20-1122
Ding, C. et al. Short-read and long-read full-length transcriptome of mouse neural stem cells across neurodevelopmental stages. Sci. Data 9, 69 (2022).
pubmed: 35236860
pmcid: 8891264
doi: 10.1038/s41597-022-01165-0
Liu, X., Andrews, M. V., Skinner, J. P., Johanson, T. M. & Chong, M. M. W. A comparison of alternative mRNA splicing in the CD4 and CD8 T cell lineages. Mol. Immunol. 133, 53–62 (2021).
pubmed: 33631555
doi: 10.1016/j.molimm.2021.02.009
Yue, F. et al. A comparative encyclopedia of DNA elements in the mouse genome. Nature 515, 355–364 (2014).
pubmed: 25409824
pmcid: 4266106
doi: 10.1038/nature13992
D’Angeli, V. et al. Polypyrimidine tract binding protein 1 regulates the activation of mouse CD8 T cells. Eur. J. Immunol. 52, 1058–1068 (2022).
pubmed: 35460072
pmcid: 9546061
doi: 10.1002/eji.202149781
Yao, F. et al. Pathologically high intraocular pressure disturbs normal iron homeostasis and leads to retinal ganglion cell ferroptosis in glaucoma. Cell Death Differ. 30, 69–81 (2023).
pubmed: 35933500
doi: 10.1038/s41418-022-01046-4
Sahlin, K. & Medvedev, P. Error correction enables use of Oxford Nanopore technology for reference-free transcriptome analysis. Nat. Commun. 12, 2 (2021).
pubmed: 33397972
pmcid: 7782715
doi: 10.1038/s41467-020-20340-8
Zhang, D. et al. Dosage sensitivity and exon shuffling shape the landscape of polymorphic duplicates in Drosophila and humans. Nat. Ecol. Evol. 6, 273–287 (2021).
pubmed: 34969986
doi: 10.1038/s41559-021-01614-w
Li, R. et al. Direct full-length RNA sequencing reveals unexpected transcriptome complexity during Caenorhabditis elegans development. Genome Res. 30, 287–298 (2020).
pubmed: 32024662
pmcid: 7050527
doi: 10.1101/gr.251512.119
Viscardi, M. J. & Arribere, J. A. Poly(a) selection introduces bias and undue noise in direct RNA-sequencing. BMC Genomics 23, 530 (2022).
pubmed: 35869428
pmcid: 9306060
doi: 10.1186/s12864-022-08762-8
Robinson, J. T. et al. Integrative genomics viewer. Nat. Biotechnol. 29, 24–26 (2011).
pubmed: 21221095
pmcid: 3346182
doi: 10.1038/nbt.1754
Deniz, Ö., Frost, J. M. & Branco, M. R. Regulation of transposable elements by DNA modifications. Nat. Rev. Genet. 20, 417–431 (2019).
pubmed: 30867571
doi: 10.1038/s41576-019-0106-6
Theunissen, T. W. et al. Molecular criteria for defining the naive human pluripotent state. Cell Stem Cell 19, 502–515 (2016).
pubmed: 27424783
pmcid: 5065525
doi: 10.1016/j.stem.2016.06.011
Pontis, J. et al. Hominoid-specific transposable elements and KZFPs facilitate human embryonic genome activation and control transcription in naive human ESCs. Cell Stem Cell 24, 724–735.e5 (2019).
pubmed: 31006620
pmcid: 6509360
doi: 10.1016/j.stem.2019.03.012
Xiang, X. et al. Human reproduction is regulated by retrotransposons derived from ancient Hominidae-specific viral infections. Nat. Commun. 13, 463 (2022).
pubmed: 35075135
pmcid: 8786967
doi: 10.1038/s41467-022-28105-1
Liao, Y., Smyth, G. K. & Shi, W. featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics 30, 923–930 (2014).
pubmed: 24227677
doi: 10.1093/bioinformatics/btt656
Vitting-Seerup, K. & Sandelin, A. IsoformSwitchAnalyzeR: analysis of changes in genome-wide patterns of alternative splicing and its functional consequences. Bioinformatics 35, 4469–4471 (2019).
pubmed: 30989184
doi: 10.1093/bioinformatics/btz247
Gao, Y. et al. ESPRESSO: Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data. Sci. Adv. 9, eabq5072 (2023).
pubmed: 36662851
pmcid: 9858503
doi: 10.1126/sciadv.abq5072
Petri, A. J. & Sahlin, K. isONform: reference-free transcriptome reconstruction from Oxford Nanopore data. Bioinformatics 39, i222–i231 (2023).
pubmed: 37387174
pmcid: 10311309
doi: 10.1093/bioinformatics/btad264
Xia, Y. et al. TAGET: a toolkit for analyzing full-length transcripts from long-read sequencing. Nat. Commun. 14, 5935 (2023).
pubmed: 37741817
pmcid: 10518008
doi: 10.1038/s41467-023-41649-0
Lienhard, M. et al. IsoTools: a flexible workflow for long-read transcriptome sequencing analysis. Bioinformatics 39, btad364 (2023).
pubmed: 37267159
pmcid: 10287928
doi: 10.1093/bioinformatics/btad364
Sims, D., Sudbery, I., Ilott, N. E., Heger, A. & Ponting, C. P. Sequencing depth and coverage: key considerations in genomic analyses. Nat. Rev. Genet. 15, 121–132 (2014).
pubmed: 24434847
doi: 10.1038/nrg3642
Garber, M., Grabherr, M. G., Guttman, M. & Trapnell, C. Computational methods for transcriptome annotation and quantification using RNA-seq. Nat. Methods 8, 469–477 (2011).
pubmed: 21623353
doi: 10.1038/nmeth.1613
Angelini, C., Canditiis, D. D. & Feis, I. D. Computational approaches for isoform detection and estimation: good and bad news. BMC Bioinforma. 15, 135 (2014).
doi: 10.1186/1471-2105-15-135
Kanitz, A. et al. Comparative assessment of methods for the computational inference of transcript isoform abundance from RNA-seq data. Genome Biol. 16, 150 (2015).
pubmed: 26201343
pmcid: 4511015
doi: 10.1186/s13059-015-0702-5
Li, H. et al. A male germ-cell-specific ribosome controls male fertility. Nature 612, 725–731 (2022).
pubmed: 36517592
doi: 10.1038/s41586-022-05508-0
Pertea, M. et al. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat. Biotechnol. 33, 290–295 (2015).
pubmed: 25690850
pmcid: 4643835
doi: 10.1038/nbt.3122
SEQC/MAQC-III Consortium. A comprehensive assessment of RNA-seq accuracy, reproducibility and information content by the Sequencing Quality Control Consortium. Nat. Biotechnol. 32, 903–914 (2014).
doi: 10.1038/nbt.2957
Li, H. et al. The sequence alignment/map format and SAMtools. Bioinformatics 25, 2078–2079 (2009).
pubmed: 19505943
pmcid: 2723002
doi: 10.1093/bioinformatics/btp352
Patro, R., Duggal, G., Love, M. I., Irizarry, R. A. & Kingsford, C. Salmon provides fast and bias-aware quantification of transcript expression. Nat. Methods 14, 417–419 (2017).
pubmed: 28263959
pmcid: 5600148
doi: 10.1038/nmeth.4197
Kiełbasa, S. M., Wan, R., Sato, K., Horton, P. & Frith, M. C. Adaptive seeds tame genomic sequence comparison. Genome Res. 21, 487–493 (2011).
pubmed: 21209072
pmcid: 3044862
doi: 10.1101/gr.113985.110
Martin, M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet J. 17, 10 (2011).
doi: 10.14806/ej.17.1.200
Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15–21 (2013).
pubmed: 23104886
doi: 10.1093/bioinformatics/bts635
Szczerbinska, I. et al. A chemically defined feeder-free system for the establishment and maintenance of the human naive pluripotent state. Stem Cell Rep. 13, 612–626 (2019).
doi: 10.1016/j.stemcr.2019.08.005
SU, Y. et al. Comprehensive assessment of mRNA isoform detection methods for long-read sequencing data, yasim. https://doi.org/10.5281/zenodo.10908532 (2024).
SU, Y. et al. Comprehensive assessment of mRNA isoform detection methods for long-read sequencing data, 2024_LRS_AS_Benchmark_Code. https://doi.org/10.5281/zenodo.10912055 (2024).