mulea: An R package for enrichment analysis using multiple ontologies and empirical false discovery rate.
False discovery rate
GMT files
Gene set enrichment
Ontologies
Overrepresentation analysis
R package
Journal
BMC bioinformatics
ISSN: 1471-2105
Titre abrégé: BMC Bioinformatics
Pays: England
ID NLM: 100965194
Informations de publication
Date de publication:
18 Oct 2024
18 Oct 2024
Historique:
received:
21
05
2024
accepted:
26
09
2024
medline:
19
10
2024
pubmed:
19
10
2024
entrez:
18
10
2024
Statut:
epublish
Résumé
Traditional gene set enrichment analyses are typically limited to a few ontologies and do not account for the interdependence of gene sets or terms, resulting in overcorrected p-values. To address these challenges, we introduce mulea, an R package offering comprehensive overrepresentation and functional enrichment analysis. mulea employs a progressive empirical false discovery rate (eFDR) method, specifically designed for interconnected biological data, to accurately identify significant terms within diverse ontologies. mulea expands beyond traditional tools by incorporating a wide range of ontologies, encompassing Gene Ontology, pathways, regulatory elements, genomic locations, and protein domains. This flexibility enables researchers to tailor enrichment analysis to their specific questions, such as identifying enriched transcriptional regulators in gene expression data or overrepresented protein domains in protein sets. To facilitate seamless analysis, mulea provides gene sets (in standardised GMT format) for 27 model organisms, covering 22 ontology types from 16 databases and various identifiers resulting in almost 900 files. Additionally, the muleaData ExperimentData Bioconductor package simplifies access to these pre-defined ontologies. Finally, mulea's architecture allows for easy integration of user-defined ontologies, or GMT files from external sources (e.g., MSigDB or Enrichr), expanding its applicability across diverse research areas. mulea is distributed as a CRAN R package downloadable from https://cran.r-project.org/web/packages/mulea/ and https://github.com/ELTEbioinformatics/mulea . It offers researchers a powerful and flexible toolkit for functional enrichment analysis, addressing limitations of traditional tools with its progressive eFDR and by supporting a variety of ontologies. Overall, mulea fosters the exploration of diverse biological questions across various model organisms.
Identifiants
pubmed: 39425047
doi: 10.1186/s12859-024-05948-7
pii: 10.1186/s12859-024-05948-7
doi:
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Pagination
334Subventions
Organisme : Biotechnology and Biological Sciences Research Council
ID : BB/CSP1720/1
Pays : United Kingdom
Organisme : Biotechnology and Biological Sciences Research Council
ID : BBS/E/F/000PR13631
Pays : United Kingdom
Organisme : Biotechnology and Biological Sciences Research Council
ID : BBS/E/F/000PR13631
Pays : United Kingdom
Organisme : Nemzeti Kutatási, Fejlesztési és Innovaciós Alap
ID : KDP-2023-C2285907
Organisme : National Research, Development and Innovation Office
ID : KKP 129814
Organisme : European Union's Horizon 2020
ID : 739593
Organisme : National Laboratory of Health Security, Hungary
ID : RRF-2.3.1-21-2022-00006
Organisme : National Laboratory of Health Security, Hungary
ID : RRF-2.3.1-21-2022-00006
Informations de copyright
© 2024. The Author(s).
Références
Dennis G, Sherman BT, Hosack DA, Yang J, Gao W, Lane HC, et al. DAVID: Database for Annotation, Visualization, and Integrated Discovery. Genome Biol. 2003;4:P3.
doi: 10.1186/gb-2003-4-5-p3
pubmed: 12734009
Zhang B, Kirov S, Snoddy J. WebGestalt: an integrated system for exploring gene sets in various biological contexts. Nucleic Acids Res. 2005;33:W741-748.
doi: 10.1093/nar/gki475
pubmed: 15980575
pmcid: 1160236
The Gene Ontology Consortium. The Gene Ontology resource: enriching a GOld mine. Nucleic Acids Res. 2021;49:D325–34.
doi: 10.1093/nar/gkaa1113
Kanehisa M, Furumichi M, Sato Y, Ishiguro-Watanabe M, Tanabe M. KEGG: integrating viruses and cellular organisms. Nucleic Acids Res. 2021;49:D545–51.
doi: 10.1093/nar/gkaa970
pubmed: 33125081
Subramanian A, Tamayo P, Mootha VK, Mukherjee S, Ebert BL, Gillette MA, et al. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proc Natl Acad Sci. 2005;102:15545–50.
doi: 10.1073/pnas.0506580102
pubmed: 16199517
pmcid: 1239896
Chen EY, Tan CM, Kou Y, Duan Q, Wang Z, Meirelles GV, et al. Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool. BMC Bioinform. 2013;14:128.
doi: 10.1186/1471-2105-14-128
Raudvere U, Kolberg L, Kuzmin I, Arak T, Adler P, Peterson H, et al. g:Profiler: a web server for functional enrichment analysis and conversions of gene lists (2019 update). Nucleic Acids Res. 2019;47:W191–8.
doi: 10.1093/nar/gkz369
pubmed: 31066453
pmcid: 6602461
Wu T, Hu E, Xu S, Chen M, Guo P, Dai Z, et al. clusterProfiler 4.0: a universal enrichment tool for interpreting omics data. Innovation. 2021;2:100141.
pubmed: 34557778
pmcid: 8454663
Krämer A, Green J, Pollard J Jr, Tugendreich S. Causal analysis approaches in ingenuity pathway analysis. Bioinformatics. 2014;30:523–30.
doi: 10.1093/bioinformatics/btt703
pubmed: 24336805
Ari E, Ölbei M, Gul L, Bohár B. muleaData: genes sets for functional enrichment analysis with the “mulea” R package. 2024.
Bonferroni CE. Teoria statistica delle classi e calcolo delle probabilità. 8th edition. Florence, Italy: Seeber; 1936.
Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Ser B Methodol. 1995;57:289–300.
doi: 10.1111/j.2517-6161.1995.tb02031.x
Hastie T, Tibshirani R, Friedman J. The elements of statistical learning: data mining, inference, and prediction. 2nd ed. New York: Springer; 2009.
doi: 10.1007/978-0-387-84858-7
Reiner A, Yekutieli D, Benjamini Y. Identifying differentially expressed genes using false discovery rate controlling procedures. Bioinformatics. 2003;19:368–75.
doi: 10.1093/bioinformatics/btf877
pubmed: 12584122
Kofler R, Schlötterer C. Gowinda: unbiased analysis of gene set enrichment for genome-wide association studies. Bioinformatics. 2012;28:2084–5.
doi: 10.1093/bioinformatics/bts315
pubmed: 22635606
pmcid: 3400962
Berriz GF, Beaver JE, Cenik C, Tasan M, Roth FP. Next generation software for functional trend analysis. Bioinformatics. 2009;25:3043–4.
doi: 10.1093/bioinformatics/btp498
pubmed: 19717575
pmcid: 2800365
Korotkevich G, Sukhov V, Budin N, Shpak B, Artyomov MN, Sergushichev A. Fast gene set enrichment analysis. bioRxiv. 2021:060012.
Rodchenkov I, Babur O, Luna A, Aksoy BA, Wong JV, Fong D, et al. Pathway Commons 2019 Update: integration, analysis and exploration of pathway data. Nucleic Acids Res. 2020;48:D489–97.
pubmed: 31647099
Jassal B, Matthews L, Viteri G, Gong C, Lorente P, Fabregat A, et al. The reactome pathway knowledgebase. Nucleic Acids Res. 2020;48:D498-503.
pubmed: 31691815
Csabai L, Fazekas D, Kadlecsik T, Szalay-Bekő M, Bohár B, Madgwick M, et al. SignaLink3: a multi-layered resource to uncover tissue-specific signaling networks. Nucleic Acids Res. 2022;50:D701–9.
doi: 10.1093/nar/gkab909
pubmed: 34634810
Martens M, Ammar A, Riutta A, Waagmeester A, Slenter DN, Hanspers K, et al. WikiPathways: connecting communities. Nucleic Acids Res. 2021;49:D613–21.
doi: 10.1093/nar/gkaa1024
pubmed: 33211851
Jin J, He K, Tang X, Li Z, Lv L, Zhao Y, et al. An Arabidopsis transcriptional regulatory map reveals distinct functional and evolutionary features of novel transcription factors. Mol Biol Evol. 2015;32:1767–73.
doi: 10.1093/molbev/msv058
pubmed: 25750178
pmcid: 4476157
Garcia-Alonso L, Holland CH, Ibrahim MM, Turei D, Saez-Rodriguez J. Benchmark and integration of resources for the estimation of human transcription factor activities. Genome Res. 2019;29:1363–75.
doi: 10.1101/gr.240663.118
pubmed: 31340985
pmcid: 6673718
Tierrafría VH, Rioualen C, Salgado H, Lara P, Gama-Castro S, Lally P, et al. RegulonDB 11.0: comprehensive high-throughput datasets on transcriptional regulation in Escherichia coli K-12. Microb Genomics. 2022;8:000833.
doi: 10.1099/mgen.0.000833
Liska O, Bohár B, Hidas A, Korcsmáros T, Papp B, Fazekas D, et al. TFLink: an integrated gateway to access transcription factor–target gene interactions for multiple species. Database. 2022;2022:baac083.
doi: 10.1093/database/baac083
pubmed: 36124642
pmcid: 9480832
Han H, Cho J-W, Lee S, Yun A, Kim H, Bae D, et al. TRRUST v2: an expanded reference database of human and mouse transcriptional regulatory interactions. Nucleic Acids Res. 2018;46:D380–6.
doi: 10.1093/nar/gkx1013
pubmed: 29087512
Teixeira MC, Monteiro PT, Palma M, Costa C, Godinho CP, Pais P, et al. YEASTRACT: an upgraded database for the analysis of transcription regulatory networks in Saccharomyces cerevisiae. Nucleic Acids Res. 2018;46:D348–53.
doi: 10.1093/nar/gkx842
pubmed: 29036684
Huang H-Y, Lin Y-C-D, Cui S, Huang Y, Tang Y, Xu J, et al. miRTarBase update an informative resource for experimentally validated miRNA–target interactions. Nucleic Acids Res. 2022;2022(50):D222–30.
doi: 10.1093/nar/gkab1079
Chintapalli VR, Wang J, Dow JAT. Using FlyAtlas to identify better Drosophila melanogaster models of human disease. Nat Genet. 2007;39:715–20.
doi: 10.1038/ng2049
pubmed: 17534367
The Modencode Consortium, Roy S, Ernst J, Kharchenko PV, Kheradpour P, Negre N, et al. Identification of functional elements and regulatory circuits by Drosophila modENCODE. Science. 2010;330:1787–97.
doi: 10.1126/science.1198374
pmcid: 3192495
Martin FJ, Amode MR, Aneja A, Austine-Orimoloye O, Azov AG, Barnes I, et al. Ensembl 2023. Nucleic Acids Res. 2023;51:D933–41.
doi: 10.1093/nar/gkac958
pubmed: 36318249
Mistry J, Chuguransky S, Williams L, Qureshi M, Salazar GA, Sonnhammer ELL, et al. Pfam: The protein families database in 2021. Nucleic Acids Res. 2021;49:D412–9.
doi: 10.1093/nar/gkaa913
pubmed: 33125078
Morgan M, Carlson M, Tenenbaum D, Arora S, Oberchain V, Morrell K, et al. ExperimentHub: client to access ExperimentHub resources. R package version 2.10.0. 2024.
Evangelista JE, Xie Z, Marino GB, Nguyen N, Clarke DJB, Maayan A. Enrichr-KG: bridging enrichment analysis across multiple libraries. Nucleic Acids Res. 2023;51:W168–79.
doi: 10.1093/nar/gkad393
pubmed: 37166973
pmcid: 10320098
Gierlinski M. fenr: Fast functional enrichment for interactive applications. R package version 1.0.5. 2022.
Grote S. GOfuncR: Gene Ontology enrichment using FUNC. R package version 1.22.2. 2024.
Falcon S, Gentleman R. Using GOstats to test gene lists for GO term association. Bioinformatics. 2007;23:257–8.
doi: 10.1093/bioinformatics/btl567
pubmed: 17098774
Zhou Y, Zhou B, Pache L, Chang M, Khodabakhshi AH, Tanaseichuk O, et al. Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat Commun. 2019;10:1523.
doi: 10.1038/s41467-019-09234-6
pubmed: 30944313
pmcid: 6447622
Bao C, Wang S, Jiang L, Fang Z, Zou K, Lin J, et al. OpenXGR: a web-server update for genomic summary data interpretation. Nucleic Acids Res. 2023;51:W387–96.
doi: 10.1093/nar/gkad357
pubmed: 37158276
pmcid: 10320191
Alexa A, Rahnenfuhrer J. topGO: enrichment analysis for Gene Ontology. R package version 2.54.0. 2023.
Liao Y, Wang J, Jaehnig EJ, Shi Z, Zhang B. WebGestalt 2019: gene set analysis toolkit with revamped UIs and APIs. Nucleic Acids Res. 2019;47:W199-205.
doi: 10.1093/nar/gkz401
pubmed: 31114916
pmcid: 6602449
Méhi O, Bogos B, Csörgő B, Pál F, Nyerges Á, Papp B, et al. Perturbation of iron homeostasis promotes the evolution of antibiotic resistance. Mol Biol Evol. 2014;31:2793–804.
doi: 10.1093/molbev/msu223
pubmed: 25063442
pmcid: 4166929