mulea: An R package for enrichment analysis using multiple ontologies and empirical false discovery rate.

False discovery rate GMT files Gene set enrichment Ontologies Overrepresentation analysis R package

Journal

BMC bioinformatics
ISSN: 1471-2105
Titre abrégé: BMC Bioinformatics
Pays: England
ID NLM: 100965194

Informations de publication

Date de publication:
18 Oct 2024
Historique:
received: 21 05 2024
accepted: 26 09 2024
medline: 19 10 2024
pubmed: 19 10 2024
entrez: 18 10 2024
Statut: epublish

Résumé

Traditional gene set enrichment analyses are typically limited to a few ontologies and do not account for the interdependence of gene sets or terms, resulting in overcorrected p-values. To address these challenges, we introduce mulea, an R package offering comprehensive overrepresentation and functional enrichment analysis. mulea employs a progressive empirical false discovery rate (eFDR) method, specifically designed for interconnected biological data, to accurately identify significant terms within diverse ontologies. mulea expands beyond traditional tools by incorporating a wide range of ontologies, encompassing Gene Ontology, pathways, regulatory elements, genomic locations, and protein domains. This flexibility enables researchers to tailor enrichment analysis to their specific questions, such as identifying enriched transcriptional regulators in gene expression data or overrepresented protein domains in protein sets. To facilitate seamless analysis, mulea provides gene sets (in standardised GMT format) for 27 model organisms, covering 22 ontology types from 16 databases and various identifiers resulting in almost 900 files. Additionally, the muleaData ExperimentData Bioconductor package simplifies access to these pre-defined ontologies. Finally, mulea's architecture allows for easy integration of user-defined ontologies, or GMT files from external sources (e.g., MSigDB or Enrichr), expanding its applicability across diverse research areas. mulea is distributed as a CRAN R package downloadable from https://cran.r-project.org/web/packages/mulea/ and https://github.com/ELTEbioinformatics/mulea . It offers researchers a powerful and flexible toolkit for functional enrichment analysis, addressing limitations of traditional tools with its progressive eFDR and by supporting a variety of ontologies. Overall, mulea fosters the exploration of diverse biological questions across various model organisms.

Identifiants

pubmed: 39425047
doi: 10.1186/s12859-024-05948-7
pii: 10.1186/s12859-024-05948-7
doi:

Types de publication

Journal Article

Langues

eng

Sous-ensembles de citation

IM

Pagination

334

Subventions

Organisme : Biotechnology and Biological Sciences Research Council
ID : BB/CSP1720/1
Pays : United Kingdom
Organisme : Biotechnology and Biological Sciences Research Council
ID : BBS/E/F/000PR13631
Pays : United Kingdom
Organisme : Biotechnology and Biological Sciences Research Council
ID : BBS/E/F/000PR13631
Pays : United Kingdom
Organisme : Nemzeti Kutatási, Fejlesztési és Innovaciós Alap
ID : KDP-2023-C2285907
Organisme : National Research, Development and Innovation Office
ID : KKP 129814
Organisme : European Union's Horizon 2020
ID : 739593
Organisme : National Laboratory of Health Security, Hungary
ID : RRF-2.3.1-21-2022-00006
Organisme : National Laboratory of Health Security, Hungary
ID : RRF-2.3.1-21-2022-00006

Informations de copyright

© 2024. The Author(s).

Références

Dennis G, Sherman BT, Hosack DA, Yang J, Gao W, Lane HC, et al. DAVID: Database for Annotation, Visualization, and Integrated Discovery. Genome Biol. 2003;4:P3.
doi: 10.1186/gb-2003-4-5-p3 pubmed: 12734009
Zhang B, Kirov S, Snoddy J. WebGestalt: an integrated system for exploring gene sets in various biological contexts. Nucleic Acids Res. 2005;33:W741-748.
doi: 10.1093/nar/gki475 pubmed: 15980575 pmcid: 1160236
The Gene Ontology Consortium. The Gene Ontology resource: enriching a GOld mine. Nucleic Acids Res. 2021;49:D325–34.
doi: 10.1093/nar/gkaa1113
Kanehisa M, Furumichi M, Sato Y, Ishiguro-Watanabe M, Tanabe M. KEGG: integrating viruses and cellular organisms. Nucleic Acids Res. 2021;49:D545–51.
doi: 10.1093/nar/gkaa970 pubmed: 33125081
Subramanian A, Tamayo P, Mootha VK, Mukherjee S, Ebert BL, Gillette MA, et al. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles. Proc Natl Acad Sci. 2005;102:15545–50.
doi: 10.1073/pnas.0506580102 pubmed: 16199517 pmcid: 1239896
Chen EY, Tan CM, Kou Y, Duan Q, Wang Z, Meirelles GV, et al. Enrichr: interactive and collaborative HTML5 gene list enrichment analysis tool. BMC Bioinform. 2013;14:128.
doi: 10.1186/1471-2105-14-128
Raudvere U, Kolberg L, Kuzmin I, Arak T, Adler P, Peterson H, et al. g:Profiler: a web server for functional enrichment analysis and conversions of gene lists (2019 update). Nucleic Acids Res. 2019;47:W191–8.
doi: 10.1093/nar/gkz369 pubmed: 31066453 pmcid: 6602461
Wu T, Hu E, Xu S, Chen M, Guo P, Dai Z, et al. clusterProfiler 4.0: a universal enrichment tool for interpreting omics data. Innovation. 2021;2:100141.
pubmed: 34557778 pmcid: 8454663
Krämer A, Green J, Pollard J Jr, Tugendreich S. Causal analysis approaches in ingenuity pathway analysis. Bioinformatics. 2014;30:523–30.
doi: 10.1093/bioinformatics/btt703 pubmed: 24336805
Ari E, Ölbei M, Gul L, Bohár B. muleaData: genes sets for functional enrichment analysis with the “mulea” R package. 2024.
Bonferroni CE. Teoria statistica delle classi e calcolo delle probabilità. 8th edition. Florence, Italy: Seeber; 1936.
Benjamini Y, Hochberg Y. Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Ser B Methodol. 1995;57:289–300.
doi: 10.1111/j.2517-6161.1995.tb02031.x
Hastie T, Tibshirani R, Friedman J. The elements of statistical learning: data mining, inference, and prediction. 2nd ed. New York: Springer; 2009.
doi: 10.1007/978-0-387-84858-7
Reiner A, Yekutieli D, Benjamini Y. Identifying differentially expressed genes using false discovery rate controlling procedures. Bioinformatics. 2003;19:368–75.
doi: 10.1093/bioinformatics/btf877 pubmed: 12584122
Kofler R, Schlötterer C. Gowinda: unbiased analysis of gene set enrichment for genome-wide association studies. Bioinformatics. 2012;28:2084–5.
doi: 10.1093/bioinformatics/bts315 pubmed: 22635606 pmcid: 3400962
Berriz GF, Beaver JE, Cenik C, Tasan M, Roth FP. Next generation software for functional trend analysis. Bioinformatics. 2009;25:3043–4.
doi: 10.1093/bioinformatics/btp498 pubmed: 19717575 pmcid: 2800365
Korotkevich G, Sukhov V, Budin N, Shpak B, Artyomov MN, Sergushichev A. Fast gene set enrichment analysis. bioRxiv. 2021:060012.
Rodchenkov I, Babur O, Luna A, Aksoy BA, Wong JV, Fong D, et al. Pathway Commons 2019 Update: integration, analysis and exploration of pathway data. Nucleic Acids Res. 2020;48:D489–97.
pubmed: 31647099
Jassal B, Matthews L, Viteri G, Gong C, Lorente P, Fabregat A, et al. The reactome pathway knowledgebase. Nucleic Acids Res. 2020;48:D498-503.
pubmed: 31691815
Csabai L, Fazekas D, Kadlecsik T, Szalay-Bekő M, Bohár B, Madgwick M, et al. SignaLink3: a multi-layered resource to uncover tissue-specific signaling networks. Nucleic Acids Res. 2022;50:D701–9.
doi: 10.1093/nar/gkab909 pubmed: 34634810
Martens M, Ammar A, Riutta A, Waagmeester A, Slenter DN, Hanspers K, et al. WikiPathways: connecting communities. Nucleic Acids Res. 2021;49:D613–21.
doi: 10.1093/nar/gkaa1024 pubmed: 33211851
Jin J, He K, Tang X, Li Z, Lv L, Zhao Y, et al. An Arabidopsis transcriptional regulatory map reveals distinct functional and evolutionary features of novel transcription factors. Mol Biol Evol. 2015;32:1767–73.
doi: 10.1093/molbev/msv058 pubmed: 25750178 pmcid: 4476157
Garcia-Alonso L, Holland CH, Ibrahim MM, Turei D, Saez-Rodriguez J. Benchmark and integration of resources for the estimation of human transcription factor activities. Genome Res. 2019;29:1363–75.
doi: 10.1101/gr.240663.118 pubmed: 31340985 pmcid: 6673718
Tierrafría VH, Rioualen C, Salgado H, Lara P, Gama-Castro S, Lally P, et al. RegulonDB 11.0: comprehensive high-throughput datasets on transcriptional regulation in Escherichia coli K-12. Microb Genomics. 2022;8:000833.
doi: 10.1099/mgen.0.000833
Liska O, Bohár B, Hidas A, Korcsmáros T, Papp B, Fazekas D, et al. TFLink: an integrated gateway to access transcription factor–target gene interactions for multiple species. Database. 2022;2022:baac083.
doi: 10.1093/database/baac083 pubmed: 36124642 pmcid: 9480832
Han H, Cho J-W, Lee S, Yun A, Kim H, Bae D, et al. TRRUST v2: an expanded reference database of human and mouse transcriptional regulatory interactions. Nucleic Acids Res. 2018;46:D380–6.
doi: 10.1093/nar/gkx1013 pubmed: 29087512
Teixeira MC, Monteiro PT, Palma M, Costa C, Godinho CP, Pais P, et al. YEASTRACT: an upgraded database for the analysis of transcription regulatory networks in Saccharomyces cerevisiae. Nucleic Acids Res. 2018;46:D348–53.
doi: 10.1093/nar/gkx842 pubmed: 29036684
Huang H-Y, Lin Y-C-D, Cui S, Huang Y, Tang Y, Xu J, et al. miRTarBase update an informative resource for experimentally validated miRNA–target interactions. Nucleic Acids Res. 2022;2022(50):D222–30.
doi: 10.1093/nar/gkab1079
Chintapalli VR, Wang J, Dow JAT. Using FlyAtlas to identify better Drosophila melanogaster models of human disease. Nat Genet. 2007;39:715–20.
doi: 10.1038/ng2049 pubmed: 17534367
The Modencode Consortium, Roy S, Ernst J, Kharchenko PV, Kheradpour P, Negre N, et al. Identification of functional elements and regulatory circuits by Drosophila modENCODE. Science. 2010;330:1787–97.
doi: 10.1126/science.1198374 pmcid: 3192495
Martin FJ, Amode MR, Aneja A, Austine-Orimoloye O, Azov AG, Barnes I, et al. Ensembl 2023. Nucleic Acids Res. 2023;51:D933–41.
doi: 10.1093/nar/gkac958 pubmed: 36318249
Mistry J, Chuguransky S, Williams L, Qureshi M, Salazar GA, Sonnhammer ELL, et al. Pfam: The protein families database in 2021. Nucleic Acids Res. 2021;49:D412–9.
doi: 10.1093/nar/gkaa913 pubmed: 33125078
Morgan M, Carlson M, Tenenbaum D, Arora S, Oberchain V, Morrell K, et al. ExperimentHub: client to access ExperimentHub resources. R package version 2.10.0. 2024.
Evangelista JE, Xie Z, Marino GB, Nguyen N, Clarke DJB, Maayan A. Enrichr-KG: bridging enrichment analysis across multiple libraries. Nucleic Acids Res. 2023;51:W168–79.
doi: 10.1093/nar/gkad393 pubmed: 37166973 pmcid: 10320098
Gierlinski M. fenr: Fast functional enrichment for interactive applications. R package version 1.0.5. 2022.
Grote S. GOfuncR: Gene Ontology enrichment using FUNC. R package version 1.22.2. 2024.
Falcon S, Gentleman R. Using GOstats to test gene lists for GO term association. Bioinformatics. 2007;23:257–8.
doi: 10.1093/bioinformatics/btl567 pubmed: 17098774
Zhou Y, Zhou B, Pache L, Chang M, Khodabakhshi AH, Tanaseichuk O, et al. Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat Commun. 2019;10:1523.
doi: 10.1038/s41467-019-09234-6 pubmed: 30944313 pmcid: 6447622
Bao C, Wang S, Jiang L, Fang Z, Zou K, Lin J, et al. OpenXGR: a web-server update for genomic summary data interpretation. Nucleic Acids Res. 2023;51:W387–96.
doi: 10.1093/nar/gkad357 pubmed: 37158276 pmcid: 10320191
Alexa A, Rahnenfuhrer J. topGO: enrichment analysis for Gene Ontology. R package version 2.54.0. 2023.
Liao Y, Wang J, Jaehnig EJ, Shi Z, Zhang B. WebGestalt 2019: gene set analysis toolkit with revamped UIs and APIs. Nucleic Acids Res. 2019;47:W199-205.
doi: 10.1093/nar/gkz401 pubmed: 31114916 pmcid: 6602449
Méhi O, Bogos B, Csörgő B, Pál F, Nyerges Á, Papp B, et al. Perturbation of iron homeostasis promotes the evolution of antibiotic resistance. Mol Biol Evol. 2014;31:2793–804.
doi: 10.1093/molbev/msu223 pubmed: 25063442 pmcid: 4166929

Auteurs

Cezary Turek (C)

Earlham Institute, Norwich Research Park, Norwich, NR4 7UZ, UK.

Márton Ölbei (M)

Earlham Institute, Norwich Research Park, Norwich, NR4 7UZ, UK.
Department of Metabolism, Digestion and Reproduction, Imperial College London, The Commonwealth Building, The Hammersmith Hospital, Du Cane Road, London, W12 0NN, UK.

Tamás Stirling (T)

Synthetic and Systems Biology Unit, Institute of Biochemistry, HUN-REN Biological Research Centre, Temesvári Krt. 62, 6726, Szeged, Hungary.
HCEMM-BRC Metabolic Systems Biology Research Group, Temesvári Krt. 62, 6726, Szeged, Hungary.
Doctoral School of Biology, University of Szeged, Közép Fasor 52, 6726, Szeged, Hungary.

Gergely Fekete (G)

Synthetic and Systems Biology Unit, Institute of Biochemistry, HUN-REN Biological Research Centre, Temesvári Krt. 62, 6726, Szeged, Hungary.
HCEMM-BRC Metabolic Systems Biology Research Group, Temesvári Krt. 62, 6726, Szeged, Hungary.

Ervin Tasnádi (E)

Synthetic and Systems Biology Unit, Institute of Biochemistry, HUN-REN Biological Research Centre, Temesvári Krt. 62, 6726, Szeged, Hungary.
Doctoral School of Computer Science, University of Szeged, Árpád Tér 2, 6720, Szeged, Hungary.

Leila Gul (L)

Earlham Institute, Norwich Research Park, Norwich, NR4 7UZ, UK.
Department of Metabolism, Digestion and Reproduction, Imperial College London, The Commonwealth Building, The Hammersmith Hospital, Du Cane Road, London, W12 0NN, UK.

Balázs Bohár (B)

Department of Metabolism, Digestion and Reproduction, Imperial College London, The Commonwealth Building, The Hammersmith Hospital, Du Cane Road, London, W12 0NN, UK.
Synthetic and Systems Biology Unit, Institute of Biochemistry, HUN-REN Biological Research Centre, Temesvári Krt. 62, 6726, Szeged, Hungary.
Department of Genetics, ELTE Eötvös Loránd University, Pázmány P. Stny. 1/C, 1117, Budapest, Hungary.

Balázs Papp (B)

Synthetic and Systems Biology Unit, Institute of Biochemistry, HUN-REN Biological Research Centre, Temesvári Krt. 62, 6726, Szeged, Hungary.
HCEMM-BRC Metabolic Systems Biology Research Group, Temesvári Krt. 62, 6726, Szeged, Hungary.

Wiktor Jurkowski (W)

Earlham Institute, Norwich Research Park, Norwich, NR4 7UZ, UK.

Eszter Ari (E)

Synthetic and Systems Biology Unit, Institute of Biochemistry, HUN-REN Biological Research Centre, Temesvári Krt. 62, 6726, Szeged, Hungary. arieszter@gmail.com.
HCEMM-BRC Metabolic Systems Biology Research Group, Temesvári Krt. 62, 6726, Szeged, Hungary. arieszter@gmail.com.
Department of Genetics, ELTE Eötvös Loránd University, Pázmány P. Stny. 1/C, 1117, Budapest, Hungary. arieszter@gmail.com.

Articles similaires

Selecting optimal software code descriptors-The case of Java.

Yegor Bugayenko, Zamira Kholmatova, Artem Kruglov et al.
1.00
Software Algorithms Programming Languages

Exploring blood-brain barrier passage using atomic weighted vector and machine learning.

Yoan Martínez-López, Paulina Phoobane, Yanaima Jauriga et al.
1.00
Blood-Brain Barrier Machine Learning Humans Support Vector Machine Software
Cephalometry Humans Anatomic Landmarks Software Internet
Humans Colorectal Neoplasms Biomarkers, Tumor Prognosis Gene Expression Regulation, Neoplastic

Classifications MeSH