The GEA pipeline for characterizing Escherichia coli and Salmonella genomes.
Journal
Scientific reports
ISSN: 2045-2322
Titre abrégé: Sci Rep
Pays: England
ID NLM: 101563288
Informations de publication
Date de publication:
10 06 2024
10 06 2024
Historique:
received:
24
01
2024
accepted:
03
06
2024
medline:
11
6
2024
pubmed:
11
6
2024
entrez:
10
6
2024
Statut:
epublish
Résumé
Salmonella enterica and Escherichia coli are major food-borne human pathogens, and their genomes are routinely sequenced for clinical surveillance. Computational pipelines designed for analyzing pathogen genomes should both utilize the most current information from annotation databases and increase the coverage of these databases over time. We report the development of the GEA pipeline to analyze large batches of E. coli and S. enterica genomes. The GEA pipeline takes as input paired Illumina raw reads files which are then assembled followed by annotation. Alternatively, assemblies can be provided as input and directly annotated. The pipeline provides predictive genome annotations for E. coli and S. enterica with a focus on the Center for Genomic Epidemiology tools. Annotation results are provided as a tab delimited text file. The GEA pipeline is designed for large-scale E. coli and S. enterica genome assembly and characterization using the Center for Genomic Epidemiology command-line tools and high-performance computing. Large scale annotation is demonstrated by an analysis of more than 14,000 Salmonella genome assemblies. Testing the GEA pipeline on E. coli raw reads demonstrates reproducibility across multiple compute environments and computational usage is optimized on high performance computers.
Identifiants
pubmed: 38858528
doi: 10.1038/s41598-024-63832-z
pii: 10.1038/s41598-024-63832-z
doi:
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Pagination
13257Informations de copyright
© 2024. This is a U.S. Government work and not under copyright protection in the US; foreign copyright protection may apply.
Références
Scallan, E. et al. Foodborne illness acquired in the United States–major pathogens. Emerg. Infect. Dis. 17, 7–15. https://doi.org/10.3201/eid1701.P11101 (2011).
doi: 10.3201/eid1701.P11101
pubmed: 21192848
pmcid: 3375761
Fjukstad, B. & Bongo, L. A. A review of scalable bioinformatics pipelines. Data Sci. Eng. 2, 245–251. https://doi.org/10.1007/s41019-017-0047-z (2017).
doi: 10.1007/s41019-017-0047-z
Leipzig, J. A review of bioinformatic pipeline frameworks. Brief. Bioinform. 18, 530–536. https://doi.org/10.1093/bib/bbw020 (2017).
doi: 10.1093/bib/bbw020
pubmed: 27013646
Center for Genomic Epidemiology Repositories [internet]. [cited 29 November 2023]. https://bitbucket.org/genomicepidemiology/workspace/repositories/
Chukamnerd, A. et al. BacSeq: A user-friendly automated pipeline for whole-genome sequence analysis of bacterial genomes. Microorganisms. 11, 1769. https://doi.org/10.3390/microorganisms11071769 (2023).
doi: 10.3390/microorganisms11071769
pubmed: 37512941
pmcid: 10385524
Couvin, D., Stattner, E., Segretier, W., Cazenave, D. & Rastogi, N. simpiTB–a pipeline designed to extract meaningful information from whole genome sequencing data of Mycobacterium tuberculosis complex, allows to combine genomic, phylogenetic and clustering analyses in existing SITVIT databases. Infect. Genet. Evol. 113, 105466. https://doi.org/10.1016/j.meegid.2023.105466 (2023).
doi: 10.1016/j.meegid.2023.105466
pubmed: 37331497
Quijada, N. M., Rodríguez-Lázaro, D., Eiros, J. M. & Hernández, M. TORMES: An automated pipeline for whole bacterial genome analysis. Bioinformatics 35, 4207–4212. https://doi.org/10.1093/bioinformatics/btz220 (2019).
doi: 10.1093/bioinformatics/btz220
pubmed: 30957837
Thomsen, M. C. F. et al. A bacterial analysis platform: An integrated system for analysing bacterial whole genome sequencing data for clinical diagnostics and surveillance. PloS One 11, e0157718. https://doi.org/10.1371/journal.pone.0157718 (2016).
doi: 10.1371/journal.pone.0157718
pubmed: 27327771
pmcid: 4915688
Xavier, B. B. et al. BacPipe: A rapid, user-friendly whole-genome sequencing pipeline for clinical diagnostic bacteriology. IScience. 23, 100769. https://doi.org/10.1016/j.isci.2019.100769 (2020).
doi: 10.1016/j.isci.2019.100769
pubmed: 31887656
Roer, L. et al. Development of a web tool for Escherichia coli subtyping based on fimH alleles. J. Clin. Microbiol. 55, 2538–2543. https://doi.org/10.1128/jcm.00737-17 (2017).
doi: 10.1128/jcm.00737-17
pubmed: 28592545
pmcid: 5527432
Larsen, M. V. et al. Multilocus sequence typing of total-genome-sequenced bacteria. J. Clin. Microbiol. 50, 1355–1361. https://doi.org/10.1128/jcm.06094-11 (2012).
doi: 10.1128/jcm.06094-11
pubmed: 22238442
pmcid: 3318499
Carattoli, A. & Hasman, H. PlasmidFinder and in silico pMLST: Identification and typing of plasmid replicons in whole-genome sequencing (WGS). Horiz. Gene Transf. 285–294. (2020). https://doi.org/10.1007/978-1-4939-9877-7_20
Bortolaia, V. et al. ResFinder 4.0 for predictions of phenotypes from genotypes. J. Antimicrob. Chemoth. 75, 3491–3500. https://doi.org/10.1093/jac/dkaa345 (2020).
doi: 10.1093/jac/dkaa345
Joensen, K. G., Tetzschner, A. M. M., Iguchi, A., Aarestrup, F. M. & Scheutz, F. Rapid and easy in silico serotyping of Escherichia coli isolates by use of whole-genome sequencing data. J. Clin. Microbiol. 53, 2410–2426. https://doi.org/10.1128/jcm.00008-15 (2015).
doi: 10.1128/jcm.00008-15
pubmed: 25972421
pmcid: 4508402
Joensen, K. G. et al. Real-time whole-genome sequencing for routine typing, surveillance, and outbreak detection of verotoxigenic Escherichia coli. J. Clin. Microbiol. 52, 1501–1510. https://doi.org/10.1128/jcm.03617-13 (2014).
doi: 10.1128/jcm.03617-13
pubmed: 24574290
pmcid: 3993690
Kurtzer, G. M., Sochat, V. & Bauer, M. W. Singularity: Scientific containers for mobility of compute. PloS One 12, e0177459. https://doi.org/10.1371/journal.pone.0177459 (2017).
doi: 10.1371/journal.pone.0177459
pubmed: 28494014
pmcid: 5426675
Grüning, B. et al. Practical computational reproducibility in the life sciences. Cell Syst. 6, 631–635. https://doi.org/10.1016/j.cels.2018.03.014 (2018).
doi: 10.1016/j.cels.2018.03.014
pubmed: 29953862
pmcid: 6263957
Sochat, V. The scientific filesystem. GigaScience giy023, 023. https://doi.org/10.1093/gigascience/giy023 (2018).
doi: 10.1093/gigascience/giy023
Guragain, M., Schmidt, J. W., Kalchayanand, N., Dickey, A. M. & Bosilevac, J. M. Characterization of Escherichia coli harboring colibactin genes (clb) isolated from beef production and processing systems. Sci. Rep. 12, 5305. https://doi.org/10.1038/s41598-022-09274-x (2022).
doi: 10.1038/s41598-022-09274-x
pubmed: 35351927
pmcid: 8964808
Guragain, M., Schmidt, J. W., Dickey, A. M. & Bosilevac, J. M. Distribution of extremely heat-resistant Escherichia coli in the beef production and processing continuum. J. Food Protect. 86, 100031. https://doi.org/10.1016/j.jfp.2022.100031 (2023).
doi: 10.1016/j.jfp.2022.100031
Schmidt, J. W. et al. Twenty-four-month longitudinal study suggests little to no horizontal gene transfer in situ between third-generation cephalosporin-resistant Salmonella and third-generation cephalosporin-resistant Escherichia coli in a beef cattle feedyard. J. Food Protect. 85, 323–335. https://doi.org/10.4315/JFP-21-371 (2022).
doi: 10.4315/JFP-21-371
SCINet Scientific Computing U.S. DEPARTMENT OF AGRICULTURE. [internet]. [cited 29 November 2023]. scinet.usda.gov
Walker, B. J. et al. Pilon: An integrated tool for comprehensive microbial variant detection and genome assembly improvement. PloS One 9, e112963. https://doi.org/10.1371/journal.pone.0112963 (2014).
doi: 10.1371/journal.pone.0112963
pubmed: 25409509
pmcid: 4237348
Souvorov, A., Agarwala, R. & Lipman, D. J. SKESA: Strategic k-mer extension for scrupulous assemblies. Genome Biol. 19, 153. https://doi.org/10.1186/s13059-018-1540-z (2018).
doi: 10.1186/s13059-018-1540-z
pubmed: 30286803
pmcid: 6172800
de Almeida, F. M., de Campos, T. A. & Pappas, G. J. Jr. Scalable and versatile container-based pipelines for de novo genome assembly and bacterial annotation. F1000Research 12, 1205. https://doi.org/10.12688/f1000research.139488.1 (2023).
doi: 10.12688/f1000research.139488.1
pubmed: 37970066
pmcid: 10646344
Steinke, K. et al. RSYD-BASIC: A bioinformatics pipeline for routine sequence analysis and data processing of bacterial isolates for clinical microbiology. Access Microbiol. https://doi.org/10.1099/acmi.0.000646.v2 (2023).
doi: 10.1099/acmi.0.000646.v2
Dykstra, D. Apptainer without setuid. arXiv 2208, 12106. [preprint] (2022). https://doi.org/10.48550/arXiv.2208.12106
Zhang, S. et al. SeqSero2: Rapid and improved Salmonella serotype determination using whole-genome sequencing data. Appl. Environ. Microbiol. 85, e01746-e1819. https://doi.org/10.1128/AEM.01746-19 (2019).
doi: 10.1128/AEM.01746-19
pubmed: 31540993
pmcid: 6856333
Camacho, C. et al. BLAST+: Architecture and applications. BMC Bioinform. 10, 44935. https://doi.org/10.1186/1471-2105-10-421 (2009).
doi: 10.1186/1471-2105-10-421
Waters, N. R., Abram, F., Brennan, F., Holmes, A. & Pritchard, L. Easy phylotyping of Escherichia coli via the EzClermont web app and command-line tool. Access Microbiol. 2, acmi000143. https://doi.org/10.1099/acmi.0.000143 (2020).
doi: 10.1099/acmi.0.000143
pubmed: 33195978
pmcid: 7656184
Bankevich, A. et al. SPAdes: A new genome assembly algorithm and its applications to single-cell sequencing. J. Comput. Biol. 19, 455–477. https://doi.org/10.1089/cmb.2012.0021 (2012).
doi: 10.1089/cmb.2012.0021
pubmed: 22506599
pmcid: 3342519
Zerbino, D. R. & Birney, E. Velvet: Algorithms for de novo short read assembly using de Bruijn graphs. Genome Res. 18, 821–829. https://doi.org/10.1101/gr.074492.107 (2008).
doi: 10.1101/gr.074492.107
pubmed: 18349386
pmcid: 2336801
Song, L., Florea, L. & Langmead, B. Lighter: Fast and memory-efficient sequencing error correction without counting. Genome Biol. 15, 44939. https://doi.org/10.1186/s13059-014-0509-9 (2014).
doi: 10.1186/s13059-014-0509-9
Magoč, T. & Salzberg, S. FLASH: Fast length adjustment of short reads to improve genome assemblies. Bioinformatics 27, 2957–2963. https://doi.org/10.1093/bioinformatics/btr507 (2011).
doi: 10.1093/bioinformatics/btr507
pubmed: 21903629
pmcid: 3198573
Danecek, P. et al. Twelve years of SAMtools and BCFtools. Gigascience 10, giab008. https://doi.org/10.1093/gigascience/giab008 (2021).
doi: 10.1093/gigascience/giab008
pubmed: 33590861
pmcid: 7931819
Li, H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv 1303, 3997 [preprint] (2013). https://doi.org/10.48550/arXiv.1303.3997
Kokot, M., Długosz, M. & Deorowicz, S. KMC 3: Counting and manipulating k-mer statistics. Bioinformatics 33, 2759–2761. https://doi.org/10.1093/bioinformatics/btx304 (2017).
doi: 10.1093/bioinformatics/btx304
pubmed: 28472236
Bolger, A. M., Lohse, M. & Usadel, B. Trimmomatic: A flexible trimmer for Illumina sequence data. Bioinformatics 30, 2114–2120. https://doi.org/10.1093/bioinformatics/btu170 (2014).
doi: 10.1093/bioinformatics/btu170
pubmed: 24695404
pmcid: 4103590