FASTAFS: file system virtualisation of random access compressed FASTA files.

FASTA FASTAFS FUSE Integrity Metadata Random access Virtualisation Zstd

Journal

BMC bioinformatics
ISSN: 1471-2105
Titre abrégé: BMC Bioinformatics
Pays: England
ID NLM: 100965194

Informations de publication

Date de publication:
01 Nov 2021
Historique:
received: 09 08 2021
accepted: 20 10 2021
entrez: 2 11 2021
pubmed: 3 11 2021
medline: 4 11 2021
Statut: epublish

Résumé

The FASTA file format, used to store polymeric sequence data, has become a bioinformatics file standard used for decades. The relatively large files require additional files, beyond the scope of the original format, to identify sequences and to provide random access. Multiple compressors have been developed to archive FASTA files back and forth, but these lack direct access to targeted content or metadata of the archive. Moreover, these solutions are not directly backwards compatible to FASTA files, resulting in limited software integration. We designed a linux based toolkit that virtualises the content of DNA, RNA and protein FASTA archives into the filesystem by using filesystem in userspace. This guarantees in-sync virtualised metadata files and offers fast random-access decompression using bit encodings plus Zstandard (zstd). The toolkit, FASTAFS, can track all its system-wide running instances, allows file integrity verification and can provide, instantly, scriptable access to sequence files and is easy to use and deploy. The file compression ratios were comparable but not superior to other state of the art archival tools, despite the innovative random access feature implemented in FASTAFS. FASTAFS is a user-friendly and easy to deploy backwards compatible generic purpose solution to store and access compressed FASTA files, since it offers file system access to FASTA files as well as in-sync metadata files through file virtualisation. Using virtual filesystems as in-between layer offers format conversion without the need to rewrite code into different programming languages while preserving compatibility.

Sections du résumé

BACKGROUND BACKGROUND
The FASTA file format, used to store polymeric sequence data, has become a bioinformatics file standard used for decades. The relatively large files require additional files, beyond the scope of the original format, to identify sequences and to provide random access. Multiple compressors have been developed to archive FASTA files back and forth, but these lack direct access to targeted content or metadata of the archive. Moreover, these solutions are not directly backwards compatible to FASTA files, resulting in limited software integration.
RESULTS RESULTS
We designed a linux based toolkit that virtualises the content of DNA, RNA and protein FASTA archives into the filesystem by using filesystem in userspace. This guarantees in-sync virtualised metadata files and offers fast random-access decompression using bit encodings plus Zstandard (zstd). The toolkit, FASTAFS, can track all its system-wide running instances, allows file integrity verification and can provide, instantly, scriptable access to sequence files and is easy to use and deploy. The file compression ratios were comparable but not superior to other state of the art archival tools, despite the innovative random access feature implemented in FASTAFS.
CONCLUSIONS CONCLUSIONS
FASTAFS is a user-friendly and easy to deploy backwards compatible generic purpose solution to store and access compressed FASTA files, since it offers file system access to FASTA files as well as in-sync metadata files through file virtualisation. Using virtual filesystems as in-between layer offers format conversion without the need to rewrite code into different programming languages while preserving compatibility.

Identifiants

pubmed: 34724897
doi: 10.1186/s12859-021-04455-3
pii: 10.1186/s12859-021-04455-3
pmc: PMC8558547
doi:

Substances chimiques

Proteins 0

Types de publication

Journal Article

Langues

eng

Sous-ensembles de citation

IM

Pagination

535

Informations de copyright

© 2021. The Author(s).

Références

Nucleic Acids Res. 2019 Jan 8;47(D1):D155-D162
pubmed: 30423142
Bioinformatics. 2018 Oct 15;34(20):3600
pubmed: 29788404
Bioinformatics. 2012 Dec 15;28(24):3211-7
pubmed: 23071270
Genome Res. 2010 Sep;20(9):1297-303
pubmed: 20644199
Nucleic Acids Res. 2013 May 1;41(10):e108
pubmed: 23558742
Bioinformatics. 2014 Aug 1;30(15):2213-5
pubmed: 24747219
Nucleic Acids Res. 2004 Jan 1;32(Database issue):D115-9
pubmed: 14681372
PLoS One. 2016 Oct 5;11(10):e0163962
pubmed: 27706213
Nat Biotechnol. 2017 Apr 11;35(4):316-319
pubmed: 28398311
Bioinformatics. 2014 Jan 1;30(1):117-8
pubmed: 24132931
Genome Res. 2012 Mar;22(3):568-76
pubmed: 22300766
Bioinformation. 2011 Jan 22;5(8):350-60
pubmed: 21383923
Nucleic Acids Res. 2013 Jan;41(Database issue):D590-6
pubmed: 23193283
Genome Biol. 2016 Apr 12;17:66
pubmed: 27072794
Mol Cell. 2010 May 28;38(4):576-89
pubmed: 20513432
Bioinformatics. 2019 Oct 1;35(19):3826-3828
pubmed: 30799504
Bioinformatics. 2013 Jan 1;29(1):15-21
pubmed: 23104886

Auteurs

Youri Hoogstrate (Y)

Department of Neurology, Erasmus University Medical Center, Dr. Molewaterplein 40, 3015 GD, Rotterdam, The Netherlands. yhoogstrate@erasmusmc.nl.
Department of Urology, Erasmus MC Cancer Institute, University Medical Center, Rotterdam, The Netherlands. yhoogstrate@erasmusmc.nl.
Cancer Computational Biology Center, Erasmus MC Cancer Institute, University Medical Center, Rotterdam, The Netherlands. yhoogstrate@erasmusmc.nl.

Guido W Jenster (GW)

Department of Urology, Erasmus MC Cancer Institute, University Medical Center, Rotterdam, The Netherlands.

Harmen J G van de Werken (HJG)

Department of Urology, Erasmus MC Cancer Institute, University Medical Center, Rotterdam, The Netherlands.
Cancer Computational Biology Center, Erasmus MC Cancer Institute, University Medical Center, Rotterdam, The Netherlands.
Department of Immunology, Erasmus MC Cancer Institute, University Medical Center, Rotterdam, The Netherlands.

Articles similaires

Selecting optimal software code descriptors-The case of Java.

Yegor Bugayenko, Zamira Kholmatova, Artem Kruglov et al.
1.00
Software Algorithms Programming Languages
Databases, Protein Protein Domains Protein Folding Proteins Deep Learning

Exploring blood-brain barrier passage using atomic weighted vector and machine learning.

Yoan Martínez-López, Paulina Phoobane, Yanaima Jauriga et al.
1.00
Blood-Brain Barrier Machine Learning Humans Support Vector Machine Software
Cephalometry Humans Anatomic Landmarks Software Internet

Classifications MeSH