FASTAFS: file system virtualisation of random access compressed FASTA files.
FASTA
FASTAFS
FUSE
Integrity
Metadata
Random access
Virtualisation
Zstd
Journal
BMC bioinformatics
ISSN: 1471-2105
Titre abrégé: BMC Bioinformatics
Pays: England
ID NLM: 100965194
Informations de publication
Date de publication:
01 Nov 2021
01 Nov 2021
Historique:
received:
09
08
2021
accepted:
20
10
2021
entrez:
2
11
2021
pubmed:
3
11
2021
medline:
4
11
2021
Statut:
epublish
Résumé
The FASTA file format, used to store polymeric sequence data, has become a bioinformatics file standard used for decades. The relatively large files require additional files, beyond the scope of the original format, to identify sequences and to provide random access. Multiple compressors have been developed to archive FASTA files back and forth, but these lack direct access to targeted content or metadata of the archive. Moreover, these solutions are not directly backwards compatible to FASTA files, resulting in limited software integration. We designed a linux based toolkit that virtualises the content of DNA, RNA and protein FASTA archives into the filesystem by using filesystem in userspace. This guarantees in-sync virtualised metadata files and offers fast random-access decompression using bit encodings plus Zstandard (zstd). The toolkit, FASTAFS, can track all its system-wide running instances, allows file integrity verification and can provide, instantly, scriptable access to sequence files and is easy to use and deploy. The file compression ratios were comparable but not superior to other state of the art archival tools, despite the innovative random access feature implemented in FASTAFS. FASTAFS is a user-friendly and easy to deploy backwards compatible generic purpose solution to store and access compressed FASTA files, since it offers file system access to FASTA files as well as in-sync metadata files through file virtualisation. Using virtual filesystems as in-between layer offers format conversion without the need to rewrite code into different programming languages while preserving compatibility.
Sections du résumé
BACKGROUND
BACKGROUND
The FASTA file format, used to store polymeric sequence data, has become a bioinformatics file standard used for decades. The relatively large files require additional files, beyond the scope of the original format, to identify sequences and to provide random access. Multiple compressors have been developed to archive FASTA files back and forth, but these lack direct access to targeted content or metadata of the archive. Moreover, these solutions are not directly backwards compatible to FASTA files, resulting in limited software integration.
RESULTS
RESULTS
We designed a linux based toolkit that virtualises the content of DNA, RNA and protein FASTA archives into the filesystem by using filesystem in userspace. This guarantees in-sync virtualised metadata files and offers fast random-access decompression using bit encodings plus Zstandard (zstd). The toolkit, FASTAFS, can track all its system-wide running instances, allows file integrity verification and can provide, instantly, scriptable access to sequence files and is easy to use and deploy. The file compression ratios were comparable but not superior to other state of the art archival tools, despite the innovative random access feature implemented in FASTAFS.
CONCLUSIONS
CONCLUSIONS
FASTAFS is a user-friendly and easy to deploy backwards compatible generic purpose solution to store and access compressed FASTA files, since it offers file system access to FASTA files as well as in-sync metadata files through file virtualisation. Using virtual filesystems as in-between layer offers format conversion without the need to rewrite code into different programming languages while preserving compatibility.
Identifiants
pubmed: 34724897
doi: 10.1186/s12859-021-04455-3
pii: 10.1186/s12859-021-04455-3
pmc: PMC8558547
doi:
Substances chimiques
Proteins
0
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Pagination
535Informations de copyright
© 2021. The Author(s).
Références
Nucleic Acids Res. 2019 Jan 8;47(D1):D155-D162
pubmed: 30423142
Bioinformatics. 2018 Oct 15;34(20):3600
pubmed: 29788404
Bioinformatics. 2012 Dec 15;28(24):3211-7
pubmed: 23071270
Genome Res. 2010 Sep;20(9):1297-303
pubmed: 20644199
Nucleic Acids Res. 2013 May 1;41(10):e108
pubmed: 23558742
Bioinformatics. 2014 Aug 1;30(15):2213-5
pubmed: 24747219
Nucleic Acids Res. 2004 Jan 1;32(Database issue):D115-9
pubmed: 14681372
PLoS One. 2016 Oct 5;11(10):e0163962
pubmed: 27706213
Nat Biotechnol. 2017 Apr 11;35(4):316-319
pubmed: 28398311
Bioinformatics. 2014 Jan 1;30(1):117-8
pubmed: 24132931
Genome Res. 2012 Mar;22(3):568-76
pubmed: 22300766
Bioinformation. 2011 Jan 22;5(8):350-60
pubmed: 21383923
Nucleic Acids Res. 2013 Jan;41(Database issue):D590-6
pubmed: 23193283
Genome Biol. 2016 Apr 12;17:66
pubmed: 27072794
Mol Cell. 2010 May 28;38(4):576-89
pubmed: 20513432
Bioinformatics. 2019 Oct 1;35(19):3826-3828
pubmed: 30799504
Bioinformatics. 2013 Jan 1;29(1):15-21
pubmed: 23104886