An improved de novo genome assembly of the common marmoset genome yields improved contiguity and increased mapping rates of sequence data.
Callithrix jacchus
Chromosome-scale scaffolds
Common marmoset
De novo assembly
Non-human primate genomics
Journal
BMC genomics
ISSN: 1471-2164
Titre abrégé: BMC Genomics
Pays: England
ID NLM: 100965258
Informations de publication
Date de publication:
02 Apr 2020
02 Apr 2020
Historique:
received:
05
03
2020
accepted:
09
03
2020
entrez:
4
4
2020
pubmed:
4
4
2020
medline:
24
10
2020
Statut:
epublish
Résumé
The common marmoset (Callithrix jacchus) is one of the most studied primate model organisms. However, the marmoset genomes available in the public databases are highly fragmented and filled with sequence gaps, hindering research advances related to marmoset genomics and transcriptomics. Here we utilize single-molecule, long-read sequence data to improve and update the existing genome assembly and report a near-complete genome of the common marmoset. The assembly is of 2.79 Gb size, with a contig N50 length of 6.37 Mb and a chromosomal scaffold N50 length of 143.91 Mb, representing the most contiguous and high-quality marmoset genome up to date. Approximately 90% of the assembled genome was represented in contigs longer than 1 Mb, with approximately 104-fold improvement in contiguity over the previously published marmoset genome. More than 98% of the gaps from the previously published genomes were filled successfully, which improved the mapping rates of genomic and transcriptomic data on to the assembled genome. Altogether the updated, high-quality common marmoset genome assembly provide improvements at various levels over the previous versions of the marmoset genome assemblies. This will allow researchers working on primate genomics to apply the genome more efficiently for their genomic and transcriptomic sequence data.
Sections du résumé
BACKGROUND
BACKGROUND
The common marmoset (Callithrix jacchus) is one of the most studied primate model organisms. However, the marmoset genomes available in the public databases are highly fragmented and filled with sequence gaps, hindering research advances related to marmoset genomics and transcriptomics.
RESULTS
RESULTS
Here we utilize single-molecule, long-read sequence data to improve and update the existing genome assembly and report a near-complete genome of the common marmoset. The assembly is of 2.79 Gb size, with a contig N50 length of 6.37 Mb and a chromosomal scaffold N50 length of 143.91 Mb, representing the most contiguous and high-quality marmoset genome up to date. Approximately 90% of the assembled genome was represented in contigs longer than 1 Mb, with approximately 104-fold improvement in contiguity over the previously published marmoset genome. More than 98% of the gaps from the previously published genomes were filled successfully, which improved the mapping rates of genomic and transcriptomic data on to the assembled genome.
CONCLUSIONS
CONCLUSIONS
Altogether the updated, high-quality common marmoset genome assembly provide improvements at various levels over the previous versions of the marmoset genome assemblies. This will allow researchers working on primate genomics to apply the genome more efficiently for their genomic and transcriptomic sequence data.
Identifiants
pubmed: 32241258
doi: 10.1186/s12864-020-6657-2
pii: 10.1186/s12864-020-6657-2
pmc: PMC7114785
doi:
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Pagination
243Subventions
Organisme : Japan Society for the Promotion of Science
ID : Kakenhi 16H06279 and 18H04127
Organisme : Ministry of Education, Culture, Sports, Science and Technology
ID : Innovative areas 221S0002
Organisme : Japan Agency for Medical Research and Development
ID : JP19kk0305008
Références
Science. 2016 Apr 1;352(6281):aae0344
pubmed: 27034376
Gigascience. 2018 Feb 1;7(2):1-7
pubmed: 29253147
BMC Bioinformatics. 2005 Feb 15;6:31
pubmed: 15713233
Science. 2018 Jun 8;360(6393):
pubmed: 29880660
Genome Res. 2017 May;27(5):722-736
pubmed: 28298431
Nat Methods. 2017 Nov;14(11):1072-1074
pubmed: 28945707
Bioinformatics. 2005 May 1;21(9):1859-75
pubmed: 15728110
Nat Methods. 2015 Jan;12(1):59-60
pubmed: 25402007
Nat Genet. 2017 Apr;49(4):643-650
pubmed: 28263316
G3 (Bethesda). 2018 May 4;8(5):1391-1398
pubmed: 29519939
Genes Cells. 2010 Sep 1;15(9):959-69
pubmed: 20670273
Nat Commun. 2019 Jan 16;10(1):260
pubmed: 30651564
Hortic Res. 2018 Aug 15;5:50
pubmed: 30131865
Gigascience. 2018 Jun 1;7(6):
pubmed: 29893829
Cell Rep. 2018 Jun 5;23(10):3078-3090
pubmed: 29874592
Brief Bioinform. 2019 May 21;20(3):866-876
pubmed: 29112696
Nat Biotechnol. 2011 May 15;29(7):644-52
pubmed: 21572440
Nature. 2017 Jun 22;546(7659):524-527
pubmed: 28605751
Nature. 2009 May 28;459(7246):523-7
pubmed: 19478777
G3 (Bethesda). 2017 Jan 5;7(1):109-117
pubmed: 27852011
Bioinformatics. 2016 Jul 15;32(14):2103-10
pubmed: 27153593
BMC Bioinformatics. 2018 Dec 14;19(1):481
pubmed: 30547739
Bioinformatics. 2003 Oct;19 Suppl 2:ii215-25
pubmed: 14534192
Bioinformatics. 2015 Oct 1;31(19):3210-2
pubmed: 26059717
Mol Ecol Resour. 2018 Nov;18(6):1188-1195
pubmed: 30035372
Nucleic Acids Res. 2003 Oct 1;31(19):5654-66
pubmed: 14500829
Nat Biotechnol. 2019 May;37(5):540-546
pubmed: 30936562
Stem Cells. 2005 Oct;23(9):1304-13
pubmed: 16109758
Gigascience. 2017 Oct 1;6(10):1-16
pubmed: 29020750
Nature. 2018 Nov;563(7732):501-507
pubmed: 30429615
PeerJ. 2018 Jun 4;6:e4958
pubmed: 29888139
Nat Methods. 2012 Mar 04;9(4):357-9
pubmed: 22388286
PLoS Comput Biol. 2019 Aug 21;15(8):e1007273
pubmed: 31433799
Bioinformatics. 2013 Jan 1;29(1):15-21
pubmed: 23104886
Gigascience. 2018 Aug 1;7(8):
pubmed: 30107523
Nat Genet. 2014 Aug;46(8):850-7
pubmed: 25038751
Nat Methods. 2020 Feb;17(2):155-158
pubmed: 31819265
Neuron. 2016 Nov 2;92(3):582-590
pubmed: 27809998
Nucleic Acids Res. 2015 Jan;43(Database issue):D737-42
pubmed: 25392405
Semin Fetal Neonatal Med. 2012 Dec;17(6):336-40
pubmed: 22871417
Sci Rep. 2015 Nov 20;5:16894
pubmed: 26586576
PLoS One. 2014 Nov 19;9(11):e112963
pubmed: 25409509
Nat Methods. 2016 Dec;13(12):1050-1054
pubmed: 27749838
Dev Growth Differ. 2014 Jan;56(1):53-62
pubmed: 24387631
Genome Biol. 2019 Oct 28;20(1):224
pubmed: 31661016
Nat Methods. 2013 Jun;10(6):563-9
pubmed: 23644548
Bioinformatics. 2018 Sep 15;34(18):3094-3100
pubmed: 29750242
Genomics. 2018 Nov;110(6):399-403
pubmed: 29665418