False gene and chromosome losses in genome assemblies caused by GC content variation and repeats.


Journal

Genome biology
ISSN: 1474-760X
Titre abrégé: Genome Biol
Pays: England
ID NLM: 100960660

Informations de publication

Date de publication:
27 09 2022
Historique:
received: 16 06 2021
accepted: 02 09 2022
entrez: 27 9 2022
pubmed: 28 9 2022
medline: 30 9 2022
Statut: epublish

Résumé

Many short-read genome assemblies have been found to be incomplete and contain mis-assemblies. The Vertebrate Genomes Project has been producing new reference genome assemblies with an emphasis on being as complete and error-free as possible, which requires utilizing long reads, long-range scaffolding data, new assembly algorithms, and manual curation. A more thorough evaluation of the recent references relative to prior assemblies can provide a detailed overview of the types and magnitude of improvements. Here we evaluate new vertebrate genome references relative to the previous assemblies for the same species and, in two cases, the same individuals, including a mammal (platypus), two birds (zebra finch, Anna's hummingbird), and a fish (climbing perch). We find that up to 11% of genomic sequence is entirely missing in the previous assemblies. In the Vertebrate Genomes Project zebra finch assembly, we identify eight new GC- and repeat-rich micro-chromosomes with high gene density. The impact of missing sequences is biased towards GC-rich 5'-proximal promoters and 5' exon regions of protein-coding genes and long non-coding RNAs. Between 26 and 60% of genes include structural or sequence errors that could lead to misunderstanding of their function when using the previous genome assemblies. Our findings reveal novel regulatory landscapes and protein coding sequences that have been greatly underestimated in previous assemblies and are now present in the Vertebrate Genomes Project reference genomes.

Sections du résumé

BACKGROUND
Many short-read genome assemblies have been found to be incomplete and contain mis-assemblies. The Vertebrate Genomes Project has been producing new reference genome assemblies with an emphasis on being as complete and error-free as possible, which requires utilizing long reads, long-range scaffolding data, new assembly algorithms, and manual curation. A more thorough evaluation of the recent references relative to prior assemblies can provide a detailed overview of the types and magnitude of improvements.
RESULTS
Here we evaluate new vertebrate genome references relative to the previous assemblies for the same species and, in two cases, the same individuals, including a mammal (platypus), two birds (zebra finch, Anna's hummingbird), and a fish (climbing perch). We find that up to 11% of genomic sequence is entirely missing in the previous assemblies. In the Vertebrate Genomes Project zebra finch assembly, we identify eight new GC- and repeat-rich micro-chromosomes with high gene density. The impact of missing sequences is biased towards GC-rich 5'-proximal promoters and 5' exon regions of protein-coding genes and long non-coding RNAs. Between 26 and 60% of genes include structural or sequence errors that could lead to misunderstanding of their function when using the previous genome assemblies.
CONCLUSIONS
Our findings reveal novel regulatory landscapes and protein coding sequences that have been greatly underestimated in previous assemblies and are now present in the Vertebrate Genomes Project reference genomes.

Identifiants

pubmed: 36167554
doi: 10.1186/s13059-022-02765-0
pii: 10.1186/s13059-022-02765-0
pmc: PMC9516821
doi:

Types de publication

Journal Article Research Support, Non-U.S. Gov't Research Support, N.I.H., Intramural Research Support, N.I.H., Extramural

Langues

eng

Sous-ensembles de citation

IM

Pagination

204

Subventions

Organisme : Howard Hughes Medical Institute
Pays : United States
Organisme : NINDS NIH HHS
ID : R03 NS115145
Pays : United States
Organisme : Intramural NIH HHS
ID : ZIA HG200398
Pays : United States
Organisme : NINDS NIH HHS
ID : R03 NS059755
Pays : United States
Organisme : Wellcome Trust
ID : WT206194
Pays : United Kingdom

Informations de copyright

© 2022. The Author(s).

Références

Genome Biol Evol. 2015 Oct 07;7(10):2913-28
pubmed: 26450849
Science. 2009 Oct 9;326(5950):289-93
pubmed: 19815776
Bioinformatics. 2006 Jan 15;22(2):134-41
pubmed: 16287941
Nature. 2020 Sep;585(7823):79-84
pubmed: 32663838
Bioinformatics. 2015 Oct 1;31(19):3210-2
pubmed: 26059717
Genome Biol. 2019 Nov 29;20(1):259
pubmed: 31783898
Genome Biol. 2011 Nov 08;12(11):R112
pubmed: 22067484
Genome Res. 2018 Jul;28(7):1029-1038
pubmed: 29884752
Bioinformatics. 2010 Mar 15;26(6):841-2
pubmed: 20110278
Nat Rev Genet. 2020 Oct;21(10):597-614
pubmed: 32504078
Curr Protoc Bioinformatics. 2004 May;Chapter 4:Unit 4.10
pubmed: 18428725
Proc Biol Sci. 2014 Jan 29;281(1779):20132780
pubmed: 24478299
Nat Biotechnol. 2012 Aug;30(8):771-6
pubmed: 22797562
Genomics. 2007 Sep;90(3):364-71
pubmed: 17590311
Genome Biol Evol. 2018 Feb 1;10(2):616-622
pubmed: 29385572
Genome Res. 2017 May;27(5):757-767
pubmed: 28381613
Bioinformatics. 2013 May 15;29(10):1341-2
pubmed: 23505295
Genome Biol. 2022 Sep 27;23(1):205
pubmed: 36167596
Nucleic Acids Res. 2007 Jan;35(Database issue):D61-5
pubmed: 17130148
Nature. 2010 Apr 1;464(7289):757-62
pubmed: 20360741
Science. 2014 Dec 12;346(6215):1256846
pubmed: 25504733
Nucleic Acids Res. 2012 May;40(10):e72
pubmed: 22323520
Gigascience. 2018 Mar 1;7(3):1-6
pubmed: 29618046
Hum Mol Genet. 2010 Oct 15;19(R2):R131-6
pubmed: 20858594
Trends Genet. 2000 Jun;16(6):276-7
pubmed: 10827456
Neuron. 2005 Apr 7;46(1):75-88
pubmed: 15820695
Genome Biol. 2015 Aug 18;16:164
pubmed: 26283656
Nucleic Acids Res. 2021 Feb 22;49(3):1497-1516
pubmed: 33450015
Genome Biol. 2017 Jun 14;18(1):113
pubmed: 28615074
G3 (Bethesda). 2017 Jan 5;7(1):109-117
pubmed: 27852011
Nature. 2004 Dec 9;432(7018):695-716
pubmed: 15592404
PLoS One. 2015 Aug 26;10(8):e0136281
pubmed: 26308360
Science. 2014 Dec 12;346(6215):1311-20
pubmed: 25504712
Mol Ecol Resour. 2018 Nov;18(6):1188-1195
pubmed: 30035372
Nature. 2021 Apr;592(7856):737-746
pubmed: 33911273
Cytogenet Genome Res. 2020;160(2):85-93
pubmed: 32235117
Nucleic Acids Res. 2021 Jan 8;49(D1):D884-D891
pubmed: 33137190
Nature. 2011 Oct 12;478(7370):476-82
pubmed: 21993624
Anim Genet. 2000 Apr;31(2):96-103
pubmed: 10782207
Nat Biotechnol. 2018 Oct 22;:
pubmed: 30346939
Mol Ecol Resour. 2021 Jan;21(1):263-286
pubmed: 32937018
Nat Genet. 2016 Oct;48(10):1204-10
pubmed: 27548311
Nature. 2021 Apr;592(7856):756-762
pubmed: 33408411
Genome Biol. 2020 Feb 12;21(1):35
pubmed: 32051000
Gigascience. 2017 Oct 1;6(10):1-16
pubmed: 29020750
Nucleic Acids Res. 2013 Jan;41(Database issue):D36-42
pubmed: 23193287
Proc Natl Acad Sci U S A. 1990 Jul;87(14):5578-82
pubmed: 2164689
Mol Biotechnol. 2013 Jul;54(3):1048-54
pubmed: 23568183
Genome Res. 2002 Jun;12(6):996-1006
pubmed: 12045153
Cytometry A. 2003 Feb;51(2):127-8; author reply 129
pubmed: 12541287
Mol Biol Evol. 2017 Dec 1;34(12):3123-3131
pubmed: 28962031
Genome Res. 2002 Apr;12(4):656-64
pubmed: 11932250
Nat Biotechnol. 2022 Sep;40(9):1332-1335
pubmed: 35332338
Neuron. 2013 Jun 5;78(5):839-54
pubmed: 23684785
Trends Genet. 2015 Dec;31(12):696-708
pubmed: 26599498
Nat Biotechnol. 2011 Jan;29(1):24-6
pubmed: 21221095
Nature. 2008 May 8;453(7192):175-83
pubmed: 18464734
J Biomol Tech. 2006 Jul;17(3):207-17
pubmed: 16870712
Nat Rev Genet. 2018 Jun;19(6):329-346
pubmed: 29599501
Front Neurosci. 2015 Oct 07;9:361
pubmed: 26500483
Chromosoma. 2016 Sep;125(4):757-68
pubmed: 26667931
Annu Rev Anim Biosci. 2016;4:45-59
pubmed: 26884102
Bioinformatics. 2009 Aug 15;25(16):2078-9
pubmed: 19505943
Development. 2010 Sep;137(18):3013-8
pubmed: 20685732
Genome Res. 2011 Sep;21(9):1512-28
pubmed: 21665927
Genomics Proteomics Bioinformatics. 2015 Oct;13(5):278-89
pubmed: 26542840
Genome Biol. 2018 Aug 24;19(1):125
pubmed: 30143029
Nat Rev Genet. 2015 Nov;16(11):627-40
pubmed: 26442640
Bioinformatics. 2018 Sep 15;34(18):3094-3100
pubmed: 29750242
Gigascience. 2020 Apr 1;9(4):
pubmed: 32242610
Cancer Inform. 2014 Jan 16;13:13-20
pubmed: 24526832
BMC Bioinformatics. 2009 Dec 15;10:421
pubmed: 20003500
PLoS One. 2008;3(10):e3440
pubmed: 18941504
Nat Biotechnol. 2000 May;18(5):505-8
pubmed: 10802616
Genome Biol. 2014;15(12):565
pubmed: 25518852
Cell Syst. 2018 Aug 22;7(2):219-226.e5
pubmed: 30138581
Proc Natl Acad Sci U S A. 2004 Nov 30;101(48):16855-60
pubmed: 15548610
RNA. 2015 Mar;21(3):333-46
pubmed: 25589248

Auteurs

Juwan Kim (J)

Interdisciplinary Program in Bioinformatics, Seoul National University, Seoul, Republic of Korea.

Chul Lee (C)

Interdisciplinary Program in Bioinformatics, Seoul National University, Seoul, Republic of Korea.

Byung June Ko (BJ)

Department of Agricultural Biotechnology and Research Institute of Agriculture and Life Sciences, Seoul National University, Seoul, Republic of Korea.

Dong Ahn Yoo (DA)

Interdisciplinary Program in Bioinformatics, Seoul National University, Seoul, Republic of Korea.

Sohyoung Won (S)

Interdisciplinary Program in Bioinformatics, Seoul National University, Seoul, Republic of Korea.

Adam M Phillippy (AM)

Genome Informatics Section, Computational and Statistical Genomics Branch, National Human Genome Research Institute, Bethesda, MD, USA.

Olivier Fedrigo (O)

Vertebrate Genome Lab, The Rockefeller University, New York City, USA.

Guojie Zhang (G)

BGI-Shenzhen, Shenzhen, 518083, China.
Villum Centre for Biodiversity Genomics, Section for Ecology and Evolution, Department of Biology, University of Copenhagen, Universitetsparken 15, 2100, Copenhagen, Denmark.
State Key Laboratory of Genetic Resources and Evolution, Kunming Institute of Zoology, Chinese Academy of Sciences, Kunming, 650223, China.
Center for Excellence in Animal Evolution and Genetics, Chinese Academy of Sciences, Kunming, 650223, China.

Kerstin Howe (K)

Wellcome Sanger Institute, Cambridge, UK.

Jonathan Wood (J)

Wellcome Sanger Institute, Cambridge, UK.

Richard Durbin (R)

Wellcome Sanger Institute, Cambridge, UK.
Department of Genetics, University of Cambridge, Cambridge, UK.

Giulio Formenti (G)

Vertebrate Genome Lab, The Rockefeller University, New York City, USA.
Laboratory of Neurogenetics of Language, The Rockefeller University, New York City, USA.

Samara Brown (S)

Laboratory of Neurogenetics of Language, The Rockefeller University, New York City, USA.

Lindsey Cantin (L)

Laboratory of Neurogenetics of Language, The Rockefeller University, New York City, USA.

Claudio V Mello (CV)

Department of Behavioral Neuroscience, Oregon Health and Science University, Portland, OR, 97239, USA.

Seoae Cho (S)

eGnome, Inc, Seoul, Republic of Korea.

Arang Rhie (A)

Genome Informatics Section, Computational and Statistical Genomics Branch, National Human Genome Research Institute, Bethesda, MD, USA.

Heebal Kim (H)

Interdisciplinary Program in Bioinformatics, Seoul National University, Seoul, Republic of Korea. heebal@snu.ac.kr.
Department of Agricultural Biotechnology and Research Institute of Agriculture and Life Sciences, Seoul National University, Seoul, Republic of Korea. heebal@snu.ac.kr.
eGnome, Inc, Seoul, Republic of Korea. heebal@snu.ac.kr.

Erich D Jarvis (ED)

Vertebrate Genome Lab, The Rockefeller University, New York City, USA. ejarvis@rockefeller.edu.
Laboratory of Neurogenetics of Language, The Rockefeller University, New York City, USA. ejarvis@rockefeller.edu.
Howard Hughes Medical Institute, Chevy Chase, MD, USA. ejarvis@rockefeller.edu.

Articles similaires

Genome, Chloroplast Phylogeny Genetic Markers Base Composition High-Throughput Nucleotide Sequencing
Robotic Surgical Procedures Animals Humans Telemedicine Models, Animal

Odour generalisation and detection dog training.

Lyn Caldicott, Thomas W Pike, Helen E Zulch et al.
1.00
Animals Odorants Dogs Generalization, Psychological Smell
Animals TOR Serine-Threonine Kinases Colorectal Neoplasms Colitis Mice

Classifications MeSH