The complete sequence of a human Y chromosome.
Humans
Base Sequence
Chromosomes, Human, Y
/ genetics
DNA, Satellite
/ genetics
Genetic Variation
/ genetics
Genetics, Population
Genomics
/ methods
Heterochromatin
/ genetics
Multigene Family
/ genetics
Reference Standards
Segmental Duplications, Genomic
/ genetics
Sequence Analysis, DNA
/ standards
Tandem Repeat Sequences
/ genetics
Telomere
/ genetics
Journal
Nature
ISSN: 1476-4687
Titre abrégé: Nature
Pays: England
ID NLM: 0410462
Informations de publication
Date de publication:
Sep 2023
Sep 2023
Historique:
received:
02
12
2022
accepted:
19
07
2023
medline:
15
9
2023
pubmed:
24
8
2023
entrez:
23
8
2023
Statut:
ppublish
Résumé
The human Y chromosome has been notoriously difficult to sequence and assemble because of its complex repeat structure that includes long palindromes, tandem repeats and segmental duplications
Identifiants
pubmed: 37612512
doi: 10.1038/s41586-023-06457-y
pii: 10.1038/s41586-023-06457-y
doi:
Substances chimiques
DNA, Satellite
0
Heterochromatin
0
TSPY1 protein, human
0
RBMY1A1 protein, human
0
DAZ1 protein, human
0
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Pagination
344-354Subventions
Organisme : NHGRI NIH HHS
ID : U41 HG010972
Pays : United States
Organisme : NHGRI NIH HHS
ID : R01 HG011274
Pays : United States
Organisme : NHGRI NIH HHS
ID : R21 HG010548
Pays : United States
Organisme : NHGRI NIH HHS
ID : U01 HG010971
Pays : United States
Organisme : NHGRI NIH HHS
ID : U41 HG006620
Pays : United States
Organisme : NHGRI NIH HHS
ID : R01 HG010169
Pays : United States
Organisme : NIGMS NIH HHS
ID : R01 GM072264
Pays : United States
Organisme : NHGRI NIH HHS
ID : R01 HG002385
Pays : United States
Informations de copyright
© 2023. This is a U.S. Government work and not under copyright protection in the US; foreign copyright protection may apply.
Références
Skaletsky, H. et al. The male-specific region of the human Y chromosome is a mosaic of discrete sequence classes. Nature 423, 825–837 (2003).
pubmed: 12815422
doi: 10.1038/nature01722
Miga, K. H. et al. Centromere reference models for human chromosomes X and Y satellite arrays. Genome Res. 24, 697–707 (2014).
pubmed: 24501022
pmcid: 3975068
doi: 10.1101/gr.159624.113
Vollger, M. R. et al. Segmental duplications and their variation in a complete human genome. Science 376, eabj6965 (2022).
pubmed: 35357917
pmcid: 8979283
doi: 10.1126/science.abj6965
Nurk, S. et al. The complete sequence of a human genome. Science 376, 44–53 (2022).
pubmed: 35357919
pmcid: 9186530
doi: 10.1126/science.abj6987
Schneider, V. A. et al. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. Genome Res. 27, 849–864 (2017).
pubmed: 28396521
pmcid: 5411779
doi: 10.1101/gr.213611.116
Gustafson, M. L. & Donahoe, P. K. Male sex determination: current concepts of male sexual differentiation. Annu. Rev. Med. 45, 505–524 (1994).
pubmed: 8198399
doi: 10.1146/annurev.med.45.1.505
Vog, P. H. et al. Human Y chromosome azoospermia factors (AZF) mapped to different subregions in Yq11. Hum. Mol. Genet. 5, 933–943 (1996).
doi: 10.1093/hmg/5.7.933
Miga, K. H. et al. Telomere-to-telomere assembly of a complete human X chromosome. Nature 585, 79–84 (2020).
pubmed: 32663838
pmcid: 7484160
doi: 10.1038/s41586-020-2547-7
Logsdon, G. A. et al. The structure, function and evolution of a complete human chromosome 8. Nature 593, 101–107 (2021).
pubmed: 33828295
pmcid: 8099727
doi: 10.1038/s41586-021-03420-7
Wenger, A. M. et al. Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome. Nat. Biotechnol. 37, 1155–1162 (2019).
pubmed: 31406327
pmcid: 6776680
doi: 10.1038/s41587-019-0217-9
Jain, M. et al. Nanopore sequencing and assembly of a human genome with ultra-long reads. Nat. Biotechnol. 36, 338–345 (2018).
pubmed: 29431738
pmcid: 5889714
doi: 10.1038/nbt.4060
Nurk, S. et al. HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads. Genome Res. 30, 1291–1305 (2020).
pubmed: 32801147
pmcid: 7545148
doi: 10.1101/gr.263566.120
Rautiainen, M. & Marschall, T. GraphAligner: rapid and versatile sequence-to-graph alignment. Genome Biol. 21, 253 (2020).
pubmed: 32972461
pmcid: 7513500
doi: 10.1186/s13059-020-02157-2
Formenti, G. et al. Merfin: improved variant filtering, assembly evaluation and polishing via k-mer validation. Nat. Methods 19, 696–704 (2022).
pubmed: 35361932
pmcid: 9745813
doi: 10.1038/s41592-022-01445-y
Kirsche, M. et al. Jasmine and Iris: population-scale structural variant comparison and analysis. Nat. Methods 20, 408–417 (2023).
pubmed: 36658279
pmcid: 10006329
doi: 10.1038/s41592-022-01753-3
Jain, C., Rhie, A., Hansen, N. F., Koren, S. & Phillippy, A. M. Long-read mapping to repetitive reference sequences using Winnowmap2. Nat. Methods 19, 705–710 (2022).
pubmed: 35365778
pmcid: 10510034
doi: 10.1038/s41592-022-01457-8
Mc Cartney, A. M. et al. Chasing perfection: validation and polishing strategies for telomere-to-telomere genome assemblies. Nat. Methods 19, 687–695 (2022).
doi: 10.1038/s41592-022-01440-3
Rhie, A., Walenz, B. P., Koren, S. & Phillippy, A. M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol. 21, 245 (2020).
pubmed: 32928274
pmcid: 7488777
doi: 10.1186/s13059-020-02134-9
Wang, T. et al. The Human Pangenome Project: a global resource to map genomic diversity. Nature 604, 437–446 (2022).
pubmed: 35444317
pmcid: 9402379
doi: 10.1038/s41586-022-04601-8
Jarvis, E. D. et al. Semi-automated assembly of high-quality diploid human reference genomes. Nature 611, 519–531 (2022).
pubmed: 36261518
pmcid: 9668749
doi: 10.1038/s41586-022-05325-5
Shumate, A. et al. Assembly and annotation of an Ashkenazi human reference genome. Genome Biol. 21, 129 (2020).
pubmed: 32487205
pmcid: 7265644
doi: 10.1186/s13059-020-02047-7
Zook, J. M. et al. Extensive sequencing of seven human genomes to characterize benchmark reference materials. Sci. Data 3, 160025 (2016).
pubmed: 27271295
pmcid: 4896128
doi: 10.1038/sdata.2016.25
Landrum, M. J. et al. ClinVar: improvements to accessing data. Nucleic Acids Res. 48, D835–D844 (2020).
pubmed: 31777943
doi: 10.1093/nar/gkz972
Welter, D. et al. The NHGRI GWAS Catalog, a curated resource of SNP-trait associations. Nucleic Acids Res. 42, D1001–D1006 (2014).
pubmed: 24316577
doi: 10.1093/nar/gkt1229
Smigielski, E. M., Sirotkin, K., Ward, M. & Sherry, S. T. dbSNP: a database of single nucleotide polymorphisms. Nucleic Acids Res. 28, 352–355 (2000).
pubmed: 10592272
pmcid: 102496
doi: 10.1093/nar/28.1.352
Karczewski, K. J. et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 581, 434–443 (2020).
pubmed: 32461654
pmcid: 7334197
doi: 10.1038/s41586-020-2308-7
Byrska-Bishop, M. et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell 185, 3426–3440 (2022).
pubmed: 36055201
pmcid: 9439720
doi: 10.1016/j.cell.2022.08.004
Mallick, S. et al. The Simons Genome Diversity Project: 300 genomes from 142 diverse populations. Nature 538, 201–206 (2016).
pubmed: 27654912
pmcid: 5161557
doi: 10.1038/nature18964
Dunham, I. et al. An integrated encyclopedia of DNA elements in the human genome. Nature 489, 57–74 (2012).
doi: 10.1038/nature11247
Ebert, P. et al. Haplotype-resolved diverse human genomes and integrated analysis of structural variation. Science 372, eabf7117 (2021).
pubmed: 33632895
pmcid: 8026704
doi: 10.1126/science.abf7117
Sanders, A. D. et al. Single-cell analysis of structural variations and complex rearrangements with tri-channel processing. Nat. Biotechnol. 38, 343–354 (2020).
pubmed: 31873213
doi: 10.1038/s41587-019-0366-x
Hallast, P. et al. Assembly of 43 human Y chromosomes reveals extensive complexity and variation. Nature https://doi.org/10.1038/s41586-023-06425-6 (2023).
Hammer, M. F. et al. Extended Y chromosome haplotypes resolve multiple and unique lineages of the Jewish priesthood. Hum. Genet. 126, 707 (2009).
pubmed: 19669163
pmcid: 2771134
doi: 10.1007/s00439-009-0727-5
Poznik, G. D. et al. Punctuated bursts in human male demography inferred from 1,244 worldwide Y-chromosome sequences. Nat. Genet. 48, 593–599 (2016).
pubmed: 27111036
pmcid: 4884158
doi: 10.1038/ng.3559
Vegesna, R., Tomaszkiewicz, M., Medvedev, P. & Makova, K. D. Dosage regulation, and variation in gene expression and copy number of human Y chromosome ampliconic genes. PLoS Genet. 15, e1008369 (2019).
pubmed: 31525193
pmcid: 6772104
doi: 10.1371/journal.pgen.1008369
NCBI RefSeq v110 Browser. Homo sapiens isolate NA24385 chromosome Y, alternate assembly T2T-CHM13v2.0. https://tinyurl.com/bdfudexn (2022).
Hoyt, S. J. et al. From telomere to telomere: the transcriptional and epigenetic state of human repeat elements. Science 376, eabk3112 (2022).
pubmed: 35357925
pmcid: 9301658
doi: 10.1126/science.abk3112
Warburton, P. E. et al. Analysis of the largest tandemly repeated DNA families in the human genome. BMC Genomics 9, 533 (2008).
pubmed: 18992157
pmcid: 2588610
doi: 10.1186/1471-2164-9-533
Halabian, R. & Makałowski, W. A map of 3′ DNA transduction variants mediated by non-LTR retroelements on 3202 human genomes. Biology 11, 1032 (2022).
pubmed: 36101413
pmcid: 9311842
doi: 10.3390/biology11071032
Weissensteiner, M. H. et al. Accurate sequencing of DNA motifs able to form alternative (non-B) structures. Genome Res. 33, 907-922 (2023).
Tyler-Smith, C., Taylor, L. & Müller, U. Structure of a hypervariable tandemly repeated DNA sequence on the short arm of the human Y chromosome. J. Mol. Biol. 203, 837–848 (1988).
pubmed: 3210241
doi: 10.1016/0022-2836(88)90110-6
Xue, Y. & Tyler-Smith, C. An exceptional gene: evolution of the TSPY gene family in humans and other great apes. Genes 2, 36–47 (2011).
pubmed: 24710137
pmcid: 3924835
doi: 10.3390/genes2010036
Saxena, R. et al. Four DAZ genes in two clusters found in the AZFc region of the human Y chromosome. Genomics 67, 256–267 (2000).
pubmed: 10936047
doi: 10.1006/geno.2000.6260
Altemose, N. et al. Complete genomic and epigenetic maps of human centromeres. Science 376, eabl4178 (2022).
pubmed: 35357911
pmcid: 9233505
doi: 10.1126/science.abl4178
Jain, M. et al. Linear assembly of a human centromere on the Y chromosome. Nat. Biotechnol. 36, 321–323 (2018).
pubmed: 29553574
pmcid: 5886786
doi: 10.1038/nbt.4109
Gershman, A. et al. Epigenetic patterns in a complete human genome. Science 376, eabj5089 (2022).
pubmed: 35357915
pmcid: 9170183
doi: 10.1126/science.abj5089
Kasinathan, S. & Henikoff, S. Non-B-form DNA is enriched at centromeres. Mol. Biol. Evol. 35, 949–962 (2018).
pubmed: 29365169
pmcid: 5889037
doi: 10.1093/molbev/msy010
Nailwal, M. & Chauhan, J. B. Azoospermia factor C subregion of the Y chromosome. J. Hum. Reprod. Sci. 10, 256 (2017).
pubmed: 29430151
pmcid: 5799928
doi: 10.4103/jhrs.JHRS_16_17
Kuroda-Kawaguchi, T. et al. The AZFc region of the Y chromosome features massive palindromes and uniform recurrent deletions in infertile men. Nat. Genet. 29, 279–286 (2001).
pubmed: 11687796
doi: 10.1038/ng757
Repping, S. et al. A family of human Y chromosomes has dispersed throughout northern Eurasia despite a 1.8-Mb deletion in the azoospermia factor c region. Genomics 83, 1046–1052 (2004).
pubmed: 15177557
doi: 10.1016/j.ygeno.2003.12.018
Porubsky, D. et al. Recurrent inversion polymorphisms in humans associate with genetic instability and genomic disorders. Cell 185, 1986–2005 (2022).
pubmed: 35525246
pmcid: 9563103
doi: 10.1016/j.cell.2022.04.017
Teitz, L. S., Pyntikova, T., Skaletsky, H. & Page, D. C. Selection has countered high mutability to preserve the ancestral copy number of Y chromosome amplicons in diverse human lineages. Am. J. Hum. Genet. 103, 261–275 (2018).
pubmed: 30075113
pmcid: 6080837
doi: 10.1016/j.ajhg.2018.07.007
Jobling, M. A. Copy number variation on the human Y chromosome. Cytogenet. Genome Res. 123, 253–262 (2008).
pubmed: 19287162
doi: 10.1159/000184715
Navarro-Costa, P., Plancha, C. E. & Gonçalves, J. Genetic dissection of the AZF regions of the human Y chromosome: thriller or filler for male (in)fertility? Biomed Res. Int. 2010, e936569 (2010).
Evans, H. J., Gosden, J. R., Mitchell, A. R. & Buckland, R. A. Location of human satellite DNAs on the Y chromosome. Nature 251, 346–347 (1974).
doi: 10.1038/251346a0
Schmid, M., Guttenbach, M., Nanda, I., Studer, R. & Epplen, J. T. Organization of DYZ2 repetitive DNA on the human Y chromosome. Genomics 6, 212–218 (1990).
pubmed: 2307465
doi: 10.1016/0888-7543(90)90559-D
Manz, E., Alkan, M., Bühler, E. & Schmidtke, J. Arrangement of DYZ1 and DYZ2 repeats on the human Y-chromosome: a case with presence of DYZ1 and absence of DYZ2. Mol. Cell. Probes 6, 257–259 (1992).
pubmed: 1406735
doi: 10.1016/0890-8508(92)90025-S
Altemose, N. A classical revival: human satellite DNAs enter the genomics era. Semin. Cell Dev. Biol. 128, 2–14 (2022).
pubmed: 35487859
doi: 10.1016/j.semcdb.2022.04.012
Gripenberg, U. Size variation and orientation of the human Y chromosome. Chromosoma 15, 618–629 (1964).
pubmed: 14333154
doi: 10.1007/BF00319995
Mathias, N., Bayés, M. & Tyler-Smith, C. Highly informative compound haplotypes for the human Y chromosome. Hum. Mol. Genet. 3, 115–123 (1994).
pubmed: 7909247
doi: 10.1093/hmg/3.1.115
Altemose, N., Miga, K. H., Maggioni, M. & Willard, H. F. Genomic characterization of large heterochromatic gaps in the human genome assembly. PLoS Comput. Biol. 10, e1003628 (2014).
pubmed: 24831296
pmcid: 4022460
doi: 10.1371/journal.pcbi.1003628
Cooke, H. Repeated sequence specific to human males. Nature 262, 182–186 (1976).
pubmed: 819844
doi: 10.1038/262182a0
Frommer, M., Prosser, J. & Vincent, P. C. Human satellite I sequences include a male specific 2.47 kb tandemly repeated unit containing one Alu family member per repeat. Nucleic Acids Res. 12, 2887–2900 (1984).
pubmed: 6324132
pmcid: 318713
doi: 10.1093/nar/12.6.2887
Babcock, M., Yatsenko, S., Stankiewicz, P., Lupski, J. R. & Morrow, B. E. AT-rich repeats associated with chromosome 22q11.2 rearrangement disorders shape human genome architecture on Yq12. Genome Res. 17, 451–460 (2007).
pubmed: 17284672
pmcid: 1832092
doi: 10.1101/gr.5651507
Webster, T. H. et al. Identifying, understanding, and correcting technical artifacts on the sex chromosomes in next-generation sequencing data. GigaScience 8, giz074 (2019).
pubmed: 31289836
pmcid: 6615978
doi: 10.1093/gigascience/giz074
Aganezov, S. et al. A complete reference genome improves analysis of human genetic variation. Science 376, eabl3533 (2022).
pubmed: 35357935
pmcid: 9336181
doi: 10.1126/science.abl3533
Bekritsky, M. A., Colombo, C. & Eberle, M. A. Identifying genomic regions with high quality single nucleotide variant calling. Illumina https://www.illumina.com/content/illumina-marketing/amr/en_US/science/genomics-research/articles/identifying-genomic-regions-with-high-quality-single-nucleotide-.html (2023).
Breitwieser, F. P., Pertea, M., Zimin, A. V. & Salzberg, S. L. Human contamination in bacterial genomes has created thousands of spurious proteins. Genome Res. 29, 954–960 (2019).
pubmed: 31064768
pmcid: 6581058
doi: 10.1101/gr.245373.118
Steinegger, M. & Salzberg, S. L. Terminating contamination: large-scale search identifies more than 2,000,000 contaminated entries in GenBank. Genome Biol. 21, 115 (2020).
pubmed: 32398145
pmcid: 7218494
doi: 10.1186/s13059-020-02023-1
Chrisman, B. et al. The human “contaminome”: bacterial, viral, and computational contamination in whole genome sequences from 1000 families. Sci. Rep. 12, 9863 (2022).
pubmed: 35701436
pmcid: 9198055
doi: 10.1038/s41598-022-13269-z
Kent, W. J. et al. The Human Genome Browser at UCSC. Genome Res. 12, 996–1006 (2002).
pubmed: 12045153
pmcid: 186604
doi: 10.1101/gr.229102
Rautiainen, M. et al. Telomere-to-telomere assembly of diploid chromosomes with Verkko. Nat. Biotechnol. https://doi.org/10.1038/s41587-023-01662-6 (2023).
Liao, W.-W. et al. A draft human pangenome reference. Nature 617, 312–324 (2023).
pubmed: 37165242
pmcid: 10172123
doi: 10.1038/s41586-023-05896-x
Jiang, Z., Hubley, R., Smit, A. & Eichler, E. E. DupMasker: a tool for annotating primate segmental duplications. Genome Res. 18, 1362–1368 (2008).
pubmed: 18502942
pmcid: 2493431
doi: 10.1101/gr.078477.108
Vollger, M. R., Kerpedjiev, P., Phillippy, A. M. & Eichler, E. E. StainedGlass: interactive visualization of massive tandem repeat structures with identity heatmaps. Bioinformatics 38, 2049–2051 (2022).
pubmed: 35020798
pmcid: 8963321
doi: 10.1093/bioinformatics/btac018
Skene, P. J. & Henikoff, S. An efficient targeted nuclease strategy for high-resolution mapping of DNA binding sites. eLife 6, e21856 (2017).
pubmed: 28079019
pmcid: 5310842
doi: 10.7554/eLife.21856
Robinson, J. T. et al. Integrative genomics viewer. Nat. Biotechnol. 29, 24–26 (2011).
pubmed: 21221095
pmcid: 3346182
doi: 10.1038/nbt.1754
Shafin, K. et al. Nanopore sequencing and the Shasta toolkit enable efficient de novo assembly of eleven human genomes. Nat. Biotechnol. 38, 1044–1053 (2020).
pubmed: 32686750
pmcid: 7483855
doi: 10.1038/s41587-020-0503-6
Koren, S. et al. De novo assembly of haplotype-resolved genomes with trio binning. Nat. Biotechnol. 36, 1174–1182 (2018).
doi: 10.1038/nbt.4277
Kolmogorov, M., Yuan, J., Lin, Y. & Pevzner, P. A. Assembly of long, error-prone reads using repeat graphs. Nat. Biotechnol. 37, 540–546 (2019).
pubmed: 30936562
doi: 10.1038/s41587-019-0072-8
Poplin, R. et al. A universal SNP and small-indel variant caller using deep neural networks. Nat. Biotechnol. 36, 983–987 (2018).
pubmed: 30247488
doi: 10.1038/nbt.4235
Shafin, K. et al. Haplotype-aware variant calling with PEPPER-Margin-DeepVariant enables high accuracy in nanopore long-reads. Nat. Methods 18, 1322–1332 (2021).
pubmed: 34725481
pmcid: 8571015
doi: 10.1038/s41592-021-01299-w
Sedlazeck, F. J. et al. Accurate detection of complex structural variations using single molecule sequencing. Nat. Methods 15, 461–468 (2018).
pubmed: 29713083
pmcid: 5990442
doi: 10.1038/s41592-018-0001-7
Jiang, T. et al. Long-read-based human genomic structural variation detection with cuteSV. Genome Biol. 21, 189 (2020).
pubmed: 32746918
pmcid: 7477834
doi: 10.1186/s13059-020-02107-y
Bzikadze, A. V., Mikheenko, A. & Pevzner, P. A. Fast and accurate mapping of long reads to complete genome assemblies with VerityMap. Genome Res. 32, 2107–2118 (2022).
pubmed: 36379716
pmcid: 9808623
doi: 10.1101/gr.276871.122
Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows–Wheeler transform. Bioinformatics 25, 1754–1760 (2009).
pubmed: 19451168
pmcid: 2705234
doi: 10.1093/bioinformatics/btp324
Porubsky, D. et al. breakpointR: an R/Bioconductor package to localize strand state changes in Strand-seq data. Bioinformatics 36, 1260–1261 (2020).
pubmed: 31504176
doi: 10.1093/bioinformatics/btz681
PacBio Revio WGS Dataset. Homo sapiens – GIAB trio HG002-4. https://downloads.pacbcloud.com/public/revio/2022Q4/ (2022).
Poznik, D. yhaplo | Identifying Y-chromosome haplogroups. GitHub https://github.com/23andMe/yhaplo (2022).
Tseng, B. et al. Y-SNP Haplogroup Hierarchy Finder: a web tool for Y-SNP haplogroup assignment. J. Hum. Genet. 67, 487–493 (2022).
pubmed: 35347230
doi: 10.1038/s10038-022-01033-0
Li, H. Minimap2: pairwise alignment for nucleotide sequences. Bioinformatics 34, 3094–3100 (2018).
pubmed: 29750242
pmcid: 6137996
doi: 10.1093/bioinformatics/bty191
Li, H. Identifying centromeric satellites with dna-brnn. Bioinformatics 35, 4408–4410 (2019).
pubmed: 30989183
pmcid: 6821349
doi: 10.1093/bioinformatics/btz264
Harris, R. S. Improved Pairwise Alignmnet of Genomic DNA (Pennsylvania State Univ., 2007).
Morgulis, A., Gertz, E. M., Schäffer, A. A. & Agarwala, R. WindowMasker: window-based masker for sequenced genomes. Bioinformatics 22, 134–141 (2006).
pubmed: 16287941
doi: 10.1093/bioinformatics/bti774
Chin, C.-S. et al. Multiscale analysis of pangenomes enables improved representation of genomic diversity for repetitive and clinically relevant genes. Nat. Methods https://doi.org/10.1038/s41592-023-01914-y (2023).
Frankish, A. et al. GENCODE 2021. Nucleic Acids Res. 49, D916–D923 (2021).
pubmed: 33270111
doi: 10.1093/nar/gkaa1087
Armstrong, J. et al. Progressive Cactus is a multiple-genome aligner for the thousand-genome era. Nature 587, 246–251 (2020).
pubmed: 33177663
pmcid: 7673649
doi: 10.1038/s41586-020-2871-y
Kovaka, S. et al. Transcriptome assembly from long-read RNA-seq alignments with StringTie2. Genome Biol. 20, 278 (2019).
pubmed: 31842956
pmcid: 6912988
doi: 10.1186/s13059-019-1910-1
Stanke, M., Diekhans, M., Baertsch, R. & Haussler, D. Using native and syntenically mapped cDNA alignments to improve de novo gene finding. Bioinformatics 24, 637–644 (2008).
pubmed: 18218656
doi: 10.1093/bioinformatics/btn013
Fiddes, I. T. et al. Comparative Annotation Toolkit (CAT)—simultaneous clade and personal genome annotation. Genome Res. 28, 1029–1038 (2018).
pubmed: 29884752
pmcid: 6028123
doi: 10.1101/gr.233460.117
Shumate, A. & Salzberg, S. L. Liftoff: accurate mapping of gene annotations. Bioinformatics 37, 1639–1643 (2021).
pubmed: 33320174
pmcid: 8289374
doi: 10.1093/bioinformatics/btaa1016
Dale, R. K., Pedersen, B. S. & Quinlan, A. R. Pybedtools: a flexible Python library for manipulating genomic datasets and annotations. Bioinformatics 27, 3423–3424 (2011).
pubmed: 21949271
pmcid: 3232365
doi: 10.1093/bioinformatics/btr539
Rhie, A. et al. Towards complete and error-free genome assemblies of all vertebrate species. Nature 592, 737–746 (2021).
pubmed: 33911273
pmcid: 8081667
doi: 10.1038/s41586-021-03451-0
Pruitt, K. D. et al. RefSeq: an update on mammalian reference sequences. Nucleic Acids Res. 42, D756–D763 (2014).
pubmed: 24259432
doi: 10.1093/nar/gkt1114
Kapustin, Y., Souvorov, A., Tatusova, T. & Lipman, D. Splign: algorithms for computing spliced alignments with identification of paralogs. Biol. Direct 3, 20 (2008).
pubmed: 18495041
pmcid: 2440734
doi: 10.1186/1745-6150-3-20
Katoh, K. & Standley, D. M. MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Mol Biol Evol. 30, 772-80 (2013).
Slater, G. S. C. & Birney, E. Automated generation of heuristics for biological sequence comparison. BMC Bioinformatics 6, 31 (2005).
pubmed: 15713233
pmcid: 553969
doi: 10.1186/1471-2105-6-31
Zook, J. M. et al. Integrating human sequence data sets provides a resource of benchmark SNP and indel genotype calls. Nat. Biotechnol. 32, 246–251 (2014).
pubmed: 24531798
doi: 10.1038/nbt.2835
Numanagić, I. et al. Fast characterization of segmental duplications in genome assemblies. Bioinformatics 34, i706–i714 (2018).
pubmed: 30423092
pmcid: 6129265
doi: 10.1093/bioinformatics/bty586
Benson, G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res. 27, 573–580 (1999).
pubmed: 9862982
pmcid: 148217
doi: 10.1093/nar/27.2.573
Arian, F. A. S., Hubley, R. & Green, P. RepeatMasker Open-4.0 2013-2015. http://www.repeatmasker.org (2015).
Storer, J., Hubley, R., Rosen, J., Wheeler, T. J. & Smit, A. F. The Dfam community resource of transposable element families, sequence models, and genome annotations. Mob. DNA 12, 2 (2021).
pubmed: 33436076
pmcid: 7805219
doi: 10.1186/s13100-020-00230-y
Olson, D. & Wheeler, T. ULTRA: a model based tool to detect tandem repeats. ACM BCB 2018, 37–46 (2018)
Quinlan, A. R. & Hall, I. M. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 26, 841–842 (2010).
pubmed: 20110278
pmcid: 2832824
doi: 10.1093/bioinformatics/btq033
Storer, J. M., Hubley, R., Rosen, J. & Smit, A. F. A. Curation guidelines for de novo generated transposable element families. Curr. Protoc. 1, e154 (2021).
pubmed: 34138525
pmcid: 9191830
doi: 10.1002/cpz1.154
Kent, W. J. BLAT—the BLAST-like alignment tool. Genome Res. 12, 656–664 (2002).
pubmed: 11932250
pmcid: 187518
Szak, S. T. et al. Molecular archeology of L1 insertions in the human genome. Genome Biol. 3, research0052.1 (2002).
doi: 10.1186/gb-2002-3-10-research0052
Altschul, S. F., Gish, W., Miller, W., Myers, E. W. & Lipman, D. J. Basic local alignment search tool. J. Mol. Biol. 215, 403–410 (1990).
pubmed: 2231712
doi: 10.1016/S0022-2836(05)80360-2
Cer, R. Z. et al. Searching for non-B DNA-forming motifs using nBMST (non-B DNA motif search tool). Curr. Protoc. Hum. Genet. 73, 18.7.1–18.7.22 (2012).
Zou, X. et al. Short inverted repeats contribute to localized mutability in human somatic cells. Nucleic Acids Res. 45, 11213–11221 (2017).
pubmed: 28977645
pmcid: 5737083
doi: 10.1093/nar/gkx731
Svetec Miklenić, M. et al. Size-dependent antirecombinogenic effect of short spacers on palindrome recombinogenicity. DNA Repair 90, 102848 (2020).
pubmed: 32388488
doi: 10.1016/j.dnarep.2020.102848
Sahakyan, A. B. et al. Machine learning model for sequence-driven DNA G-quadruplex formation. Sci. Rep. 7, 14535 (2017).
pubmed: 29109402
pmcid: 5673958
doi: 10.1038/s41598-017-14017-4
Hao, Z. et al. RIdeogram: drawing SVG graphics to visualize and map genome-wide data on the idiograms. PeerJ Comput. Sci. 6, e251 (2020).
pubmed: 33816903
pmcid: 7924719
doi: 10.7717/peerj-cs.251
Dotmatics. GraphPad Prism v.9.1.0 for Windows. https://www.graphpad.com (16 March 2021).
Vollger, M. R. SafFire. GitHub https://github.com/mrvollger/SafFire (2022).
Pendleton, A. L. et al. Comparison of village dog and wolf genomes highlights the role of the neural crest in dog domestication. BMC Biol. 16, 64 (2018).
pubmed: 29950181
pmcid: 6022502
doi: 10.1186/s12915-018-0535-2
Hach, F. et al. mrsFAST: a cache-oblivious algorithm for short-read mapping. Nat. Methods 7, 576–577 (2010).
pubmed: 20676076
pmcid: 3115707
doi: 10.1038/nmeth0810-576
Escalona, M. et al. Whole-genome sequence and assembly of the Javan gibbon (Hylobates moloch). J. Hered. 114, 35–43 (2023).
pubmed: 36146896
doi: 10.1093/jhered/esac043
Cortez, D. et al. Origins and functional evolution of Y chromosomes across mammals. Nature 508, 488–493 (2014).
pubmed: 24759410
doi: 10.1038/nature13151
Stamatakis, A. RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies. Bioinformatics 30, 1312–1313 (2014).
pubmed: 24451623
pmcid: 3998144
doi: 10.1093/bioinformatics/btu033
Dotmatics. Geneious v2019.2.3. https://www.geneious.com/ (2019).
Rambaut et al. FigTree v1.4.4. http://tree.bio.ed.ac.uk/software/figtree/ (2018).
Tyler-Smith, C. & Brown, W. R. A. Structure of the major block of alphoid satellite DNA on the human Y chromosome. J. Mol. Biol. 195, 457–470 (1987).
pubmed: 2821279
doi: 10.1016/0022-2836(87)90175-6
Shepelev, V. A. et al. Annotation of suprachromosomal families reveals uncommon types of alpha satellite organization in pericentromeric regions of hg38 human genome assembly. Genomics Data 5, 139–146 (2015).
pubmed: 26167452
pmcid: 4496801
doi: 10.1016/j.gdata.2015.05.035
Lee, I. et al. Simultaneous profiling of chromatin accessibility and methylation on human cell lines with nanopore sequencing. Nat. Methods 17, 1191–1199 (2020).
pubmed: 33230324
pmcid: 7704922
doi: 10.1038/s41592-020-01000-7
Krumsiek, J., Arnold, R. & Rattei, T. Gepard: a rapid and sensitive tool for creating dotplots on genome scale. Bioinformatics 23, 1026–1028 (2007).
pubmed: 17309896
doi: 10.1093/bioinformatics/btm039
Rice, P., Longden, I. & Bleasby, A. EMBOSS: The European Molecular Biology Open Software Suite. Trends Genet. 16, 276–277 (2000).
pubmed: 10827456
doi: 10.1016/S0168-9525(00)02024-2
Sun, C. et al. Deletion of azoospermia factor a (AZFa) region of human Y chromosome caused by recombination between HERV15 proviruses. Hum. Mol. Genet. 9, 2291–2296 (2000).
pubmed: 11001932
doi: 10.1093/oxfordjournals.hmg.a018920
Lassmann, T. Kalign 3: multiple sequence alignment of large datasets. Bioinformatics 36, 1928–1929 (2020).
doi: 10.1093/bioinformatics/btz795
Wheeler, T. J. & Eddy, S. R. nhmmer: DNA homology search with profile HMMs. Bioinformatics 29, 2487–2489 (2013).
pubmed: 23842809
pmcid: 3777106
doi: 10.1093/bioinformatics/btt403
Stephens, Z. D. et al. Simulating next-generation sequencing datasets from empirical mutation and sequencing models. PLoS ONE 11, e0167047 (2016).
pubmed: 27893777
pmcid: 5125660
doi: 10.1371/journal.pone.0167047
Bushnell, B. BBMap: a fast, accurate, splice-aware aligner. OSTI.gov https://www.osti.gov/biblio/1241166 (2017).
Aken, B. L. et al. Ensembl 2017. Nucleic Acids Res. 45, D635–D642 (2017).
pubmed: 27899575
doi: 10.1093/nar/gkw1104
Poznik, G. D. et al. Sequencing Y chromosomes resolves discrepancy in time to common ancestor of males versus females. Science 341, 562–565 (2013).
pubmed: 23908239
pmcid: 4032117
doi: 10.1126/science.1237619
McKenna, A. et al. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res. 20, 1297–1303 (2010).
pubmed: 20644199
pmcid: 2928508
doi: 10.1101/gr.107524.110
Schatz, M. C. et al. Inverting the model of genomics data sharing with the NHGRI Genomic Data Science Analysis, Visualization, and Informatics Lab-space. Cell Genomics 2, 100085 (2022).
pubmed: 35199087
pmcid: 8863334
doi: 10.1016/j.xgen.2021.100085
Danecek, P. et al. Twelve years of SAMtools and BCFtools. GigaScience 10, giab008 (2021).
pubmed: 33590861
pmcid: 7931819
doi: 10.1093/gigascience/giab008
Talenti, A. & Prendergast, J. nf-LO: a scalable, containerized workflow for genome-to-genome lift over. Genome Biol. Evol. 13, evab183 (2021).
pubmed: 34383887
pmcid: 8412297
doi: 10.1093/gbe/evab183
Guarracino, A., Mwaniki, N., Marco-Sola, S., & Garrison, E. wfmash: whole-chromosome pairwise alignment using the hierarchical wavefront algorithm. GitHub https://github.com/ekg/wfmash (2021).
Sherry, S. T., Ward, M. & Sirotkin, K. dbSNP—database for single nucleotide polymorphisms and other classes of minor genetic variation. Genome Res. 9, 677–679 (1999).
pubmed: 10447503
doi: 10.1101/gr.9.8.677
Landrum, M. J. et al. ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Res. 46, D1062–D1067 (2018).
pubmed: 29165669
doi: 10.1093/nar/gkx1153
Buniello, A. et al. The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019. Nucleic Acids Res. 47, D1005–D1012 (2019).
pubmed: 30445434
doi: 10.1093/nar/gky1120
Van der Auwera G. A. & O’Connor B. D. Genomics in the Cloud: Using Docker, GATK, and WDL in Terra (O’Reilly Media, 2020).
Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat. Methods 9, 357–359 (2012).
pubmed: 22388286
pmcid: 3322381
doi: 10.1038/nmeth.1923
Ramírez, F. et al. deepTools2: a next generation web server for deep-sequencing data analysis. Nucleic Acids Res. 44, W160–W165 (2016).
pubmed: 27079975
pmcid: 4987876
doi: 10.1093/nar/gkw257
Zhao, H. et al. CrossMap: a versatile tool for coordinate conversion between genome assemblies. Bioinformatics 30, 1006–1007 (2014).
pubmed: 24351709
doi: 10.1093/bioinformatics/btt730
Marçais, G. et al. MUMmer4: a fast and versatile genome alignment system. PLoS Comput. Biol. 14, e1005944 (2018).
pubmed: 29373581
pmcid: 5802927
doi: 10.1371/journal.pcbi.1005944
Ondov, B. D., Bergman, N. H. & Phillippy, A. M. Interactive metagenomic visualization in a Web browser. BMC Bioinformatics 12, 385 (2011).
pubmed: 21961884
pmcid: 3190407
doi: 10.1186/1471-2105-12-385
Rhie, A. Repositories for the analysis of T2T-Y and T2T-CHM13v2.0. Zenodo https://doi.org/10.5281/zenodo.8136598 (2023).
Falconer, E. et al. DNA template strand sequencing of single-cells maps genomic rearrangements at high resolution. Nat. Methods 9, 1107–1112 (2012).
pubmed: 23042453
pmcid: 3580294
doi: 10.1038/nmeth.2206