Chromosome-level genome assembly of the diploid oat species Avena longiglumis.
Journal
Scientific data
ISSN: 2052-4463
Titre abrégé: Sci Data
Pays: England
ID NLM: 101640192
Informations de publication
Date de publication:
22 Apr 2024
22 Apr 2024
Historique:
received:
28
12
2023
accepted:
10
04
2024
medline:
23
4
2024
pubmed:
23
4
2024
entrez:
22
4
2024
Statut:
epublish
Résumé
Diploid wild oat Avena longiglumis has nutritional and adaptive traits which are valuable for common oat (A. sativa) breeding. The combination of Illumina, Nanopore and Hi-C data allowed us to assemble a high-quality chromosome-level genome of A. longiglumis (ALO), evidenced by contig N50 of 12.68 Mb with 99% BUSCO completeness for the assembly size of 3,960.97 Mb. A total of 40,845 protein-coding genes were annotated. The assembled genome was composed of 87.04% repetitive DNA sequences. Dotplots of the genome assembly (PI657387) with two published ALO genomes were compared to indicate the conservation of gene order and equal expansion of all syntenic blocks among three genome assemblies. Two recent whole-genome duplication events were characterized in genomes of diploid Avena species. These findings provide new knowledge for the genomic features of A. longiglumis, give information about the species diversity, and will accelerate the functional genomics and breeding studies in oat and related cereal crops.
Identifiants
pubmed: 38649380
doi: 10.1038/s41597-024-03248-6
pii: 10.1038/s41597-024-03248-6
doi:
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Pagination
412Subventions
Organisme : National Natural Science Foundation of China (National Science Foundation of China)
ID : 32070359
Organisme : National Natural Science Foundation of China (National Science Foundation of China)
ID : 32370402
Informations de copyright
© 2024. The Author(s).
Références
Grundy, M. M. L., Fardet, A., Tosh, S. M., Rich, G. T. & Wilde, P. J. Processing of oat: the impact on oat’s cholesterol lowering effect. Food Funct. 9, 1328–1343 (2018).
pubmed: 29431835
pmcid: 5885279
doi: 10.1039/C7FO02006F
Liu, K. S. Comparison of lipid content and fatty acid composition and their distribution within seeds of 5 small grain species. J. Food Sci. 76, C334–C342 (2011).
pubmed: 21535754
doi: 10.1111/j.1750-3841.2010.02038.x
White, D. A., Fisk, I. D. & Gray, D. A. Characterisation of oat (Avena sativa L.) oil bodies and intrinsically associated E-vitamers. J. Cereal Sci. 43, 244–249 (2006).
doi: 10.1016/j.jcs.2005.10.002
Yang, Z. et al. Oat: current state and challenges in plant-based food applications. Trends Food Sci. Technol. 134, 56–71 (2023).
doi: 10.1016/j.tifs.2023.02.017
Kamal, N. et al. The mosaic oat genome gives insights into a uniquely healthy cereal crop. Nature 606, 113–119 (2022).
pubmed: 35585233
pmcid: 9159951
doi: 10.1038/s41586-022-04732-y
Ouyang, S. et al. The TIGR rice genome annotation resource: improvements and new features. Nucleic Acids Res. 35, D883–D887 (2007).
pubmed: 17145706
doi: 10.1093/nar/gkl976
McCormick, R. F. et al. The Sorghum bicolor reference genome: improved assembly, gene annotations, a transcriptome atlas, and signatures of genome organization. Plant J. 93, 338–354 (2018).
pubmed: 29161754
doi: 10.1111/tpj.13781
Yang, Z. R. et al. A mini foxtail millet with an Arabidopsis-like life cycle as a C
pubmed: 32868891
doi: 10.1038/s41477-020-0747-7
Peng, Y. Y. et al. Reference genome assemblies reveal the origin and evolution of allohexaploid oat. Nat. Genet. 54, 1248–1258 (2022).
pubmed: 35851189
pmcid: 9355876
doi: 10.1038/s41588-022-01127-7
Liu, Q. et al. Genome-wide expansion and reorganization during grass evolution: from 30 Mb chromosomes in rice and Brachypodium to 550 Mb in Avena. BMC Plant Biol. 23, 627 (2023).
pubmed: 38062402
pmcid: 10704644
doi: 10.1186/s12870-023-04644-7
Saini, P. et al. Disease Resistance in Crop Plants: Molecular, Genetic and Genomic Perspectives (ed. Wani, S. H.) Ch. 9 (Springer Nature, 2019).
Simão, F. A., Waterhouse, R. M., Ioannidis, P., Kriventseva, E. V. & Zdobnov, E. M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210–3212 (2015).
pubmed: 26059717
doi: 10.1093/bioinformatics/btv351
Pruitt, K. D., Tatusova, T. & Maglott, D. R. NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts, and proteins. Nucleic Acids Res. 33, D501–D504 (2005).
pubmed: 15608248
doi: 10.1093/nar/gki025
Huerta-Cepas, J. et al. eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Res. 44, D309–D314 (2019).
doi: 10.1093/nar/gky1085
Finn, R. D. et al. The Pfam protein family’s database. Nucleic Acids Res. 36, D281–D288 (2014).
doi: 10.1093/nar/gkm960
Kristensen, D. M. et al. A low-polynomial algorithm for assembling clusters of orthologous groups from intergenomic symmetric best matches. Bioinformatics 26, 1481–1487 (2010).
pubmed: 20439257
pmcid: 2881409
doi: 10.1093/bioinformatics/btq229
Bairoch, A. & Apweiler, R. The SWISS-PROT protein sequence database and its supplement TrEMBL. Nucleic Acids Res. 28, 45–48 (2000).
pubmed: 10592178
pmcid: 102476
doi: 10.1093/nar/28.1.45
Ashburner, M. et al. Gene Ontology: tool for the unification of biology. Nat Genet. 25, 25–29 (2001).
doi: 10.1038/75556
Tatusov, R. L. et al. The COG database: an updated version includes eukaryotes. BMC Bioinformatics 4, 41 (2003).
pubmed: 12969510
pmcid: 222959
doi: 10.1186/1471-2105-4-41
Kanehisa, M. et al. KEGG for taxonomy-based analysis of pathways and genomes. Nucleic Acids Res. 51, D587–D592 (2023).
pubmed: 36300620
doi: 10.1093/nar/gkac963
Jin, J. et al. PlantTFDB 4.0: toward a central hub for transcription factors and regulatory interactions in plants. Nucleic Acids Res. 45, D1040–D1045 (2016).
pubmed: 27924042
pmcid: 5210657
doi: 10.1093/nar/gkw982
Levasseur, A., Drula, E., Lombard, V., Coutinho, P. M. & Henrissat, B. Expansion of the enzymatic repertoire of the CAZy database to integrate auxiliary redox enzymes. Biotechnol. Biofuels 6, 41 (2013).
pubmed: 23514094
pmcid: 3620520
doi: 10.1186/1754-6834-6-41
Chen, S., Zhou, Y., Chen, Y. & Gu, J. Fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, i884–i890 (2018).
pubmed: 30423086
pmcid: 6129281
doi: 10.1093/bioinformatics/bty560
Luo, R. et al. SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. GigaScience 1, 18 (2012).
pubmed: 23587118
pmcid: 3626529
doi: 10.1186/2047-217X-1-18
Marcais, G. & Kingsford, C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics 27, 764 (2011).
pubmed: 21217122
pmcid: 3051319
doi: 10.1093/bioinformatics/btr011
Ranallo-Benavidez, T. R., Jaron, K. S. & Schatz, M. C. GenomeScope 2.0 and Smudgeplot for reference-free profiling of polyploid genomes. Nat. Commun. 11, 1–10 (2020).
doi: 10.1038/s41467-020-14998-3
Ruan, J. & Li, H. Fast and accurate long-read assembly with wtdbg2. Nat. Methods 17, 155–158 (2020).
pubmed: 31819265
doi: 10.1038/s41592-019-0669-3
Liu, H., Wu, S., Li, A. & Ruan, J. SMARTdenovo: a de novo assembler using long noisy reads. GigaByte 15, 1–9 (2021).
doi: 10.46471/gigabyte.15
Hu, J., Fan, J. P., Sun, Z. Y. & Liu, S. L. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 36, 2253–2255 (2020).
pubmed: 31778144
doi: 10.1093/bioinformatics/btz891
Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat. Methods 9, 357–359 (2012).
pubmed: 22388286
pmcid: 3322381
doi: 10.1038/nmeth.1923
Durand, N. C. et al. Juicer provides a one-click system for analyzing loop-resolution Hi-C experiments. Cell Syst. 3, 95–98 (2016).
pubmed: 27467249
pmcid: 5846465
doi: 10.1016/j.cels.2016.07.002
Servant, N. et al. HiC-Pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol. 16, 259 (2015).
pubmed: 26619908
pmcid: 4665391
doi: 10.1186/s13059-015-0831-x
Burton, J. N. et al. Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions. Nat. Biotechnol. 31, 1119–1125 (2013).
pubmed: 24185095
pmcid: 4117202
doi: 10.1038/nbt.2727
Ou, S. J. & Jiang, N. LTR_FINDER_parallel: parallelization of LTR_FINDER enabling rapid identification of long terminal repeat retrotransposons. Mob. DNA 10, 48 (2019).
pubmed: 31857828
pmcid: 6909508
doi: 10.1186/s13100-019-0193-0
Ellinghaus, D., Kurtz, S. & Willhoeft, U. LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons. BMC Bioinformatics 9, 18 (2008).
pubmed: 18194517
pmcid: 2253517
doi: 10.1186/1471-2105-9-18
Ou, S. J. & Jiang, N. LTR_retriever: a highly accurate and sensitive program for identification of long terminal repeat retrotransposons. Plant Physiol. 176, 1410–1422 (2018).
pubmed: 29233850
doi: 10.1104/pp.17.01310
Shi, J. M. & Liang, C. Generic repeat finder: a high-sensitivity tool for genome-wide de novo repeat detection. Plant Physiol. 180, 1803–1815 (2019).
pubmed: 31152127
pmcid: 6670090
doi: 10.1104/pp.19.00386
Su, W., Gu, X. & Peterson, T. TIR-Learner, a new ensemble method for TIR transposable element annotation, provides evidence for abundant new transposable elements in the maize genome. Mol. Plant 12, 447–460 (2016).
doi: 10.1016/j.molp.2019.02.008
Xiong, W., He, L. M., Lai, J. S., Dooner, H. K. & Du, C. G. HelitronScanner uncovers a large overlooked cache of Helitron transposons in many plant genomes. Proc. Natl. Acad. Sci. USA 111, 10263–10268 (2014).
pubmed: 24982153
pmcid: 4104883
doi: 10.1073/pnas.1410068111
Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc. Natl. Acad. Sci. USA 117, 9451–9457 (2020).
pubmed: 32300014
pmcid: 7196820
doi: 10.1073/pnas.1921046117
Tarailo-Graovac, M. & Chen, N. S. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr. Protoc. Bioinformatics 4, 1–14 (2009).
Zhang, R. G. et al. TEsorter: an accurate and fast method to classify LTR-retrotransposons in plant genomes. Hortic. Res. 9, uhac017 (2022).
pubmed: 35184178
pmcid: 9002660
doi: 10.1093/hr/uhac017
Yandell, M. & Ence, D. A beginner’s guide to eukaryotic genome annotation. Nat. Rev. Genet. 13, 329–342 (2012).
pubmed: 22510764
doi: 10.1038/nrg3174
Stanke, M. et al. Augustus: ab initio prediction of alternative transcripts. Nucleic Acids Res. 34, W435–W439 (2006).
pubmed: 16845043
pmcid: 1538822
doi: 10.1093/nar/gkl200
Keilwagen, J., Hartung, F. & Grau, J. GeMoMa: homology-based gene prediction utilizing intron position conservation and RNA-seq data. Methods Mol. Biol. 1962, 161–177 (2019).
pubmed: 31020559
doi: 10.1007/978-1-4939-9173-0_9
The Arabidopsis Genome Initiative. Analysis of the genome sequence of the flowering plant Arabidopsis thaliana. Nature 408, 796–815 (2000).
doi: 10.1038/35048692
International Brachypodium Initiative. Genome sequencing and analysis of the model grass Brachypodium distachyon. Nature 463, 763–768 (2010).
doi: 10.1038/nature08747
Mascher, M. et al. A chromosome conformation capture ordered sequence of the barley genome. Nature 544, 427–433 (2017).
pubmed: 28447635
doi: 10.1038/nature22043
International Wheat Genome Sequencing Consortium (IWGSC). Shifting the limits in wheat research and breeding using a fully annotated reference genome. Science 361, eaar7191 (2018).
doi: 10.1126/science.aar7191
Schnable, P. S. et al. The B73 maize genome: complexity, diversity, and dynamics. Science 326, 1112–1115 (2009).
pubmed: 19965430
doi: 10.1126/science.1178534
Haas, B. J. et al. De novo transcript sequence reconstruction from RNA-seq using the Trinity platform for reference generation and analysis. Nat. Protoc. 8, 1494–1512 (2013).
pubmed: 23845962
doi: 10.1038/nprot.2013.084
Haas, B. J. et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the program to assemble spliced alignments. Genome Biol. 9, 1–22 (2008).
doi: 10.1186/gb-2008-9-1-r7
Lavigne, R., Seto, D., Mahadevan, P., Ackermann, H. W. & Kropinski, A. M. Unifying classical and molecular taxonomic classification: analysis of the Podoviridae using BLASTP-based tools. Res. Microbiol. 159, 406–414 (2008).
pubmed: 18555669
doi: 10.1016/j.resmic.2008.03.005
Conesa, A. et al. Blast2GO: a universal tool for annotation, visualization and analysis in functional genomics research. Bioinformatics 21, 3674–3676 (2005).
pubmed: 16081474
doi: 10.1093/bioinformatics/bti610
Griffiths-Jones, S. et al. Rfam: annotating non-coding RNAs in complete genomes. Nucleic Acids Res. 33, D121–D124 (2005).
pubmed: 15608160
doi: 10.1093/nar/gki081
Chan, P. P., Lin, B. Y., Mak, A. J. & Lowe, T. M. tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic Acids Res. 49, 9077–9096 (2021).
pubmed: 34417604
pmcid: 8450103
doi: 10.1093/nar/gkab688
Li, H., Feng, X. & Chu, C. The design and construction of reference pangenome graphs with minigraph. Genome Biol. 21, 1–19 (2020).
doi: 10.1186/s13059-020-02168-z
Cabanettes, F. & Klopp, C. D-GENIES: dot plot large genomes in an interactive, efficient and simple way. Peer J. 6, e4958 (2018).
pubmed: 29888139
pmcid: 5991294
doi: 10.7717/peerj.4958
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRP375311 (2022).
NCBI RNA Sequencing Data https://identifiers.org/ncbi/insdc.sra:SRP433645 (2023).
NCBI Assembly https://identifiers.org/ncbi/insdc.gca:GCA_030063025.1 (2023).