Chromosome-level genome assembly of the diploid oat species Avena longiglumis.


Journal

Scientific data
ISSN: 2052-4463
Titre abrégé: Sci Data
Pays: England
ID NLM: 101640192

Informations de publication

Date de publication:
22 Apr 2024
Historique:
received: 28 12 2023
accepted: 10 04 2024
medline: 23 4 2024
pubmed: 23 4 2024
entrez: 22 4 2024
Statut: epublish

Résumé

Diploid wild oat Avena longiglumis has nutritional and adaptive traits which are valuable for common oat (A. sativa) breeding. The combination of Illumina, Nanopore and Hi-C data allowed us to assemble a high-quality chromosome-level genome of A. longiglumis (ALO), evidenced by contig N50 of 12.68 Mb with 99% BUSCO completeness for the assembly size of 3,960.97 Mb. A total of 40,845 protein-coding genes were annotated. The assembled genome was composed of 87.04% repetitive DNA sequences. Dotplots of the genome assembly (PI657387) with two published ALO genomes were compared to indicate the conservation of gene order and equal expansion of all syntenic blocks among three genome assemblies. Two recent whole-genome duplication events were characterized in genomes of diploid Avena species. These findings provide new knowledge for the genomic features of A. longiglumis, give information about the species diversity, and will accelerate the functional genomics and breeding studies in oat and related cereal crops.

Identifiants

pubmed: 38649380
doi: 10.1038/s41597-024-03248-6
pii: 10.1038/s41597-024-03248-6
doi:

Types de publication

Journal Article

Langues

eng

Sous-ensembles de citation

IM

Pagination

412

Subventions

Organisme : National Natural Science Foundation of China (National Science Foundation of China)
ID : 32070359
Organisme : National Natural Science Foundation of China (National Science Foundation of China)
ID : 32370402

Informations de copyright

© 2024. The Author(s).

Références

Grundy, M. M. L., Fardet, A., Tosh, S. M., Rich, G. T. & Wilde, P. J. Processing of oat: the impact on oat’s cholesterol lowering effect. Food Funct. 9, 1328–1343 (2018).
pubmed: 29431835 pmcid: 5885279 doi: 10.1039/C7FO02006F
Liu, K. S. Comparison of lipid content and fatty acid composition and their distribution within seeds of 5 small grain species. J. Food Sci. 76, C334–C342 (2011).
pubmed: 21535754 doi: 10.1111/j.1750-3841.2010.02038.x
White, D. A., Fisk, I. D. & Gray, D. A. Characterisation of oat (Avena sativa L.) oil bodies and intrinsically associated E-vitamers. J. Cereal Sci. 43, 244–249 (2006).
doi: 10.1016/j.jcs.2005.10.002
Yang, Z. et al. Oat: current state and challenges in plant-based food applications. Trends Food Sci. Technol. 134, 56–71 (2023).
doi: 10.1016/j.tifs.2023.02.017
Kamal, N. et al. The mosaic oat genome gives insights into a uniquely healthy cereal crop. Nature 606, 113–119 (2022).
pubmed: 35585233 pmcid: 9159951 doi: 10.1038/s41586-022-04732-y
Ouyang, S. et al. The TIGR rice genome annotation resource: improvements and new features. Nucleic Acids Res. 35, D883–D887 (2007).
pubmed: 17145706 doi: 10.1093/nar/gkl976
McCormick, R. F. et al. The Sorghum bicolor reference genome: improved assembly, gene annotations, a transcriptome atlas, and signatures of genome organization. Plant J. 93, 338–354 (2018).
pubmed: 29161754 doi: 10.1111/tpj.13781
Yang, Z. R. et al. A mini foxtail millet with an Arabidopsis-like life cycle as a C
pubmed: 32868891 doi: 10.1038/s41477-020-0747-7
Peng, Y. Y. et al. Reference genome assemblies reveal the origin and evolution of allohexaploid oat. Nat. Genet. 54, 1248–1258 (2022).
pubmed: 35851189 pmcid: 9355876 doi: 10.1038/s41588-022-01127-7
Liu, Q. et al. Genome-wide expansion and reorganization during grass evolution: from 30 Mb chromosomes in rice and Brachypodium to 550 Mb in Avena. BMC Plant Biol. 23, 627 (2023).
pubmed: 38062402 pmcid: 10704644 doi: 10.1186/s12870-023-04644-7
Saini, P. et al. Disease Resistance in Crop Plants: Molecular, Genetic and Genomic Perspectives (ed. Wani, S. H.) Ch. 9 (Springer Nature, 2019).
Simão, F. A., Waterhouse, R. M., Ioannidis, P., Kriventseva, E. V. & Zdobnov, E. M. BUSCO: assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210–3212 (2015).
pubmed: 26059717 doi: 10.1093/bioinformatics/btv351
Pruitt, K. D., Tatusova, T. & Maglott, D. R. NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts, and proteins. Nucleic Acids Res. 33, D501–D504 (2005).
pubmed: 15608248 doi: 10.1093/nar/gki025
Huerta-Cepas, J. et al. eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Res. 44, D309–D314 (2019).
doi: 10.1093/nar/gky1085
Finn, R. D. et al. The Pfam protein family’s database. Nucleic Acids Res. 36, D281–D288 (2014).
doi: 10.1093/nar/gkm960
Kristensen, D. M. et al. A low-polynomial algorithm for assembling clusters of orthologous groups from intergenomic symmetric best matches. Bioinformatics 26, 1481–1487 (2010).
pubmed: 20439257 pmcid: 2881409 doi: 10.1093/bioinformatics/btq229
Bairoch, A. & Apweiler, R. The SWISS-PROT protein sequence database and its supplement TrEMBL. Nucleic Acids Res. 28, 45–48 (2000).
pubmed: 10592178 pmcid: 102476 doi: 10.1093/nar/28.1.45
Ashburner, M. et al. Gene Ontology: tool for the unification of biology. Nat Genet. 25, 25–29 (2001).
doi: 10.1038/75556
Tatusov, R. L. et al. The COG database: an updated version includes eukaryotes. BMC Bioinformatics 4, 41 (2003).
pubmed: 12969510 pmcid: 222959 doi: 10.1186/1471-2105-4-41
Kanehisa, M. et al. KEGG for taxonomy-based analysis of pathways and genomes. Nucleic Acids Res. 51, D587–D592 (2023).
pubmed: 36300620 doi: 10.1093/nar/gkac963
Jin, J. et al. PlantTFDB 4.0: toward a central hub for transcription factors and regulatory interactions in plants. Nucleic Acids Res. 45, D1040–D1045 (2016).
pubmed: 27924042 pmcid: 5210657 doi: 10.1093/nar/gkw982
Levasseur, A., Drula, E., Lombard, V., Coutinho, P. M. & Henrissat, B. Expansion of the enzymatic repertoire of the CAZy database to integrate auxiliary redox enzymes. Biotechnol. Biofuels 6, 41 (2013).
pubmed: 23514094 pmcid: 3620520 doi: 10.1186/1754-6834-6-41
Chen, S., Zhou, Y., Chen, Y. & Gu, J. Fastp: an ultra-fast all-in-one FASTQ preprocessor. Bioinformatics 34, i884–i890 (2018).
pubmed: 30423086 pmcid: 6129281 doi: 10.1093/bioinformatics/bty560
Luo, R. et al. SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler. GigaScience 1, 18 (2012).
pubmed: 23587118 pmcid: 3626529 doi: 10.1186/2047-217X-1-18
Marcais, G. & Kingsford, C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics 27, 764 (2011).
pubmed: 21217122 pmcid: 3051319 doi: 10.1093/bioinformatics/btr011
Ranallo-Benavidez, T. R., Jaron, K. S. & Schatz, M. C. GenomeScope 2.0 and Smudgeplot for reference-free profiling of polyploid genomes. Nat. Commun. 11, 1–10 (2020).
doi: 10.1038/s41467-020-14998-3
Ruan, J. & Li, H. Fast and accurate long-read assembly with wtdbg2. Nat. Methods 17, 155–158 (2020).
pubmed: 31819265 doi: 10.1038/s41592-019-0669-3
Liu, H., Wu, S., Li, A. & Ruan, J. SMARTdenovo: a de novo assembler using long noisy reads. GigaByte 15, 1–9 (2021).
doi: 10.46471/gigabyte.15
Hu, J., Fan, J. P., Sun, Z. Y. & Liu, S. L. NextPolish: a fast and efficient genome polishing tool for long-read assembly. Bioinformatics 36, 2253–2255 (2020).
pubmed: 31778144 doi: 10.1093/bioinformatics/btz891
Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat. Methods 9, 357–359 (2012).
pubmed: 22388286 pmcid: 3322381 doi: 10.1038/nmeth.1923
Durand, N. C. et al. Juicer provides a one-click system for analyzing loop-resolution Hi-C experiments. Cell Syst. 3, 95–98 (2016).
pubmed: 27467249 pmcid: 5846465 doi: 10.1016/j.cels.2016.07.002
Servant, N. et al. HiC-Pro: an optimized and flexible pipeline for Hi-C data processing. Genome Biol. 16, 259 (2015).
pubmed: 26619908 pmcid: 4665391 doi: 10.1186/s13059-015-0831-x
Burton, J. N. et al. Chromosome-scale scaffolding of de novo genome assemblies based on chromatin interactions. Nat. Biotechnol. 31, 1119–1125 (2013).
pubmed: 24185095 pmcid: 4117202 doi: 10.1038/nbt.2727
Ou, S. J. & Jiang, N. LTR_FINDER_parallel: parallelization of LTR_FINDER enabling rapid identification of long terminal repeat retrotransposons. Mob. DNA 10, 48 (2019).
pubmed: 31857828 pmcid: 6909508 doi: 10.1186/s13100-019-0193-0
Ellinghaus, D., Kurtz, S. & Willhoeft, U. LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons. BMC Bioinformatics 9, 18 (2008).
pubmed: 18194517 pmcid: 2253517 doi: 10.1186/1471-2105-9-18
Ou, S. J. & Jiang, N. LTR_retriever: a highly accurate and sensitive program for identification of long terminal repeat retrotransposons. Plant Physiol. 176, 1410–1422 (2018).
pubmed: 29233850 doi: 10.1104/pp.17.01310
Shi, J. M. & Liang, C. Generic repeat finder: a high-sensitivity tool for genome-wide de novo repeat detection. Plant Physiol. 180, 1803–1815 (2019).
pubmed: 31152127 pmcid: 6670090 doi: 10.1104/pp.19.00386
Su, W., Gu, X. & Peterson, T. TIR-Learner, a new ensemble method for TIR transposable element annotation, provides evidence for abundant new transposable elements in the maize genome. Mol. Plant 12, 447–460 (2016).
doi: 10.1016/j.molp.2019.02.008
Xiong, W., He, L. M., Lai, J. S., Dooner, H. K. & Du, C. G. HelitronScanner uncovers a large overlooked cache of Helitron transposons in many plant genomes. Proc. Natl. Acad. Sci. USA 111, 10263–10268 (2014).
pubmed: 24982153 pmcid: 4104883 doi: 10.1073/pnas.1410068111
Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc. Natl. Acad. Sci. USA 117, 9451–9457 (2020).
pubmed: 32300014 pmcid: 7196820 doi: 10.1073/pnas.1921046117
Tarailo-Graovac, M. & Chen, N. S. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr. Protoc. Bioinformatics 4, 1–14 (2009).
Zhang, R. G. et al. TEsorter: an accurate and fast method to classify LTR-retrotransposons in plant genomes. Hortic. Res. 9, uhac017 (2022).
pubmed: 35184178 pmcid: 9002660 doi: 10.1093/hr/uhac017
Yandell, M. & Ence, D. A beginner’s guide to eukaryotic genome annotation. Nat. Rev. Genet. 13, 329–342 (2012).
pubmed: 22510764 doi: 10.1038/nrg3174
Stanke, M. et al. Augustus: ab initio prediction of alternative transcripts. Nucleic Acids Res. 34, W435–W439 (2006).
pubmed: 16845043 pmcid: 1538822 doi: 10.1093/nar/gkl200
Keilwagen, J., Hartung, F. & Grau, J. GeMoMa: homology-based gene prediction utilizing intron position conservation and RNA-seq data. Methods Mol. Biol. 1962, 161–177 (2019).
pubmed: 31020559 doi: 10.1007/978-1-4939-9173-0_9
The Arabidopsis Genome Initiative. Analysis of the genome sequence of the flowering plant Arabidopsis thaliana. Nature 408, 796–815 (2000).
doi: 10.1038/35048692
International Brachypodium Initiative. Genome sequencing and analysis of the model grass Brachypodium distachyon. Nature 463, 763–768 (2010).
doi: 10.1038/nature08747
Mascher, M. et al. A chromosome conformation capture ordered sequence of the barley genome. Nature 544, 427–433 (2017).
pubmed: 28447635 doi: 10.1038/nature22043
International Wheat Genome Sequencing Consortium (IWGSC). Shifting the limits in wheat research and breeding using a fully annotated reference genome. Science 361, eaar7191 (2018).
doi: 10.1126/science.aar7191
Schnable, P. S. et al. The B73 maize genome: complexity, diversity, and dynamics. Science 326, 1112–1115 (2009).
pubmed: 19965430 doi: 10.1126/science.1178534
Haas, B. J. et al. De novo transcript sequence reconstruction from RNA-seq using the Trinity platform for reference generation and analysis. Nat. Protoc. 8, 1494–1512 (2013).
pubmed: 23845962 doi: 10.1038/nprot.2013.084
Haas, B. J. et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the program to assemble spliced alignments. Genome Biol. 9, 1–22 (2008).
doi: 10.1186/gb-2008-9-1-r7
Lavigne, R., Seto, D., Mahadevan, P., Ackermann, H. W. & Kropinski, A. M. Unifying classical and molecular taxonomic classification: analysis of the Podoviridae using BLASTP-based tools. Res. Microbiol. 159, 406–414 (2008).
pubmed: 18555669 doi: 10.1016/j.resmic.2008.03.005
Conesa, A. et al. Blast2GO: a universal tool for annotation, visualization and analysis in functional genomics research. Bioinformatics 21, 3674–3676 (2005).
pubmed: 16081474 doi: 10.1093/bioinformatics/bti610
Griffiths-Jones, S. et al. Rfam: annotating non-coding RNAs in complete genomes. Nucleic Acids Res. 33, D121–D124 (2005).
pubmed: 15608160 doi: 10.1093/nar/gki081
Chan, P. P., Lin, B. Y., Mak, A. J. & Lowe, T. M. tRNAscan-SE 2.0: improved detection and functional classification of transfer RNA genes. Nucleic Acids Res. 49, 9077–9096 (2021).
pubmed: 34417604 pmcid: 8450103 doi: 10.1093/nar/gkab688
Li, H., Feng, X. & Chu, C. The design and construction of reference pangenome graphs with minigraph. Genome Biol. 21, 1–19 (2020).
doi: 10.1186/s13059-020-02168-z
Cabanettes, F. & Klopp, C. D-GENIES: dot plot large genomes in an interactive, efficient and simple way. Peer J. 6, e4958 (2018).
pubmed: 29888139 pmcid: 5991294 doi: 10.7717/peerj.4958
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRP375311 (2022).
NCBI RNA Sequencing Data https://identifiers.org/ncbi/insdc.sra:SRP433645 (2023).
NCBI Assembly https://identifiers.org/ncbi/insdc.gca:GCA_030063025.1 (2023).

Auteurs

Qing Liu (Q)

State Key Laboratory of Plant Diversity and Specialty Crops / Guangdong Provincial Key Laboratory of Applied Botany, South China Botanical Garden, Chinese Academy of Sciences, Guangzhou, 510650, China. liuqing@scib.ac.cn.
South China National Botanical Garden, Guangzhou, China. liuqing@scib.ac.cn.

Gui Xiong (G)

State Key Laboratory of Plant Diversity and Specialty Crops / Guangdong Provincial Key Laboratory of Applied Botany, South China Botanical Garden, Chinese Academy of Sciences, Guangzhou, 510650, China.
University of Chinese Academy of Sciences, Beijing, China.

Ziwei Wang (Z)

School of Biology and Agriculture, Shaoguan University, Shaoguan, China.

Yongxing Wu (Y)

College of Agriculture, South China Agricultural University, Guangzhou, China.

Tieyao Tu (T)

State Key Laboratory of Plant Diversity and Specialty Crops / Guangdong Provincial Key Laboratory of Applied Botany, South China Botanical Garden, Chinese Academy of Sciences, Guangzhou, 510650, China.
South China National Botanical Garden, Guangzhou, China.

Trude Schwarzacher (T)

State Key Laboratory of Plant Diversity and Specialty Crops / Guangdong Provincial Key Laboratory of Applied Botany, South China Botanical Garden, Chinese Academy of Sciences, Guangzhou, 510650, China.
South China National Botanical Garden, Guangzhou, China.
University of Leicester, Department of Genetics and Genome Biology, Institute for Environmental Futures, Leicester, UK.

John Seymour Heslop-Harrison (JS)

State Key Laboratory of Plant Diversity and Specialty Crops / Guangdong Provincial Key Laboratory of Applied Botany, South China Botanical Garden, Chinese Academy of Sciences, Guangzhou, 510650, China. phh4@le.ac.uk.
South China National Botanical Garden, Guangzhou, China. phh4@le.ac.uk.
University of Leicester, Department of Genetics and Genome Biology, Institute for Environmental Futures, Leicester, UK. phh4@le.ac.uk.

Classifications MeSH