Chromosome-level genome assembly of Euphorbia tirucalli (Euphorbiaceae), a highly stress-tolerant oil plant.


Journal

Scientific data
ISSN: 2052-4463
Titre abrégé: Sci Data
Pays: England
ID NLM: 101640192

Informations de publication

Date de publication:
21 Jun 2024
Historique:
received: 15 02 2024
accepted: 10 06 2024
medline: 22 6 2024
pubmed: 22 6 2024
entrez: 21 6 2024
Statut: epublish

Résumé

Euphorbia, one of the largest genera of flowering plants, is well-known for containing many biofuel crops. Euphorbia tirucalli, an evergreen succulent mainly native to the Africa continent but cultivated worldwide, is a promising petroleum plant with high tolerance to drought and salt stress. However, the exploration of such an important plant resource is severely hampered by the lack of a reference genome. Here, we present the chromosome-level genome assembly of E. tirucalli using PacBio HiFi sequencing and Hi-C technology. Its genome size was approximately 745.62 Mb, with a contig N50 of 74.16 Mb. A total of 743.63 Mb (99.73%) of the assembled sequences were anchored to 10 chromosomes with a complete BUSCO score of 97.80%. Genome annotation revealed 26,304 protein-coding genes, and 76.37% of the genome was identified as repeat elements. The high-quality genome provides valuable genetic resources that would be useful for unraveling the genetic mechanisms of biofuel synthesis and evolutionary adaptation of E. tirucalli.

Identifiants

pubmed: 38906925
doi: 10.1038/s41597-024-03503-w
pii: 10.1038/s41597-024-03503-w
doi:

Types de publication

Journal Article Dataset

Langues

eng

Sous-ensembles de citation

IM

Pagination

658

Subventions

Organisme : Guangdong Science and Technology Department (Science and Technology Department, Guangdong Province)
ID : 2023B0303050001
Organisme : National Natural Science Foundation of China (National Science Foundation of China)
ID : 31970360

Informations de copyright

© 2024. The Author(s).

Références

Bruyns, P. V., Mapaya, R. J. & Hedderson, T. J. J. T. A new subgeneric classification for Euphorbia (Euphorbiaceae) in southern Africa based on ITS and psbA-trnH sequence data. Taxon 55, 397–420 (2006).
doi: 10.2307/25065587
Duke, J. Euphorbia tirucalli L., handbook of energy crops. Purdue University centre for new crops and plant products (1983).
Loke, J., Mesa, L. A. & Franken, J. Y. Euphorbia tirucalli biology manual: Feedstock production, bioenergy conversion, application, economics Version 2. FACT. (2011).
Hastilestari, B. R. et al. Euphorbia tirucalli L.–Comprehensive characterization of a drought tolerant plant with a potential as biofuel source. PLoS one 8, e63501 (2013).
pubmed: 23658836 pmcid: 3643915 doi: 10.1371/journal.pone.0063501
Duy Khang, N. V., Nhi Nguyen, T. & Quan, T. L. Potential biofuel exploitation from two common Vietnamese Euphorbia plants (Euphorbiaceae). Biofuel. Bioprod. Biorefin. 17, 1315–1327 (2023).
doi: 10.1002/bbb.2472
Li, Z., Huang, J. & Li, L. Polyploid induction in Euphorbia tirucalli L. with Colchicine. J Southwest 29, 106–110 (2007). (in Chinese).
Van Damme, P. L. Euphorbia tirucalli for high biomass production. In Combating desertification with plants. (Boston, MA: Springer US, 2001).
Saigo, R. H. & Saigo, B. W. Botany, principles and applications. (Pren-tice-Hall, Englewood CliVs, 1983).
Calvin, M. Petroleum plantations for fuel and materials. Bioscience. 29, 533–538 (1979).
doi: 10.2307/1307721
Calvin, M. Fuel oils from euphorbs and other plants. Bot. J. Linnean Soc. 94, 97–110 (1987).
doi: 10.1111/j.1095-8339.1987.tb01040.x
Van Damme, P. Het traditioneel gebruik van Euphorbia tirucalli. Afrika Focus 5, 176–193 (1989).
doi: 10.1163/2031356X-0050304006
Aylward, J. H. & Parsons, P. G. Treatment of prostate cancer. (Peplin Research May, US, 2008).
Kumar, A. Some potential plants for medicine from India. (Ayurvedic medicines, University of Rajasthan, Rajasthan, 1999).
Schmelzer, G. H. & Gurib-Fakim, A. Medicinal plants. (Plant Resources of Tropical Africa, 2008).
Abuelsoud, W., Hirschmann, F. & Papenbrock, J. Sulfur metabolism and drought stress tolerance in plants. Vol. 1: Physiology and Biochemistry 227–249 (2016).
Cushman, J. C. & Borland, A. Induction of Crassulacean acid metabolism by water limitation. Plant Cell Environ. 25, 295–310 (2002).
pubmed: 11841671 doi: 10.1046/j.0016-8025.2001.00760.x
Janssens, M. J., Keutgen, N. & Pohlan, J. The role of bio-productivity on bio-energy yields. J. Agr. Rural Dev. Trop. 110, 41–48 (2009).
Zika, M. & Erb, K. H. The global loss of net primary production resulting from human-induced soil degradation in drylands. Ecol. Econ. 69, 310–318 (2009).
doi: 10.1016/j.ecolecon.2009.06.014
Wang, R. et al. Production and selected fuel properties of biodiesel from promising non-edible oils: Euphorbia lathyris L., Sapium sebiferum L. and Jatropha curcas L. Bioresour. Technol. 102, 1194–1199 (2011).
pubmed: 20951029 doi: 10.1016/j.biortech.2010.09.066
Dai, A. Increasing drought under global warming in observations and models. Nat. Clim. Change 3, 52–58 (2013).
doi: 10.1038/nclimate1633
Taparia, T., Mvss, M., Mehrotra, R., Shukla, P. & Mehrotra, S. Developments and challenges in biodiesel production from microalgae: A review. Biotechnol. Appl. Biochem. 63, 715–726 (2016).
pubmed: 26178774 doi: 10.1002/bab.1412
Demirbas, A. Competitive liquid biofuels from biomass. Appl. Energy. 88, 17–28 (2011).
doi: 10.1016/j.apenergy.2010.07.016
Sharma, R., Wungrampha, S., Singh, V., Pareek, A. & Sharma, M. K. Halophytes as bioenergy crops. Front. Plant Sci. 7, 1372 (2016).
pubmed: 27679645 pmcid: 5020079 doi: 10.3389/fpls.2016.01372
Winnepenninckx, B., Backeljau, T. & De Wachter, R. Extraction of high molecular weight DNA from molluscs. Trends Genet. 9, 407 (1993).
pubmed: 8122306 doi: 10.1016/0168-9525(93)90102-N
Liu, B. et al. Estimation of genomic characteristics by analyzing k-mer frequency in de novo genome projects. arXiv.org, arXiv: 1308.2012 (2013).
Cheng, H., Concepcion, G. T., Feng, X., Zhang, H. & Li, H. Haplotype-resolved de novo assembly using phased assembly graphs with hifiasm. Nat. Methods 18, 170–175 (2021).
pubmed: 33526886 pmcid: 7961889 doi: 10.1038/s41592-020-01056-5
Simão, F. A., Waterhouse, R. M., Ioannidis, P., Kriventseva, E. V. & Zdobnov, E. M. BUSCO: Assessing genome assembly and annotation completeness with single-copy orthologs. Bioinformatics 31, 3210–3212 (2015).
pubmed: 26059717 doi: 10.1093/bioinformatics/btv351
Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25, 1754-1760.
Durand, N. C. et al. Juicer provides a one-click system for analyzing loop-resolution Hi-C experiments. Cell Syst. 3, 95–98 (2016).
pubmed: 27467249 pmcid: 5846465 doi: 10.1016/j.cels.2016.07.002
Dudchenko, O. et al. De novo assembly of the Aedes aegypti genome using Hi-C yields chromosome-length scaffolds. Science 356, 92–95 (2017).
pubmed: 28336562 pmcid: 5635820 doi: 10.1126/science.aal3327
Durand, N. C. et al. Juicebox provides a visualization system for Hi-C contact maps with unlimited zoom. Cell Syst. 3, 99–101 (2016).
pubmed: 27467250 pmcid: 5596920 doi: 10.1016/j.cels.2015.07.012
Ou, S., Chen, J. & Jiang, N. Assessing genome assembly quality using the LTR Assembly Index (LAI). Nucleic Acids Res. 46, e126–e126 (2018).
pubmed: 30107434 pmcid: 6265445
Benson, G. Tandem repeats finder: a program to analyze DNA sequences. Nucleic Acids Res. 27, 573–580 (1999).
pubmed: 9862982 pmcid: 148217 doi: 10.1093/nar/27.2.573
Chen, N. Using Repeat Masker to identify repetitive elements in genomic sequences. Curr. Protoc. Bioinformatics 5, 4.10.11–14.10.14 (2004).
doi: 10.1002/0471250953.bi0410s05
Xu, Z. & Wang, H. LTR_FINDER: an efficient tool for the prediction of full-length LTR retrotransposons. Nucleic Acids Res. 35, W265–W268 (2007).
pubmed: 17485477 pmcid: 1933203 doi: 10.1093/nar/gkm286
Ellinghaus, D., Kurtz, S. & Willhoeft, U. LTRharvest, an efficient and flexible software for de novo detection of LTR retrotransposons. Bmc. Bioinformatics 9, 1–114 (2008).
doi: 10.1186/1471-2105-9-18
Ou, S. & Jiang, N. LTR_retriever: A highly accurate and sensitive program for identification of long terminal repeat retrotransposons. Plant Physiol. 176, 1410–1422 (2018).
pubmed: 29233850 doi: 10.1104/pp.17.01310
Han, Y. & Wessler, S. R. MITE-Hunter: a program for discovering miniature inverted-repeat transposable elements from genomic sequences. Nucleic Acids Res. 38, e199 (2010).
pubmed: 20880995 pmcid: 3001096 doi: 10.1093/nar/gkq862
Flynn, J. M. et al. RepeatModeler2 for automated genomic discovery of transposable element families. Proc. Natl Acad. Sci. USA 117, 9451–9457 (2020).
pubmed: 32300014 pmcid: 7196820 doi: 10.1073/pnas.1921046117
Grabherr, M. G. et al. Full-length transcriptome assembly from RNA-Seq data without a reference genome. Nat. Biotechnol. 29, 644–652 (2011).
pubmed: 21572440 pmcid: 3571712 doi: 10.1038/nbt.1883
Kim, D., Langmead, B. & Salzberg, S. L. HISAT: A fast spliced aligner with low memory requirements. Nat. Methods 12, 357–360 (2015).
pubmed: 25751142 pmcid: 4655817 doi: 10.1038/nmeth.3317
Li, H. et al. The sequence alignment/map format and SAMtools. Bioinformatics 25, 2078–2079 (2009).
pubmed: 19505943 pmcid: 2723002 doi: 10.1093/bioinformatics/btp352
Pertea, M. et al. StringTie enables improved reconstruction of a transcriptome from RNA-seq reads. Nat. Biotechnol. 33, 290–295 (2015).
pubmed: 25690850 pmcid: 4643835 doi: 10.1038/nbt.3122
Haas, B. J. et al. Improving the Arabidopsis genome annotation using maximal transcript alignment assemblies. Nucleic Acids Res. 31, 5654–5666 (2003).
pubmed: 14500829 pmcid: 206470 doi: 10.1093/nar/gkg770
Johnson, A. D. et al. SNAP: a web-based tool for identification and annotation of proxy SNPs using HapMap. Bioinformatics 24, 2938–2939 (2008).
pubmed: 18974171 pmcid: 2720775 doi: 10.1093/bioinformatics/btn564
Guigó, R., Knudsen, S., Drake, N. & Smith, T. Prediction of gene structure. J. Mol. Biol. 226, 141–157 (1992).
pubmed: 1619647 doi: 10.1016/0022-2836(92)90130-C
Majoros, W. H., Pertea, M. & Salzberg, S. L. TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders. Bioinformatics 20, 2878–2879 (2004).
pubmed: 15145805 doi: 10.1093/bioinformatics/bth315
Lomsadze, A., Burns, P. D. & Borodovsky, M. Integration of mapped RNA-Seq reads into automatic training of eukaryotic gene finding algorithm. Nucleic Acids Res. 42, e119–e119 (2014).
pubmed: 24990371 pmcid: 4150757 doi: 10.1093/nar/gku557
Stanke, M. et al. AUGUSTUS: ab initio prediction of alternative transcripts. Nucleic Acids Res. 34, W435–439 (2006).
pubmed: 16845043 pmcid: 1538822 doi: 10.1093/nar/gkl200
Keilwagen, J., Hartung, F. & Grau, J. GeMoMa: homology-based gene prediction utilizing intron position conservation and RNA-seq data. Methods Mol. Biol. 1962, 161–177 (2019).
pubmed: 31020559 doi: 10.1007/978-1-4939-9173-0_9
Haas, B. J. et al. Automated eukaryotic gene structure annotation using EVidenceModeler and the program to assemble spliced alignments. Genome Biol. 9, R7 (2008).
pubmed: 18190707 pmcid: 2395244 doi: 10.1186/gb-2008-9-1-r7
Blum, M. et al. The InterPro protein families and domains database: 20 years on. Nucleic Acids Res. 49, D344–D354 (2021).
pubmed: 33156333 doi: 10.1093/nar/gkaa977
Cantalapiedra, C. P., Hernández-Plaza, A., Letunic, I., Bork, P. & Huerta-Cepas, J. Eggnog-mapper v2: Functional annotation, orthology assignments, and domain prediction at the metagenomic scale. Mol. Biol. Evol. 38, 5825–5829 (2021).
pubmed: 34597405 pmcid: 8662613 doi: 10.1093/molbev/msab293
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR27885842 (2024).
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR27885834 (2024).
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR27885835 (2024).
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR27885836 (2024).
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR27885837 (2024).
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR27885838 (2024).
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR27885839 (2024).
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR27885840 (2024).
NCBI Sequence Read Archive https://identifiers.org/ncbi/insdc.sra:SRR27885841 (2024).
NCBI GenBank https://identifiers.org/ncbi/insdc:JAZDXJ000000000 (2024).
Wang, J. et al. Chromosome-level genome assembly of Euphorbia tirucalli (Euphorbiaceae), a highly stress-tolerant oil plant. Figshare https://doi.org/10.6084/m9.figshare.25224737 (2024).
Rhie, A., Walenz, B. P., Koren, S. & Phillippy, A. M. Merqury: reference-free quality, completeness, and phasing assessment for genome assemblies. Genome Biol. 21, 1–27 (2020).
doi: 10.1186/s13059-020-02134-9

Auteurs

Zuoying Wei (Z)

State Key Laboratory of Plant Diversity and Specialty Crops, South China Botanical Garden, Chinese Academy of Sciences, Guangzhou, 510650, China.
University of Chinese Academy of Sciences, Beijing, China.
South China National Botanical Garden, Chinese Academy of Sciences (CAS), Guangzhou, China.

Chao Feng (C)

State Key Laboratory of Plant Diversity and Specialty Crops, South China Botanical Garden, Chinese Academy of Sciences, Guangzhou, 510650, China.
South China National Botanical Garden, Chinese Academy of Sciences (CAS), Guangzhou, China.
Key Laboratory of National Forestry and Grassland Administration on Plant Conservation and Utilization in Southern China, Guangzhou, 510650, China.

Jiayun Xu (J)

State Key Laboratory of Plant Diversity and Specialty Crops, South China Botanical Garden, Chinese Academy of Sciences, Guangzhou, 510650, China.
University of Chinese Academy of Sciences, Beijing, China.
South China National Botanical Garden, Chinese Academy of Sciences (CAS), Guangzhou, China.

Xizuo Shi (X)

State Key Laboratory of Plant Diversity and Specialty Crops, South China Botanical Garden, Chinese Academy of Sciences, Guangzhou, 510650, China.
University of Chinese Academy of Sciences, Beijing, China.
South China National Botanical Garden, Chinese Academy of Sciences (CAS), Guangzhou, China.

Ming Kang (M)

State Key Laboratory of Plant Diversity and Specialty Crops, South China Botanical Garden, Chinese Academy of Sciences, Guangzhou, 510650, China.
South China National Botanical Garden, Chinese Academy of Sciences (CAS), Guangzhou, China.
Key Laboratory of National Forestry and Grassland Administration on Plant Conservation and Utilization in Southern China, Guangzhou, 510650, China.

Jing Wang (J)

State Key Laboratory of Plant Diversity and Specialty Crops, South China Botanical Garden, Chinese Academy of Sciences, Guangzhou, 510650, China. wjing@scbg.ac.cn.
South China National Botanical Garden, Chinese Academy of Sciences (CAS), Guangzhou, China. wjing@scbg.ac.cn.
Key Laboratory of National Forestry and Grassland Administration on Plant Conservation and Utilization in Southern China, Guangzhou, 510650, China. wjing@scbg.ac.cn.

Articles similaires

Drought Resistance Gene Expression Profiling Gene Expression Regulation, Plant Gossypium Multigene Family
Fragaria Light Plant Leaves Osmosis Stress, Physiological
Genome Size Genome, Plant Magnoliopsida Evolution, Molecular Arabidopsis
Citrus Phenylalanine Ammonia-Lyase Stress, Physiological Multigene Family Phylogeny

Classifications MeSH