Biological factors in the synthetic construction of overlapping genes.


Journal

BMC genomics
ISSN: 1471-2164
Titre abrégé: BMC Genomics
Pays: England
ID NLM: 100965258

Informations de publication

Date de publication:
11 Dec 2021
Historique:
received: 08 11 2020
accepted: 17 11 2021
entrez: 13 12 2021
pubmed: 14 12 2021
medline: 15 12 2021
Statut: epublish

Résumé

Overlapping genes (OLGs) with long protein-coding overlapping sequences are disallowed by standard genome annotation programs, outside of viruses. Recently however they have been discovered in Archaea, diverse Bacteria, and Mammals. The biological factors underlying life's ability to create overlapping genes require more study, and may have important applications in understanding evolution and in biotechnology. A previous study claimed that protein domains from viruses were much better suited to forming overlaps than those from other cellular organisms - in this study we assessed this claim, in order to discover what might underlie taxonomic differences in the creation of gene overlaps. After overlapping arbitrary Pfam domain pairs and evaluating them with Hidden Markov Models we find OLG construction to be much less constrained than expected. For instance, close to 10% of the constructed sequences cannot be distinguished from typical sequences in their protein family. Most are also indistinguishable from natural protein sequences regarding identity and secondary structure. Surprisingly, contrary to a previous study, virus domains were much less suitable for designing OLGs than bacterial or eukaryotic domains were. In general, the amount of amino acid change required to force a domain to overlap is approximately equal to the variation observed within a typical domain family. The resulting high similarity between natural sequences and those altered so as to overlap is mostly due to the combination of high redundancy in the genetic code and the evolutionary exchangeability of many amino acids. Synthetic overlapping genes which closely resemble natural gene sequences, as measured by HMM profiles, are remarkably easy to construct, and most arbitrary domain pairs can be altered so as to overlap while retaining high similarity to the original sequences. Future work however will need to assess important factors not considered such as intragenic interactions which affect protein folding. While the analysis here is not sufficient to guarantee functional folding proteins, further analysis of constructed OLGs will improve our understanding of the origin of these remarkable genetic elements across life and opens up exciting possibilities for synthetic biology.

Sections du résumé

BACKGROUND BACKGROUND
Overlapping genes (OLGs) with long protein-coding overlapping sequences are disallowed by standard genome annotation programs, outside of viruses. Recently however they have been discovered in Archaea, diverse Bacteria, and Mammals. The biological factors underlying life's ability to create overlapping genes require more study, and may have important applications in understanding evolution and in biotechnology. A previous study claimed that protein domains from viruses were much better suited to forming overlaps than those from other cellular organisms - in this study we assessed this claim, in order to discover what might underlie taxonomic differences in the creation of gene overlaps.
RESULTS RESULTS
After overlapping arbitrary Pfam domain pairs and evaluating them with Hidden Markov Models we find OLG construction to be much less constrained than expected. For instance, close to 10% of the constructed sequences cannot be distinguished from typical sequences in their protein family. Most are also indistinguishable from natural protein sequences regarding identity and secondary structure. Surprisingly, contrary to a previous study, virus domains were much less suitable for designing OLGs than bacterial or eukaryotic domains were. In general, the amount of amino acid change required to force a domain to overlap is approximately equal to the variation observed within a typical domain family. The resulting high similarity between natural sequences and those altered so as to overlap is mostly due to the combination of high redundancy in the genetic code and the evolutionary exchangeability of many amino acids.
CONCLUSIONS CONCLUSIONS
Synthetic overlapping genes which closely resemble natural gene sequences, as measured by HMM profiles, are remarkably easy to construct, and most arbitrary domain pairs can be altered so as to overlap while retaining high similarity to the original sequences. Future work however will need to assess important factors not considered such as intragenic interactions which affect protein folding. While the analysis here is not sufficient to guarantee functional folding proteins, further analysis of constructed OLGs will improve our understanding of the origin of these remarkable genetic elements across life and opens up exciting possibilities for synthetic biology.

Identifiants

pubmed: 34895142
doi: 10.1186/s12864-021-08181-1
pii: 10.1186/s12864-021-08181-1
pmc: PMC8665328
doi:

Substances chimiques

Biological Factors 0

Types de publication

Journal Article

Langues

eng

Sous-ensembles de citation

IM

Pagination

888

Informations de copyright

© 2021. The Author(s).

Références

Proc Natl Acad Sci U S A. 1992 Nov 15;89(22):10915-9
pubmed: 1438297
Mol Cell. 2007 Mar 23;25(6):851-62
pubmed: 17386262
Angew Chem Int Ed Engl. 2018 May 14;57(20):5674-5678
pubmed: 29512300
Nat Nanotechnol. 2018 Apr;13(4):309-315
pubmed: 29133926
J Biol Chem. 2016 Nov 4;291(45):23830-23831
pubmed: 27815453
Nucleic Acids Res. 2019 Jul 2;47(W1):W632-W635
pubmed: 31114895
BMC Genet. 2020 Mar 6;21(1):25
pubmed: 32138667
Genetics. 2018 Sep;210(1):303-313
pubmed: 30026186
Proteins. 2019 Dec;87(12):1011-1020
pubmed: 31589781
Sci Rep. 2017 Nov 20;7(1):15873
pubmed: 29158504
J Biol Chem. 1994 Feb 11;269(6):4523-31
pubmed: 8308022
Genome Res. 2007 Oct;17(10):1496-504
pubmed: 17785537
Mol Biol Evol. 2021 Sep 27;38(10):4301-4309
pubmed: 34043802
Mol Biol Evol. 2013 Apr;30(4):772-80
pubmed: 23329690
Biosystems. 2019 Nov;185:104023
pubmed: 31520875
Proc Natl Acad Sci U S A. 2016 Oct 11;113(41):11537-11542
pubmed: 27681623
Elife. 2020 Oct 01;9:
pubmed: 33001029
Nature. 1976 Nov 4;264(5581):34-41
pubmed: 1004533
Proc Natl Acad Sci U S A. 2020 Mar 17;117(11):5907-5912
pubmed: 32127487
J Mol Evol. 1998 Sep;47(3):238-48
pubmed: 9732450
Trends Biotechnol. 1990 Jun;8(6):140-4
pubmed: 1369993
Proc Natl Acad Sci U S A. 2019 Jan 29;116(5):1733-1738
pubmed: 30635413
Virus Evol. 2020 Feb 13;6(1):veaa009
pubmed: 32071766
Proteins. 1997 Jul;28(3):405-20
pubmed: 9223186
Curr Opin Struct Biol. 2021 Jun;68:142-148
pubmed: 33529785
Cell. 2016 Dec 15;167(7):1762-1773.e12
pubmed: 27984726
Orig Life Evol Biosph. 1995 Dec;25(6):565-89
pubmed: 7494636
Proteins. 2021 Dec;89(12):1607-1617
pubmed: 34533838
J Theor Biol. 2017 Feb 21;415:90-101
pubmed: 27737786
Biol Direct. 2016 May 21;11:26
pubmed: 27209091
Front Microbiol. 2018 May 14;9:931
pubmed: 29867840
Protein Eng. 1999 Feb;12(2):85-94
pubmed: 10195279
Nat Rev Genet. 2012 Jun 12;13(7):455-68
pubmed: 22688678
PLoS Comput Biol. 2021 Oct 8;17(10):e1009475
pubmed: 34624014
Nature. 2021 Aug;596(7873):583-589
pubmed: 34265844
Biosystems. 2018 Feb;164:199-216
pubmed: 29107641
Science. 1977 Jun 10;196(4295):1187-8
pubmed: 17787080
Proc Natl Acad Sci U S A. 2020 Oct 6;117(40):24936-24946
pubmed: 32958672
Science. 2019 Aug 9;365(6453):595-598
pubmed: 31395784
PLoS Comput Biol. 2013;9(8):e1003162
pubmed: 23966842
Proc Natl Acad Sci U S A. 1992 Oct 15;89(20):9489-93
pubmed: 1329098
Nucleic Acids Res. 2020 Jun 4;48(10):5201-5216
pubmed: 32382758
J Mol Evol. 2008 Nov;67(5):510-6
pubmed: 18855039
Nucleic Acids Res. 2017 Aug 21;45(14):8484-8492
pubmed: 28582582
Proc Natl Acad Sci U S A. 1951 Apr;37(4):205-11
pubmed: 14816373
Phys Rev E Stat Nonlin Soft Matter Phys. 2009 Jun;79(6 Pt 1):060901
pubmed: 19658466
J Theor Biol. 2017 Apr 21;419:266-268
pubmed: 28167103
Sci Rep. 2019 Aug 26;9(1):12374
pubmed: 31451723
Bioinformatics. 1998;14(9):755-63
pubmed: 9918945
Mol Biol Evol. 2009 Feb;26(2):445-50
pubmed: 19037009
Nucleic Acids Res. 2011 Jan;39(Database issue):D411-9
pubmed: 21071423
Biopolymers. 1983 Dec;22(12):2577-637
pubmed: 6667333
Biophys J. 2017 Oct 17;113(8):1719-1730
pubmed: 29045866
Front Mol Biosci. 2020 Aug 14;7:187
pubmed: 32923454
J Biol Chem. 1994 Feb 11;269(6):4513-22
pubmed: 8308021
Mol Biol Evol. 2012 Dec;29(12):3767-80
pubmed: 22821011
Nature. 1978 Apr 6;272(5653):532-5
pubmed: 692657
J Theor Biol. 1979 Sep 7;80(1):21-6
pubmed: 94642
Biol Direct. 2014 Jun 14;9:11
pubmed: 24927791
Nucleic Acids Res. 2014 Jul;42(Web Server issue):W337-43
pubmed: 24799431
Microbiol Spectr. 2018 Jul;6(4):
pubmed: 30003865
Virus Evol. 2020 Feb 10;6(1):veaa007
pubmed: 32064120
Trends Biochem Sci. 1990 Jul;15(7):257-61
pubmed: 2200170
Sci Rep. 2018 Dec 14;8(1):17875
pubmed: 30552341
BMC Bioinformatics. 2010 Mar 15;11:131
pubmed: 20230630
Front Microbiol. 2020 Mar 20;11:377
pubmed: 32265854
Nucleic Acids Res. 2004 Mar 19;32(5):1792-7
pubmed: 15034147

Auteurs

Stefan Wichmann (S)

Chair of Microbial Ecology, Department of Molecular Life Sciences, Technical University of Munich, Freising, Germany.

Siegfried Scherer (S)

Chair of Microbial Ecology, Department of Molecular Life Sciences, Technical University of Munich, Freising, Germany.

Zachary Ardern (Z)

Chair of Microbial Ecology, Department of Molecular Life Sciences, Technical University of Munich, Freising, Germany. zachary.ardern@sanger.ac.uk.
Wellcome Sanger Institute, Wellcome Genome Campus, Hinxton, UK. zachary.ardern@sanger.ac.uk.

Articles similaires

Robotic Surgical Procedures Animals Humans Telemedicine Models, Animal

Odour generalisation and detection dog training.

Lyn Caldicott, Thomas W Pike, Helen E Zulch et al.
1.00
Animals Odorants Dogs Generalization, Psychological Smell
Animals TOR Serine-Threonine Kinases Colorectal Neoplasms Colitis Mice
Animals Tail Swine Behavior, Animal Animal Husbandry

Classifications MeSH