Creating De Novo Overlapped Genes.

Deep learning Generative model Genome compression Machine learning Markov random field Multiple sequence alignments Overlapping genes Protein design Synthetic biology Synthetic genomes

Journal

Methods in molecular biology (Clifton, N.J.)
ISSN: 1940-6029
Titre abrégé: Methods Mol Biol
Pays: United States
ID NLM: 9214969

Informations de publication

Date de publication:
2023
Historique:
entrez: 13 10 2022
pubmed: 14 10 2022
medline: 18 10 2022
Statut: ppublish

Résumé

Future applications of synthetic biology will rely on deploying engineered cells outside of lab environments for long periods of time. Currently, a significant roadblock to this application is the potential for deactivating mutations in engineered genes. A recently developed method to protect engineered coding sequences from mutation is called Constraining Adaptive Mutations using Engineered Overlapping Sequences (CAMEOS). In this chapter we provide a workflow for utilizing CAMEOS to create synthetic overlaps between two genes, one essential (infA) and one non-essential (aroB), to protect the non-essential gene from mutation and loss of protein function. In this workflow we detail the methods to collect large numbers of related protein sequences, produce multiple sequence alignments (MSAs), use the MSAs to generate hidden Markov models and Markov random field models, and finally generate a library of overlapping coding sequences through CAMEOS scripts. To assist practitioners with basic coding skills to try out the CAMEOS method, we have created a virtual machine containing all the required packages already installed that can be downloaded and run locally.

Identifiants

pubmed: 36227541
doi: 10.1007/978-1-0716-2617-7_6
doi:

Substances chimiques

Proteins 0

Types de publication

Journal Article Research Support, Non-U.S. Gov't

Langues

eng

Sous-ensembles de citation

IM

Pagination

95-120

Informations de copyright

© 2023. The Author(s), under exclusive license to Springer Science+Business Media, LLC, part of Springer Nature.

Références

Tang T-C et al (2021) Materials design by synthetic biology. Nat Rev Mater 6(4):332–350
doi: 10.1038/s41578-020-00265-w
Brophy JAN, Voigt CA (2014) Principles of genetic circuit design. Nat Methods 11(5):508–520
doi: 10.1038/nmeth.2926
Chan LY, Kosuri S, Endy D (2005) Refactoring bacteriophage T7. Mol Syst Biol 1(1):2005.0018
doi: 10.1038/msb4100025
Jaschke PR et al (2012) A fully decompressed synthetic bacteriophage øX174 genome assembled and archived in yeast. Virology 434(2):278–284
doi: 10.1016/j.virol.2012.09.020
Logel DY, Trofimova E, Jaschke PR (2022) Codon-restrained method for both eliminating and creating intragenic bacterial promoters. ACS Synth Biol 11(2):689–699
doi: 10.1021/acssynbio.1c00359
Temme K, Zhao D, Voigt CA (2012) Refactoring the nitrogen fixation gene cluster from Klebsiella oxytoca. Proc Natl Acad Sci 109(18):7085–7090
doi: 10.1073/pnas.1120788109
Wright BW, Molloy MP, Jaschke PR (2022) Overlapping genes in natural and engineered genomes. Nat Rev Genet 23(3):154–168
doi: 10.1038/s41576-021-00417-w
Song M et al (2017) Control of type III protein secretion using a minimal genetic system. Nat Commun 8:14737
doi: 10.1038/ncomms14737
Springman R et al (2012) Evolutionary stability of a refactored phage genome. ACS Synth Biol 1(9):425–430
doi: 10.1021/sb300040v
Borkowski O et al (2016) Overloaded and stressed: whole-cell considerations for bacterial synthetic biology. Curr Opin Microbiol 33:123–130
doi: 10.1016/j.mib.2016.07.009
Ellis T (2019) Predicting how evolution will beat us. Microb Biotechnol 12(1):41
doi: 10.1111/1751-7915.13327
Gallup O, Ming H, Ellis T (2021) Ten future challenges for synthetic biology. Eng Biol 5(3):51–59
doi: 10.1049/enb2.12011
Clark M, Maselko M (2020) Transgene biocontainment strategies for molecular farming. Front Plant Sci 11:1–11
Maselko M et al (2017) Engineering species-like barriers to sexual reproduction. Nat Commun 8(1):883
doi: 10.1038/s41467-017-01007-3
Ryu M-H et al (2020) Control of nitrogen fixation in bacteria that associate with cereals. Nat Microbiol 5(2):314–330
doi: 10.1038/s41564-019-0631-2
Rylott EL, Bruce NC (2020) How synthetic biology can help bioremediation. Curr Opin Chem Biol 58:86–95
doi: 10.1016/j.cbpa.2020.07.004
Voigt CA (2020) Synthetic biology 2020–2030: six commercially-available products that are changing our world. Nat Commun 11(1):6379
doi: 10.1038/s41467-020-20122-2
Jaschke PR et al (2019) Definitive demonstration by synthesis of genome annotation completeness. Proc Natl Acad Sci 116(48):24206–24213
doi: 10.1073/pnas.1905990116
Wright BW et al (2020) Genome modularization reveals overlapped gene topology is necessary for efficient viral reproduction. ACS Synth Biol 9(11):3079–3090
doi: 10.1021/acssynbio.0c00323
Blazejewski T, Ho H-I, Wang HH (2019) Synthetic sequence entanglement augments stability and containment of genetic information in cells. Science 365(6453):595–598
doi: 10.1126/science.aav5477
Chowdhury B, Garai G (2017) A review on multiple sequence alignment from the perspective of genetic algorithm. Genomics 109(5):419–431
doi: 10.1016/j.ygeno.2017.06.007
Thompson JD, Higgins DG, Gibson TJ (1994) CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice. Nucleic Acids Res 22(22):4673–4680
doi: 10.1093/nar/22.22.4673
Deorowicz S, Debudaj-Grabysz A, Gudyś A (2016) FAMSA: fast and accurate multiple sequence alignment of huge protein families. Sci Rep 6(1):33964
doi: 10.1038/srep33964
Katoh K et al (2002) MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform. Nucleic Acids Res 30(14):3059–3066
doi: 10.1093/nar/gkf436
Eddy SR (2004) What is a hidden Markov model? Nat Biotechnol 22(10):1315–1316
doi: 10.1038/nbt1004-1315
Thomas J, Ramakrishnan N, Bailey-Kellogg C (2008) Graphical models of residue coupling in protein families. IEEE/ACM Trans Comput Biol Bioinform 5(2):183–197
doi: 10.1109/TCBB.2007.70225
Steinegger M et al (2019) HH-suite3 for fast remote homology detection and deep protein annotation. BMC Bioinformatics 20(1):473
doi: 10.1186/s12859-019-3019-7
Seemayer S, Gruber M, Söding J (2014) CCMpred--fast and precise prediction of protein residue-residue contacts from correlated mutations. Bioinformatics 30(21):3128–3130
doi: 10.1093/bioinformatics/btu500
Finn RD, Clements J, Eddy SR (2011) HMMER web server: interactive sequence similarity searching. Nucleic Acids Res 39(Web Server issue):W29–W37
doi: 10.1093/nar/gkr367
Finn RD et al (2014) Pfam: the protein families database. Nucleic Acids Res 42(D1):D222–D230
doi: 10.1093/nar/gkt1223
Blum M et al (2021) The InterPro protein families and domains database: 20 years on. Nucleic Acids Res 49(D1):D344–D354
doi: 10.1093/nar/gkaa977
Bhagwat M, Aravind L (2007) PSI-BLAST tutorial. In: Bergman NH (ed) Comparative genomics. Humana Press, Totowa
Jehl P, Sievers F, Higgins DG (2015) OD-seq: outlier detection in multiple sequence alignments. BMC Bioinformatics 16(1):1–11
doi: 10.1186/s12859-015-0702-1
Bailey TL et al (2015) The MEME suite. Nucleic Acids Res 43(W1):W39–W49
doi: 10.1093/nar/gkv416
Continuum A (2015) Anaconda software distribution computer software vers 2-2.4. 0.
Bezanson J et al (2012) Julia: A fast dynamic language for technical computing. arXiv preprint arXiv:1209.5145
Jumper J et al (2021) Highly accurate protein structure prediction with AlphaFold. Nature 596(7873):583–589
doi: 10.1038/s41586-021-03819-2
Varadi M et al (2022) AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Res 50(D1):D439–D444
doi: 10.1093/nar/gkab1061
Sette M et al (1997) The structure of the translational initiation factor IF1 from E. coli contains an oligomer-binding motif. EMBO J 16(6):1436–1443
doi: 10.1093/emboj/16.6.1436
Reis AC, Salis HM (2020) An automated model test system for systematic development and improvement of gene expression models. ACS Synth Biol 9(11):3145–3156
doi: 10.1021/acssynbio.0c00394
Keseler IM et al (2017) The EcoCyc database: reflecting new knowledge about Escherichia coli K-12. Nucleic Acids Res 45(D1):D543–D550
doi: 10.1093/nar/gkw1003

Auteurs

Dominic Y Logel (DY)

School of Natural Sciences, Macquarie University, Sydney, NSW, Australia.
ARC Centre of Excellence in Synthetic Biology, Macquarie University, Sydney, NSW, Australia.

Paul R Jaschke (PR)

School of Natural Sciences, Macquarie University, Sydney, NSW, Australia. paul.jaschke@mq.edu.au.
ARC Centre of Excellence in Synthetic Biology, Macquarie University, Sydney, NSW, Australia. paul.jaschke@mq.edu.au.

Articles similaires

Databases, Protein Protein Domains Protein Folding Proteins Deep Learning
Animals Hemiptera Insect Proteins Phylogeny Insecticides
Cicer Germination Proteolysis Seeds Plant Proteins

Mutational analysis of Phanerochaete chrysosporium´s purine transporter.

Mariana Barraco-Vega, Manuel Sanguinetti, Gabriela da Rosa et al.
1.00
Phanerochaete Fungal Proteins Purines Aspergillus nidulans DNA Mutational Analysis

Classifications MeSH