A survey on protein-DNA-binding sites in computational biology.

DNA–protein-binding sites bioinformatics convolutional neural network deep learning machine learning recurrent neural networks transcription factor binding site

Journal

Briefings in functional genomics
ISSN: 2041-2657
Titre abrégé: Brief Funct Genomics
Pays: England
ID NLM: 101528229

Informations de publication

Date de publication:
16 09 2022
Historique:
received: 02 02 2022
revised: 07 04 2022
accepted: 22 04 2022
pubmed: 3 6 2022
medline: 21 9 2022
entrez: 2 6 2022
Statut: ppublish

Résumé

Transcription factors are important cellular components of the process of gene expression control. Transcription factor binding sites are locations where transcription factors specifically recognize DNA sequences, targeting gene-specific regions and recruiting transcription factors or chromatin regulators to fine-tune spatiotemporal gene regulation. As the common proteins, transcription factors play a meaningful role in life-related activities. In the face of the increase in the protein sequence, it is urgent how to predict the structure and function of the protein effectively. At present, protein-DNA-binding site prediction methods are based on traditional machine learning algorithms and deep learning algorithms. In the early stage, we usually used the development method based on traditional machine learning algorithm to predict protein-DNA-binding sites. In recent years, methods based on deep learning to predict protein-DNA-binding sites from sequence data have achieved remarkable success. Various statistical and machine learning methods used to predict the function of DNA-binding proteins have been proposed and continuously improved. Existing deep learning methods for predicting protein-DNA-binding sites can be roughly divided into three categories: convolutional neural network (CNN), recursive neural network (RNN) and hybrid neural network based on CNN-RNN. The purpose of this review is to provide an overview of the computational and experimental methods applied in the field of protein-DNA-binding site prediction today. This paper introduces the methods of traditional machine learning and deep learning in protein-DNA-binding site prediction from the aspects of data processing characteristics of existing learning frameworks and differences between basic learning model frameworks. Our existing methods are relatively simple compared with natural language processing, computational vision, computer graphics and other fields. Therefore, the summary of existing protein-DNA-binding site prediction methods will help researchers better understand this field.

Identifiants

pubmed: 35652477
pii: 6596853
doi: 10.1093/bfgp/elac009
doi:

Substances chimiques

Chromatin 0
DNA-Binding Proteins 0
Transcription Factors 0
DNA 9007-49-2

Types de publication

Journal Article Review Research Support, Non-U.S. Gov't

Langues

eng

Sous-ensembles de citation

IM

Pagination

357-375

Informations de copyright

© The Author(s) 2022. Published by Oxford University Press. All rights reserved. For Permissions, please email: journals.permissions@oup.com.

Auteurs

Articles similaires

Selecting optimal software code descriptors-The case of Java.

Yegor Bugayenko, Zamira Kholmatova, Artem Kruglov et al.
1.00
Software Algorithms Programming Languages
1.00
Humans Magnetic Resonance Imaging Brain Infant, Newborn Infant, Premature
Adenosine Triphosphate Adenosine Diphosphate Mitochondrial ADP, ATP Translocases Binding Sites Mitochondria
Humans Colorectal Neoplasms Biomarkers, Tumor Prognosis Gene Expression Regulation, Neoplastic

Classifications MeSH