Automatic consistency assurance for literature-based gene ontology annotation.
Biological database quality
Gene ontology annotation
Text mining
Journal
BMC bioinformatics
ISSN: 1471-2105
Titre abrégé: BMC Bioinformatics
Pays: England
ID NLM: 100965194
Informations de publication
Date de publication:
25 Nov 2021
25 Nov 2021
Historique:
received:
12
04
2021
accepted:
15
11
2021
entrez:
26
11
2021
pubmed:
27
11
2021
medline:
30
11
2021
Statut:
epublish
Résumé
Literature-based gene ontology (GO) annotation is a process where expert curators use uniform expressions to describe gene functions reported in research papers, creating computable representations of information about biological systems. Manual assurance of consistency between GO annotations and the associated evidence texts identified by expert curators is reliable but time-consuming, and is infeasible in the context of rapidly growing biological literature. A key challenge is maintaining consistency of existing GO annotations as new studies are published and the GO vocabulary is updated. In this work, we introduce a formalisation of biological database annotation inconsistencies, identifying four distinct types of inconsistency. We propose a novel and efficient method using state-of-the-art text mining models to automatically distinguish between consistent GO annotation and the different types of inconsistent GO annotation. We evaluate this method using a synthetic dataset generated by directed manipulation of instances in an existing corpus, BC4GO. We provide detailed error analysis for demonstrating that the method achieves high precision on more confident predictions. Two models built using our method for distinct annotation consistency identification tasks achieved high precision and were robust to updates in the GO vocabulary. Our approach demonstrates clear value for human-in-the-loop curation scenarios.
Sections du résumé
BACKGROUND
BACKGROUND
Literature-based gene ontology (GO) annotation is a process where expert curators use uniform expressions to describe gene functions reported in research papers, creating computable representations of information about biological systems. Manual assurance of consistency between GO annotations and the associated evidence texts identified by expert curators is reliable but time-consuming, and is infeasible in the context of rapidly growing biological literature. A key challenge is maintaining consistency of existing GO annotations as new studies are published and the GO vocabulary is updated.
RESULTS
RESULTS
In this work, we introduce a formalisation of biological database annotation inconsistencies, identifying four distinct types of inconsistency. We propose a novel and efficient method using state-of-the-art text mining models to automatically distinguish between consistent GO annotation and the different types of inconsistent GO annotation. We evaluate this method using a synthetic dataset generated by directed manipulation of instances in an existing corpus, BC4GO. We provide detailed error analysis for demonstrating that the method achieves high precision on more confident predictions.
CONCLUSIONS
CONCLUSIONS
Two models built using our method for distinct annotation consistency identification tasks achieved high precision and were robust to updates in the GO vocabulary. Our approach demonstrates clear value for human-in-the-loop curation scenarios.
Identifiants
pubmed: 34823464
doi: 10.1186/s12859-021-04479-9
pii: 10.1186/s12859-021-04479-9
pmc: PMC8620237
doi:
Types de publication
Journal Article
Langues
eng
Sous-ensembles de citation
IM
Pagination
565Subventions
Organisme : Australian Research Council
ID : DP190101350
Informations de copyright
© 2021. The Author(s).
Références
BMC Bioinformatics. 2007 May 22;8:170
pubmed: 17519041
Methods Mol Biol. 2017;1446:97-109
pubmed: 27812938
Genome Biol. 2019 Nov 19;20(1):244
pubmed: 31744546
Methods Mol Biol. 2017;1446:189-205
pubmed: 27812944
J Biomed Inform. 2013 Oct;46(5):914-20
pubmed: 23906817
Methods Mol Biol. 2017;1446:161-173
pubmed: 27812942
Brief Bioinform. 2011 Nov;12(6):723-35
pubmed: 21330331
Database (Oxford). 2014 Aug 25;2014:
pubmed: 25157073
Sci Data. 2016 May 24;3:160035
pubmed: 27219127
Bioinformatics. 2002 Dec;18(12):1641-9
pubmed: 12490449
Methods Mol Biol. 2017;1446:55-67
pubmed: 27812935
PLoS Comput Biol. 2013;9(5):e1003063
pubmed: 23737737
Methods Mol Biol. 2017;1446:15-24
pubmed: 27812932
BMC Bioinformatics. 2012 Jul 09;13:161
pubmed: 22776079
Sci Rep. 2018 Jan 22;8(1):1362
pubmed: 29358745
PLoS Comput Biol. 2012 May;8(5):e1002533
pubmed: 22693439
Nat Genet. 2000 May;25(1):25-9
pubmed: 10802651
Nucleic Acids Res. 2015 Jan;43(Database issue):D1049-56
pubmed: 25428369
BMC Med Inform Decis Mak. 2018 Jun 25;18(1):46
pubmed: 29940927
Nat Genet. 2004 May;36(5):431-2
pubmed: 15118671
Database (Oxford). 2014 Jul 28;2014:
pubmed: 25070993
Bioinformatics. 2009 Nov 15;25(22):3045-6
pubmed: 19744993
Database (Oxford). 2013 Jul 09;2013:bat054
pubmed: 23842463
Nucleic Acids Res. 2017 Jan 4;45(D1):D331-D338
pubmed: 27899567
Methods Mol Biol. 2017;1446:69-84
pubmed: 27812936
Methods Mol Biol. 2017;1446:41-54
pubmed: 27812934
BMC Bioinformatics. 2014 Feb 26;15:59
pubmed: 24571547
Bioinformatics. 2017 Jul 15;33(14):i49-i58
pubmed: 28881973
Database (Oxford). 2013 Jul 09;2013:bat041
pubmed: 23842461