Home LiteratureArticle Details
PMID: 17646325 Published · ppublish English Journal Article Research Support, N.I.H., Extramural

Manual curation is not sufficient for annotation of genomic databases.

Bioinformatics (Oxford, England) ·Vol. 23 ·No. 13 ·2007-07-01 ·Pages i41-8

Baumgartner WA, Cohen KB, Fox LM, Acquaah-Mensah G, Hunter L

Abstract

Knowledge base construction has been an area of intense activity and great importance in the growth of computational biology. However, there is little or no history of work on the subject of evaluation of knowledge bases, either with respect to their contents or with respect to the processes by which they are constructed. This article proposes the application of a metric from software engineering known as the found/fixed graph to the problem of evaluating the processes by which genomic knowledge bases are built, as well as the completeness of their contents. Well-understood patterns of change in the found/fixed graph are found to occur in two large publicly available knowledge bases. These patterns suggest that the current manual curation processes will take far too long to complete the annotations of even just the most important model organisms, and that at their current rate of production, they will never be sufficient for completing the annotation of all currently available proteomes.

MeSH Terms
Chromosome Mapping/methods Databases, Protein Documentation/methods Genomics/methods Proteins/chemistry,genetics Sequence Analysis, Protein/methods
Chemicals
Proteins
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Baumgartner William A
Center for Computational Pharmacology, University of Colorado School of Medicine, USA.
Cohen K Bretonnel
Fox Lynne M
Acquaah-Mensah George
Hunter Lawrence
References (32)
32 references, click to expand
  1. Gene indexing: characterization and analysis of NLM's GeneRIFs.
    AMIA Annu Symp Proc. 2003;:460-4 PMID: 14728215
  2. Extraction of transcript diversity from scientific literature.
    PLoS Comput Biol. 2005 Jun;1(1):e10 PMID: 16103899
  3. xGDB: open-source computational infrastructure for the integrated evaluation and analysis of genome features.
    Genome Biol. 2006;7(11):R111 PMID: 17116260
  4. GeneRIF quality assurance as summary revision.
    Pac Symp Biocomput. 2007;:269-80 PMID: 17990498
  5. Entrez Gene: gene-centered information at NCBI.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D54-8 PMID: 15608257
  6. The database revolution.
    Nature. 2007 Jan 18;445(7125):229-30 PMID: 17230150
  7. Finding GeneRIFs via gene ontology annotations.
    Pac Symp Biocomput. 2006;:52-63 PMID: 17094227
  8. ASAP, a systematic annotation package for community analysis of genomes.
    Nucleic Acids Res. 2003 Jan 1;31(1):147-51 PMID: 12519969
  9. Complete genome sequence of Pseudomonas aeruginosa PAO1, an opportunistic pathogen.
    Nature. 2000 Aug 31;406(6799):959-64 PMID: 10984043
  10. Creating the gene ontology resource: design and implementation.
    Genome Res. 2001 Aug;11(8):1425-33 PMID: 11483584
  11. A biocurator perspective: annotation at the Research Collaboratory for Structural Bioinformatics Protein Data Bank.
    PLoS Comput Biol. 2006 Oct 27;2(10):e99 PMID: 17069453
  12. The Gene Ontology Annotation (GOA) Database: sharing knowledge in Uniprot with Gene Ontology.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D262-6 PMID: 14681408
  13. Sequencing solution: use volunteer annotators organized via Internet.
    Nature. 2000 Aug 31;406(6799):933 PMID: 10984027
  14. Genome re-annotation: a wiki solution?
    Genome Biol. 2007;8(1):102 PMID: 17274839
  15. Automated extraction of mutation data from the literature: application of MuteXt to G protein-coupled receptors and nuclear hormone receptors.
    Bioinformatics. 2004 Mar 1;20(4):557-68 PMID: 14990452
  16. Investigating semantic similarity measures across the Gene Ontology: the relationship between sequence and annotation.
    Bioinformatics. 2003 Jul 1;19(10):1275-83 PMID: 12835272
  17. MILANO--custom annotation of microarray results using automatic literature searches.
    BMC Bioinformatics. 2005;6:12 PMID: 15661078
  18. GO PaD: the Gene Ontology Partition Database.
    Nucleic Acids Res. 2007 Jan;35(Database issue):D322-7 PMID: 17098937
  19. yrGATE: a web-based gene-structure annotation tool for the identification and dissemination of eukaryotic genes.
    Genome Biol. 2006;7(7):R58 PMID: 16859520
  20. Publishing perishing? Towards tomorrow's information architecture.
    BMC Bioinformatics. 2007;8:17 PMID: 17239245
  21. Mistakes in medical ontologies: where do they come from and how can they be detected?
    Stud Health Technol Inform. 2004;102:145-63 PMID: 15853269
  22. PharmGKB: the Pharmacogenetics Knowledge Base.
    Nucleic Acids Res. 2002 Jan 1;30(1):163-5 PMID: 11752281
  23. RIBOWEB: linking structural computations to a knowledge base of published experimental data.
    Proc Int Conf Intell Syst Mol Biol. 1997;5:84-7 PMID: 9322019
  24. Semantic similarity measures as tools for exploring the gene ontology.
    Pac Symp Biocomput. 2003;:601-12 PMID: 12603061
  25. Building large knowledge bases in molecular biology.
    Proc Int Conf Intell Syst Mol Biol. 1993;1:345-53 PMID: 7584356
  26. Consistency across the hierarchies of the UMLS Semantic Network and Metathesaurus.
    J Biomed Inform. 2003 Dec;36(6):450-61 PMID: 14759818
  27. Key biology databases go wiki.
    Nature. 2007 Feb 15;445(7129):691 PMID: 17301755
  28. Quality control for terms and definitions in ontologies and taxonomies.
    BMC Bioinformatics. 2006;7:212 PMID: 16623942
  29. Gene-function wiki would let biologists pool worldwide resources.
    Nature. 2006 Feb 2;439(7076):534 PMID: 16452957
  30. Evaluation of long-term maintenance of a large medical knowledge base.
    J Am Med Inform Assoc. 1995 Sep-Oct;2(5):297-306 PMID: 7496879
  31. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003.
    Nucleic Acids Res. 2003 Jan 1;31(1):365-70 PMID: 12520024
  32. Community-based gene structure annotation.
    Trends Plant Sci. 2005 Jan;10(1):9-14 PMID: 15642518
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2007-07-01
Pages
i41-8
Language
English
Region
England
NLM ID
9808944
PMCID
PMC2516305
Subset
IM
Grants
NLM NIH HHS · R01 LM008111 · United States
NLM NIH HHS · R01 LM008111-03 · United States
NLM NIH HHS · R01-LM008111 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]