Home LiteratureArticle Details
PMID: 23424147 Published · ppublish English Journal Article

Detection of protein catalytic sites in the biomedical literature.

Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing ·2013-00-00 ·Pages 433-44

Verspoor K, Mackinlay A, Cohn JD, Wall ME

Abstract

This paper explores the application of text mining to the problem of detecting protein functional sites in the biomedical literature, and specifically considers the task of identifying catalytic sites in that literature. We provide strong evidence for the need for text mining techniques that address residue-level protein function annotation through an analysis of two corpora in terms of their coverage of curated data sources. We also explore the viability of building a text-based classifier for identifying protein functional sites, identifying the low coverage of curated data sources and the potential ambiguity of information about protein functional sites as challenges that must be addressed. Nevertheless we produce a simple classifier that achieves a reasonable ∼69% F-score on our full text silver corpus on the first attempt to address this classification task. The work has application in computational prediction of the functional significance of protein sites as well as in curation workflows for databases that capture this information.

MeSH Terms
Amino Acids/chemistry Artificial Intelligence Binding Sites Catalytic Domain Computational Biology Data Mining/statistics & numerical data Databases, Protein/statistics & numerical data Ligands Natural Language Processing Proteins/chemistry,classification,metabolism
Chemicals
Amino Acids Ligands Proteins
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Verspoor Karin
National ICT Australia, Victoria Research Lab, Parkville, VIC 3010, Australia. [email protected]
Mackinlay Andrew
Cohn Judith D
Wall Michael E
References (13)
13 references, click to expand
  1. The structural and content aspects of abstracts versus bodies of full text journal articles are different.
    BMC Bioinformatics. 2010 Sep 29;11:492 PMID: 20920264
  2. Automated extraction and semantic analysis of mutation impacts from the biomedical literature.
    BMC Genomics. 2012 Jun 18;13 Suppl 4:S10 PMID: 22759648
  3. A reappraisal of sentence and token splitting for life sciences documents.
    Stud Health Technol Inform. 2007;129(Pt 1):524-8 PMID: 17911772
  4. The Catalytic Site Atlas: a resource of catalytic sites and residues identified in enzymes using structural data.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D129-33 PMID: 14681376
  5. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  6. BioLemmatizer: a lemmatization tool for morphological processing of biomedical text.
    J Biomed Semantics. 2012 Apr 01;3:3 PMID: 22464129
  7. Text mining improves prediction of protein functional sites.
    PLoS One. 2012;7(2):e32171 PMID: 22393388
  8. Manual curation is not sufficient for annotation of genomic databases.
    Bioinformatics. 2007 Jul 1;23(13):i41-8 PMID: 17646325
  9. Literature mining of protein-residue associations with graph rules learned through distant supervision.
    J Biomed Semantics. 2012 Oct 5;3 Suppl 3:S2 PMID: 23046792
  10. The Protein Data Bank.
    Nucleic Acids Res. 2000 Jan 1;28(1):235-42 PMID: 10592235
  11. Binding MOAD, a high-quality protein-ligand database.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D674-8 PMID: 18055497
  12. Annotation of protein residues based on a literature analysis: cross-validation against UniProtKb.
    BMC Bioinformatics. 2009 Aug 27;10 Suppl 8:S4 PMID: 19758468
  13. Beyond annotation transfer by homology: novel protein-function prediction methods to assist drug discovery.
    Drug Discov Today. 2005 Nov 1;10(21):1475-82 PMID: 16243268
Article Info
Journal
Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing
Abbr.
Pac Symp Biocomput
ISSN
2335-6936
Published
2013-00-00
Pages
433-44
Language
English
Region
United States
NLM ID
9711271
PMCID
PMC3664919
Subset
IM
Grants
NLM NIH HHS · R01 LM010120 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]