Home LiteratureArticle Details
PMID: 12424124 Published · ppublish English Evaluation Study Journal Article Validation Study

Finding relevant references to genes and proteins in Medline using a Bayesian approach.

Bioinformatics (Oxford, England) ·Vol. 18 ·No. 11 ·2002-11-00 ·Pages 1515-22

Leonard JE, Colombe JB, Levy JL

Abstract

Mining the biomedical literature for references to genes and proteins always involves a tradeoff between high precision with false negatives, and high recall with false positives. Having a reliable method for assessing the relevance of literature mining results is crucial to finding ways to balance precision and recall, and for subsequently building automated systems to analyze these results. We hypothesize that abstracts and titles that discuss the same gene or protein use similar words. To validate this hypothesis, we built a dictionary- and rule-based system to mine Medline for references to genes and proteins, and used a Bayesian metric for scoring the relevance of each reference assignment. We analyzed the entire set of Medline records from 1966 to late 2001, and scored each gene and protein reference using a Bayesian estimated probability (EP) based on word frequency in a training set of 137837 known assignments from 30594 articles to 36197 gene and protein symbols. Two test sets of 148 and 150 randomly chosen assignments, respectively, were hand-validated and categorized as either good or bad. The distributions of EP values, when plotted on a log-scale histogram, are shown to markedly differ between good and bad assignments. Using EP values, recall was 100% at 61% precision (EP=2 x 10(-5)), 63% at 88% precision (EP=0.008), and 10% at 100% precision (EP=0.1). These results show that Medline entries discussing the same gene or protein have similar word usage, and that our method of assessing this similarity using EP values is valid, and enables an EP cutoff value to be determined that accurately and reproducibly balances precision and recall, allowing automated analysis of literature mining results. .

MeSH Terms
Abstracting and Indexing/methods Algorithms Bayes Theorem Database Management Systems Dictionaries as Topic False Negative Reactions False Positive Reactions Genes Information Storage and Retrieval/methods MEDLINE Models, Statistical National Library of Medicine (U.S.) Natural Language Processing Pattern Recognition, Automated Proteins Subject Headings United States
Chemicals
Proteins
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Leonard Julie E
Incellico Inc, 2327 Englert Dr, Durham, NC 27713, USA. [email protected]
Colombe Jeffrey B
Levy Joshua L
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4803
Published
2002-11-00
Pages
1515-22
Language
English
Region
England
NLM ID
9808944
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]