Abstract
The analysis of large-scale genomic information (such as sequence data or expression patterns) frequently involves grouping genes on the basis of common experimental features. Often, as with gene expression clustering, there are too many groups to easily identify the functionally relevant ones. One valuable source of information about gene function is the published literature. We present a method, neighbor divergence, for assessing whether the genes within a group share a common biological function based on their associated scientific literature. The method uses statistical natural language processing techniques to interpret biological text. It requires only a corpus of documents relevant to the genes being studied (e.g., all genes in an organism) and an index connecting the documents to appropriate genes. Given a group of genes, neighbor divergence assigns a numerical score indicating how "functionally coherent" the gene group is from the perspective of the published literature. We evaluate our method by testing its ability to distinguish 19 known functional gene groups from 1900 randomly assembled groups. Neighbor divergence achieves 79% sensitivity at 100% specificity, comparing favorably to other tested methods. We also apply neighbor divergence to previously published gene expression clusters to assess its ability to recognize gene groups that had been manually identified as representative of a common function.
MeSH Terms
Algorithms
Artificial Intelligence
Cluster Analysis
Computational Biology/methods,trends
Databases, Genetic/statistics & numerical data
Discriminant Analysis
Gene Expression Profiling/methods,statistics & numerical data,trends
Genes, Fungal/physiology
Genome, Fungal
Information Services
Natural Language Processing
Research Design/statistics & numerical data,trends
Saccharomyces cerevisiae/genetics
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Raychaudhuri Soumya
Department of Genetics, Stanford Medical Informatics, University, Stanford, California 94305-5479, USA.
Schütze Hinrich
Altman Russ B
References (30)
30 references, click to expand
-
Automatic annotation for biological sequences by extraction of keywords from MEDLINE abstracts. Development of a prototype system.
Proc Int Conf Intell Syst Mol Biol. 1997;5:25-32
PMID: 9322011
-
SGD: Saccharomyces Genome Database.
Nucleic Acids Res. 1998 Jan 1;26(1):73-9
PMID: 9399804
-
FlyBase: a Drosophila database. The FlyBase consortium.
Nucleic Acids Res. 1997 Jan 1;25(1):63-6
PMID: 9045212
-
Structural and functional analyses of APG5, a gene involved in autophagy in yeast.
Gene. 1996 Oct 31;178(1-2):139-43
PMID: 8921905
-
Resolution of subunit interactions and cytoplasmic subcomplexes of the yeast vacuolar proton-translocating ATPase.
J Biol Chem. 1996 Apr 26;271(17):10397-404
PMID: 8626613
-
Basic local alignment search tool.
J Mol Biol. 1990 Oct 5;215(3):403-10
PMID: 2231712
-
Predicting the sub-cellular location of proteins from text using support vector machines.
Pac Symp Biocomput. 2002;:374-85
PMID: 11928491
-
Associating genes with gene ontology codes using a maximum entropy analysis of biomedical literature.
Genome Res. 2002 Jan;12(1):203-14
PMID: 11779846
-
The Mouse Genome Database (MGD): the model organism database for the laboratory mouse.
Nucleic Acids Res. 2002 Jan 1;30(1):113-5
PMID: 11752269
-
A literature network of human genes for high-throughput analysis of gene expression.
Nat Genet. 2001 May;28(1):21-8
PMID: 11326270
-
Use of keyword hierarchies to interpret gene expression patterns.
Bioinformatics. 2001 Apr;17(4):319-26
PMID: 11301300
-
Basic microarray analysis: grouping and feature reduction.
Trends Biotechnol. 2001 May;19(5):189-93
PMID: 11301132
-
Information access. Building a "GenBank" of the published literature.
Science. 2001 Mar 23;291(5512):2318-9
PMID: 11269300
-
Detecting gene relations from Medline abstracts.
Pac Symp Biocomput. 2001;:483-95
PMID: 11262966
-
Including biological literature improves homology search.
Pac Symp Biocomput. 2001;:374-83
PMID: 11262956
-
MIPS: a database for genomes and protein sequences.
Nucleic Acids Res. 2000 Jan 1;28(1):37-40
PMID: 10592176
-
Genes, themes and microarrays: using information retrieval for large-scale gene analysis.
Proc Int Conf Intell Syst Mol Biol. 2000;8:317-28
PMID: 10977093
-
Functional discovery via a compendium of expression profiles.
Cell. 2000 Jul 7;102(1):109-26
PMID: 10929718
-
Automatic extraction of protein interactions from scientific abstracts.
Pac Symp Biocomput. 2000;:541-52
PMID: 10902201
-
SAWTED: structure assignment with text description--enhanced detection of remote homologues with automated SWISS-PROT annotation comparisons.
Bioinformatics. 2000 Feb;16(2):125-9
PMID: 10842733
-
Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
Nat Genet. 2000 May;25(1):25-9
PMID: 10802651
-
Evaluation of human-readable annotation in biomolecular sequence databases with biological rule libraries.
Bioinformatics. 1999 Jul-Aug;15(7-8):528-35
PMID: 10487860
-
A comprehensive analysis of protein-protein interactions in Saccharomyces cerevisiae.
Nature. 2000 Feb 10;403(6770):623-7
PMID: 10688190
-
Automatic extraction of biological information from scientific text: protein-protein interactions.
Proc Int Conf Intell Syst Mol Biol. 1999;:60-7
PMID: 10786287
-
Functional characterization of the S. cerevisiae genome by gene deletion and parallel analysis.
Science. 1999 Aug 6;285(5429):901-6
PMID: 10436161
-
A novel method for automatic functional annotation of proteins.
Bioinformatics. 1999 Mar;15(3):228-33
PMID: 10222410
-
The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1999.
Nucleic Acids Res. 1999 Jan 1;27(1):49-54
PMID: 9847139
-
Cluster analysis and display of genome-wide expression patterns.
Proc Natl Acad Sci U S A. 1998 Dec 8;95(25):14863-8
PMID: 9843981
-
EUCLID: automatic classification of proteins in functional classes by their database annotations.
Bioinformatics. 1998;14(6):542-3
PMID: 9694995
-
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
PMID: 9254694