Home LiteratureArticle Details
PMID: 12368251 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S. Research Support, U.S. Gov't, P.H.S.

Using text analysis to identify functionally coherent gene groups.

Genome research ·Vol. 12 ·No. 10 ·2002-10-00 ·Pages 1582-90

Raychaudhuri S, Schütze H, Altman RB

Abstract

The analysis of large-scale genomic information (such as sequence data or expression patterns) frequently involves grouping genes on the basis of common experimental features. Often, as with gene expression clustering, there are too many groups to easily identify the functionally relevant ones. One valuable source of information about gene function is the published literature. We present a method, neighbor divergence, for assessing whether the genes within a group share a common biological function based on their associated scientific literature. The method uses statistical natural language processing techniques to interpret biological text. It requires only a corpus of documents relevant to the genes being studied (e.g., all genes in an organism) and an index connecting the documents to appropriate genes. Given a group of genes, neighbor divergence assigns a numerical score indicating how "functionally coherent" the gene group is from the perspective of the published literature. We evaluate our method by testing its ability to distinguish 19 known functional gene groups from 1900 randomly assembled groups. Neighbor divergence achieves 79% sensitivity at 100% specificity, comparing favorably to other tested methods. We also apply neighbor divergence to previously published gene expression clusters to assess its ability to recognize gene groups that had been manually identified as representative of a common function.

MeSH Terms
Algorithms Artificial Intelligence Cluster Analysis Computational Biology/methods,trends Databases, Genetic/statistics & numerical data Discriminant Analysis Gene Expression Profiling/methods,statistics & numerical data,trends Genes, Fungal/physiology Genome, Fungal Information Services Natural Language Processing Research Design/statistics & numerical data,trends Saccharomyces cerevisiae/genetics
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Raychaudhuri Soumya
Department of Genetics, Stanford Medical Informatics, University, Stanford, California 94305-5479, USA.
Schütze Hinrich
Altman Russ B
References (30)
30 references, click to expand
  1. Automatic annotation for biological sequences by extraction of keywords from MEDLINE abstracts. Development of a prototype system.
    Proc Int Conf Intell Syst Mol Biol. 1997;5:25-32 PMID: 9322011
  2. SGD: Saccharomyces Genome Database.
    Nucleic Acids Res. 1998 Jan 1;26(1):73-9 PMID: 9399804
  3. FlyBase: a Drosophila database. The FlyBase consortium.
    Nucleic Acids Res. 1997 Jan 1;25(1):63-6 PMID: 9045212
  4. Structural and functional analyses of APG5, a gene involved in autophagy in yeast.
    Gene. 1996 Oct 31;178(1-2):139-43 PMID: 8921905
  5. Resolution of subunit interactions and cytoplasmic subcomplexes of the yeast vacuolar proton-translocating ATPase.
    J Biol Chem. 1996 Apr 26;271(17):10397-404 PMID: 8626613
  6. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  7. Predicting the sub-cellular location of proteins from text using support vector machines.
    Pac Symp Biocomput. 2002;:374-85 PMID: 11928491
  8. Associating genes with gene ontology codes using a maximum entropy analysis of biomedical literature.
    Genome Res. 2002 Jan;12(1):203-14 PMID: 11779846
  9. The Mouse Genome Database (MGD): the model organism database for the laboratory mouse.
    Nucleic Acids Res. 2002 Jan 1;30(1):113-5 PMID: 11752269
  10. A literature network of human genes for high-throughput analysis of gene expression.
    Nat Genet. 2001 May;28(1):21-8 PMID: 11326270
  11. Use of keyword hierarchies to interpret gene expression patterns.
    Bioinformatics. 2001 Apr;17(4):319-26 PMID: 11301300
  12. Basic microarray analysis: grouping and feature reduction.
    Trends Biotechnol. 2001 May;19(5):189-93 PMID: 11301132
  13. Information access. Building a "GenBank" of the published literature.
    Science. 2001 Mar 23;291(5512):2318-9 PMID: 11269300
  14. Detecting gene relations from Medline abstracts.
    Pac Symp Biocomput. 2001;:483-95 PMID: 11262966
  15. Including biological literature improves homology search.
    Pac Symp Biocomput. 2001;:374-83 PMID: 11262956
  16. MIPS: a database for genomes and protein sequences.
    Nucleic Acids Res. 2000 Jan 1;28(1):37-40 PMID: 10592176
  17. Genes, themes and microarrays: using information retrieval for large-scale gene analysis.
    Proc Int Conf Intell Syst Mol Biol. 2000;8:317-28 PMID: 10977093
  18. Functional discovery via a compendium of expression profiles.
    Cell. 2000 Jul 7;102(1):109-26 PMID: 10929718
  19. Automatic extraction of protein interactions from scientific abstracts.
    Pac Symp Biocomput. 2000;:541-52 PMID: 10902201
  20. SAWTED: structure assignment with text description--enhanced detection of remote homologues with automated SWISS-PROT annotation comparisons.
    Bioinformatics. 2000 Feb;16(2):125-9 PMID: 10842733
  21. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  22. Evaluation of human-readable annotation in biomolecular sequence databases with biological rule libraries.
    Bioinformatics. 1999 Jul-Aug;15(7-8):528-35 PMID: 10487860
  23. A comprehensive analysis of protein-protein interactions in Saccharomyces cerevisiae.
    Nature. 2000 Feb 10;403(6770):623-7 PMID: 10688190
  24. Automatic extraction of biological information from scientific text: protein-protein interactions.
    Proc Int Conf Intell Syst Mol Biol. 1999;:60-7 PMID: 10786287
  25. Functional characterization of the S. cerevisiae genome by gene deletion and parallel analysis.
    Science. 1999 Aug 6;285(5429):901-6 PMID: 10436161
  26. A novel method for automatic functional annotation of proteins.
    Bioinformatics. 1999 Mar;15(3):228-33 PMID: 10222410
  27. The SWISS-PROT protein sequence data bank and its supplement TrEMBL in 1999.
    Nucleic Acids Res. 1999 Jan 1;27(1):49-54 PMID: 9847139
  28. Cluster analysis and display of genome-wide expression patterns.
    Proc Natl Acad Sci U S A. 1998 Dec 8;95(25):14863-8 PMID: 9843981
  29. EUCLID: automatic classification of proteins in functional classes by their database annotations.
    Bioinformatics. 1998;14(6):542-3 PMID: 9694995
  30. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1088-9051
Published
2002-10-00
Pages
1582-90
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC187532
Subset
IM
Grants
NIGMS NIH HHS · R24 GM061374 · United States
NIGMS NIH HHS · T32 GM007365 · United States
NLM NIH HHS · LM06244 · United States
NIGMS NIH HHS · GM61374 · United States
NIGMS NIH HHS · GM-07365 · United States
NIGMS NIH HHS · U01 GM061374 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]