Home LiteratureArticle Details
PMID: 11262956 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S. Research Support, U.S. Gov't, P.H.S.

Including biological literature improves homology search.

Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing ·2001-00-00 ·Pages 374-83

Chang JT, Raychaudhuri S, Altman RB

Abstract

Annotating the tremendous amount of sequence information being generated requires accurate automated methods for recognizing homology. Although sequence similarity is only one of many indicators of evolutionary homology, it is often the only one used. Here we find that supplementing sequence similarity with information from biomedical literature is successful in increasing the accuracy of homology search results. We modified the PSI-BLAST algorithm to use literature similarity in each iteration of its database search. The modified algorithm is evaluated and compared to standard PSI-BLAST in searching for homologous proteins. The performance of the modified algorithm achieved 32% recall with 95% precision, while the original one achieved 33% recall with 84% precision; the literature similarity requirement preserved the sensitive characteristic of the PSI-BLAST algorithm while improving the precision.

MeSH Terms
Algorithms Computational Biology Databases, Factual Natural Language Processing Periodicals as Topic Sensitivity and Specificity Sequence Alignment/methods,statistics & numerical data Software
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Chang J T
Stanford Medical Informatics, Stanford University, 251 Campus Drive, MSOB X-215, Stanford, CA 94305-5479, USA. [email protected]
Raychaudhuri S
Altman R B
References (19)
19 references, click to expand
  1. Identification of related proteins on family, superfamily and fold level.
    J Mol Biol. 2000 Jan 21;295(3):613-25 PMID: 10623551
  2. Assessing annotation transfer for genomics: quantifying the relations between protein sequence, structure and function through traditional and probabilistic scores.
    J Mol Biol. 2000 Mar 17;297(1):233-49 PMID: 10704319
  3. Position-specific annotation of protein function based on multiple homologs.
    Proc Int Conf Intell Syst Mol Biol. 1999;:28-33 PMID: 10786283
  4. Large-scale comparison of protein sequence alignment algorithms with structure alignments.
    Proteins. 2000 Jul 1;40(1):6-22 PMID: 10813826
  5. SAWTED: structure assignment with text description--enhanced detection of remote homologues with automated SWISS-PROT annotation comparisons.
    Bioinformatics. 2000 Feb;16(2):125-9 PMID: 10842733
  6. The PSIPRED protein structure prediction server.
    Bioinformatics. 2000 Apr;16(4):404-5 PMID: 10869041
  7. A general method applicable to the search for similarities in the amino acid sequence of two proteins.
    J Mol Biol. 1970 Mar;48(3):443-53 PMID: 5420325
  8. The Protein Data Bank: a computer-based archival file for macromolecular structures.
    J Mol Biol. 1977 May 25;112(3):535-42 PMID: 875032
  9. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  10. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  11. The SWISS-PROT protein sequence data bank.
    Nucleic Acids Res. 1991 Apr 25;19 Suppl:2247-9 PMID: 2041811
  12. SCOP: a structural classification of proteins database for the investigation of sequences and structures.
    J Mol Biol. 1995 Apr 7;247(4):536-40 PMID: 7723011
  13. An analysis of statistical term strength and its use in the indexing and retrieval of molecular biology texts.
    Comput Biol Med. 1996 May;26(3):209-22 PMID: 8725772
  14. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  15. Automatic annotation for biological sequences by extraction of keywords from MEDLINE abstracts. Development of a prototype system.
    Proc Int Conf Intell Syst Mol Biol. 1997;5:25-32 PMID: 9322011
  16. Homology-based fold predictions for Mycoplasma genitalium proteins.
    J Mol Biol. 1998 Jul 17;280(3):323-6 PMID: 9665839
  17. Characterization of single-nucleotide polymorphisms in coding regions of human genes.
    Nat Genet. 1999 Jul;22(3):231-8 PMID: 10391209
  18. Patterns of single-nucleotide polymorphisms in candidate genes for blood-pressure homeostasis.
    Nat Genet. 1999 Jul;22(3):239-47 PMID: 10391210
  19. Comparative modeling of CASP3 targets using PSI-BLAST and SCWRL.
    Proteins. 1999;Suppl 3:81-7 PMID: 10526356
Article Info
Journal
Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing
Abbr.
Pac Symp Biocomput
ISSN
2335-6928
Published
2001-00-00
Pages
374-83
Language
English
Region
United States
NLM ID
9711271
PMCID
PMC2671075
Subset
IM
Grants
NIGMS NIH HHS · T32 GM007365 · United States
NIGMS NIH HHS · GM07365 · United States
NIGMS NIH HHS · U01 GM061374 · United States
NIGMS NIH HHS · 1U01-GM61374-01 · United States
NIGMS NIH HHS · T32 GM007365-23 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]