Home LiteratureArticle Details
PMID: 10627144 Published · ppublish English Journal Article

Significance of Z-value statistics of Smith-Waterman scores for protein alignments.

Computers & chemistry ·Vol. 23 ·No. 3-4 ·1999-06-15 ·Pages 317-31

Comet JP, Aude JC, Glémet E, Risler JL, Hénaut A, Slonimski PP, Codani JJ

Abstract

The Z-value is an attempt to estimate the statistical significance of a Smith-Waterman dynamic alignment score (SW-score) through the use of a Monte-Carlo process. It partly reduces the bias induced by the composition and length of the sequences. This paper is not a theoretical study on the distribution of SW-scores and Z-values. Rather, it presents a statistical analysis of Z-values on large datasets of protein sequences, leading to a law of probability that the experimental Z-values follow. First, we determine the relationships between the computed Z-value, an estimation of its variance and the number of randomizations in the Monte-Carlo process. Then, we illustrate that Z-values are less correlated to sequence lengths than SW-scores. Then we show that pairwise alignments, performed on 'quasi-real' sequences (i.e., randomly shuffled sequences of the same length and amino acid composition as the real ones) lead to Z-value distributions that statistically fit the extreme value distribution, more precisely the Gumbel distribution (global EVD, Extreme Value Distribution). However, for real protein sequences, we observe an over-representation of high Z-values. We determine first a cutoff value which separates these overestimated Z-values from those which follow the global EVD. We then show that the interesting part of the tail of distribution of Z-values can be approximated by another EVD (i.e., an EVD which differs from the global EVD) or by a Pareto law. This has been confirmed for all proteins analysed so far, whether extracted from individual genomes, or from the ensemble of five complete microbial genomes comprising altogether 16956 protein sequences.

MeSH Terms
Computing Methodologies Escherichia coli/genetics Genome, Bacterial Genome, Fungal Mathematics Monte Carlo Method Saccharomyces cerevisiae/genetics Sequence Alignment
Authors & Affiliations
7 authors, click to expand affiliations / ORCID
Comet J P
INRIA Rocquencourt, Le-Chesnay Cedex, France. [email protected]
Aude J C
Glémet E
Risler J L
Hénaut A
Slonimski P P
Codani J J
Article Info
Journal
Computers & chemistry
Abbr.
Comput Chem
ISSN
0097-8485
Published
1999-06-15
Pages
317-31
Language
English
Region
England
NLM ID
7607706
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]