Home LiteratureArticle Details
PMID: 15751111 Published · ppublish English Comparative Study Evaluation Study Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, P.H.S. Validation Study

Optimizing long intrinsic disorder predictors with protein evolutionary information.

Journal of bioinformatics and computational biology ·Vol. 3 ·No. 1 ·2005-02-00 ·Pages 35-60

Peng K, Vucetic S, Radivojac P, Brown CJ, Dunker AK, Obradovic Z

Abstract

Protein existing as an ensemble of structures, called intrinsically disordered, has been shown to be responsible for a wide variety of biological functions and to be common in nature. Here we focus on improving sequence-based predictions of long (>30 amino acid residues) regions lacking specific 3-D structure by means of four new neural-network-based Predictors Of Natural Disordered Regions (PONDRs): VL3, VL3H, VL3P, and VL3E. PONDR VL3 used several features from a previously introduced PONDR VL2, but benefitted from optimized predictor models and a slightly larger (152 vs. 145) set of disordered proteins that were cleaned of mislabeling errors found in the smaller set. PONDR VL3H utilized homologues of the disordered proteins in the training stage, while PONDR VL3P used attributes derived from sequence profiles obtained by PSI-BLAST searches. The measure of accuracy was the average between accuracies on disordered and ordered protein regions. By this measure, the 30-fold cross-validation accuracies of VL3, VL3H, and VL3P were, respectively, 83.6 +/- 1.4%, 85.3 +/- 1.4%, and 85.2 +/- 1.5%. By combining VL3H and VL3P, the resulting PONDR VL3E achieved an accuracy of 86.7 +/- 1.4%. This is a significant improvement over our previous PONDRs VLXT (71.6 +/- 1.3%) and VL2 (80.9 +/- 1.4%). The new disorder predictors with the corresponding datasets are freely accessible through the web server at http://www.ist.temple.edu/disprot.

MeSH Terms
Algorithms Computer Simulation Conserved Sequence Evolution, Molecular Models, Chemical Models, Genetic Models, Statistical Neural Networks, Computer Proteins/chemistry,genetics Sequence Alignment/methods Sequence Analysis, Protein/methods Sequence Homology, Amino Acid
Chemicals
Proteins
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Peng Kang
Center for Information Science and Technology, Temple University, Philadelphia, PA 19122, USA.
Vucetic Slobodan
Radivojac Predrag
Brown Celeste J
Dunker A Keith
Obradovic Zoran
Article Info
Journal
Journal of bioinformatics and computational biology
Abbr.
J Bioinform Comput Biol
ISSN
0219-7200
Published
2005-02-00
Pages
35-60
Language
English
Region
Singapore
NLM ID
101187344
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]