Abstract
We evaluate tertiary structure predictions on medium to large size proteins by TASSER, a new algorithm that assembles protein structures through rearranging the rigid fragments from threading templates guided by a reduced Calpha and side-chain based potential consistent with threading based tertiary restraints. Predictions were generated for 745 proteins 201-300 residues in length that cover the Protein Data Bank (PDB) at the level of 35% sequence identity. With homologous proteins excluded, in 365 cases, the templates identified by our threading program, PROSPECTOR_3, have a root-mean-square deviation (RMSD) to native < 6.5 angstroms, with >70% alignment coverage. After TASSER assembly, in 408 cases the best of the top five full-length models has a RMSD < 6.5 angstroms. Among the 745 targets are 18 membrane proteins, with one-third having a predicted RMSD < 5.5 A. For all representative proteins less than or equal to 300 residues that have corresponding multiple NMR structures in the Protein Data Bank, approximately 20% of the models generated by TASSER are closer to the NMR structure centroid than the farthest individual NMR model. These results suggest that reasonable structure predictions for nonhomologous large size proteins can be automatically generated on a proteomic scale, and the application of this approach to structural as well as functional genomics represent promising applications of TASSER.
MeSH Terms
Algorithms
Amino Acid Sequence
Benchmarking/methods
Computer Simulation
Databases, Protein
Magnetic Resonance Spectroscopy
Models, Molecular
Molecular Sequence Data
Protein Folding
Protein Structure, Tertiary
Proteins/analysis,chemistry,standards
Reference Standards
Reproducibility of Results
Sensitivity and Specificity
Sequence Analysis, Protein/methods,standards
Sequence Homology, Amino Acid
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Zhang Yang
Center of Excellence in Bioinformatics, University at Buffalo, Buffalo, New York 14203, USA.
Skolnick Jeffrey
References (30)
30 references, click to expand
-
The Protein Data Bank.
Nucleic Acids Res. 2000 Jan 1;28(1):235-42
PMID: 10592235
-
Membrane protein folding and stability: physical principles.
Annu Rev Biophys Biomol Struct. 1999;28:319-65
PMID: 10410805
-
Structural genomics and its importance for gene function analysis.
Nat Biotechnol. 2000 Mar;18(3):283-7
PMID: 10700142
-
Do more complex organisms have a greater proportion of membrane proteins in their genomes?
Proteins. 2000 Jun 1;39(4):417-20
PMID: 10813823
-
Comparative protein structure modeling of genes and genomes.
Annu Rev Biophys Biomol Struct. 2000;29:291-325
PMID: 10940251
-
Modeling of loops in protein structures.
Protein Sci. 2000 Sep;9(9):1753-73
PMID: 11045621
-
Structure determination of membrane-associated proteins from nuclear magnetic resonance data.
Anal Biochem. 2001 Jan 1;288(1):1-15
PMID: 11141300
-
Defrosting the frozen approximation: PROSPECTOR--a new approach to threading.
Proteins. 2001 Feb 15;42(3):319-31
PMID: 11151004
-
Prospects for ab initio protein structural genomics.
J Mol Biol. 2001 Mar 9;306(5):1191-9
PMID: 11237627
-
Two-dimensional crystallization of membrane proteins: the lipid layer strategy.
FEBS Lett. 2001 Aug 31;504(3):187-93
PMID: 11532452
-
Protein structure prediction and structural genomics.
Science. 2001 Oct 5;294(5540):93-6
PMID: 11588250
-
GTOP: a database of protein structures predicted from genome sequences.
Nucleic Acids Res. 2002 Jan 1;30(1):294-8
PMID: 11752318
-
Critical assessment of methods of protein structure prediction (CASP): round IV.
Proteins. 2001;Suppl 5:2-7
PMID: 11835476
-
Local energy landscape flattening: parallel hyperbolic Monte Carlo sampling of protein folding.
Proteins. 2002 Aug 1;48(2):192-201
PMID: 12112688
-
The PEDANT genome database.
Nucleic Acids Res. 2003 Jan 1;31(1):207-11
PMID: 12519983
-
TMPDB: a database of experimentally-characterized transmembrane topologies.
Nucleic Acids Res. 2003 Jan 1;31(1):406-9
PMID: 12520035
-
Improving the performance of DomainParser for structural domain partition using neural network.
Nucleic Acids Res. 2003 Feb 1;31(3):944-52
PMID: 12560490
-
TOUCHSTONE II: a new approach to ab initio protein structure prediction.
Biophys J. 2003 Aug;85(2):1145-64
PMID: 12885659
-
Critical assessment of methods of protein structure prediction (CASP)-round V.
Proteins. 2003;53 Suppl 6:334-9
PMID: 14579322
-
SPICKER: a clustering approach to identify near-native protein folds.
J Comput Chem. 2004 Apr 30;25(6):865-71
PMID: 15011258
-
Automated structure prediction of weakly homologous proteins on a genomic scale.
Proc Natl Acad Sci U S A. 2004 May 18;101(20):7594-9
PMID: 15126668
-
Development and large scale benchmark testing of the PROSPECTOR_3 threading algorithm.
Proteins. 2004 Aug 15;56(3):502-18
PMID: 15229883
-
A general method applicable to the search for similarities in the amino acid sequence of two proteins.
J Mol Biol. 1970 Mar;48(3):443-53
PMID: 5420325
-
Mapping the protein universe.
Science. 1996 Aug 2;273(5275):595-603
PMID: 8662544
-
What is the probability of a chance prediction of a protein structure with an rmsd of 6 A?
Fold Des. 1998;3(2):141-7
PMID: 9565758
-
Assembly of protein structure from sparse experimental data: an efficient Monte Carlo model.
Proteins. 1998 Sep 1;32(4):475-94
PMID: 9726417
-
Hidden Markov models for detecting remote protein homologies.
Bioinformatics. 1998;14(10):846-56
PMID: 9927713
-
Protein structure prediction by global optimization of a potential energy function.
Proc Natl Acad Sci U S A. 1999 May 11;96(10):5482-5
PMID: 10318909
-
Protein secondary structure prediction based on position-specific scoring matrices.
J Mol Biol. 1999 Sep 17;292(2):195-202
PMID: 10493868
-
Derivation of protein-specific pair potentials based on weak sequence fragment similarity.
Proteins. 2000 Jan 1;38(1):3-16
PMID: 10651034