Home LiteratureArticle Details
PMID: 18094477 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

The JCSG MR pipeline: optimized alignments, multiple models and parallel searches.

Acta crystallographica. Section D, Biological crystallography ·Vol. 64 ·No. Pt 1 ·2008-01-00 ·Pages 133-40

Schwarzenbacher R, Godzik A, Jaroszewski L

Abstract

The success rate of molecular replacement (MR) falls considerably when search models share less than 35% sequence identity with their templates, but can be improved significantly by using fold-recognition methods combined with exhaustive MR searches. Models based on alignments calculated with fold-recognition algorithms are more accurate than models based on conventional alignment methods such as FASTA or BLAST, which are still widely used for MR. In addition, by designing MR pipelines that integrate phasing and automated refinement and allow parallel processing of such calculations, one can effectively increase the success rate of MR. Here, updated results from the JCSG MR pipeline are presented, which to date has solved 33 MR structures with less than 35% sequence identity to the closest homologue of known structure. By using difficult MR problems as examples, it is demonstrated that successful MR phasing is possible even in cases where the similarity between the model and the template can only be detected with fold-recognition algorithms. In the first step, several search models are built based on all homologues found in the PDB by fold-recognition algorithms. The models resulting from this process are used in parallel MR searches with different combinations of input parameters of the MR phasing algorithm. The putative solutions are subjected to rigid-body and restrained crystallographic refinement and ranked based on the final values of free R factor, figure of merit and deviations from ideal geometry. Finally, crystal packing and electron-density maps are checked to identify the correct solution. If this procedure does not yield a solution with interpretable electron-density maps, then even more alternative models are prepared. The structurally variable regions of a protein family are identified based on alignments of sequences and known structures from that family and appropriate trimmings of the models are proposed. All combinations of these trimmings are applied to the search models and the resulting set of models is used in the MR pipeline. It is estimated that with the improvements in model building and exhaustive parallel searches with existing phasing algorithms, MR can be successful for more than 50% of recognizable homologues of known structures below the threshold of 35% sequence identity. This implies that about one-third of the proteins in a typical bacterial proteome are potential MR targets.

MeSH Terms
Algorithms Amino Acid Sequence Computational Biology/methods Computer Simulation Crystallography, X-Ray Models, Molecular Molecular Sequence Data Proteins/chemistry,genetics Sequence Alignment/methods Sequence Homology, Amino Acid
Chemicals
Proteins
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Schwarzenbacher Robert
University of Salzburg, Structural Biology, Billrothstrasse 11, 5020 Salzburg, Austria.
Godzik Adam
Jaroszewski Lukasz
References (39)
39 references, click to expand
  1. A stochastic approach to molecular replacement.
    Acta Crystallogr D Biol Crystallogr. 2000 Feb;56(Pt 2):169-74 PMID: 10666596
  2. On the potential of normal-mode analysis for solving difficult molecular-replacement problems.
    Acta Crystallogr D Biol Crystallogr. 2004 Apr;60(Pt 4):796-9 PMID: 15039589
  3. BALBES: a molecular-replacement pipeline.
    Acta Crystallogr D Biol Crystallogr. 2008 Jan;64(Pt 1):125-32 PMID: 18094476
  4. A database of protein structure families with common folding motifs.
    Protein Sci. 1992 Dec;1(12):1691-8 PMID: 1304898
  5. Molecular replacement--historical background.
    Acta Crystallogr D Biol Crystallogr. 2001 Oct;57(Pt 10):1360-6 PMID: 11567146
  6. Profile hidden Markov models.
    Bioinformatics. 1998;14(9):755-63 PMID: 9918945
  7. Solution solution: using NMR models for molecular replacement.
    Acta Crystallogr D Biol Crystallogr. 2001 Oct;57(Pt 10):1457-61 PMID: 11567160
  8. CaspR: a web server for automated molecular replacement using homology modelling.
    Nucleic Acids Res. 2004 Jul 1;32(Web Server issue):W606-9 PMID: 15215460
  9. SCOP: a structural classification of proteins database for the investigation of sequences and structures.
    J Mol Biol. 1995 Apr 7;247(4):536-40 PMID: 7723011
  10. The CCP4 suite: programs for protein crystallography.
    Acta Crystallogr D Biol Crystallogr. 1994 Sep 1;50(Pt 5):760-3 PMID: 15299374
  11. An approach to multi-copy search in molecular replacement.
    Acta Crystallogr D Biol Crystallogr. 2000 Dec;56(Pt 12):1622-4 PMID: 11092928
  12. Protein homology detection by HMM-HMM comparison.
    Bioinformatics. 2005 Apr 1;21(7):951-60 PMID: 15531603
  13. Structural genomics of the Thermotoga maritima proteome implemented in a high-throughput structure determination pipeline.
    Proc Natl Acad Sci U S A. 2002 Sep 3;99(18):11664-9 PMID: 12193646
  14. Implementation of molecular replacement in AMoRe.
    Acta Crystallogr D Biol Crystallogr. 2001 Oct;57(Pt 10):1367-72 PMID: 11567147
  15. MrBUMP: an automated pipeline for molecular replacement.
    Acta Crystallogr D Biol Crystallogr. 2008 Jan;64(Pt 1):119-24 PMID: 18094475
  16. FUGUE: sequence-structure homology recognition using environment-specific substitution tables and structure-dependent gap penalties.
    J Mol Biol. 2001 Jun 29;310(1):243-57 PMID: 11419950
  17. Refinement of macromolecular structures by the maximum-likelihood method.
    Acta Crystallogr D Biol Crystallogr. 1997 May 1;53(Pt 3):240-55 PMID: 15299926
  18. Rapid automated molecular replacement by evolutionary search.
    Acta Crystallogr D Biol Crystallogr. 1999 Feb;55(Pt 2):484-91 PMID: 10089360
  19. Crystallography & NMR system: A new software suite for macromolecular structure determination.
    Acta Crystallogr D Biol Crystallogr. 1998 Sep 1;54(Pt 5):905-21 PMID: 9757107
  20. The Protein Data Bank.
    Nucleic Acids Res. 2000 Jan 1;28(1):235-42 PMID: 10592235
  21. An assessment of amino acid exchange matrices in aligning protein sequences: the twilight zone revisited.
    J Mol Biol. 1995 Jun 16;249(4):816-31 PMID: 7602593
  22. WHAT IF: a molecular modeling and drug design program.
    J Mol Graph. 1990 Mar;8(1):52-6, 29 PMID: 2268628
  23. The ASTRAL Compendium in 2004.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D189-92 PMID: 14681391
  24. Parameter-space screening: a powerful tool for high-throughput crystal structure determination.
    Acta Crystallogr D Biol Crystallogr. 2005 May;61(Pt 5):520-7 PMID: 15858261
  25. The relation between the divergence of sequence and structure in proteins.
    EMBO J. 1986 Apr;5(4):823-6 PMID: 3709526
  26. Evaluating the potential of using fold-recognition models for molecular replacement.
    Acta Crystallogr D Biol Crystallogr. 2001 Oct;57(Pt 10):1428-34 PMID: 11567156
  27. Hidden Markov models for detecting remote protein homologies.
    Bioinformatics. 1998;14(10):846-56 PMID: 9927713
  28. Hybrid fold recognition: combining sequence derived properties with evolutionary information.
    Pac Symp Biocomput. 2000;:119-30 PMID: 10902162
  29. Synergistic effects of substrate-induced conformational changes in phosphoglycerate kinase activation.
    Nature. 1997 Jan 16;385(6613):275-8 PMID: 9000079
  30. The importance of alignment accuracy for molecular replacement.
    Acta Crystallogr D Biol Crystallogr. 2004 Jul;60(Pt 7):1229-36 PMID: 15213384
  31. Protein threading using PROSPECT: design and evaluation.
    Proteins. 2000 Aug 15;40(3):343-54 PMID: 10861926
  32. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  33. Likelihood-enhanced fast rotation functions.
    Acta Crystallogr D Biol Crystallogr. 2004 Mar;60(Pt 3):432-8 PMID: 14993666
  34. A method for finding candidate conformations for molecular replacement using relative rotation between domains of a known structure.
    Acta Crystallogr D Biol Crystallogr. 2006 Apr;62(Pt 4):398-409 PMID: 16552141
  35. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  36. Comparison of sequence profiles. Strategies for structural predictions using sequence information.
    Protein Sci. 2000 Feb;9(2):232-41 PMID: 10716175
  37. Structural diversity of domain superfamilies in the CATH database.
    J Mol Biol. 2006 Jul 14;360(3):725-41 PMID: 16780872
  38. Multiple flexible structure alignment using partial order graphs.
    Bioinformatics. 2005 May 15;21(10):2362-9 PMID: 15746292
  39. FFAS03: a server for profile--profile sequence alignments.
    Nucleic Acids Res. 2005 Jul 1;33(Web Server issue):W284-8 PMID: 15980471
Article Info
Journal
Acta crystallographica. Section D, Biological crystallography
Abbr.
Acta Crystallogr D Biol Crystallogr
ISSN
0907-4449
Published
2008-01-00
Epub
2007-00-05
Pages
133-40
Language
English
Region
United States
NLM ID
9305878
PMCID
PMC2394805
Subset
IM
Grants
NIGMS NIH HHS · U54 GM074898 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]