Home LiteratureArticle Details
PMID: 19472362 Published · ppublish English Comparative Study Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

Evolutionary constraints on structural similarity in orthologs and paralogs.

Protein science : a publication of the Protein Society ·Vol. 18 ·No. 6 ·2009-06-00 ·Pages 1306-15

Peterson ME, Chen F, Saven JG, Roos DS, Babbitt PC, Sali A

Abstract

Although a quantitative relationship between sequence similarity and structural similarity has long been established, little is known about the impact of orthology on the relationship between protein sequence and structure. Among homologs, orthologs (derived by speciation) more frequently have similar functions than paralogs (derived by duplication). Here, we hypothesize that an orthologous pair will tend to exhibit greater structural similarity than a paralogous pair at the same level of sequence similarity. To test this hypothesis, we used 284,459 pairwise structure-based alignments of 12,634 unique domains from SCOP as well as orthology and paralogy assignments from OrthoMCL DB. We divided the comparisons by sequence identity and determined whether the sequence-structure relationship differed between the orthologs and paralogs. We found that at levels of sequence identity between 30 and 70%, orthologous domain pairs indeed tend to be significantly more structurally similar than paralogous pairs at the same level of sequence identity. An even larger difference is found when comparing ligand binding residues instead of whole domains. These differences between orthologs and paralogs are expected to be useful for selecting template structures in comparative modeling and target proteins in structural genomics.

MeSH Terms
Amino Acid Sequence Databases, Protein Evolution, Molecular Protein Structure, Tertiary Proteins/chemistry Sequence Alignment Structure-Activity Relationship
Chemicals
Proteins
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Peterson Mark E
Department of Bioengineering and Therapeutic Sciences, University of California-San Francisco, 1700 4th Street, San Francisco, CA 94158, USA. [email protected]
Chen Feng
Saven Jeffery G
Roos David S
Babbitt Patricia C
Sali Andrej
References (64)
64 references, click to expand
  1. Assessing performance of orthology detection strategies applied to eukaryotic genomes.
    PLoS One. 2007 Apr 18;2(4):e383 PMID: 17440619
  2. History of the enzyme nomenclature system.
    Bioinformatics. 2000 Jan;16(1):34-40 PMID: 10812475
  3. On a Mirkin-Muchnik-Smith conjecture for comparing molecular phylogenies.
    J Comput Biol. 1997 Summer;4(2):177-87 PMID: 9228616
  4. How well is enzyme function conserved as a function of pairwise sequence identity?
    J Mol Biol. 2003 Oct 31;333(4):863-82 PMID: 14568541
  5. A biologically consistent model for comparing molecular phylogenies.
    J Comput Biol. 1995 Winter;2(4):493-507 PMID: 8634901
  6. The relation between the divergence of sequence and structure in proteins.
    EMBO J. 1986 Apr;5(4):823-6 PMID: 3709526
  7. Practical limits of function prediction.
    Proteins. 2000 Oct 1;41(1):98-107 PMID: 10944397
  8. BLOSUM62 miscalculations improve search performance.
    Nat Biotechnol. 2008 Mar;26(3):274-5 PMID: 18327232
  9. Evolution of protein sequences and structures.
    J Mol Biol. 1999 Aug 27;291(4):977-95 PMID: 10452901
  10. A practical and robust sequence search strategy for structural genomics target selection.
    Bioinformatics. 2004 Sep 22;20(14):2288-95 PMID: 15201178
  11. A comprehensive evolutionary classification of proteins encoded in complete eukaryotic genomes.
    Genome Biol. 2004;5(2):R7 PMID: 14759257
  12. SCOP: a structural classification of proteins database for the investigation of sequences and structures.
    J Mol Biol. 1995 Apr 7;247(4):536-40 PMID: 7723011
  13. Amino acid substitution matrices from protein blocks.
    Proc Natl Acad Sci U S A. 1992 Nov 15;89(22):10915-9 PMID: 1438297
  14. Protein structure prediction: inroads to biology.
    Mol Cell. 2005 Dec 22;20(6):811-9 PMID: 16364908
  15. Update on the pfam5000 strategy for selection of structural genomics targets.
    Conf Proc IEEE Eng Med Biol Soc. 2005;2006:751-5 PMID: 17282292
  16. Comparison of solvent-inaccessible cores of homologous proteins: definitions useful for protein modelling.
    Protein Eng. 1987 Jun;1(3):159-71 PMID: 3507702
  17. Enzyme function less conserved than anticipated.
    J Mol Biol. 2002 Apr 26;318(2):595-608 PMID: 12051862
  18. The COG database: an updated version includes eukaryotes.
    BMC Bioinformatics. 2003 Sep 11;4:41 PMID: 12969510
  19. Comparative modeling for protein structure prediction.
    Curr Opin Struct Biol. 2006 Apr;16(2):172-7 PMID: 16510277
  20. Comparative protein modelling by satisfaction of spatial restraints.
    J Mol Biol. 1993 Dec 5;234(3):779-815 PMID: 8254673
  21. Orthologs, paralogs, and evolutionary genomics.
    Annu Rev Genet. 2005;39:309-38 PMID: 16285863
  22. UCSF Chimera--a visualization system for exploratory research and analysis.
    J Comput Chem. 2004 Oct;25(13):1605-12 PMID: 15264254
  23. Definitions of enzyme function for the structural genomics era.
    Curr Opin Chem Biol. 2003 Apr;7(2):230-7 PMID: 12714057
  24. Evolution of function in protein superfamilies, from a structural perspective.
    J Mol Biol. 2001 Apr 6;307(4):1113-43 PMID: 11286560
  25. Twilight zone of protein sequence alignments.
    Protein Eng. 1999 Feb;12(2):85-94 PMID: 10195279
  26. The ASTRAL Compendium in 2004.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D189-92 PMID: 14681391
  27. Virtual screening of chemical libraries.
    Nature. 2004 Dec 16;432(7019):862-5 PMID: 15602552
  28. Probing protein fold space with a simplified model.
    J Mol Biol. 2008 Jan 25;375(4):920-33 PMID: 18054792
  29. Quantifying structure-function uncertainty: a graph theoretical exploration into the origins and limitations of protein annotation.
    J Mol Biol. 2004 Apr 2;337(4):933-49 PMID: 15033362
  30. LigBase: a database of families of aligned ligand binding sites in known protein sequences and structures.
    Bioinformatics. 2002 Jan;18(1):200-1 PMID: 11836232
  31. A genomic perspective on protein families.
    Science. 1997 Oct 24;278(5338):631-7 PMID: 9381173
  32. Molecular mechanics methods for predicting protein-ligand binding.
    Phys Chem Chem Phys. 2006 Nov 28;8(44):5166-77 PMID: 17203140
  33. Physically realistic homology models built with ROSETTA can be more accurate than their templates.
    Proc Natl Acad Sci U S A. 2006 Apr 4;103(14):5361-6 PMID: 16567638
  34. From gene to organismal phylogeny: reconciled trees and the gene tree/species tree problem.
    Mol Phylogenet Evol. 1997 Apr;7(2):231-40 PMID: 9126565
  35. Protein structure modeling with MODELLER.
    Methods Mol Biol. 2008;426:145-59 PMID: 18542861
  36. Orthology, paralogy and proposed classification for paralog subtypes.
    Trends Genet. 2002 Dec;18(12):619-20 PMID: 12446146
  37. An integrated approach to the analysis and modeling of protein sequences and structures. II. On the relationship between sequence and structural similarity for proteins that are not obviously related in sequence.
    J Mol Biol. 2000 Aug 18;301(3):679-89 PMID: 10966777
  38. Assessing annotation transfer for genomics: quantifying the relations between protein sequence, structure and function through traditional and probabilistic scores.
    J Mol Biol. 2000 Mar 17;297(1):233-49 PMID: 10704319
  39. TreeFam: a curated database of phylogenetic trees of animal gene families.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D572-80 PMID: 16381935
  40. Protein structure alignment by incremental combinatorial extension (CE) of the optimal path.
    Protein Eng. 1998 Sep;11(9):739-47 PMID: 9796821
  41. Sequence variations within protein families are linearly related to structural variations.
    J Mol Biol. 2002 Oct 25;323(3):551-62 PMID: 12381308
  42. Shining a light on structural genomics.
    Nat Struct Biol. 1998 Aug;5 Suppl:643-5 PMID: 9699614
  43. OrthoMCL: identification of ortholog groups for eukaryotic genomes.
    Genome Res. 2003 Sep;13(9):2178-89 PMID: 12952885
  44. Duplication-based measures of difference between gene and species trees.
    J Comput Biol. 1998 Spring;5(1):135-48 PMID: 9541877
  45. The impact of structural genomics: expectations and outcomes.
    Science. 2006 Jan 20;311(5759):347-51 PMID: 16424331
  46. The relationship between protein structure and function: a comprehensive survey with application to the yeast genome.
    J Mol Biol. 1999 Apr 23;288(1):147-64 PMID: 10329133
  47. Progress of structural genomics initiatives: an analysis of solved target structures.
    J Mol Biol. 2005 May 20;348(5):1235-60 PMID: 15854658
  48. Protein structure prediction and structural genomics.
    Science. 2001 Oct 5;294(5540):93-6 PMID: 11588250
  49. Progress over the first decade of CASP experiments.
    Proteins. 2005;61 Suppl 7:225-36 PMID: 16187365
  50. Recognition of analogous and homologous protein folds: analysis of sequence and structure conservation.
    J Mol Biol. 1997 Jun 13;269(3):423-39 PMID: 9199410
  51. Distinguishing homologous from analogous proteins.
    Syst Zool. 1970 Jun;19(2):99-113 PMID: 5449325
  52. Protein structure modeling for structural genomics.
    Nat Struct Biol. 2000 Nov;7 Suppl:986-90 PMID: 11104007
  53. Comparison of conformational characteristics in structurally similar protein pairs.
    Protein Sci. 1993 Nov;2(11):1811-26 PMID: 8268794
  54. Protein folds, functions and evolution.
    J Mol Biol. 1999 Oct 22;293(2):333-42 PMID: 10529349
  55. Quantitative sequence-function relationships in proteins based on gene ontology.
    BMC Bioinformatics. 2007 Aug 08;8:294 PMID: 17686158
  56. Quantitative assessment of relationship between sequence similarity and function similarity.
    BMC Genomics. 2007 Jul 09;8:222 PMID: 17620139
  57. Large-scale protein structure modeling of the Saccharomyces cerevisiae genome.
    Proc Natl Acad Sci U S A. 1998 Nov 10;95(23):13597-602 PMID: 9811845
  58. Target selection and deselection at the Berkeley Structural Genomics Center.
    Proteins. 2006 Feb 1;62(2):356-70 PMID: 16276528
  59. TM-align: a protein structure alignment algorithm based on the TM-score.
    Nucleic Acids Res. 2005 Apr 22;33(7):2302-9 PMID: 15849316
  60. OrthoMCL-DB: querying a comprehensive multi-species collection of ortholog groups.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D363-8 PMID: 16381887
  61. Large-scale comparison of protein sequence alignment algorithms with structure alignments.
    Proteins. 2000 Jul 1;40(1):6-22 PMID: 10813826
  62. Prediction of protein function from protein sequence and structure.
    Q Rev Biophys. 2003 Aug;36(3):307-40 PMID: 15029827
  63. Physics-based methods for studying protein-ligand interactions.
    Curr Opin Drug Discov Devel. 2007 May;10(3):325-31 PMID: 17554859
  64. Benchmarking ortholog identification methods using functional genomics data.
    Genome Biol. 2006;7(4):R31 PMID: 16613613
Article Info
Journal
Protein science : a publication of the Protein Society
Abbr.
Protein Sci
ISSN
1469-896X
Published
2009-06-00
Pages
1306-15
Language
English
Region
United States
NLM ID
9211750
PMCID
PMC2774440
Subset
IM
Grants
NIGMS NIH HHS · P01 GM071790 · United States
NIGMS NIH HHS · P01 GM 71790 · United States
NIGMS NIH HHS · R01 GM 54762 · United States
NIGMS NIH HHS · R01 GM 60595 · United States
NIGMS NIH HHS · R01 GM 61267 · United States
NIGMS NIH HHS · U54 GM 074945 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]