Home LiteratureArticle Details
PMID: 16914055 Published · epublish English Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, Non-P.H.S.

MultiSeq: unifying sequence and structure data for evolutionary analysis.

BMC bioinformatics ·Vol. 7 ·2006-08-16 ·Pages 382

Roberts E, Eargle J, Wright D, Luthey-Schulten Z

Abstract

Since the publication of the first draft of the human genome in 2000, bioinformatic data have been accumulating at an overwhelming pace. Currently, more than 3 million sequences and 35 thousand structures of proteins and nucleic acids are available in public databases. Finding correlations in and between these data to answer critical research questions is extremely challenging. This problem needs to be approached from several directions: information science to organize and search the data; information visualization to assist in recognizing correlations; mathematics to formulate statistical inferences; and biology to analyze chemical and physical properties in terms of sequence and structure changes. Here we present MultiSeq, a unified bioinformatics analysis environment that allows one to organize, display, align and analyze both sequence and structure data for proteins and nucleic acids. While special emphasis is placed on analyzing the data within the framework of evolutionary biology, the environment is also flexible enough to accommodate other usage patterns. The evolutionary approach is supported by the use of predefined metadata, adherence to standard ontological mappings, and the ability for the user to adjust these classifications using an electronic notebook. MultiSeq contains a new algorithm to generate complete evolutionary profiles that represent the topology of the molecular phylogenetic tree of a homologous group of distantly related proteins. The method, based on the multidimensional QR factorization of multiple sequence and structure alignments, removes redundancy from the alignments and orders the protein sequences by increasing linear dependence, resulting in the identification of a minimal basis set of sequences that spans the evolutionary space of the homologous group of proteins. MultiSeq is a major extension of the Multiple Alignment tool that is provided as part of VMD, a structural visualization program for analyzing molecular dynamics simulations. Both are freely distributed by the NIH Resource for Macromolecular Modeling and Bioinformatics and MultiSeq is included with VMD starting with version 1.8.5. The MultiSeq website has details on how to download and use the software: http://www.scs.uiuc.edu/~schulten/multiseq/

MeSH Terms
Algorithms Computational Biology/methods Computer Graphics Computer Simulation DNA/chemistry Databases, Genetic Evolution, Molecular Models, Molecular Nucleic Acid Conformation Phylogeny Protein Conformation Proteins/chemistry,classification,genetics Sequence Alignment Sequence Analysis, DNA Sequence Analysis, Protein Sequence Homology, Nucleic Acid Software Structural Homology, Protein
Chemicals
Proteins DNA
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Roberts Elijah
Center for Biophysics and Computational Biology, University of Illinois at Urbana-Champaign, Urbana, IL, USA. [email protected]
Eargle John
Wright Dan
Luthey-Schulten Zaida
References (57)
57 references, click to expand
  1. Aminoacyl-tRNA synthetases, the genetic code, and the evolutionary process.
    Microbiol Mol Biol Rev. 2000 Mar;64(1):202-36 PMID: 10704480
  2. Knowledge-based protein secondary structure assignment.
    Proteins. 1995 Dec;23(4):566-79 PMID: 8749853
  3. Comparative protein structure modeling of genes and genomes.
    Annu Rev Biophys Biomol Struct. 2000;29:291-325 PMID: 10940251
  4. T-Coffee: A novel method for fast and accurate multiple sequence alignment.
    J Mol Biol. 2000 Sep 8;302(1):205-17 PMID: 10964570
  5. Which craft is best in bioinformatics?
    Comput Chem. 2001 Jul;25(4):329-39 PMID: 11459349
  6. A standard reference frame for the description of nucleic acid base-pair geometry.
    J Mol Biol. 2001 Oct 12;313(1):229-37 PMID: 11601858
  7. Combining multiple structure and sequence alignments to improve sequence detection and alignment: application to the SH2 domains of Janus kinases.
    Proc Natl Acad Sci U S A. 2001 Dec 18;98(26):14796-801 PMID: 11752426
  8. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003.
    Nucleic Acids Res. 2003 Jan 1;31(1):365-70 PMID: 12520024
  9. The Ribosomal Database Project (RDP-II): previewing a new autoaligner that allows regular updates and the new prokaryotic taxonomy.
    Nucleic Acids Res. 2003 Jan 1;31(1):442-3 PMID: 12520046
  10. Protein family annotation in a multiple alignment viewer.
    Bioinformatics. 2003 Mar 1;19(4):544-5 PMID: 12611813
  11. MrBayes 3: Bayesian phylogenetic inference under mixed models.
    Bioinformatics. 2003 Aug 12;19(12):1572-4 PMID: 12912839
  12. A graph-theory algorithm for rapid protein side-chain prediction.
    Protein Sci. 2003 Sep;12(9):2001-14 PMID: 12930999
  13. A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood.
    Syst Biol. 2003 Oct;52(5):696-704 PMID: 14530136
  14. The comparative RNA web (CRW) site: an online database of comparative sequence and structure information for ribosomal, intron, and other RNAs.
    BMC Bioinformatics. 2002;3:2 PMID: 11869452
  15. Protein modelling for all.
    Trends Biochem Sci. 1999 Sep;24(9):364-7 PMID: 10470037
  16. The Ribosomal Database Project (RDP-II): sequences and tools for high-throughput rRNA analysis.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D294-6 PMID: 15608200
  17. GenBank.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D34-8 PMID: 15608212
  18. NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D501-4 PMID: 15608248
  19. Evolutionary profiles derived from the QR factorization of multiple structural alignments gives an economy of information.
    J Mol Biol. 2005 Feb 25;346(3):875-94 PMID: 15713469
  20. Evolutionary profiles from the QR factorization of multiple sequence alignments.
    Proc Natl Acad Sci U S A. 2005 Mar 15;102(11):4045-50 PMID: 15741270
  21. MollDE: a homology modeling framework you can click with.
    Bioinformatics. 2005 Jun 15;21(12):2914-6 PMID: 15845657
  22. Multiple sequence alignments.
    Curr Opin Struct Biol. 2005 Jun;15(3):261-6 PMID: 15963889
  23. The Diamond STING server.
    Nucleic Acids Res. 2005 Jul 1;33(Web Server issue):W29-35 PMID: 15980473
  24. Friend, an integrated analytical front-end application for bioinformatics.
    Bioinformatics. 2005 Sep 15;21(18):3677-8 PMID: 16076889
  25. Evolutionary information for specifying a protein fold.
    Nature. 2005 Sep 22;437(7058):512-8 PMID: 16177782
  26. Assessment of protein distance measures and tree-building methods for phylogenetic tree reconstruction.
    Mol Biol Evol. 2005 Nov;22(11):2257-64 PMID: 16049194
  27. The evolutionary history of Cys-tRNACys formation.
    Proc Natl Acad Sci U S A. 2005 Dec 27;102(52):19003-8 PMID: 16380427
  28. EMBL Nucleotide Sequence Database: developments in 2005.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D10-5 PMID: 16381823
  29. The integrated microbial genomes (IMG) system.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D344-8 PMID: 16381883
  30. DDBJ in preparation for overview of research activities behind data submissions.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D6-9 PMID: 16381940
  31. Multiple Alignment of protein structures and sequences for VMD.
    Bioinformatics. 2006 Feb 15;22(4):504-6 PMID: 16339280
  32. tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence.
    Nucleic Acids Res. 1997 Mar 1;25(5):955-64 PMID: 9023104
  33. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  34. CATH--a hierarchic classification of protein domain structures.
    Structure. 1997 Aug 15;5(8):1093-108 PMID: 9309224
  35. The CLUSTAL_X windows interface: flexible strategies for multiple sequence alignment aided by quality analysis tools.
    Nucleic Acids Res. 1997 Dec 15;25(24):4876-82 PMID: 9396791
  36. Compilation of tRNA sequences and sequences of tRNA genes.
    Nucleic Acids Res. 1998 Jan 1;26(1):148-53 PMID: 9399820
  37. CINEMA--a novel colour INteractive editor for multiple alignments.
    Gene. 1998 Oct 9;221(1):GC57-63 PMID: 9852962
  38. Profile hidden Markov models.
    Bioinformatics. 1998;14(9):755-63 PMID: 9918945
  39. Database resources of the National Center for Biotechnology Information.
    Nucleic Acids Res. 2000 Jan 1;28(1):10-4 PMID: 10592169
  40. The Protein Data Bank.
    Nucleic Acids Res. 2000 Jan 1;28(1):235-42 PMID: 10592235
  41. On the evolution of structure in aminoacyl-tRNA synthetases.
    Microbiol Mol Biol Rev. 2003 Dec;67(4):550-73 PMID: 14665676
  42. The ASTRAL Compendium in 2004.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D189-92 PMID: 14681391
  43. SCOP database in 2004: refinements integrate structure and sequence family data.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D226-9 PMID: 14681400
  44. The Jalview Java alignment editor.
    Bioinformatics. 2004 Feb 12;20(3):426-7 PMID: 14960472
  45. MEGA3: Integrated software for Molecular Evolutionary Genetics Analysis and sequence alignment.
    Brief Bioinform. 2004 Jun;5(2):150-63 PMID: 15260895
  46. UCSF Chimera--a visualization system for exploratory research and analysis.
    J Comput Chem. 2004 Oct;25(13):1605-12 PMID: 15264254
  47. Evolutionary trees from DNA sequences: a maximum likelihood approach.
    J Mol Evol. 1981;17(6):368-76 PMID: 7288891
  48. The relation between the divergence of sequence and structure in proteins.
    EMBO J. 1986 Apr;5(4):823-6 PMID: 3709526
  49. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  50. Multiple protein sequence alignment from tertiary structure comparison: assignment of global and residue confidence levels.
    Proteins. 1992 Oct;14(2):309-23 PMID: 1409577
  51. The nucleic acid database. A comprehensive relational database of three-dimensional structures of nucleic acids.
    Biophys J. 1992 Sep;63(3):751-9 PMID: 1384741
  52. Summary: the modified nucleosides of RNA.
    Nucleic Acids Res. 1994 Jun 25;22(12):2183-96 PMID: 7518580
  53. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  54. SCOP: a structural classification of proteins database for the investigation of sequences and structures.
    J Mol Biol. 1995 Apr 7;247(4):536-40 PMID: 7723011
  55. RASMOL: biomolecular graphics for all.
    Trends Biochem Sci. 1995 Sep;20(9):374 PMID: 7482707
  56. VMD: visual molecular dynamics.
    J Mol Graph. 1996 Feb;14(1):33-8, 27-8 PMID: 8744570
  57. Cn3D: sequence and structure views for Entrez.
    Trends Biochem Sci. 2000 Jun;25(6):300-2 PMID: 10838572
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2006-08-16
Epub
2006-00-16
Pages
382
Language
English
Region
England
NLM ID
100965194
PMCID
PMC1586216
Subset
IM
Grants
NCRR NIH HHS · P41 RR005969 · United States
NIGMS NIH HHS · T32 GM008276 · United States
NCRR NIH HHS · PHS2P41RR05969 · United States
NIGMS NIH HHS · PHS5T32GM08276 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]