Home LiteratureArticle Details
PMID: 20195253 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Review

Visualization of multiple alignments, phylogenies and gene family evolution.

Nature methods ·Vol. 7 ·No. 3 Suppl ·2010-03-00 ·Pages S16-25

Procter JB, Thompson J, Letunic I, Creevey C, Jossinet F, Barton GJ

Abstract

Software for visualizing sequence alignments and trees are essential tools for life scientists. In this review, we describe the major features and capabilities of a selection of stand-alone and web-based applications useful when investigating the function and evolution of a gene family. These range from simple viewers, to systems that provide sophisticated editing and analysis functions. We conclude with a discussion of the challenges that these tools now face due to the flood of next generation sequence data and the increasingly complex network of bioinformatics information sources.

MeSH Terms
Amino Acid Sequence Biological Evolution Internet Molecular Sequence Data Multigene Family Phylogeny Proteins/chemistry Sequence Alignment Sequence Homology, Amino Acid
Chemicals
Proteins
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Procter James B
University of Dundee, UK. [email protected]
Thompson Julie
Letunic Ivica
Creevey Chris
Jossinet Fabrice
Barton Geoffrey J
References (83)
83 references, click to expand
  1. Annotation and visualization of endogenous retroviral sequences using the Distributed Annotation System (DAS) and eBioX.
    BMC Bioinformatics. 2009 Jun 16;10 Suppl 6:S18 PMID: 19534743
  2. JOY: protein sequence-structure representation and analysis.
    Bioinformatics. 1998;14(7):617-23 PMID: 9730927
  3. NOTUNG: a program for dating gene duplications and optimizing gene family trees.
    J Comput Biol. 2000;7(3-4):429-47 PMID: 11108472
  4. Characterization of pairwise and multiple sequence alignment errors.
    Gene. 2009 Jul 15;441(1-2):141-7 PMID: 18614299
  5. Dendroscope: An interactive viewer for large phylogenetic trees.
    BMC Bioinformatics. 2007 Nov 22;8:460 PMID: 18034891
  6. The GOA database in 2009--an integrated Gene Ontology Annotation resource.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D396-403 PMID: 18957448
  7. ConStruct: Improved construction of RNA consensus structures.
    BMC Bioinformatics. 2008 Apr 28;9:219 PMID: 18442401
  8. Ensemble approach to predict specificity determinants: benchmarking and validation.
    BMC Bioinformatics. 2009 Jul 02;10:207 PMID: 19573245
  9. The ConSurf-DB: pre-calculated evolutionary conservation profiles of protein structures.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D323-7 PMID: 18971256
  10. Integrating sequence and structural biology with DAS.
    BMC Bioinformatics. 2007 Sep 12;8:333 PMID: 17850653
  11. Protein sequence alignments: a strategy for the hierarchical analysis of residue conservation.
    Comput Appl Biosci. 1993 Dec;9(6):745-56 PMID: 8143162
  12. INTREPID--INformation-theoretic TREe traversal for Protein functional site IDentification.
    Bioinformatics. 2008 Nov 1;24(21):2445-52 PMID: 18776193
  13. Sequence to Structure (S2S): display, manipulate and interconnect RNA data from sequence to structure.
    Bioinformatics. 2005 Aug 1;21(15):3320-1 PMID: 15905274
  14. The classification of amino acid conservation.
    J Theor Biol. 1986 Mar 21;119(2):205-18 PMID: 3461222
  15. CHROMA: consensus-based colouring of multiple alignments for publication.
    Bioinformatics. 2001 Sep;17(9):845-6 PMID: 11590103
  16. The nucleic acid database. A comprehensive relational database of three-dimensional structures of nucleic acids.
    Biophys J. 1992 Sep;63(3):751-9 PMID: 1384741
  17. Evolutionary trees from DNA sequences: a maximum likelihood approach.
    J Mol Evol. 1981;17(6):368-76 PMID: 7288891
  18. JEvTrace: refinement and variations of the evolutionary trace in JAVA.
    Genome Biol. 2002;3(12):RESEARCH0077 PMID: 12537566
  19. VISSA: a program to visualize structural features from structure sequence alignment.
    Bioinformatics. 2006 Apr 1;22(7):887-8 PMID: 16434438
  20. PFAAT version 2.0: a tool for editing, annotating, and analyzing multiple sequence alignments.
    BMC Bioinformatics. 2007 Oct 11;8:381 PMID: 17931421
  21. TRANSFAC and its module TRANSCompel: transcriptional gene regulation in eukaryotes.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D108-10 PMID: 16381825
  22. Visualizing large hierarchical clusters in hyperbolic space.
    Bioinformatics. 2000 Jul;16(7):660-1 PMID: 11038340
  23. ALSCRIPT: a tool to format multiple sequence alignments.
    Protein Eng. 1993 Jan;6(1):37-40 PMID: 8433969
  24. Visualising biological data: a semantic approach to tool and database integration.
    BMC Bioinformatics. 2009 Jun 16;10 Suppl 6:S19 PMID: 19534744
  25. The Sequence Ontology: a tool for the unification of genome annotations.
    Genome Biol. 2005;6(5):R44 PMID: 15892872
  26. Treevolution: visual analysis of phylogenetic trees.
    Bioinformatics. 2009 Aug 1;25(15):1970-1 PMID: 19470585
  27. T-Coffee: A novel method for fast and accurate multiple sequence alignment.
    J Mol Biol. 2000 Sep 8;302(1):205-17 PMID: 10964570
  28. Database resources of the National Center for Biotechnology Information.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D5-15 PMID: 18940862
  29. Application of phylogenetic networks in evolutionary studies.
    Mol Biol Evol. 2006 Feb;23(2):254-67 PMID: 16221896
  30. The Protein Data Bank.
    Nucleic Acids Res. 2000 Jan 1;28(1):235-42 PMID: 10592235
  31. Twenty Years of Delila and Molecular Information Theory: The Altenberg-Austin Workshop in Theoretical Biology Biological Information, Beyond Metaphor: Causality, Explanation, and Unification Altenberg, Austria, 11-14 July 2002.
    Biol Theory. 2006;1(3):250-260 PMID: 18084638
  32. PhyloWidget: web-based visualizations for the tree of life.
    Bioinformatics. 2008 Jul 15;24(14):1641-2 PMID: 18487241
  33. Detecting species-site dependencies in large multiple sequence alignments.
    Nucleic Acids Res. 2009 Oct;37(18):5959-68 PMID: 19661281
  34. Universally conserved positions in protein folds: reading evolutionary signals about stability, folding kinetics and function.
    J Mol Biol. 1999 Aug 6;291(1):177-96 PMID: 10438614
  35. Circos: an information aesthetic for comparative genomics.
    Genome Res. 2009 Sep;19(9):1639-45 PMID: 19541911
  36. A novel method for multiple alignment of sequences with repeated and shuffled elements.
    Genome Res. 2004 Nov;14(11):2336-46 PMID: 15520295
  37. CINEMA-MX: a modular multiple alignment editor.
    Bioinformatics. 2002 Oct;18(10):1402-3 PMID: 12376388
  38. The TRANSFAC project as an example of framework technology that supports the analysis of genomic regulation.
    Brief Bioinform. 2008 Jul;9(4):326-32 PMID: 18436575
  39. CTree: comparison of clusters between phylogenetic trees made easy.
    Bioinformatics. 2007 Nov 1;23(21):2952-3 PMID: 17717036
  40. Analysis and prediction of functionally important sites in proteins.
    Protein Sci. 2007 Jan;16(1):4-13 PMID: 17192586
  41. Interactive Tree Of Life (iTOL): an online tool for phylogenetic tree display and annotation.
    Bioinformatics. 2007 Jan 1;23(1):127-8 PMID: 17050570
  42. Visualising very large phylogenetic trees in three dimensional hyperbolic space.
    BMC Bioinformatics. 2004 Apr 29;5:48 PMID: 15117420
  43. TOPALi v2: a rich graphical interface for evolutionary analyses of multiple alignments on HPC clusters and multi-core desktops.
    Bioinformatics. 2009 Jan 1;25(1):126-7 PMID: 18984599
  44. Scoring residue conservation.
    Proteins. 2002 Aug 1;48(2):227-41 PMID: 12112692
  45. Multiple sequence alignment.
    Curr Opin Struct Biol. 2006 Jun;16(3):368-73 PMID: 16679011
  46. TreeDyn: towards dynamic graphics and annotations for analyses of trees.
    BMC Bioinformatics. 2006 Oct 10;7:439 PMID: 17032440
  47. The 20 years of PROSITE.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D245-9 PMID: 18003654
  48. Reconciliation with non-binary species trees.
    J Comput Biol. 2008 Oct;15(8):981-1006 PMID: 18808330
  49. The RNA structure alignment ontology.
    RNA. 2009 Sep;15(9):1623-31 PMID: 19622678
  50. The Pfam protein families database.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D281-8 PMID: 18039703
  51. Jalview Version 2--a multiple sequence alignment editor and analysis workbench.
    Bioinformatics. 2009 May 1;25(9):1189-91 PMID: 19151095
  52. A method to predict functional residues in proteins.
    Nat Struct Biol. 1995 Feb;2(2):171-8 PMID: 7749921
  53. The Universal Protein Resource (UniProt) 2009.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D169-74 PMID: 18836194
  54. OXBench: a benchmark for evaluation of protein multiple sequence alignment accuracy.
    BMC Bioinformatics. 2003 Oct 10;4:47 PMID: 14552658
  55. Multiple sequence alignment using ClustalW and ClustalX.
    Curr Protoc Bioinformatics. 2002 Aug;Chapter 2:Unit 2.3 PMID: 18792934
  56. Profile hidden Markov models.
    Bioinformatics. 1998;14(9):755-63 PMID: 9918945
  57. ATV: display and manipulation of annotated phylogenetic trees.
    Bioinformatics. 2001 Apr;17(4):383-4 PMID: 11301314
  58. Sequence logos: a new way to display consensus sequences.
    Nucleic Acids Res. 1990 Oct 25;18(20):6097-100 PMID: 2172928
  59. Bayesian inference of phylogeny and its impact on evolutionary biology.
    Science. 2001 Dec 14;294(5550):2310-4 PMID: 11743192
  60. Joint evolutionary trees: a large-scale method to predict protein interfaces based on sequence sampling.
    PLoS Comput Biol. 2009 Jan;5(1):e1000267 PMID: 19165315
  61. The Gene Ontology's Reference Genome Project: a unified framework for functional annotation across species.
    PLoS Comput Biol. 2009 Jul;5(7):e1000431 PMID: 19578431
  62. Vector NTI, a balanced all-in-one sequence analysis suite.
    Brief Bioinform. 2004 Dec;5(4):378-88 PMID: 15606974
  63. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  64. Prediction of protein secondary structure and active sites using the alignment of homologous sequences.
    J Mol Biol. 1987 Jun 20;195(4):957-61 PMID: 3656439
  65. Synchronous visual analysis and editing of RNA sequence and secondary structure alignments using 4SALE.
    BMC Res Notes. 2008 Oct 14;1:91 PMID: 18854023
  66. The Protein Feature Ontology: a tool for the unification of protein feature annotations.
    Bioinformatics. 2008 Dec 1;24(23):2767-72 PMID: 18936051
  67. ESPript/ENDscript: Extracting and rendering sequence and 3D information from atomic structures of proteins.
    Nucleic Acids Res. 2003 Jul 1;31(13):3320-3 PMID: 12824317
  68. WWW-query: an on-line retrieval system for biological sequence banks.
    Biochimie. 1996;78(5):364-9 PMID: 8905155
  69. Approaches to comparative sequence analysis: towards a functional view of vertebrate genomes.
    Nat Rev Genet. 2008 Apr;9(4):303-13 PMID: 18347593
  70. HotSwap for bioinformatics: a STRAP tutorial.
    BMC Bioinformatics. 2006 Feb 09;7:64 PMID: 16469097
  71. SEAVIEW and PHYLO_WIN: two graphic tools for sequence alignment and molecular phylogeny.
    Comput Appl Biosci. 1996 Dec;12(6):543-8 PMID: 9021275
  72. Semiautomated improvement of RNA alignments.
    RNA. 2007 Nov;13(11):1850-9 PMID: 17804647
  73. NCBI BLAST: a better web interface.
    Nucleic Acids Res. 2008 Jul 1;36(Web Server issue):W5-9 PMID: 18440982
  74. Structural interpretation of mutations and SNPs using STRAP-NT.
    Protein Sci. 2006 Jan;15(1):208-10 PMID: 16322575
  75. MacVector. Integrated sequence analysis for the Macintosh.
    Methods Mol Biol. 2000;132:47-69 PMID: 10547831
  76. MEGA: a biologist-centric software for evolutionary analysis of DNA and protein sequences.
    Brief Bioinform. 2008 Jul;9(4):299-306 PMID: 18417537
  77. MACSIMS: multiple alignment of complete sequences information management system.
    BMC Bioinformatics. 2006 Jun 23;7:318 PMID: 16792820
  78. Visualizing profile-profile alignment: pairwise HMM logos.
    Bioinformatics. 2005 Jun 15;21(12):2912-3 PMID: 15827079
  79. Amino acid encoding schemes from protein structure alignments: multi-dimensional vectors to describe residue types.
    J Theor Biol. 2002 Jun 7;216(3):361-65 PMID: 12183124
  80. Correlated substitution analysis and the prediction of amino acid structural contacts.
    Brief Bioinform. 2008 Jan;9(1):46-56 PMID: 18000015
  81. TreeView: an application to display phylogenetic trees on personal computers.
    Comput Appl Biosci. 1996 Aug;12(4):357-8 PMID: 8902363
  82. Phylogeny estimation: traditional and Bayesian approaches.
    Nat Rev Genet. 2003 Apr;4(4):275-84 PMID: 12671658
  83. MEME SUITE: tools for motif discovery and searching.
    Nucleic Acids Res. 2009 Jul;37(Web Server issue):W202-8 PMID: 19458158
Article Info
Journal
Nature methods
Abbr.
Nat Methods
ISSN
1548-7105
Published
2010-03-00
Pages
S16-25
Language
English
Region
United States
NLM ID
101215604
Subset
IM
Grants
Biotechnology and Biological Sciences Research Council · BB/G022682/1 · United Kingdom
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]