Home LiteratureArticle Details
PMID: 15985156 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

Multiple sequence alignments of partially coding nucleic acid sequences.

BMC bioinformatics ·Vol. 6 ·2005-06-28 ·Pages 160

Stocsits RR, Hofacker IL, Fried C, Stadler PF

Abstract

High quality sequence alignments of RNA and DNA sequences are an important prerequisite for the comparative analysis of genomic sequence data. Nucleic acid sequences, however, exhibit a much larger sequence heterogeneity compared to their encoded protein sequences due to the redundancy of the genetic code. It is desirable, therefore, to make use of the amino acid sequence when aligning coding nucleic acid sequences. In many cases, however, only a part of the sequence of interest is translated. On the other hand, overlapping reading frames may encode multiple alternative proteins, possibly with intermittent non-coding parts. Examples are, in particular, RNA virus genomes. The standard scoring scheme for nucleic acid alignments can be extended to incorporate simultaneously information on translation products in one or more reading frames. Here we present a multiple alignment tool, codaln, that implements a combined nucleic acid plus amino acid scoring model for pairwise and progressive multiple alignments that allows arbitrary weighting for almost all scoring parameters. Resource requirements of codaln are comparable with those of standard tools such as ClustalW. We demonstrate the applicability of codaln to various biologically relevant types of sequences (bacteriophage Levivirus and Vertebrate Hox clusters) and show that the combination of nucleic acid and amino acid sequence information leads to improved alignments. These, in turn, increase the performance of analysis tools that depend strictly on good input alignments such as methods for detecting conserved RNA secondary structure elements.

MeSH Terms
Algorithms Amino Acid Sequence Codon Conserved Sequence Homeodomain Proteins/chemistry Levivirus/genetics Models, Molecular Open Reading Frames Protein Structure, Secondary RNA/chemistry Reproducibility of Results Sequence Alignment/methods Sequence Analysis, DNA/methods Sequence Homology, Nucleic Acid
Chemicals
Codon Homeodomain Proteins RNA
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Stocsits Roman R
Interdisciplinary Centre for Bioinformatics, University of Leipzig, Haertelstrasse 16-18, D-04107 Leipzig, Germany. [email protected]
Hofacker Ivo L
Fried Claudia
Stadler Peter F
References (34)
34 references, click to expand
  1. Ancient origin of the Hox gene cluster.
    Nat Rev Genet. 2001 Jan;2(1):33-8 PMID: 11253066
  2. Expression of alternatively spliced FGF-2 antisense RNA transcripts in the central nervous system: regulation of FGF-2 mRNA translation.
    Mol Cell Endocrinol. 2000 Dec 22;170(1-2):233-42 PMID: 11162906
  3. Reconstructing the origins of human hepatitis viruses.
    Philos Trans R Soc Lond B Biol Sci. 2001 Jul 29;356(1411):1013-26 PMID: 11516379
  4. Conserved RNA secondary structures in Picornaviridae genomes.
    Nucleic Acids Res. 2001 Dec 15;29(24):5079-89 PMID: 11812840
  5. Purifying and directional selection in overlapping prokaryotic genes.
    Trends Genet. 2002 May;18(5):228-32 PMID: 12047938
  6. Influenza virus still surprises.
    Curr Opin Microbiol. 2002 Aug;5(4):414-8 PMID: 12160862
  7. Computational discovery of sense-antisense transcription in the human and mouse genomes.
    Genome Biol. 2002 Aug 22;3(9):RESEARCH0044 PMID: 12225583
  8. The virological and clinical significance of mutations in the overlapping envelope and polymerase genes of hepatitis B virus.
    J Clin Virol. 2002 Aug;25(2):97-106 PMID: 12367644
  9. Widespread occurrence of antisense transcription in the human genome.
    Nat Biotechnol. 2003 Apr;21(4):379-86 PMID: 12640466
  10. Molecular biology of umbraviruses: phantom warriors.
    J Gen Virol. 2003 Aug;84(Pt 8):1951-60 PMID: 12867625
  11. Noncoding RNA gene detection using comparative sequence analysis.
    BMC Bioinformatics. 2001;2:8 PMID: 11801179
  12. Gene fusion and overlapping reading frames in the mammalian genes for 4E-BP3 and MASK.
    J Biol Chem. 2003 Dec 26;278(52):52290-7 PMID: 14557257
  13. Conserved RNA secondary structures in Flaviviridae genomes.
    J Gen Virol. 2004 May;85(Pt 5):1113-24 PMID: 15105528
  14. Conserved RNA secondary structures in viral genomes: a survey.
    Bioinformatics. 2004 Jul 10;20(10):1495-9 PMID: 15231541
  15. A comparative method for finding and folding RNA secondary structures within protein-coding regions.
    Nucleic Acids Res. 2004;32(16):4925-36 PMID: 15448187
  16. Panhandles and hairpin structures at the termini of germiston virus RNAs (Bunyavirus).
    Virology. 1982 Oct 15;122(1):191-7 PMID: 7135833
  17. An improved algorithm for matching biological sequences.
    J Mol Biol. 1982 Dec 15;162(3):705-8 PMID: 7166760
  18. Conserved elements in the 3' untranslated region of flavivirus RNAs and potential cyclization sequences.
    J Mol Biol. 1987 Nov 5;198(1):33-41 PMID: 2828633
  19. Two steps in the evolution of Antennapedia-class vertebrate homeobox genes.
    Proc Natl Acad Sci U S A. 1989 Jul;86(14):5459-63 PMID: 2568634
  20. Homeobox genes and axial patterning.
    Cell. 1992 Jan 24;68(2):283-302 PMID: 1346368
  21. In vitro recombination and terminal elongation of RNA by Q beta replicase.
    EMBO J. 1992 Dec;11(13):5129-35 PMID: 1281452
  22. Sequence analysis of RNA species synthesized by Q beta replicase without template.
    Biochemistry. 1993 May 11;32(18):4848-54 PMID: 7683911
  23. An algorithm combining DNA and protein alignment.
    J Theor Biol. 1994 Mar 21;167(2):169-74 PMID: 8207946
  24. From sequences to shapes and back: a case study in RNA secondary structures.
    Proc Biol Sci. 1994 Mar 22;255(1344):279-84 PMID: 7517565
  25. Archetypal organization of the amphioxus Hox gene cluster.
    Nature. 1994 Aug 18;370(6490):563-6 PMID: 7914353
  26. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  27. Combined DNA and protein alignment.
    Methods Enzymol. 1996;266:402-18 PMID: 8743696
  28. Automatic detection of conserved RNA structure elements in complete RNA virus genomes.
    Nucleic Acids Res. 1998 Aug 15;26(16):3825-36 PMID: 9685502
  29. BAliBASE: a benchmark alignment database for the evaluation of multiple alignment programs.
    Bioinformatics. 1999 Jan;15(1):87-8 PMID: 10068696
  30. Automatic detection of conserved base pairing patterns in RNA virus genomes.
    Comput Chem. 1999 Jun 15;23(3-4):401-14 PMID: 10404627
  31. Properties of overlapping genes are conserved across microbial genomes.
    Genome Res. 2004 Nov;14(11):2268-72 PMID: 15520290
  32. Fast and reliable prediction of noncoding RNAs.
    Proc Natl Acad Sci U S A. 2005 Feb 15;102(7):2454-9 PMID: 15665081
  33. The hepatitis B virus pregenome: prediction of RNA structure and implications for the emergence of deletions.
    Intervirology. 2000;43(3):154-64 PMID: 11044809
  34. Two overlapping reading frames in a single exon encode interacting proteins--a novel way of gene usage.
    EMBO J. 2001 Jul 16;20(14):3849-60 PMID: 11447126
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2005-06-28
Epub
2005-00-28
Pages
160
Language
English
Region
England
NLM ID
100965194
PMCID
PMC1182351
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]