Home LiteratureArticle Details
PMID: 18687127 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

Performing statistical analyses on quantitative data in Taverna workflows: an example using R and maxdBrowse to identify differentially-expressed genes from microarray data.

BMC bioinformatics ·Vol. 9 ·2008-08-07 ·Pages 334

Li P, Castrillo JI, Velarde G, Wassink I, Soiland-Reyes S, Owen S, Withers D, Oinn T, Pocock MR, Goble CA, Oliver SG, Kell DB

Abstract

There has been a dramatic increase in the amount of quantitative data derived from the measurement of changes at different levels of biological complexity during the post-genomic era. However, there are a number of issues associated with the use of computational tools employed for the analysis of such data. For example, computational tools such as R and MATLAB require prior knowledge of their programming languages in order to implement statistical analyses on data. Combining two or more tools in an analysis may also be problematic since data may have to be manually copied and pasted between separate user interfaces for each tool. Furthermore, this transfer of data may require a reconciliation step in order for there to be interoperability between computational tools. Developments in the Taverna workflow system have enabled pipelines to be constructed and enacted for generic and ad hoc analyses of quantitative data. Here, we present an example of such a workflow involving the statistical identification of differentially-expressed genes from microarray data followed by the annotation of their relationships to cellular processes. This workflow makes use of customised maxdBrowse web services, a system that allows Taverna to query and retrieve gene expression data from the maxdLoad2 microarray database. These data are then analysed by R to identify differentially-expressed genes using the Taverna RShell processor which has been developed for invoking this tool when it has been deployed as a service using the RServe library. In addition, the workflow uses Beanshell scripts to reconcile mismatches of data between services as well as to implement a form of user interaction for selecting subsets of microarray data for analysis as part of the workflow execution. A new plugin system in the Taverna software architecture is demonstrated by the use of renderers for displaying PDF files and CSV formatted data within the Taverna workbench. Taverna can be used by data analysis experts as a generic tool for composing ad hoc analyses of quantitative data by combining the use of scripts written in the R programming language with tools exposed as services in workflows. When these workflows are shared with colleagues and the wider scientific community, they provide an approach for other scientists wanting to use tools such as R without having to learn the corresponding programming language to analyse their own data.

MeSH Terms
Data Interpretation, Statistical Databases, Genetic Gene Expression Profiling/statistics & numerical data Information Storage and Retrieval Oligonucleotide Array Sequence Analysis/statistics & numerical data Programming Languages Software
Authors & Affiliations
12 authors, click to expand affiliations / ORCID
Li Peter
Manchester Centre for Integrative Systems Biology and School of Chemistry, Manchester Interdisciplinary Biocentre, University of Manchester, 131 Princess St, Manchester, M1 7DN, UK. [email protected]
Castrillo Juan I
Velarde Giles
Wassink Ingo
Soiland-Reyes Stian
Owen Stuart
Withers David
Oinn Tom
Pocock Matthew R
Goble Carole A
Oliver Stephen G
Kell Douglas B
References (15)
15 references, click to expand
  1. Multiple high-throughput analyses monitor the response of E. coli to perturbations.
    Science. 2007 Apr 27;316(5824):593-7 PMID: 17379776
  2. maxdLoad2 and maxdBrowse: standards-compliant tools for microarray experimental annotation, data management and dissemination.
    BMC Bioinformatics. 2005 Nov 03;6:264 PMID: 16269077
  3. Here is the evidence, now what is the hypothesis? The complementary roles of inductive and hypothesis-driven science in the post-genomic era.
    Bioessays. 2004 Jan;26(1):99-105 PMID: 14696046
  4. State of the nation in data integration for bioinformatics.
    J Biomed Inform. 2008 Oct;41(5):687-93 PMID: 18358788
  5. GO::TermFinder--open source software for accessing Gene Ontology information and finding significantly enriched Gene Ontology terms associated with a list of genes.
    Bioinformatics. 2004 Dec 12;20(18):3710-5 PMID: 15297299
  6. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  7. Taverna: a tool for the composition and enactment of bioinformatics workflows.
    Bioinformatics. 2004 Nov 22;20(17):3045-54 PMID: 15201187
  8. oneChannelGUI: a graphical interface to Bioconductor tools, designed for life scientists who are not familiar with R language.
    Bioinformatics. 2007 Dec 15;23(24):3406-8 PMID: 17875544
  9. Minimum information about a microarray experiment (MIAME)-toward standards for microarray data.
    Nat Genet. 2001 Dec;29(4):365-71 PMID: 11726920
  10. GenePublisher: Automated analysis of DNA microarray data.
    Nucleic Acids Res. 2003 Jul 1;31(13):3471-6 PMID: 12824347
  11. Growth control of the eukaryote cell: a systems biology study in yeast.
    J Biol. 2007;6(2):4 PMID: 17439666
  12. BioMoby extensions to the Taverna workflow management and enactment software.
    BMC Bioinformatics. 2006 Nov 30;7:523 PMID: 17137515
  13. Taverna: a tool for building and running workflows of services.
    Nucleic Acids Res. 2006 Jul 1;34(Web Server issue):W729-32 PMID: 16845108
  14. A systematic strategy for large-scale analysis of genotype phenotype correlations: identification of candidate genes involved in African trypanosomiasis.
    Nucleic Acids Res. 2007;35(16):5625-33 PMID: 17709344
  15. BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis.
    Bioinformatics. 2005 Aug 15;21(16):3439-40 PMID: 16082012
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2008-08-07
Epub
2008-00-07
Pages
334
Language
English
Region
England
NLM ID
100965194
PMCID
PMC2528018
Subset
IM
Grants
Biotechnology and Biological Sciences Research Council · G18886 · United Kingdom
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]