Abstract
Dramatic increases in the throughput of nucleotide sequencing machines, and the promise of ever greater performance, have thrust bioinformatics into the era of petabyte-scale data sets. Sequence repositories, which provide the feed for these data sets into the worldwide computational infrastructure, are challenged by the impact of these data volumes. The European Nucleotide Archive (ENA; http://www.ebi.ac.uk/embl), comprising the EMBL Nucleotide Sequence Database and the Ensembl Trace Archive, has identified challenges in the storage, movement, analysis, interpretation and visualization of petabyte-scale data sets. We present here our new repository for next generation sequence data, a brief summary of contents of the ENA and provide details of major developments to submission pipelines, high-throughput rule-based validation infrastructure and data integration approaches.
MeSH Terms
Databases, Nucleic Acid
Internet
Sequence Analysis/trends
Systems Integration
Authors & Affiliations
27 authors, click to expand affiliations / ORCID
Cochrane Guy
EMBL-European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge CB10 1SD, UK.
[email protected]
Akhtar Ruth
Bonfield James
Bower Lawrence
Demiralp Fehmi
Faruque Nadeem
Gibson Richard
Hoad Gemma
Hubbard Tim
Hunter Christopher
Jang Mikyung
Juhos Szilveszter
Leinonen Rasko
Leonard Steven
Lin Quan
Lopez Rodrigo
Lorenc Dariusz
McWilliam Hamish
Mukherjee Gaurab
Plaister Sheila
Radhakrishnan Rajesh
Robinson Stephen
Sobhany Siamak
Hoopen Petra Ten
Vaughan Robert
Zalunin Vadim
Birney Ewan
References (12)
12 references, click to expand
-
Comparative analysis of Acinetobacters: three genomes for three lifestyles.
PLoS One. 2008;3(3):e1805
PMID: 18350144
-
The minimum information about a genome sequence (MIGS) specification.
Nat Biotechnol. 2008 May;26(5):541-7
PMID: 18464787
-
High-throughput sequencing provides insights into genome variation and evolution in Salmonella Typhi.
Nat Genet. 2008 Aug;40(8):987-93
PMID: 18660809
-
ArrayExpress--a public database of microarray experiments and gene expression profiles.
Nucleic Acids Res. 2007 Jan;35(Database issue):D747-50
PMID: 17132828
-
DDBJ with new system and face.
Nucleic Acids Res. 2008 Jan;36(Database issue):D22-4
PMID: 17962300
-
The HGNC Database in 2008: a resource for the human genome.
Nucleic Acids Res. 2008 Jan;36(Database issue):D445-8
PMID: 17984084
-
Ensembl 2008.
Nucleic Acids Res. 2008 Jan;36(Database issue):D707-14
PMID: 18000006
-
Priorities for nucleotide trace, sequence and annotation data capture at the Ensembl Trace Archive and the EMBL Nucleotide Sequence Database.
Nucleic Acids Res. 2008 Jan;36(Database issue):D5-12
PMID: 18039715
-
The universal protein resource (UniProt).
Nucleic Acids Res. 2008 Jan;36(Database issue):D190-5
PMID: 18045787
-
Database resources of the National Center for Biotechnology Information.
Nucleic Acids Res. 2008 Jan;36(Database issue):D13-21
PMID: 18045790
-
GenBank.
Nucleic Acids Res. 2008 Jan;36(Database issue):D25-30
PMID: 18073190
-
The Mouse Genome Database (MGD): mouse biology and model systems.
Nucleic Acids Res. 2008 Jan;36(Database issue):D724-8
PMID: 18158299