Home LiteratureArticle Details
PMID: 15130933 Published · ppublish English Comparative Study Evaluation Study Journal Article Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S.

A statistical framework for genomic data fusion.

Bioinformatics (Oxford, England) ·Vol. 20 ·No. 16 ·2004-11-01 ·Pages 2626-35

Lanckriet GR, De Bie T, Cristianini N, Jordan MI, Noble WS

Abstract

During the past decade, the new focus on genomics has highlighted a particular challenge: to integrate the different views of the genome that are provided by various types of experimental data. This paper describes a computational framework for integrating and drawing inferences from a collection of genome-wide measurements. Each dataset is represented via a kernel function, which defines generalized similarity relationships between pairs of entities, such as genes or proteins. The kernel representation is both flexible and efficient, and can be applied to many different types of data. Furthermore, kernel functions derived from different types of data can be combined in a straightforward fashion. Recent advances in the theory of kernel methods have provided efficient algorithms to perform such combinations in a way that minimizes a statistical loss function. These methods exploit semidefinite programming techniques to reduce the problem of finding optimizing kernel combinations to a convex optimization problem. Computational experiments performed using yeast genome-wide datasets, including amino acid sequences, hydropathy profiles, gene expression data and known protein-protein interactions, demonstrate the utility of this approach. A statistical learning algorithm trained from all of these data to recognize particular classes of proteins--membrane proteins and ribosomal proteins--performs significantly better than the same algorithm trained on any single type of data. Supplementary data at http://noble.gs.washington.edu/proj/sdp-svm

MeSH Terms
Algorithms Artificial Intelligence Chromosome Mapping/methods Databases, Genetic Databases, Protein Fungal Proteins/chemistry,genetics Gene Expression Profiling/methods Genomics/methods Information Storage and Retrieval/methods Membrane Proteins/genetics,metabolism Models, Genetic Models, Statistical Pattern Recognition, Automated Proteins/analysis,chemistry,classification,genetics Ribosomal Proteins/chemistry,genetics Sequence Alignment Sequence Analysis, Protein/methods Sequence Homology, Amino Acid Systems Integration
Chemicals
Fungal Proteins Membrane Proteins Proteins Ribosomal Proteins
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Lanckriet Gert R G
Department of Electrical Engineering and Computer Science, University of California, Berkeley 94720, USA.
De Bie Tijl
Cristianini Nello
Jordan Michael I
Noble William Stafford
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4803
Published
2004-11-01
Epub
2004-00-06
Pages
2626-35
Language
English
Region
England
NLM ID
9808944
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]