Home LiteratureArticle Details
PMID: 27716030 Published · epublish English Journal Article

Handling missing rows in multi-omics data integration: multiple imputation in multiple factor analysis framework.

BMC bioinformatics ·Vol. 17 ·No. 1 ·2016-10-03 ·Pages 402

Voillet V, Besse P, Liaubet L, San Cristobal M, González I

Abstract

In omics data integration studies, it is common, for a variety of reasons, for some individuals to not be present in all data tables. Missing row values are challenging to deal with because most statistical methods cannot be directly applied to incomplete datasets. To overcome this issue, we propose a multiple imputation (MI) approach in a multivariate framework. In this study, we focus on multiple factor analysis (MFA) as a tool to compare and integrate multiple layers of information. MI involves filling the missing rows with plausible values, resulting in M completed datasets. MFA is then applied to each completed dataset to produce M different configurations (the matrices of coordinates of individuals). Finally, the M configurations are combined to yield a single consensus solution. We assessed the performance of our method, named MI-MFA, on two real omics datasets. Incomplete artificial datasets with different patterns of missingness were created from these data. The MI-MFA results were compared with two other approaches i.e., regularized iterative MFA (RI-MFA) and mean variable imputation (MVI-MFA). For each configuration resulting from these three strategies, the suitability of the solution was determined against the true MFA configuration obtained from the original data and a comprehensive graphical comparison showing how the MI-, RI- or MVI-MFA configurations diverge from the true configuration was produced. Two approaches i.e., confidence ellipses and convex hulls, to visualize and assess the uncertainty due to missing values were also described. We showed how the areas of ellipses and convex hulls increased with the number of missing individuals. A free and easy-to-use code was proposed to implement the MI-MFA method in the R statistical environment. We believe that MI-MFA provides a useful and attractive method for estimating the coordinates of individuals on the first MFA components despite missing rows. MI-MFA configurations were close to the true configuration even when many individuals were missing in several data tables. This method takes into account the uncertainty of MI-MFA configurations induced by the missing rows, thereby allowing the reliability of the results to be evaluated.

Keywords
Hot-deck imputation Missing individuals Multiple imputation Multiple omics data integration Multivariate factor analysis
MeSH Terms
Acetaminophen/toxicity Analgesics, Non-Narcotic/toxicity Animals Chemical and Drug Induced Liver Injury/etiology,metabolism,pathology Data Interpretation, Statistical Factor Analysis, Statistical Gene Expression Regulation/drug effects Genomics/methods Humans Male Multivariate Analysis Neoplasms/genetics,metabolism Proteomics/methods Rats Rats, Wistar Reproducibility of Results Tumor Cells, Cultured
Chemicals
Analgesics, Non-Narcotic Acetaminophen
Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Voillet Valentin
Université de Toulouse, INRA, INPT, INP-ENVT, UMR1388, GenPhySE, Castanet-Tolosan, F-31326, France.
Besse Philippe
Université de Toulouse INSA, UMR5219 Institut de Mathématiques, Toulouse, F-31077, France.
Liaubet Laurence
Université de Toulouse, INRA, INPT, INP-ENVT, UMR1388, GenPhySE, Castanet-Tolosan, F-31326, France.
San Cristobal Magali
Université de Toulouse, INRA, INPT, INP-ENVT, UMR1388, GenPhySE, Castanet-Tolosan, F-31326, France. | Université de Toulouse INSA, UMR5219 Institut de Mathématiques, Toulouse, F-31077, France.
González Ignacio ORCID
INRAUR875 Mathématiques et Informatiques Appliquées, F-31326, Castanet-Tolosan, France. [email protected].
References (10)
10 references, click to expand
  1. Proteomic profiling of the NCI-60 cancer cell lines using new high-density reverse-phase lysate microarrays.
    Proc Natl Acad Sci U S A. 2003 Nov 25;100(24):14229-34 PMID: 14623978
  2. Simultaneous clustering of gene expression data with clinical chemistry and pathological evaluations reveals phenotypic prototypes.
    BMC Syst Biol. 2007 Feb 23;1:15 PMID: 17408499
  3. Multiple imputation of discrete and continuous data by fully conditional specification.
    Stat Methods Med Res. 2007 Jun;16(3):219-42 PMID: 17621469
  4. Missing inaction: the dangers of ignoring missing data.
    Trends Ecol Evol. 2008 Nov;23(11):592-6 PMID: 18823677
  5. mRNA and microRNA expression profiles of the NCI-60 integrated with drug activities.
    Mol Cancer Ther. 2010 May;9(5):1080-91 PMID: 20442302
  6. A Review of Hot Deck Imputation for Survey Non-response.
    Int Stat Rev. 2010 Apr;78(1):40-64 PMID: 21743766
  7. CellMiner: a web-based suite of genomic and pharmacologic tools to explore transcript and drug patterns in the NCI-60 cell line set.
    Cancer Res. 2012 Jul 15;72(14):3499-511 PMID: 22802077
  8. A multivariate approach to the integration of multi-omics datasets.
    BMC Bioinformatics. 2014 May 29;15:162 PMID: 24884486
  9. Data integration in the era of omics: current and future challenges.
    BMC Syst Biol. 2014;8 Suppl 2:I1 PMID: 25032990
  10. Generalized canonical correlation analysis of matrices with missing rows: a simulation study.
    Psychometrika. 2006 Jun;71(2):323-331 PMID: 28197957
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2016-10-03
Epub
2016-00-03
Pages
402
Language
English
Region
England
NLM ID
100965194
PMCID
PMC5048483
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]