Missing Values Estimation for Time Course Gene Expression Data Using the Sequential Partial Least Squares Regression Fitting

Fitting Cox models in a big data context -on a massive scale in terms of volume, intensity, and complexity exceeding the capacity of usual analytic tools-is often challenging. If some data are missing, it is even more difficult. We proposed algorithms that were able to fit Cox models in high dimensional settings using extensions of partial least squares regression to the Cox models. Some of them were able to cope with missing data. We were recently able to extend our most recent algorithms to big data, thus allowing to fit Cox model for big data with missing values. When cross-validating standard or extended Cox models, the commonly used criterion is the cross-validated partial loglikelihood using a naive or a van Houwelingen scheme —to make efficient use of the death times of the left out data in relation to the death times of all the data. Quite astonishingly, we will show, using a strong simulation study involving three different data simulation algorithms, that these two cross-validation methods fail with the extensions, either straightforward or more involved ones, of partial least squares regression to the Cox model. This is quite an interesting result for at least two reasons. Firstly, several nice features of PLS based models, including regularization, interpretability of the components, missing data support, data visualization thanks to biplots of individuals and variables —and even parsimony or group parsimony for Sparse partial least squares or sparse group SPLS based models, account for a common use of these extensions by statisticians who usually select their hyperparameters using cross-validation. Secondly, they are almost always featured in benchmarking studies to assess the performance of a new estimation technique used in a high dimensional or big data context and often show poor statistical properties. We carried out a vast simulation study to evaluate more than a dozen of potential cross-validation criteria, either AUC or prediction error based. Several of them lead to the selection of a reasonable number of components. Using these newly found cross-validation criteria to fit extensions of partial least squares regression to the Cox model, we performed a benchmark reanalysis that showed enhanced performances of these techniques. In addition, we proposed sparse group extensions of our algorithms and defined a new robust measure based on the Schmid score and the R coefficient of determination for least absolute deviation: the integrated R Schmid Score weighted. The R-package used in this article is available on the CRAN, http://cran.r-project.org/web/packages/plsRcox/index.html. The R package bigPLS will soon be available on the CRAN and, until then, is available on Github https://github.com/fbertran/bigPLS.

Download Full-text

Tumor classification by partial least squares using microarray gene expression data

Bioinformatics ◽

10.1093/bioinformatics/18.1.39 ◽

2002 ◽

Vol 18 (1) ◽

pp. 39-50 ◽

Cited By ~ 505

Author(s):

D. V. Nguyen ◽

D. M. Rocke

Keyword(s):

Gene Expression ◽

Least Squares ◽

Partial Least Squares ◽

Gene Expression Data ◽

Microarray Gene Expression Data ◽

Tumor Classification ◽

Expression Data ◽

Microarray Gene Expression ◽

Microarray Gene

Download Full-text

Partial least squares and logistic regression random-effects estimates for gene selection in supervised classification of gene expression data

Journal of Biomedical Informatics ◽

10.1016/j.jbi.2013.05.008 ◽

2013 ◽

Vol 46 (4) ◽

pp. 697-709 ◽

Cited By ~ 6

Author(s):

Arief Gusnanto ◽

Alexander Ploner ◽

Farag Shuweihdi ◽

Yudi Pawitan

Keyword(s):

Gene Expression ◽

Logistic Regression ◽

Least Squares ◽

Partial Least Squares ◽

Random Effects ◽

Gene Expression Data ◽

Supervised Classification ◽

Gene Selection ◽

Expression Data

Download Full-text

ITERATED LOCAL LEAST SQUARES MICROARRAY MISSING VALUE IMPUTATION

Journal of Bioinformatics and Computational Biology ◽

10.1142/s0219720006002302 ◽

2006 ◽

Vol 04 (05) ◽

pp. 935-957 ◽

Cited By ~ 51

Author(s):

ZHIPENG CAI ◽

MAYSAM HEYDARI ◽

GUOHUI LIN

Keyword(s):

Gene Expression ◽

Data Analysis ◽

Least Squares ◽

Gene Expression Data ◽

Missing Values ◽

Target Genes ◽

Accurate Estimation ◽

Expression Data ◽

Microarray Gene Expression ◽

Missing Value

Microarray gene expression data often contains multiple missing values due to various reasons. However, most of gene expression data analysis algorithms require complete expression data. Therefore, accurate estimation of the missing values is critical to further data analysis. In this paper, an Iterated Local Least Squares Imputation (ILLSimpute) method is proposed for estimating missing values. Two unique features of ILLSimpute method are: ILLSimpute method does not fix a common number of coherent genes for target genes for estimation purpose, but defines coherent genes as those within a distance threshold to the target genes. Secondly, in ILLSimpute method, estimated values in one iteration are used for missing value estimation in the next iteration and the method terminates after certain iterations or the imputed values converge. Experimental results on six real microarray datasets showed that ILLSimpute method performed at least as well as, and most of the time much better than, five most recent imputation methods.

Download Full-text