IDEAS home Printed from https://ideas.repec.org/a/eee/csdana/v53y2009i5p1906-1922.html
   My bibliography  Save this article

Exploration of distributional models for a novel intensity-dependent normalization procedure in censored gene expression data

Author

Listed:
  • Lama, Nicola
  • Boracchi, Patrizia
  • Biganzoli, Elia

Abstract

Current gene intensity-dependent normalization methods, based on regression smoothing techniques, usually approach the two problems of reducing location bias and data rescaling without taking into account the censoring that is characteristic of certain gene expressions, produced by experimental measurement constraints or by previous normalization steps. Moreover, control of normalization procedures for balancing bias versus variance is often left to the user's experience. An approximate maximum likelihood procedure for fitting a model smoothing the dependences of log-fold gene expression differences on average gene intensities is presented. Central tendency and scaling factor are modeled by means of the B-spline smoothing technique. As an alternative to the outlier theory and robust methods, the approach presented looks for suitable distributional models, possibly generalizing the classical Gaussian and Laplacian assumptions, controlling for different types of censoring. The Bayesian information criterion is adopted for model selection. Distributional assumptions are tested using goodness-of-fit statistics and Monte Carlo evaluation. Randomization quantiles are proposed to produce normally distributed adjusted data. Three publicly available data sets are analyzed for demonstration purposes. Student's t error models reveal best performances in all of the data sets considered. More validating evidence is needed to evaluate the Asymmetric Laplace distribution, which showed interesting results in one data set.

Suggested Citation

  • Lama, Nicola & Boracchi, Patrizia & Biganzoli, Elia, 2009. "Exploration of distributional models for a novel intensity-dependent normalization procedure in censored gene expression data," Computational Statistics & Data Analysis, Elsevier, vol. 53(5), pages 1906-1922, March.
  • Handle: RePEc:eee:csdana:v:53:y:2009:i:5:p:1906-1922
    as

    Download full text from publisher

    File URL: http://www.sciencedirect.com/science/article/pii/S0167-9473(08)00563-X
    Download Restriction: Full text for ScienceDirect subscribers only.
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    References listed on IDEAS

    as
    1. Mineo, Angelo & Ruggieri, Mariantonietta, 2005. "A Software Tool for the Exponential Power Distribution: The normalp Package," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 12(i04).
    2. Laura J. van 't Veer & Hongyue Dai & Marc J. van de Vijver & Yudong D. He & Augustinus A. M. Hart & Mao Mao & Hans L. Peterse & Karin van der Kooy & Matthew J. Marton & Anke T. Witteveen & George J. S, 2002. "Gene expression profiling predicts clinical outcome of breast cancer," Nature, Nature, vol. 415(6871), pages 530-536, January.
    3. Tibshirani Robert J. & Efron Brad, 2002. "Pre-validation and inference in microarrays," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 1(1), pages 1-20, August.
    4. Cui Xiangqin & Kerr M. Kathleen & Churchill Gary A., 2003. "Transformations for cDNA Microarray Data," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 2(1), pages 1-22, June.
    5. Purdom Elizabeth & Holmes Susan P, 2005. "Error Distribution for Gene Expression Data," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 4(1), pages 1-35, July.
    6. Huber Wolfgang & von Heydebreck Anja & Sueltmann Holger & Poustka Annemarie & Vingron Martin, 2003. "Parameter estimation for the calibration and variance stabilization of microarray data," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 2(1), pages 1-24, April.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Kelmansky Diana M. & Martínez Elena J. & Leiva Víctor, 2013. "A new variance stabilizing transformation for gene expression data analysis," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 12(6), pages 653-666, December.
    2. Ambroise Jérôme & Bearzatto Bertrand & Robert Annie & Macq Benoit & Gala Jean-Luc, 2012. "Combining Multiple Laser Scans of Spotted Microarrays by Means of a Two-Way ANOVA Model," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 11(3), pages 1-20, February.
    3. Mehmet Niyazi Çankaya & Abdullah Yalçınkaya & Ömer Altındaǧ & Olcay Arslan, 2019. "On the robustness of an epsilon skew extension for Burr III distribution on the real line," Computational Statistics, Springer, vol. 34(3), pages 1247-1273, September.
    4. Punathumparambath, Bindu & Kulathinal, Sangita & George, Sebastian, 2012. "Asymmetric type II compound Laplace distribution and its application to microarray gene expression," Computational Statistics & Data Analysis, Elsevier, vol. 56(6), pages 1396-1404.
    5. Tibshirani Robert J., 2009. "Univariate Shrinkage in the Cox Model for High Dimensional Data," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 8(1), pages 1-18, April.
    6. Jing Zhang & Qihua Wang & Xuan Wang, 2022. "Surrogate-variable-based model-free feature screening for survival data under the general censoring mechanism," Annals of the Institute of Statistical Mathematics, Springer;The Institute of Statistical Mathematics, vol. 74(2), pages 379-397, April.
    7. Gaorong Li & Liugen Xue & Heng Lian, 2012. "SCAD-penalised generalised additive models with non-polynomial dimensionality," Journal of Nonparametric Statistics, Taylor & Francis Journals, vol. 24(3), pages 681-697.
    8. Zemin Zheng & Jie Zhang & Yang Li, 2022. "L 0 -Regularized Learning for High-Dimensional Additive Hazards Regression," INFORMS Journal on Computing, INFORMS, vol. 34(5), pages 2762-2775, September.
    9. Lian, Heng & Du, Pang & Li, YuanZhang & Liang, Hua, 2014. "Partially linear structure identification in generalized additive models with NP-dimensionality," Computational Statistics & Data Analysis, Elsevier, vol. 80(C), pages 197-208.
    10. Huixia Judy Wang & Leonard A. Stefanski & Zhongyi Zhu, 2012. "Corrected-loss estimation for quantile regression with covariate measurement errors," Biometrika, Biometrika Trust, vol. 99(2), pages 405-421.
    11. Jan, Budczies & Kosztyla, Daniel & von Törne, Christian & Stenzinger, Albrecht & Darb-Esfahani, Silvia & Dietel, Manfred & Denkert, Carsten, 2014. "cancerclass: An R Package for Development and Validation of Diagnostic Tests from High-Dimensional Molecular Data," Journal of Statistical Software, Foundation for Open Access Statistics, vol. 59(i01).
    12. Zhaoliang Wang & Liugen Xue & Gaorong Li & Fei Lu, 2019. "Spline estimator for ultra-high dimensional partially linear varying coefficient models," Annals of the Institute of Statistical Mathematics, Springer;The Institute of Statistical Mathematics, vol. 71(3), pages 657-677, June.
    13. Tu, Shiyi & Wang, Min & Sun, Xiaoqian, 2016. "Bayesian analysis of two-piece location–scale models under reference priors with partial information," Computational Statistics & Data Analysis, Elsevier, vol. 96(C), pages 133-144.
    14. Lian, I.B. & Chang, C.J. & Liang, Y.J. & Yang, M.J. & Fann, C.S.J., 2007. "Identifying differentially expressed genes in dye-swapped microarray experiments of small sample size," Computational Statistics & Data Analysis, Elsevier, vol. 51(5), pages 2602-2620, February.
    15. Reiner Franke, 2015. "How Fat-Tailed is US Output Growth?," Metroeconomica, Wiley Blackwell, vol. 66(2), pages 213-242, May.
    16. Grace Y. Yi & Wenqing He & Raymond. J. Carroll, 2022. "Feature screening with large‐scale and high‐dimensional survival data," Biometrics, The International Biometric Society, vol. 78(3), pages 894-907, September.
    17. Olcay Arslan, 2010. "An alternative multivariate skew Laplace distribution: properties and estimation," Statistical Papers, Springer, vol. 51(4), pages 865-887, December.
    18. Massimiliano Giacalone & Demetrio Panarello & Raffaele Mattera, 2018. "Multicollinearity in regression: an efficiency comparison between Lp-norm and least squares estimators," Quality & Quantity: International Journal of Methodology, Springer, vol. 52(4), pages 1831-1859, July.
    19. Mattera, Raffaele, 2017. "A GED-based regression to fit the actual data distribution," MPRA Paper 80501, University Library of Munich, Germany.
    20. Khan Md Hasinur Rahaman & Bhadra Anamika & Howlader Tamanna, 2019. "Stability selection for lasso, ridge and elastic net implemented with AFT models," Statistical Applications in Genetics and Molecular Biology, De Gruyter, vol. 18(5), pages 1-14, October.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:csdana:v:53:y:2009:i:5:p:1906-1922. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.elsevier.com/locate/csda .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.