The cluster graphical lasso for improved estimation of Gaussian graphical models

The cluster graphical lasso for improved estimation of Gaussian graphical models

Author

Listed:

Tan, Kean Ming
Witten, Daniela
Shojaie, Ali

Abstract

The task of estimating a Gaussian graphical model in the high-dimensional setting is considered. The graphical lasso, which involves maximizing the Gaussian log likelihood subject to a lasso penalty, is a well-studied approach for this task. A surprising connection between the graphical lasso and hierarchical clustering is introduced: the graphical lasso in effect performs a two-step procedure, in which (1) single linkage hierarchical clustering is performed on the variables in order to identify connected components, and then (2) a penalized log likelihood is maximized on the subset of variables within each connected component. Thus, the graphical lasso determines the connected components of the estimated network via single linkage clustering. The single linkage clustering is known to perform poorly in certain finite-sample settings. Therefore, the cluster graphical lasso, which involves clustering the features using an alternative to single linkage clustering, and then performing the graphical lasso on the subset of variables within each cluster, is proposed. Model selection consistency for this technique is established, and its improved performance relative to the graphical lasso is demonstrated in a simulation study, as well as in applications to a university webpage and a gene expression data sets.

Suggested Citation

Tan, Kean Ming & Witten, Daniela & Shojaie, Ali, 2015. "The cluster graphical lasso for improved estimation of Gaussian graphical models," Computational Statistics & Data Analysis, Elsevier, vol. 85(C), pages 23-36.

Handle: RePEc:eee:csdana:v:85:y:2015:i:c:p:23-36
DOI: 10.1016/j.csda.2014.11.015

Download full text from publisher

As the access to this document is restricted, you may want to

for a different version of it.

References listed on IDEAS

Zou, Hui, 2006. "The Adaptive Lasso and Its Oracle Properties," Journal of the American Statistical Association, American Statistical Association, vol. 101, pages 1418-1429, December.
Nicolai Meinshausen & Peter Bühlmann, 2010. "Stability selection," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 72(4), pages 417-473, September.
Jian Guo & Elizaveta Levina & George Michailidis & Ji Zhu, 2011. "Joint estimation of multiple graphical models," Biometrika, Biometrika Trust, vol. 98(1), pages 1-15.
Robert Tibshirani & Guenther Walther & Trevor Hastie, 2001. "Estimating the number of clusters in a data set via the gap statistic," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 63(2), pages 411-423.
Lam, Clifford & Fan, Jianqing, 2009. "Sparsistency and rates of convergence in large covariance matrix estimation," LSE Research Online Documents on Economics 31540, London School of Economics and Political Science, LSE Library.
Peng, Jie & Wang, Pei & Zhou, Nengfeng & Zhu, Ji, 2009. "Partial Correlation Estimation by Joint Sparse Regression Models," Journal of the American Statistical Association, American Statistical Association, vol. 104(486), pages 735-746.
Glenn Milligan & Martha Cooper, 1985. "An examination of procedures for determining the number of clusters in a data set," Psychometrika, Springer;The Psychometric Society, vol. 50(2), pages 159-179, June.
Ming Yuan & Yi Lin, 2007. "Model selection and estimation in the Gaussian graphical model," Biometrika, Biometrika Trust, vol. 94(1), pages 19-35.
Beatrix Jones & Mike West, 2005. "Covariance decomposition in undirected Gaussian graphical models," Biometrika, Biometrika Trust, vol. 92(4), pages 779-786, December.
Cai, Tony & Liu, Weidong & Luo, Xi, 2011. "A Constrained â„“1 Minimization Approach to Sparse Precision Matrix Estimation," Journal of the American Statistical Association, American Statistical Association, vol. 106(494), pages 594-607.

Full references (including those not matched with items on IDEAS)

Citations

Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.

Cited by:

Chen, Shuo & Kang, Jian & Xing, Yishi & Zhao, Yunpeng & Milton, Donald K., 2018. "Estimating large covariance matrix with network topology for high-dimensional biomedical data," Computational Statistics & Data Analysis, Elsevier, vol. 127(C), pages 82-95.
Jos'e Vin'icius de Miranda Cardoso & Jiaxi Ying & Daniel Perez Palomar, 2020. "Algorithms for Learning Graphs in Financial Markets," Papers 2012.15410, arXiv.org.
Ines Wilms & Jacob Bien, 2021. "Tree-based Node Aggregation in Sparse Graphical Models," Papers 2101.12503, arXiv.org.
Mariaelena Bottazzi Schenone & Roberto Rocci & Maurizio Vichi, 2025. "Generalized Reduced K–Means," Computational Statistics, Springer, vol. 40(4), pages 1753-1778, April.

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

Ziqi Chen & Chenlei Leng, 2016. "Dynamic Covariance Models," Journal of the American Statistical Association, Taylor & Francis Journals, vol. 111(515), pages 1196-1207, July.
Wang, Ke & Franks, Alexander & Oh, Sang-Yun, 2023. "Learning Gaussian graphical models with latent confounders," Journal of Multivariate Analysis, Elsevier, vol. 198(C).
Bailey, Natalia & Pesaran, M. Hashem & Smith, L. Vanessa, 2019. "A multiple testing approach to the regularisation of large sample correlation matrices," Journal of Econometrics, Elsevier, vol. 208(2), pages 507-534.
- Natalia Bailey & M. Hashem Pesaran & L. Vanessa Smith, 2014. "A Multiple Testing Approach to the Regularisation of Large Sample Correlation Matrices," CESifo Working Paper Series 4834, CESifo.
- Natalia Bailey & M. Hashem Pesaran & L. Vanessa Smith, 2015. "A Multiple Testing Approach to the Regularisation of Large Sample Correlation Matrices," Working Papers 764, Queen Mary University of London, School of Economics and Finance.
- Natalia Bailey & Vanessa Smith & M. Hashem Pesaran, 2014. "A multiple testing approach to the regularisation of large sample correlation matrices," Cambridge Working Papers in Economics 1413, Faculty of Economics, University of Cambridge.
Lafit, Ginette & Nogales Martín, Francisco Javier & Zamar, Rubén, 2015. "Ranking Edges and Model Selection in High-Dimensional Graphs," DES - Working Papers. Statistics and Econometrics. WS ws1511, Universidad Carlos III de Madrid. Departamento de EstadÃstica.
Xiao Guo & Hai Zhang, 2020. "Sparse directed acyclic graphs incorporating the covariates," Statistical Papers, Springer, vol. 61(5), pages 2119-2148, October.
Banerjee, Sayantan & Ghosal, Subhashis, 2015. "Bayesian structure learning in graphical models," Journal of Multivariate Analysis, Elsevier, vol. 136(C), pages 147-162.
Yin, Jianxin & Li, Hongzhe, 2012. "Model selection and estimation in the matrix normal graphical model," Journal of Multivariate Analysis, Elsevier, vol. 107(C), pages 119-140.
Lin Zhang & Andrew DiLernia & Karina Quevedo & Jazmin Camchong & Kelvin Lim & Wei Pan, 2021. "A random covariance model for bi‐level graphical modeling with application to resting‐state fMRI data," Biometrics, The International Biometric Society, vol. 77(4), pages 1385-1396, December.
Lafit, Ginette & Nogales Martín, Francisco Javier, 2017. "Robust and sparse estimation of high-dimensional precision matrices via bivariate outlier detection," DES - Working Papers. Statistics and Econometrics. WS 24534, Universidad Carlos III de Madrid. Departamento de EstadÃstica.
Jianqing Fan & Yuan Liao & Han Liu, 2016. "An overview of the estimation of large covariance and precision matrices," Econometrics Journal, Royal Economic Society, vol. 19(1), pages 1-32, February.
Li‐Pang Chen, 2024. "Estimation of Graphical Models: An Overview of Selected Topics," International Statistical Review, International Statistical Institute, vol. 92(2), pages 194-245, August.
Hirose, Kei & Fujisawa, Hironori & Sese, Jun, 2017. "Robust sparse Gaussian graphical modeling," Journal of Multivariate Analysis, Elsevier, vol. 161(C), pages 172-190.
Banerjee, Sayantan & Akbani, Rehan & Baladandayuthapani, Veerabhadran, 2019. "Spectral clustering via sparse graph structure learning with application to proteomic signaling networks in cancer," Computational Statistics & Data Analysis, Elsevier, vol. 132(C), pages 46-69.
Pan, Yuqing & Mai, Qing, 2020. "Efficient computation for differential network analysis with applications to quadratic discriminant analysis," Computational Statistics & Data Analysis, Elsevier, vol. 144(C).
Fan, Xinyan & Zhang, Qingzhao & Ma, Shuangge & Fang, Kuangnan, 2021. "Conditional score matching for high-dimensional partial graphical models," Computational Statistics & Data Analysis, Elsevier, vol. 153(C).
Avagyan, Vahe & Alonso Fernández, Andrés Modesto & Nogales, Francisco J., 2015. "D-trace Precision Matrix Estimation Using Adaptive Lasso Penalties," DES - Working Papers. Statistics and Econometrics. WS 21775, Universidad Carlos III de Madrid. Departamento de EstadÃstica.
Yang, Yihe & Zhou, Jie & Pan, Jianxin, 2021. "Estimation and optimal structure selection of high-dimensional Toeplitz covariance matrix," Journal of Multivariate Analysis, Elsevier, vol. 184(C).
Lee, Wonyul & Liu, Yufeng, 2012. "Simultaneous multiple response regression and inverse covariance matrix estimation via penalized Gaussian maximum likelihood," Journal of Multivariate Analysis, Elsevier, vol. 111(C), pages 241-255.
Ines Wilms & Jacob Bien, 2021. "Tree-based Node Aggregation in Sparse Graphical Models," Papers 2101.12503, arXiv.org.
Dong Liu & Changwei Zhao & Yong He & Lei Liu & Ying Guo & Xinsheng Zhang, 2023. "Simultaneous cluster structure learning and estimation of heterogeneous graphs for matrix‐variate fMRI data," Biometrics, The International Biometric Society, vol. 79(3), pages 2246-2259, September.

More about this item

Keywords

; ; ; ; ;

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:csdana:v:85:y:2015:i:c:p:23-36. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.elsevier.com/locate/csda .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

The cluster graphical lasso for improved estimation of Gaussian graphical models

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Citations

Most related items

More about this item

Keywords

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data