IDEAS home Printed from https://ideas.repec.org/p/arx/papers/1910.06677.html
   My bibliography  Save this paper

Matrix Completion, Counterfactuals, and Factor Analysis of Missing Data

Author

Listed:
  • Jushan Bai
  • Serena Ng

Abstract

This paper proposes an imputation procedure that uses the factors estimated from a tall block along with the re-rotated loadings estimated from a wide block to impute missing values in a panel of data. Assuming that a strong factor structure holds for the full panel of data and its sub-blocks, it is shown that the common component can be consistently estimated at four different rates of convergence without requiring regularization or iteration. An asymptotic analysis of the estimation error is obtained. An application of our analysis is estimation of counterfactuals when potential outcomes have a factor structure. We study the estimation of average and individual treatment effects on the treated and establish a normal distribution theory that can be useful for hypothesis testing.

Suggested Citation

  • Jushan Bai & Serena Ng, 2019. "Matrix Completion, Counterfactuals, and Factor Analysis of Missing Data," Papers 1910.06677, arXiv.org, revised Aug 2021.
  • Handle: RePEc:arx:papers:1910.06677
    as

    Download full text from publisher

    File URL: http://arxiv.org/pdf/1910.06677
    File Function: Latest version
    Download Restriction: no
    ---><---

    Other versions of this item:

    References listed on IDEAS

    as
    1. Giannone, Domenico & Reichlin, Lucrezia & Small, David, 2008. "Nowcasting: The real-time informational content of macroeconomic data," Journal of Monetary Economics, Elsevier, vol. 55(4), pages 665-676, May.
    2. Laurent Gobillon & Thierry Magnac, 2016. "Regional Policy Evaluation: Interactive Fixed Effects and Synthetic Controls," The Review of Economics and Statistics, MIT Press, vol. 98(3), pages 535-551, July.
    3. Jushan Bai & Serena Ng, 2002. "Determining the Number of Factors in Approximate Factor Models," Econometrica, Econometric Society, vol. 70(1), pages 191-221, January.
    4. Horton, Nicholas J. & Kleinman, Ken P., 2007. "Much Ado About Nothing: A Comparison of Missing Data Methods and Software to Fit Incomplete Data Regression Models," The American Statistician, American Statistical Association, vol. 61, pages 79-90, February.
    5. Bai, Jushan & Ng, Serena, 2019. "Rank regularized estimation of approximate factor models," Journal of Econometrics, Elsevier, vol. 212(1), pages 78-96.
    6. J. B. Taylor & Harald Uhlig (ed.), 2016. "Handbook of Macroeconomics," Handbook of Macroeconomics, Elsevier, edition 1, volume 2, number 2.
    7. Jungbacker, B. & Koopman, S.J. & van der Wel, M., 2011. "Maximum likelihood estimation for dynamic factor models with missing data," Journal of Economic Dynamics and Control, Elsevier, vol. 35(8), pages 1358-1368, August.
    8. R. H. Shumway & D. S. Stoffer, 1982. "An Approach To Time Series Smoothing And Forecasting Using The Em Algorithm," Journal of Time Series Analysis, Wiley Blackwell, vol. 3(4), pages 253-264, July.
    9. Cheng Hsiao & H. Steve Ching & Shui Ki Wan, 2012. "A Panel Data Approach For Program Evaluation: Measuring The Benefits Of Political And Economic Integration Of Hong Kong With Mainland China," Journal of Applied Econometrics, John Wiley & Sons, Ltd., vol. 27(5), pages 705-740, August.
    10. Marta Bańbura & Michele Modugno, 2014. "Maximum Likelihood Estimation Of Factor Models On Datasets With Arbitrary Pattern Of Missing Data," Journal of Applied Econometrics, John Wiley & Sons, Ltd., vol. 29(1), pages 133-160, January.
    11. Carl Eckart & Gale Young, 1936. "The approximation of one matrix by another of lower rank," Psychometrika, Springer;The Psychometric Society, vol. 1(3), pages 211-218, September.
    12. Jin, Sainan & Miao, Ke & Su, Liangjun, 2021. "On factor models with random missing: EM estimation, inference, and cross validation," Journal of Econometrics, Elsevier, vol. 222(1), pages 745-777.
    13. Xu, Yiqing, 2017. "Generalized Synthetic Control Method: Causal Inference with Interactive Fixed Effects Models," Political Analysis, Cambridge University Press, vol. 25(1), pages 57-76, January.
    14. Jushan Bai, 2009. "Panel Data Models With Interactive Fixed Effects," Econometrica, Econometric Society, vol. 77(4), pages 1229-1279, July.
    15. Domenico Giannone & Lucrezia Reichlin & David Small, 2008. "Nowcasting: the real time informational content of macroeconomic data releases," ULB Institutional Repository 2013/6409, ULB -- Universite Libre de Bruxelles.
    16. Stock J.H. & Watson M.W., 2002. "Forecasting Using Principal Components From a Large Number of Predictors," Journal of the American Statistical Association, American Statistical Association, vol. 97, pages 1167-1179, December.
    17. Domenico Giannone & Lucrezia Reichlin & David H. Small, 2005. "Nowcasting GDP and inflation: the real-time informational content of macroeconomic data releases," Finance and Economics Discussion Series 2005-42, Board of Governors of the Federal Reserve System (U.S.).
    18. Jushan Bai, 2003. "Inferential Theory for Factor Models of Large Dimensions," Econometrica, Econometric Society, vol. 71(1), pages 135-171, January.
    19. James Honaker & Gary King, 2010. "What to Do about Missing Values in Time‐Series Cross‐Section Data," American Journal of Political Science, John Wiley & Sons, vol. 54(2), pages 561-581, April.
    Full references (including those not matched with items on IDEAS)

    Most related items

    These are the items that most often cite the same works as this one and are cited by the same works as this one.
    1. Xiong, Ruoxuan & Pelger, Markus, 2023. "Large dimensional latent factor modeling with missing observations and applications to causal inference," Journal of Econometrics, Elsevier, vol. 233(1), pages 271-301.
    2. Matteo Barigozzi & Matteo Luciani, 2019. "Quasi Maximum Likelihood Estimation and Inference of Large Approximate Dynamic Factor Models via the EM algorithm," Papers 1910.03821, arXiv.org, revised Sep 2024.
    3. Cahan, Ercument & Bai, Jushan & Ng, Serena, 2023. "Factor-based imputation of missing values and covariances in panel data of large dimensions," Journal of Econometrics, Elsevier, vol. 233(1), pages 113-131.
    4. Poncela, Pilar & Ruiz, Esther & Miranda, Karen, 2021. "Factor extraction using Kalman filter and smoothing: This is not just another survey," International Journal of Forecasting, Elsevier, vol. 37(4), pages 1399-1425.
    5. Catherine Doz & Peter Fuleky, 2019. "Dynamic Factor Models," Working Papers 2019-4, University of Hawaii Economic Research Organization, University of Hawaii at Manoa.
    6. Stock, J.H. & Watson, M.W., 2016. "Dynamic Factor Models, Factor-Augmented Vector Autoregressions, and Structural Vector Autoregressions in Macroeconomics," Handbook of Macroeconomics, in: J. B. Taylor & Harald Uhlig (ed.), Handbook of Macroeconomics, edition 1, volume 2, chapter 0, pages 415-525, Elsevier.
    7. Kaufmann, Daniel & Scheufele, Rolf, 2017. "Business tendency surveys and macroeconomic fluctuations," International Journal of Forecasting, Elsevier, vol. 33(4), pages 878-893.
    8. Jin, Sainan & Miao, Ke & Su, Liangjun, 2021. "On factor models with random missing: EM estimation, inference, and cross validation," Journal of Econometrics, Elsevier, vol. 222(1), pages 745-777.
    9. Pilar Poncela & Esther Ruiz, 2016. "Small- Versus Big-Data Factor Extraction in Dynamic Factor Models: An Empirical Assessment," Advances in Econometrics, in: Dynamic Factor Models, volume 35, pages 401-434, Emerald Group Publishing Limited.
    10. Marcellino, Massimiliano & Sivec, Vasja, 2016. "Monetary, fiscal and oil shocks: Evidence based on mixed frequency structural FAVARs," Journal of Econometrics, Elsevier, vol. 193(2), pages 335-348.
    11. Monica Defend & Aleksey Min & Lorenzo Portelli & Franz Ramsauer & Francesco Sandrini & Rudi Zagst, 2021. "Quantifying Drivers of Forecasted Returns Using Approximate Dynamic Factor Models for Mixed-Frequency Panel Data," Forecasting, MDPI, vol. 3(1), pages 1-35, February.
    12. Matteo Barigozzi & Matteo Luciani, 2017. "Common Factors, Trends, and Cycles in Large Datasets," Finance and Economics Discussion Series 2017-111, Board of Governors of the Federal Reserve System (U.S.).
    13. Ma, Tao & Zhou, Zhou & Antoniou, Constantinos, 2018. "Dynamic factor model for network traffic state forecast," Transportation Research Part B: Methodological, Elsevier, vol. 118(C), pages 281-317.
    14. Jonas Krampe & Luca Margaritella, 2021. "Factor Models with Sparse VAR Idiosyncratic Components," Papers 2112.07149, arXiv.org, revised May 2022.
    15. Hindrayanto, Irma & Koopman, Siem Jan & de Winter, Jasper, 2016. "Forecasting and nowcasting economic growth in the euro area using factor models," International Journal of Forecasting, Elsevier, vol. 32(4), pages 1284-1305.
    16. Alvarez, Rocio & Camacho, Maximo & Perez-Quiros, Gabriel, 2016. "Aggregate versus disaggregate information in dynamic factor models," International Journal of Forecasting, Elsevier, vol. 32(3), pages 680-694.
    17. Modugno, Michele & Soybilgen, Barış & Yazgan, Ege, 2016. "Nowcasting Turkish GDP and news decomposition," International Journal of Forecasting, Elsevier, vol. 32(4), pages 1369-1384.
    18. Daniel J. Lewis & Karel Mertens & James H. Stock & Mihir Trivedi, 2022. "Measuring real activity using a weekly economic index," Journal of Applied Econometrics, John Wiley & Sons, Ltd., vol. 37(4), pages 667-687, June.
    19. Bańbura, Marta & Giannone, Domenico & Lenza, Michele, 2015. "Conditional forecasts and scenario analysis with vector autoregressions for large cross-sections," International Journal of Forecasting, Elsevier, vol. 31(3), pages 739-756.
    20. Matteo Luciani & Madhavi Pundit & Arief Ramayandi & Giovanni Veronese, 2018. "Nowcasting Indonesia," Empirical Economics, Springer, vol. 55(2), pages 597-619, September.

    More about this item

    NEP fields

    This paper has been announced in the following NEP Reports:

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:arx:papers:1910.06677. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: arXiv administrators (email available below). General contact details of provider: http://arxiv.org/ .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.