IDEAS home Printed from https://ideas.repec.org/a/wly/quante/v9y2018i1p395-440.html

Identification, data combination, and the risk of disclosure

Author

Listed:
  • Tatiana Komarova
  • Denis Nekipelov
  • Evgeny Yakovlev

Abstract

It is commonplace that the data needed for econometric inference are not contained in a single source. In this paper we analyze the problem of parametric inference from combined individual‐level data when data combination is based on personal and demographic identifiers such as name, age, or address. Our main question is the identification of the econometric model based on the combined data when the data do not contain exact individual identifiers and no parametric assumptions are imposed on the joint distribution of information that is common across the combined data set. We demonstrate the conditions on the observable marginal distributions of data in individual data sets that can and cannot guarantee identification of the parameters of interest. We also note that the data combination procedure is essential in a semiparametric setting such as ours. Provided that the (nonparametric) data combination procedure can only be defined in finite samples, we introduce a new notion of identification based on the concept of limits of statistical experiments. Our results apply to the setting where the individual data used for inferences are sensitive and their combination may lead to a substantial increase in the data sensitivity or lead to a “de‐anonymization” of the previously “anonymized” information. We demonstrate that the point identification of an econometric model from combined data is incompatible with restrictions on the risk of individual disclosure. If the data combination procedure guarantees a bound on the risk of individual disclosure, then the information available from the combined data set allows one to identify the parameter of interest only partially, and the size of the identification region is inversely related to the upper bound guarantee for the disclosure risk. This result is new in the context of data combination as we notice that the quality of links that need to be used in the combined data to assure point identification may be much higher than the average link quality in the entire data set, and thus point inference requires the use of the most sensitive subset of the data. Our results provide important insights into the ongoing discourse on the empirical analysis of merged administrative records as well as discussions on the “disclosive” nature of policies implemented by the data‐driven companies (such as internet services companies and medical companies using individual patient records for policy decisions).

Suggested Citation

  • Tatiana Komarova & Denis Nekipelov & Evgeny Yakovlev, 2018. "Identification, data combination, and the risk of disclosure," Quantitative Economics, Econometric Society, vol. 9(1), pages 395-440, March.
  • Handle: RePEc:wly:quante:v:9:y:2018:i:1:p:395-440
    DOI: 10.3982/QE568
    as

    Download full text from publisher

    File URL: https://doi.org/10.3982/QE568
    Download Restriction: no

    File URL: https://libkey.io/10.3982/QE568?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    Other versions of this item:

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Tatiana Komarova & Denis Nekipelov & Ahnaf Al Rafi & Evgeny Yakovlev, 2017. "K-anonymity: A note on the trade-off between data utility and data security," Applied Econometrics, Russian Presidential Academy of National Economy and Public Administration (RANEPA), vol. 48, pages 44-62.
    2. Tatiana Komarova & Denis Nekipelov, 2020. "Privacy-aware identification," Papers 2006.14732, arXiv.org, revised Nov 2025.
    3. David Pacini, 2012. "Least Square Linear Prediction with Two-Sample Data," Bristol Economics Discussion Papers 12/631, School of Economics, University of Bristol, UK.

    More about this item

    JEL classification:

    • C13 - Mathematical and Quantitative Methods - - Econometric and Statistical Methods and Methodology: General - - - Estimation: General
    • C14 - Mathematical and Quantitative Methods - - Econometric and Statistical Methods and Methodology: General - - - Semiparametric and Nonparametric Methods: General
    • C25 - Mathematical and Quantitative Methods - - Single Equation Models; Single Variables - - - Discrete Regression and Qualitative Choice Models; Discrete Regressors; Proportions; Probabilities
    • C35 - Mathematical and Quantitative Methods - - Multiple or Simultaneous Equation Models; Multiple Variables - - - Discrete Regression and Qualitative Choice Models; Discrete Regressors; Proportions

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:wly:quante:v:9:y:2018:i:1:p:395-440. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Wiley Content Delivery (email available below). General contact details of provider: https://edirc.repec.org/data/essssea.html .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.