Regression analysis under incomplete linkage

Regression analysis under incomplete linkage

Author

Listed:

Kim, Gunky
Chambers, Raymond

Abstract

Most probability-based methods used to link records from two distinct data sets corresponding to the same target population do not lead to perfect linkage, i.e. there are linkage errors in the merged data. Further, the linkage is often incomplete, in the sense that many records in the two data sets remain unmatched at the completion of the linkage process. This paper introduces methods that correct for the biases due to linkage errors and incomplete linkage when carrying out regression analysis using linked data. In particular, it focuses on the case where one of the linked data sets is a sample from the target population and the other is a register, i.e. it covers the entire target population.

Suggested Citation

Kim, Gunky & Chambers, Raymond, 2012. "Regression analysis under incomplete linkage," Computational Statistics & Data Analysis, Elsevier, vol. 56(9), pages 2756-2770.

Handle: RePEc:eee:csdana:v:56:y:2012:i:9:p:2756-2770
DOI: 10.1016/j.csda.2012.02.026

Download full text from publisher

As the access to this document is restricted, you may want to

for a different version of it.

References listed on IDEAS

P. Lahiri & Michael D. Larsen, 2005. "Regression Analysis With Linked Data," Journal of the American Statistical Association, American Statistical Association, vol. 100, pages 222-230, March.
Jixian Wang & Peter Donnan, 2002. "Adjusting for missing record linkage in outcome studies," Journal of Applied Statistics, Taylor & Francis Journals, vol. 29(6), pages 873-884.

Full references (including those not matched with items on IDEAS)

Citations

Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.

Cited by:

Li‐Chun Zhang & Tiziana Tuoto, 2021. "Linkage‐data linear regression," Journal of the Royal Statistical Society Series A, Royal Statistical Society, vol. 184(2), pages 522-547, April.
Han Ying, 2020. "Discussion of “Small area estimation: its evolution in five decades”, by Malay Ghosh," Statistics in Transition New Series, Statistics Poland, vol. 21(4), pages 30-34, August.
N. Salvati & E. Fabrizi & M. G. Ranalli & R. L. Chambers, 2021. "Small area estimation with linked data," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 83(1), pages 78-107, February.
Ray Chambers & Andrea Diniz da Silva, 2020. "Improved secondary analysis of linked data: a framework and an illustration," Journal of the Royal Statistical Society Series A, Royal Statistical Society, vol. 183(1), pages 37-59, January.
Hendrik van Broekhuizen, 2016. "Graduate unemployment and Higher Education Institutions in South Africa," Working Papers 08/2016, Stellenbosch University, Department of Economics.
Tatiana Komarova & Denis Nekipelov & Evgeny Yakovlev, 2018. "Identification, data combination, and the risk of disclosure," Quantitative Economics, Econometric Society, vol. 9(1), pages 395-440, March.
- Tatiana V. Komarova & Denis Nekipelov & Evgeny Yakovlev, 2011. "Identification, data combination and the risk of disclosure," CeMMAP working papers CWP38/11, Centre for Microdata Methods and Practice, Institute for Fiscal Studies.
- Komarova, Tatiana & Nekipelov, Denis & Yakovlev, Evgeny, 2018. "Identification, data combination and the risk of disclosure," LSE Research Online Documents on Economics 79384, London School of Economics and Political Science, LSE Library.
Vo, Thanh Huan & Chauvet, Guillaume & Happe, André & Oger, Emmanuel & Paquelet, Stéphane & Garès, Valérie, 2023. "Extending the Fellegi-Sunter record linkage model for mixed-type data with application to the French national health data system," Computational Statistics & Data Analysis, Elsevier, vol. 179(C).
Chenarides, Lauren & Hanks, Andrew S. & Berard, Jake & Carlson, Andrea C. & Davis, George & Finaret, Amelia B., 2025. "A review of data linkages for policy-informing research in food and agricultural economics," Food Policy, Elsevier, vol. 137(C).
Ying Han, 2020. "Discussion of "Small area estimation: its evolution in five decades", by Malay Ghosh," Statistics in Transition New Series, Polish Statistical Association, vol. 21(4), pages 30-34, August.
Angelo Moretti & Natalie Shlomo, 2023. "Improving Probabilistic Record Linkage Using Statistical Prediction Models," International Statistical Review, International Statistical Institute, vol. 91(3), pages 368-394, December.
Ben Powell & Paul A. Smith, 2020. "Computing expectations and marginal likelihoods for permutations," Computational Statistics, Springer, vol. 35(2), pages 871-891, June.

Most related items

These are the items that most often cite the same works as this one and are cited by the same works as this one.

Al-Kandari Noriah M. & Lahiri Partha, 2016. "Prediction of a Function of Misclassified Binary Data," Statistics in Transition New Series, Statistics Poland, vol. 17(3), pages 429-447, September.
Ben Powell & Paul A. Smith, 2020. "Computing expectations and marginal likelihoods for permutations," Computational Statistics, Springer, vol. 35(2), pages 871-891, June.
Han Ying, 2020. "Discussion of “Small area estimation: its evolution in five decades”, by Malay Ghosh," Statistics in Transition New Series, Statistics Poland, vol. 21(4), pages 30-34, August.
Durrant, Gabriele B. & D'Arrigo, Julia & Steele, Fiona, 2011. "Using field process data to predict best times of contact conditioning on household and interviewer influences," LSE Research Online Documents on Economics 52201, London School of Economics and Political Science, LSE Library.
Ying Han, 2020. "Discussion of "Small area estimation: its evolution in five decades", by Malay Ghosh," Statistics in Transition New Series, Polish Statistical Association, vol. 21(4), pages 30-34, August.
Noriah M. Al-Kandari & Partha Lahiri, 2016. "Prediction Of A Function Of Misclassified Binary Data," Statistics in Transition New Series, Polish Statistical Association, vol. 17(3), pages 429-447, September.
Shurong Lin & Eric D. Kolaczyk, 2026. "Data Privacy for Record Linkage and Beyond," NBER Chapters, in: Data Privacy Protection and the Conduct of Applied Research: Methods, Approaches and New Findings, National Bureau of Economic Research, Inc.
John M. Abowd & Joelle Abramowitz & Margaret C. Levenstein & Kristin McCue & Dhiren Patki & Trivellore Raghunathan & Ann M. Rodgers & Matthew D. Shapiro & Nada Wasi & Dawn Zinsser, 2021. "Finding Needles in Haystacks: Multiple-Imputation Record Linkage Using Machine Learning," Working Papers 21-35, Center for Economic Studies, U.S. Census Bureau.
- John M. Abowd & Joelle Hillary Abramowitz & Margaret Catherine Levenstein & Kristin McCue & Dhiren Patki & Trivellore Raghunathan & Ann Michelle Rodgers & Matthew D. Shapiro & Nada Wasi & Dawn Zinsser, 2021. "Finding Needles in Haystacks: Multiple-Imputation Record Linkage Using Machine Learning," Working Papers 22-11, Federal Reserve Bank of Boston.
Sarah Tahamont & Zubin Jelveh & Aaron Chalfin & Shi Yan & Benjamin Hansen, 2019. "Administrative Data Linking and Statistical Power Problems in Randomized Experiments," NBER Working Papers 25657, National Bureau of Economic Research, Inc.
Sarah Tahamont & Zubin Jelveh & Aaron Chalfin & Shi Yan & Benjamin Hansen, 2021. "Dude, Where’s My Treatment Effect? Errors in Administrative Data Linking and the Destruction of Statistical Power in Randomized Experiments," Journal of Quantitative Criminology, Springer, vol. 37(3), pages 715-749, September.
D. H. Judson, 2007. "Information integration for constructing social statistics: history, theory and ideas towards a research programme," Journal of the Royal Statistical Society Series A, Royal Statistical Society, vol. 170(2), pages 483-501, March.
Loredana Di Consiglio & Tiziana Tuoto, 2020. "A comparison of area level and unit level small area models in the presence of linkage errors," Statistics in Transition New Series, Polish Statistical Association, vol. 21(4), pages 103-122, August.
John M. Abowd & Joelle Abramowitz & Margaret C. Levenstein & Kristin McCue & Dhiren Patki & Trivellore Raghunathan & Ann M. Rodgers & Matthew D. Shapiro & Nada Wasi, 2019. "Optimal Probabilistic Record Linkage: Best Practice for Linking Employers in Survey and Administrative Data," Working Papers 19-08, Center for Economic Studies, U.S. Census Bureau.
Kreuter Frauke, 2013. "Discussion," Journal of Official Statistics, Sciendo, vol. 29(1), pages 165-169, March.
Dasylva Abel, 2018. "Design-Based Estimation with Record-Linked Administrative Files and a Clerical Review Sample," Journal of Official Statistics, Sciendo, vol. 34(1), pages 41-54, March.
Afshin Fallah & Mohsen Mohammadzadeh, 2010. "Bayesian regression analysis with linked data using mixture normal distributions," Statistical Papers, Springer, vol. 51(2), pages 421-430, June.
Tatiana Komarova & Denis Nekipelov & Evgeny Yakovlev, 2018. "Identification, data combination, and the risk of disclosure," Quantitative Economics, Econometric Society, vol. 9(1), pages 395-440, March.
- Tatiana V. Komarova & Denis Nekipelov & Evgeny Yakovlev, 2011. "Identification, data combination and the risk of disclosure," CeMMAP working papers CWP38/11, Centre for Microdata Methods and Practice, Institute for Fiscal Studies.
- Komarova, Tatiana & Nekipelov, Denis & Yakovlev, Evgeny, 2018. "Identification, data combination and the risk of disclosure," LSE Research Online Documents on Economics 79384, London School of Economics and Political Science, LSE Library.
Vo, Thanh Huan & Chauvet, Guillaume & Happe, André & Oger, Emmanuel & Paquelet, Stéphane & Garès, Valérie, 2023. "Extending the Fellegi-Sunter record linkage model for mixed-type data with application to the French national health data system," Computational Statistics & Data Analysis, Elsevier, vol. 179(C).
Deborah Wagner & Mary Lane, 2014. "The Person Identification Validation System (PVS): Applying the Center for Administrative Records Research and Applications’ (CARRA) Record Linkage Software," CARRA Working Papers 2014-01, Center for Economic Studies, U.S. Census Bureau.
N. Salvati & E. Fabrizi & M. G. Ranalli & R. L. Chambers, 2021. "Small area estimation with linked data," Journal of the Royal Statistical Society Series B, Royal Statistical Society, vol. 83(1), pages 78-107, February.

More about this item

Keywords

; ; ; ;

Statistics

Access and download statistics

Corrections

All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:eee:csdana:v:56:y:2012:i:9:p:2756-2770. See general information about how to correct material in RePEc.

If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

If CitEc recognized a bibliographic reference but did not link an item in RePEc to it, you can help with this form .

If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Catherine Liu (email available below). General contact details of provider: http://www.elsevier.com/locate/csda .

Please note that corrections may take a couple of weeks to filter through the various RePEc services.

IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.

Browse Econ Literature

More features

Regression analysis under incomplete linkage

Author

Abstract

Suggested Citation

Download full text from publisher

References listed on IDEAS

Citations

Most related items

More about this item

Keywords

Statistics

Corrections

More services and features

MyIDEAS

Author registration

Rankings

RePEc Genealogy

RePEc Biblio

MPRA

New papers by email

EconAcademics

Plagiarism

About RePEc

RePEc home

Blog

Help/FAQ

RePEc team

Participating archives

Privacy statement

Help us

Corrections

Volunteers

Get papers listed

Open a RePEc archive

Get RePEc data