Harmonizing and Combining Large Datasets – An Application to Firm-Level Patent and Accounting Data
AbstractThis paper discusses methods for the harmonization and combination of large-scale patent and trademark datasets with each other and other sources of data. Dictionary- and rule-based approaches to the consolidation of applicant names in patent data are presented and shown to have both benefits and drawbacks in isolation. We combine the two methods and develop a set of rules and dictionaries to consolidate European, Patent Cooperation Treaty (PCT) and US patent data with firm accounting data. The resulting data encompass about 131,000 patent applicant names from 46 countries, covering 58.8 percent of EPO applications and 50.6 percent of PCT applications by business organizations during the time period from 1979 to 2008. For US data, the resulting dataset includes around 54,000 assignee names and 51.3 percent of US granted patents during approximately the same time period.
Download InfoIf you experience problems downloading a file, check if you have the proper application to view it first. In case of further problems read the IDEAS help page. Note that these files are not on the IDEAS site. Please be patient as the files may be large.
Bibliographic InfoPaper provided by National Bureau of Economic Research, Inc in its series NBER Working Papers with number 15851.
Date of creation: Mar 2010
Date of revision:
Contact details of provider:
Postal: National Bureau of Economic Research, 1050 Massachusetts Avenue Cambridge, MA 02138, U.S.A.
Web page: http://www.nber.org
More information through EDIRC
Find related papers by JEL classification:
- C81 - Mathematical and Quantitative Methods - - Data Collection and Data Estimation Methodology; Computer Programs - - - Methodology for Collecting, Estimating, and Organizing Microeconomic Data
- O34 - Economic Development, Technological Change, and Growth - - Technological Change; Research and Development; Intellectual Property Rights - - - Intellectual Property Rights
You can help add them by filling out this form.
CitEc Project, subscribe to its RSS feed for this item.
- Justus Baron & Yann Ménière & Tim Pohlmann, 2012. "Joint innovation in ICT standards: How consortia drive the volume of patent filings," Working Papers hal-00707291, HAL.
- Cristiano Antonelli & Alessandra Colombelli, 2013.
"Knowledge Cumulability and Complementarity in the Knowledge Generation Function,"
GREDEG Working Papers
2013-08, Groupe de REcherche en Droit, Économie, Gestion (GREDEG CNRS), University of Nice Sophia Antipolis.
- Antonelli Cristiano & Colombelli Alessandra, 2013. "Knowledge cumulability and complementarity in the knowledge generation function," Department of Economics and Statistics Cognetti de Martiis. Working Papers 201305, University of Turin.
- Antonelli,Cristiano & Colombelli, Alessandra, 2013. "Knowledge cumulability and complementarity in the knowledge generation function," Department of Economics and Statistics Cognetti de Martiis LEI & BRICK - Laboratory of Economics of Innovation "Franco Momigliano", Bureau of Research in Innovation, Complexity and Knowledge, Collegio 201303, University of Turin.
- Alessandra Colombelli & Jackie Krafft & Francesco Quatraro, 2013.
"Properties of knowledge base and firm survival: Evidence from a sample of French manufacturing firms,"
- Colombelli Alessandra & Krafft.Jackie & Quatraro Francesco, 2012. "Properties of knowledge base and firm survival: Evidence from a sample of French manufacturing firms," Department of Economics and Statistics Cognetti de Martiis LEI & BRICK - Laboratory of Economics of Innovation "Franco Momigliano", Bureau of Research in Innovation, Complexity and Knowledge, Collegio 201209, University of Turin.
- Michele PEZZONI (University of Milano-Bicocca - KiTES-Università Bocconi - Observatoire des Sciences et des Techniques) & Francesco LISSONI (GREThA, CNRS, UMR 5113 - KiTES) & Gianluca TARASCONI (KiTE, 2012. "How To Kill Inventors: Testing The Massacrator© Algorithm For Inventor Disambiguation," Cahiers du GREThA 2012-29, Groupe de Recherche en Economie Théorique et Appliquée.
- Niels Bosma & Frank van Oort, 2012. "Agglomeration Economies, Inventors and Entrepreneurs as Engines of European Regional Productivity," Working Papers 12-20, Utrecht School of Economics.
- Christoph Ernst & Katharina Richter & Nadine Riedel, 2013. "Corporate taxation and the quality of research & development," Working Papers 1301, Oxford University Centre for Business Taxation.
- Isabel Tecu, 2013. "The Location of Industrial Innovation: Does Manufacturing Matter?," Working Papers 13-09, Center for Economic Studies, U.S. Census Bureau.
- repec:hal:wpaper:hal-00686007 is not listed on IDEAS
- Markus Eberhardt & Christian Helmers & Zhihong Yu, 2011.
"Is the Dragon Learning to Fly? An Analysis of the Chinese Patent Explosion,"
CSAE Working Paper Series
2011-15, Centre for the Study of African Economies, University of Oxford.
- Markus Eberhardt & Christian Helmers & Zhihong Yu, . "Is the Dragon Learning to Fly? An Analysis of the Chinese Patent Explosion," Discussion Papers 11/16, University of Nottingham, GEP.
- Markus Eberhardt & Christian Helmers, 2011. "Is the Dragon Learning to Fly? An Analysis of the Chinese Patent Explosion," Economics Series Working Papers WPS/2011-15, University of Oxford, Department of Economics.
- Brown, James R. & Martinsson, Gustav & Petersen, Bruce C., 2012. "Do financing constraints matter for R&D?," European Economic Review, Elsevier, vol. 56(8), pages 1512-1529.
- Patrick Llerena & Valentine Millot, 2013. "Are Trade Marks and Patents Complementary or Substitute Protections for Innovation," Working Papers of BETA 2013-01, Bureau d'Economie Théorique et Appliquée, UDS, Strasbourg.
- Alessandra Colombelli & Francesco Quatraro, 2013. "The persistence of firms' knowledge base: A quantile approach to Italian data," Working Papers hal-00867132, HAL.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: ().
If references are entirely missing, you can add them using this form.