Privacy preservation in data mining using hybrid perturbation methods: an application to bankruptcy prediction in banks
AbstractToday, the data related to business, finance and healthcare pose problems for Privacy-Preserving Data Mining (PPDM). Privacy regulations and concerns prevent data owners from sharing data for mining purposes. To circumvent this problem, data owners must design strategies to meet privacy requirements and ensure valid data mining results. This paper proposes the hybridisation of the random projection and random rotation methods for privacy-preserving classification. The hybrid method is tested on six benchmark data sets and four bank bankruptcy data sets. These methods ensure the privacy and secrecy of bank data and the resulting data set is mined without a considerable loss of accuracy. A multilayer perceptron, decision tree J48 and logistic regression are used as classifiers. The results of a tenfold cross-validation and t-test indicate improved average accuracies for the hybrid privacy preservation method compared to when random projection is used alone. The reasons for the superior performance of the hybrid privacy preservation method are also highlighted.
Download InfoIf you experience problems downloading a file, check if you have the proper application to view it first. In case of further problems read the IDEAS help page. Note that these files are not on the IDEAS site. Please be patient as the files may be large.
As the access to this document is restricted, you may want to look for a different version under "Related research" (further below) or search for a different version of it.
Bibliographic InfoArticle provided by Inderscience Enterprises Ltd in its journal Int. J. of Data Analysis Techniques and Strategies.
Volume (Year): 1 (2009)
Issue (Month): 4 ()
Contact details of provider:
Web page: http://www.inderscience.com/browse/index.php?journalID=282
privacy preservation; data mining; PPDM; random projection; random rotation; stress function; bankruptcy prediction; classification; multilayer perceptron; decision tree J48; logistic regression; banks.;
Find related papers by JEL classification:
- dec - - - - - -
- tre - - - - - -
- J48 - Labor and Demographic Economics - - Particular Labor Markets - - - Particular Labor Markets; Public Policy
You can help add them by filling out this form.
reading list or among the top items on IDEAS.Access and download statisticsgeneral information about how to correct material in RePEc.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: (Graham Langley) or (Christopher F. Baum).
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
If references are entirely missing, you can add them using this form.
If the full references list an item that is present in RePEc, but the system did not link to it, you can help with this form.
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your profile, as there may be some citations waiting for confirmation.
Please note that corrections may take a couple of weeks to filter through the various RePEc services.