IDEAS home Printed from https://ideas.repec.org/a/taf/japsta/v47y2020i3p568-581.html
   My bibliography  Save this article

A descriptive study of variable discretization and cost-sensitive logistic regression on imbalanced credit data

Author

Listed:
  • Lili Zhang
  • Herman Ray
  • Jennifer Priestley
  • Soon Tan

Abstract

Training classification models on imbalanced data tends to result in bias towards the majority class. In this paper, we demonstrate how variable discretization and cost-sensitive logistic regression help mitigate this bias on an imbalanced credit scoring dataset, and further show the application of the variable discretization technique on the data from other domains, demonstrating its potential as a generic technique for classifying imbalanced data beyond credit socring. The performance measurements include ROC curves, Area under ROC Curve (AUC), Type I Error, Type II Error, accuracy, and F1 score. The results show that proper variable discretization and cost-sensitive logistic regression with the best class weights can reduce the model bias and/or variance. From the perspective of the algorithm, cost-sensitive logistic regression is beneficial for increasing the value of predictors even if they are not in their optimized forms while maintaining monotonicity. From the perspective of predictors, the variable discretization performs better than cost-sensitive logistic regression, provides more reasonable coefficient estimates for predictors which have nonlinear relationships against their empirical logit, and is robust to penalty weights on misclassifications of events and non-events determined by their apriori proportions.

Suggested Citation

  • Lili Zhang & Herman Ray & Jennifer Priestley & Soon Tan, 2020. "A descriptive study of variable discretization and cost-sensitive logistic regression on imbalanced credit data," Journal of Applied Statistics, Taylor & Francis Journals, vol. 47(3), pages 568-581, February.
  • Handle: RePEc:taf:japsta:v:47:y:2020:i:3:p:568-581
    DOI: 10.1080/02664763.2019.1643829
    as

    Download full text from publisher

    File URL: http://hdl.handle.net/10.1080/02664763.2019.1643829
    Download Restriction: Access to full text is restricted to subscribers.

    File URL: https://libkey.io/10.1080/02664763.2019.1643829?utm_source=ideas
    LibKey link: if access is restricted and if your library uses this service, LibKey will redirect you to where you can use your library subscription to access this item
    ---><---

    As the access to this document is restricted, you may want to search for a different version of it.

    Citations

    Citations are extracted by the CitEc Project, subscribe to its RSS feed for this item.
    as


    Cited by:

    1. Fisnik Doko & Slobodan Kalajdziski & Igor Mishkovski, 2021. "Credit Risk Model Based on Central Bank Credit Registry Data," JRFM, MDPI, vol. 14(3), pages 1-17, March.

    More about this item

    Statistics

    Access and download statistics

    Corrections

    All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:taf:japsta:v:47:y:2020:i:3:p:568-581. See general information about how to correct material in RePEc.

    If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.

    We have no bibliographic references for this item. You can help adding them by using this form .

    If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.

    For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Chris Longhurst (email available below). General contact details of provider: http://www.tandfonline.com/CJAS20 .

    Please note that corrections may take a couple of weeks to filter through the various RePEc services.

    IDEAS is a RePEc service. RePEc uses bibliographic data supplied by the respective publishers.