Author
Abstract
Diabetes, a pervasive metabolic disorder, presents a significant global health challenge, with its escalating prevalence contributing to millions of annual fatalities. This study aims to develop machine learning techniques to enhance the accuracy of early diabetes diagnosis and to address the class imbalance in medical datasets, which often skews machine learning model performance. The research question focuses on the effectiveness of machine learning techniques and class imbalance rectification methods in improving diabetes prediction accuracy. Using the Pima Indians Diabetes Database, this study applies the Instance Hardness Threshold (IHT) undersampling technique and the Synthetic Minority Over-sampling Technique (SMOTE) to address class imbalances. Data preprocessing steps include imputing missing values, rebalancing the dataset, normalizing data, and selecting relevant features. Recursive Feature Elimination identifies vital features such as glucose, BMI, skin thickness, insulin, diabetes pedigree function, and age. Decision Trees, K-Nearest Neighbour, Support Vector Machines, Random Forest, and Artificial Neural Networks classification algorithms have been evaluated and compared. The Random Forest algorithm emerges as the most effective, achieving an accuracy of 97.21%, precision of 95.5%, recall of 98.8%, F1-score of 97.1%, and AUC of 99.0% with the IHT undersampling technique. Other algorithms also show enhanced performance metrics following the application of resampling techniques. The study underscores the importance of addressing class imbalances in datasets for diabetes prediction. The proposed framework enhances early diabetes diagnosis and contributes to improved healthcare efficiency. Future research should explore larger datasets and incorporate advanced deep-learning methods to refine predictive performance further.
Suggested Citation
Abdoul Malik & Cengiz Tepe, 2025.
"Improving Early Diabetes Diagnosis : A Machine Learning Approach for Class Imbalance Mitigation,"
International Journal of Scientific Research in Computer Science, Engineering and Information Technology, International Journal of Scientific Research in Computer Science, Engineering and Information Technology, vol. 11(3), pages 861-874, June.
Handle:
RePEc:jbh:ijsrcs:v11:y2025:i3:id:1534
DOI: 10.32628/CSEIT25113322
Note: Article URL: https://ijsrcseit.com/home/article/view/CSEIT25113322
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:jbh:ijsrcs:v11:y2025:i3:id:1534. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Pankaj Sharma (USA) (email available below). General contact details of provider: https://ijsrcseit.com/home .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.