SMOTETomek-DNN: A Machine Learning Framework for Credit Risk Prediction with an Imbalanced Dataset

Authors

  • Tiruneh Kebede Dubale Wolaita Sodo University, Ethiopia
  • Siraj Sebhatu Seyoum Wolaita Sodo University, Ethiopia

DOI:

https://doi.org/10.64539/sjcs.v2i2.2026.520

Keywords:

Credit risk prediction, Microfinance institutions, Imbalanced learning, Machine learning, Sampling techniques

Abstract

Accurate credit risk prediction plays a critical role in strengthening the financial stability of microfinance institutions, especially in developing economies, where increasing loan defaults and imbalanced borrower records create significant challenges for reliable decision-making. Although machine learning approaches have improved credit assessment practices, existing models often favor majority-class borrowers and fail to detect high-risk default cases effectively because of severe class imbalance. This limitation highlights the need for more robust and imbalanced-sensitive predictive frameworks. This study aims to develop an effective machine learning-based credit risk prediction framework by integrating data balancing strategies with ensemble and deep learning models. This study systematically investigates the impact of baseline learning and multiple resampling techniques, including oversampling, undersampling, and hybrid methods, when applied to Random Forest, XGBoost, LightGBM, CatBoost, and Deep Neural Network classifiers. The effectiveness of the proposed models was assessed using imbalance-aware evaluation measures, particularly ROC-AUC and Geometric Mean, along with conventional classification metrics. The experimental findings demonstrate that incorporating resampling techniques substantially improved the default detection performance. The DNN model combined with SMOTETomek achieved the best results, obtaining 94.9% F1-score, ROC-AUC 98%, and 97.2% of G-Mean. CatBoost also exhibited consistent competitiveness across different sampling configurations. These findings suggest that hybrid sampling integrated with advanced learning architectures can provide a reliable and practical solution for managing credit risk in imbalanced microfinance datasets, supporting improved lending decisions and sustainable financial operations in the future.

References

[1] M. T. Molla, “Association of risk management practices and financial performance of microfinance institutions in Ethiopia : A two-step system generalized method of moments approach,” PLOS One, pp. 1–18, 2025. https://doi.org/10.1371/journal.pone.0321415.

[2] R. Sifrain, “Factors Influencing Loan Portfolio Quality of Microfinance Institutions in Haiti,” Journal of Financial Risk Management, vol. 11, no. 1, pp. 95–115, 2022. https://doi.org/10.4236/jfrm.2022.111005.

[3] H. A. Legass and H. A. Roba, “The Impact of Credit Risk Management on Financial Performance of Commercial Banks in Ethiopia,” Journal of Finance and Accounting, vol. 12, no. 6, pp. 142–155, 2024. https://doi.org/10.11648/j.jfa.20241206.11.

[4] V. García, A. I. Marqués, and J. S. Sánchez, “Exploring the synergetic effects of sample types on the performance of ensembles for credit risk and corporate bankruptcy prediction,” Inf. Fusion, vol. 47, pp. 88–101, 2019. https://doi.org/10.1016/j.inffus.2018.07.004.

[5] F. Barboza, H. Kimura, and E. Altman, “Machine learning models and bankruptcy prediction,” Expert Syst. Appl., vol. 83, pp. 405–417, 2017. https://doi.org/10.1016/j.eswa.2017.04.006.

[6] S. Shi, R. Tse, W. Luo, S. D’Addona, and G. Pau, “Machine learning-driven credit risk: a systemic review,” Neural Comput. Appl., vol. 34, no. 17, pp. 14327–14339, 2022. https://doi.org/10.1007/s00521-022-07472-2.

[7] S. Zhang, J. Tay, and P. Baiz, “The Effects of Data Imbalance Under a Federated Learning Approach for Credit Risk Forecasting,” arXiv preprint arXiv:2401.07234, 2024. https://doi.org/10.48550/arXiv.2401.07234.

[8] A. Namvar, M. Siami, F. Rabhi, and M. Naderpour, “Credit risk prediction in an imbalanced social lending environment,” International Journal of Computational Intelligence Systems, vol. 11, pp. 925–935, 2018. https://doi.org/10.2991/ijcis.11.1.70.

[9] E. B. Abdelghafour, C. Mohamed, A. Noura, and B. Abdelhamid, “Enhancing Credit Card Fraud Detection Using a Stacking Model Approach and Hyperparameter Optimization,” International Journal of Advanced Computer Science and Applications (IJACSA), vol. 15, no. 10, 2024. https://doi.org/10.14569/ijacsa.2024.01510110.

[10] H. Wang, “Credit Risk Management of Consumer Finance Based on Big Data,” Mobile Information Systems, vol. 2021, 2021. https://doi.org/10.1155/2021/8189255.

[11] L. Gao and J. Xiao, “Big Data Credit Report in Credit Risk Management of Consumer Finance,” Wireless Communications and Mobile Computing, vol. 2021, 2021. https://doi.org/10.1155/2021/4811086.

[12] A. Niu, B. Cai, and S. Cai, “Big Data Analytics for Complex Credit Risk Assessment of Network Lending Based on SMOTE Algorithm,” Complexity, vol. 2020, 2020. https://doi.org/10.1155/2020/8563030.

[13] J. P. Noriega, L. A. Rivera, and J. A. Herrera, “Machine Learning for Credit Risk Prediction: A Systematic Literature Review,” Data, vol. 8, no. 11, 2023. https://doi.org/10.3390/data8110169.

[14] N. Chai, M. Zoynul, L. Yang, and B. Shi, “Farmers’ credit risk evaluation with an explainable hybrid ensemble approach : A closer look in microfinance,” Pacific-Basin Financ. J., vol. 89, no. June 2024, p. 102612, 2025. https://doi.org/10.1016/j.pacfin.2024.102612.

[15] L. Marceau, L. Qiu, N. Vandewiele, E. Charton, “A comparison of Deep Learning performances with other machine learning algorithms on credit scoring unbalanced data,” arXiv preprint arXiv:1907.12363, 2019. https://doi.org/10.48550/arXiv.1907.12363.

[16] A. Ampountolas, T. N. Nde, P. Date, and C. Constantinescu, “A Machine Learning Approach for Micro-Credit Scoring,” Risks, vol. 9, no. 3, 2021. https://doi.org/10.3390/risks9030050.

[17] M. Z. Abedin, C. Guotai, P. Hajek, and T. Zhang, “Combining weighted SMOTE with ensemble learning for the class-imbalanced prediction of small business credit risk,” Complex Intell. Syst., vol. 9, no. 4, pp. 3559–3579, 2023. https://doi.org/10.1007/s40747-021-00614-4.

[18] I. Aruleba and Y. Sun, “Effective Credit Risk Prediction Using Ensemble Classifiers With Model Explanation,” IEEE Access, pp. 115015 – 115025, vol. 12, 2024. https://doi.org/10.1109/ACCESS.2024.3445308.

[19] X. Liu, Z. Zhang, and D. Wang, “Classification of Imbalanced Credit scoring data sets Based on Ensemble Method with the Weighted-Hybrid-Sampling,” arXiv preprint arXiv:2102.04721, 2021. https://doi.org/10.48550/arXiv.2102.04721.

[20] E. Alfredo, S. Alvarez, M. Angel, and C. Lengua, “Design of a machine learning model for predicting credit risk in microfinance using environmental data,” Bulletin of Electrical Engineering and Informatics (BEEI), vol. 14, no. 6, pp. 4578–4589, 2025. https://doi.org/10.11591/eei.v14i6.9848.

[21] C. Yu, Y. Jin, Q. Xing, Y. Zhang, S. Guo, S. Meng, “Advanced User Credit Risk Prediction Model Using LightGBM, XGBoost and Tabnet with SMOTEENN,” in 2024 IEEE 6th International Conference on Power, Intelligent Computing and Systems (ICPICS), 2024. https://doi.org/10.1109/ICPICS62053.2024.10796247.

[22] N. A. Siagian, S. P. Sipayung, A. Rikki, and N. Marbun, “Integrating SMOTE with XGBoost for Robust Classification on Imbalanced Datasets: A Dual-Domain Evaluation,” Sinkron: Jurnal dan Penelitian Teknik Informatika, vol. 9, no. 3, pp. 1094–1107, 2025. https://doi.org/10.33395/sinkron.v9i3.15029.

[23] T. Sodnomdavaa and M. Sandagsuren, “Temporal and Cost-Sensitive Evaluation Framework for Credit Risk Modeling Under Distributional Shifts,” Risks, vol. 14, no. 4, art. No. 95, 2026. https://doi.org/10.3390/risks14040095.

[24] Z. Zhao, T. Cui, S. Ding, J. Li, and A. G. Bellotti, “Resampling Techniques Study on Class Imbalance Problem in Credit Risk Prediction,” Mathematics, vol. 12, no. 5, art. No. 701, 2024. https://doi.org/10.3390/math12050701.

[25] A. B. Shaik and S. Srinivasan, “A Brief Survey on Random Forest Ensembles in Classification Model,” in International Conference on Innovative Computing and Communications, pp. 253–260, 2018. https://doi.org/10.1007/978-981-13-2354-6_27.

[26] T. Chen and C. Guestrin, “XGBoost: A Scalable Tree Boosting System,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 785–794, 2016. https://doi.org/10.1145/2939672.2939785.

[27] A. H. Ginting, R. W. Sembiring, and E. M. Zamzami, “Comparison Analysis: Logistic Regression, Random Forest, XGBoost, and CatBoost in Credit Scoring,” in Proceedings of the 2024 Brawijaya International Conference (BIC 2024), 2025. https://doi.org/10.2991/978-94-6463-854-7_16.

[28] H. A. Salman, A. Kalakech, and A. Steiti, “Random Forest Algorithm Overview,” Babylonian Journal of Machine Learning, vol. 2024, pp. 69–79, 2024. https://doi.org/10.58496/BJML/2024/007.

[29] A. O. Kuyoro, O. A. Ogunyolu, T. G. Ayanwola, and F. Yetunde, “Dynamic Effectiveness of Random Forest Algorithm in Financial Credit Risk Management for Improving Output Accuracy and Loan Classification Prediction,” Ingénierie des Systèmes d’Information, vol. 27, no. 5, pp. 815–821, 2022. https://doi.org/10.18280/isi.270515.

[30] C.-H. Chen, J.-P. Lai, Y.-M. Chang, C.-J. Lai, and P.-F. Pai, “A Study of Optimization in Deep Neural Networks for Regression,” Electronics, vol. 12, no. 14, art. No. 3071, 2023. https://doi.org/10.3390/electronics12143071.

[31] M. Schmitt, “Deep Learning vs. Gradient Boosting: Benchmarking state-of-the-art machine learning algorithms for credit scoring,” arXiv preprint arXiv:2205.10535, 2022. https://doi.org/10.48550/arXiv.2205.10535.

[32] O. Radović, S. Marinković, and J. Radojičić, “Credit Scoring with an Ensemble Deep Learning Classification Methods – Comparison with Traditional Methods,” Facta Universitatis, Series: Economics and Organization, vol. 18, no. 1, pp. 29–43, 2021. https://doi.org/10.22190/FUEO201028001R.

[33] Y. Chen, R. Calabrese, and B. Martin-barragan, “Interpretable machine learning for imbalanced credit scoring datasets,” Eur. J. Oper. Res., vol. 312, no. 1, pp. 357–372, 2024. https://doi.org/10.1016/j.ejor.2023.06.036.

[34] L. Wang, “Imbalanced credit risk prediction based on SMOTE and multi-kernel FCM improved by particle swarm optimization,” Appl. Soft Comput., vol. 114, p. 108153, 2022. https://doi.org/10.1016/j.asoc.2021.108153.

[35] M. Song, H. Ma, Y. Zhu, M. Zhang, “Credit Risk Prediction Based on Improved ADASYN Sampling and Optimized LightGBM,” Journal of Social Computing, vol. 5, no. 3, pp. 232-241, 2024. https://doi.org/10.23919/JSC.2024.0019.

[36] L. H. Chia, “Finding the Sweet Spot: Optimal Data Augmentation Ratio for Imbalanced Credit Scoring Using ADASYN,” arXiv preprint arXiv:2510.18252, 2025. https://doi.org/10.48550/arXiv.2510.18252.

[37] J. Guo, H. Wu, X. Chen, and W. Lin, “Adaptive SV-Borderline SMOTE-SVM algorithm for imbalanced data classification,” Appl. Soft Comput., vol. 150, p. 110986, 2024. https://doi.org/10.1016/j.asoc.2023.110986.

[38] A. S. Almajid and R. Arifudin, “Multilayer Perceptron Optimization on Imbalanced Data Using SVM-SMOTE and One-Hot Encoding for Credit Card Default Prediction,” Journal of Advances in Information Systems and Technology, vol. 3, no. 2, pp. 67–74, 2021. https://doi.org/10.15294/jaist.v3i2.57061.

[39] B. Krawczyk, “Learning from imbalanced data: open challenges and future directions,” Prog. Artif. Intell., vol. 5, no. 4, pp. 221–232, 2016. https://doi.org/10.1007/s13748-016-0094-0.

[40] G. Hou, D. L. Tong, S. Y. Liew, and P. Y. Choo “Comparative Analysis of Resampling Techniques for Class Imbalance in Financial Distress Prediction Using XGBoost,” Mathematics, vol. 13, no. 13, art. No. 2186, 2025. https://doi.org/10.3390/math13132186.

[41] O.-A. Ampomah, E. Agyemang, K. Acheampong, and L. Agyekum, “Enhancing Credit Default Prediction Using Boruta Feature Selection and DBSCAN Algorithm with Different Resampling Techniques,” arXiv preprint arXiv:2509.19408, 2025. https://doi.org/10.48550/arXiv.2509.19408.

[42] C. Vairetti, J. L. Assadi, and S. Maldonado, “Efficient hybrid oversampling and intelligent undersampling for imbalanced big data classification,” Expert Systems with Applications, vol. 246, 2024. https://doi.org/10.1016/j.eswa.2024.123149.

[43] I. Aruleba and Y. Sun, “Machine Learning with Applications Enhanced credit risk prediction using deep learning and SMOTE-ENN resampling,” Mach. Learn. with Appl., vol. 21, no. July 2024, p. 100692, 2025. https://doi.org/10.1016/j.mlwa.2025.100692.

[44] G. E. A. P. A. Batista, R. C. Prati, and M. C. Monard, “A study of the behavior of several methods for balancing machine learning training data,” ACM SIGKDD Explorations Newsletter, vol. 6, no. 1, pp. 20 – 29, 2004. https://doi.org/10.1145/1007730.1007735.

[45] N. V Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic Minority Over-sampling Technique,” Journal of Artificial Intelligence Research (JAIR), vol. 16, pp. 321–357, 2002. https://doi.org/10.1613/jair.953.

[46] Z. Zhao and V. Aumeboonsuke, “Imbalanced Credit Risk Prediction in Ensemble Learning Classifiers: A Comparative Analysis of SMOTE, ADASYN, SMOTETomek, and Cluster Centroids,” Arts of Management Journal, vol. 7, no. 3, pp. 959–984, 2023. https://so02.tci-thaijo.org/index.php/jam/article/view/264093.

Downloads

Published

2026-08-12

How to Cite

Dubale, T. K., & Seyoum, S. S. (2026). SMOTETomek-DNN: A Machine Learning Framework for Credit Risk Prediction with an Imbalanced Dataset. Scientific Journal of Computer Science, 2(2), 355–373. https://doi.org/10.64539/sjcs.v2i2.2026.520

Similar Articles

<< < 1 2 3 > >> 

You may also start an advanced similarity search for this article.