Tóm tắt:
Mục tiêu nghiên cứu: Nghiên cứu này phân tích và so sánh hiệu quả của bốn mô hình dựa trên cây gồm cây quyết định, rừng ngẫu nhiên, cây quyết định khả vi và rừng quyết định nơ-ron sâu trong dự báo rủi ro vỡ nợ (RRVN) của doanh nghiệp nhỏ và vừa (SMEs) tại Việt Nam.
Thiết kế nghiên cứu/phương pháp/tiếp cận: Trên cơ sở bộ dữ liệu gồm nhiều chỉ số tài chính doanh nghiệp, nghiên cứu triển khai quy trình huấn luyện và đánh giá thực nghiệm để đối chiếu hiệu suất dự báo giữa các mô hình theo những tiêu chí phân loại và khả năng giải thích kết quả.
Kết quả nghiên cứu chính: Kết quả cho thấy mô hình cây quyết định khả vi đạt hiệu suất tốt nhất, thể hiện ưu thế trong nhận diện quan hệ phi tuyến và phân loại RRVN. Đồng thời, các biến phản ánh khả năng thanh khoản, cấu trúc nợ và mức độ chịu đựng chi phí tài chính được xác định là những nhân tố có giá trị dự báo nổi bật.
Giá trị đóng góp mới: Nghiên cứu bổ sung bằng chứng thực nghiệm về ứng dụng các mô hình dựa trên cây và mô hình lai cây–học sâu trong bối cảnh SMEs tại Việt Nam. Đồng thời, bài viết cung cấp hàm ý thực tiễn cho việc lựa chọn mô hình phục vụ quản trị rủi ro tín dụng tại các tổ chức tài chính.
Tài liệu tham khảo:
- Altman, E. I. (1968). Financial ratios, discriminant analysis and the prediction of corporate bankruptcy. Journal of Finance.
- Balestriero, R. (2017). Neural decision trees. arXiv preprint arXiv:1702.07360.
- Bian, L., Qin, X., Zhang, C., Guo, P., & Wu, H. (2023). Application, interpretability and prediction of machine learning method combined with LSTM and LightGBM-a case study for runoff simulation in an arid area. Journal of Hydrology, 625. https://doi.org/10.1016/j.jhydrol.2023.130091.
- Cui, S., Wang, Y., Yin, Y., Cheng, T. C. E., Wang, D., & Zhai, M. (2021). A cluster-based intelligence ensemble learning method for classification problems. Information Sciences, 560, 386-409. https://doi.org/10.1016/j.ins.2021.01.061.
- Dapogny, A., & Bailly, K. (2018). Face alignment with cascaded semi-parametric deep greedy neural forests. Pattern recognition letters, 102, 75-81. https://doi.org/10.1016/j.patrec.2017.12.010.
- Detthamrong, U., Chansanam, W., Boongoen, T., & Iam-On, N. (2024). Enhancing fraud detection in banking using advanced machine learning techniques. International Journal of Economics and Financial Issues, 14(5), 177-184. https://doi.org/10.32479/ijefi.16613.
- Farhangi, F. (2022). Investigating the role of data preprocessing, hyperparameters tuning, and type of machine learning algorithm in the improvement of drowsy EEG signal modeling. Intelligent Systems with Applications, 15, 200100. https://doi.org/10.1016/j.iswa.2022.200100.
- Gajdosikova, D., Valaskova, K., & Durana, P. (2026). Cross-National Benchmarking of Bankruptcy Prediction Models Across V4 Economies. International Journal of Economic Sciences, 1-19. https://doi.org/10.31181/ijes1512026223.
- Hernandez Tinoco, M., & Wilson, N. (2013). Financial distress and bankruptcy prediction among listed companies using accounting, market and macroeconomic variables.
- Irsoy, O., Yıldız, O. T., & Alpaydın, E. (2012). Soft decision trees. In Proceedings of the 21st international conference on pattern recognition (ICPR2012) (pp. 1819-1822). IEEE.
- Kalra, A., & Brown, D. S. (2022). Interpretable reward learning via differentiable decision trees. In NeurIPS ML Safety Workshop.
- Kaya, O. (2022). Determinants and consequences of SME insolvency risk during the pandemic. Economic Modelling, 115, 105958.
- Lao, Z., He, D., Wei, Z., Shang, H., Jin, Z., Miao, J., & Ren, C. (2023). Intelligent fault diagnosis for rail transit switch machine based on adaptive feature selection and improved LightGBM. Engineering Failure Analysis, 148. https://doi.org/10.1016/j.engfailanal.2023.107219.
- Lapuschkin, S., Wäldchen, S., Binder, A., Montavon, G., Samek, W., & Müller, K. R. (2019). Unmasking Clever Hans predictors and assessing what machines really learn. Nature communications, 10(1), 1096. https://doi.org/10.1038/s41467-019-08987-4.
- Lee, S., Choi, K., & Yoo, D. (2023). Building a core rule-based decision tree to explain the causes of insolvency in small and medium-sized enterprises more easily. Humanities and Social Sciences Communications, 10(1), 1-16. https://doi.org/10.1057/s41599-023-02382-7.
- Manokaran, J., & Vairavel, G. (2023). GIWRF-SMOTE: Gini impurity-based weighted random forest with SMOTE for effective malware attack and anomaly detection in IoT-Edge. Smart Science, 11(2), 276-292.
- Maturo, F., & Verde, R. (2022). Pooling random forest and functional data analysis for biomedical signals supervised classification: Theory and application to electrocardiogram data. Statistics in Medicine, 41(12), 2247-2275. https://doi.org/10.1002/sim.9353.
- Nguyễn Minh Nhật, & Ngô Hoàng Khánh Duy (2025). Lựa chọn đặc trưng và dự báo rủi ro vỡ nợ doanh nghiệp: thực nghiệm với mô hình học máy. Tạp chí Nghiên cứu Tài chính - Marketing, 90 (Tập 16, kỳ 3), 74-85. https://doi.org/10.52932/jfmr.v16i3.
- Ohlson, J. A. (1980). Financial ratios and the probabilistic prediction of bankruptcy. Journal of Accounting Research.
- Ozili, P. K. (2025). Bank non-performing loans research around the world. Asian Journal of Economics and Banking, 9(3), 437–462. https://doi.org/10.1108/AJEB-09-2024-0103
- Papíková, L., & Papík, M. (2022). Effects of classification, feature selection, and resampling methods on bankruptcy prediction of small and medium‐sized enterprises. Intelligent Systems in Accounting, Finance and Management, 29(4), 254-281. https://doi.org/10.1002/isaf.1521.
- Platt, H. D., & Platt, M. B. (2002). Predicting corporate financial distress: Reflections on choice-based sample bias.
- Qiu, W., Rudkin, S., & Dłotko, P. (2020). Refining understanding of corporate failure through a topological data analysis mapping of Altman’s Z-score model. Expert Systems with Applications, 156, 113475.
- Reddy, K. S., Yadav, R. S., & Agarwal, A. (2025). Credit exposures and systemic risk in Indian banks. Asian Journal of Economics and Banking, 9(2), 222–239. https://doi.org/10.1108/AJEB-12-2024-0148
- Saarela, M., & Jauhiainen, S. (2021). Comparison of feature importance measures as explanations for classification models. SN Applied Sciences, 3(2), 272. https://doi.org/10.1007/s42452-021-04148-9.
- Salih, A. M. (2024). Explainable Artificial Intelligence for Dependent Features: Additive Effects of Collinearity. In Proceedings of the 2024 8th International Conference on Advances in Artificial Intelligence (pp. 94-99). https://doi.org/10.1145/3704137.3704152.
- Sandhu, G., Singh, A., Bedi, P. S., Lamba, P. S., & Chaudhary, G. (2024). Improvement of Random Forest Ensembling Algorithm Efficiency Through Cardinal Tuning of n_estimators Parameter. Journal of Multiple-Valued Logic & Soft Computing, 42.
- Silva, A., Killian, T., Jimenez, I., Son, S. H., & Gombolay, M. (2020). Optimization methods for interpretable differentiable decision trees applied to reinforcement learning. In International conference on artificial intelligence and statistics, (pp. 1855-1865).
- Wang, J., Zheng, Y., Li, X., Yu, C., Kodaka, K., & Li, K. (2015). Driving risk assessment using near-crash database through data mining of tree-based model. Accident Analysis & Prevention, 84, 54-64.
- Yao, J., Levy-Chapira, M., & Margaryan, M. (2017). Checking account activity and credit default risk of enterprises: An application of statistical learning methods. arXiv preprint arXiv:1707.00757.
Abstract:
Purpose: This study compares the predictive performance of four tree-based models, namely Decision Tree, Random Forest, Differentiable Decision Trees, and Deep Neural Decision Forest, in forecasting default risk among small and medium-sized enterprises in Vietnam.
Design/methodology/approach: Using a dataset of firm-level financial indicators, the study employs an empirical training and evaluation framework to examine model performance in terms of classification effectiveness and interpretability.
Findings: The findings show that Differentiable Decision Trees deliver the strongest predictive performance, indicating a superior ability to capture nonlinear relationships and classify default risk. In addition, variables related to liquidity, debt structure, and the capacity to absorb financing costs emerge as the most influential predictors.
Originality/value: The study contributes additional empirical evidence on the use of tree-based and hybrid tree–deep learning models in SME default prediction in Vietnam and provides practical implications for model selection in credit risk management.