ABSTRACT
Objective: To characterize, via a predictive model using real-world data, patients with diabetes with a heightened probability of hospitalization.
Methods: At the Endocrinology Unit of a tertiary public hospital in Rio Grande do Sul, Brazil, a retrospective cohort study analyzed initial consultations from January 1, 2015, to December 31, 2017, focusing on 617 patients with diabetes. Within this group, 82.98% (512 patients) did not require hospitalization, while 17.02% (105 patients) were hospitalized at least once. Multiple machine learning algorithms were tested, and the combination of XGBoost and Instance Hardness Threshold models displayed the best predictive performance. The SHapley Additive exPlanations method was used for result interpretation.
Results: The most optimal performance was observed by combining the XGBoost and Instance Hardness Threshold models, resulting in the highest sensitivity (0.93) in accurately classifying hospitalization events, with an acceptable area under the curve of 0.72. Key predictive features included the number of outpatient visits, amplitude of estimated glomerular filtration rate, and age (individuals below 24 years old and between 65 to 70 years old had higher hospitalization likelihood).
Conclusion: The proposed model demonstrated high predictive capability and may help to identify patients with diabetes who should be more closely monitored to reduce their risk of hospitalization.
Keywords
Diabetes mellitus; Hospitalization; Machine learning
INTRODUCTION
Diabetes mellitus is a disease characterized by chronic hyperglycemia due to the impaired release and action of insulin, as well as a failure to regulate hepatic glucose production (1). The two most prevalent types are type 1 and type 2, which account for about 10% and nearly 90% of all cases, respectively (2). The prevalence of diabetes mellitus has been increasing steadily in recent years and, by 2021, about 537 million people were estimated to have this condition (3). In Brazil, 12% of the population is diagnosed with diabetes (4), the sixth country with the highest number of adults with diabetes in the world (3).
Patients with diabetes are at greater risk of developing chronic complications and related diseases and require more access to health services than those without diabetes (5,6), representing one of the leading causes for hospital admissions and outpatient visits. It was estimated that the economic burden in Brazil reached US$ 2.15 billion in 2016, of which 70.6% are indirect costs related to premature deaths, absenteeism, and early retirement. If the growth rate of diabetes prevalence continues in Brazil, the direct and indirect costs of diabetes will be more than double by 2030 (an increase of 133.4% or 6.2% per year) (7).
Artificial intelligence (AI) is a growing field and its applications to diabetes could reshape the management of this chronic condition. Artificial intelligence algorithms have been used to develop predictive models for the risk of developing diabetes and its complications and to optimize the use of healthcare resources (8,9,10,11,12).
Electronic medical records (EMRs) allow the consistent and homogeneous gathering of data, enabling the repository use to train and develop algorithms (10,13). Such records have been used in medical studies with several objectives, including the prediction of hospitalization using the patient’s first record in the emergency room (14). The possibility of predicting which patients with diabetes are at greater risk of hospitalization and mortality via their characteristics would lead to early interventions that reduce risk, optimize treatments, and better prepare hospital resources to provide adequate care. Few studies evaluate the prediction of hospitalization of patients with diabetes (15,16,17,18).
This study aimed to characterize, via a predictive model using real-world data, patients with diabetes with a heightened probability of hospitalization.
MATERIALS AND METHODS
Study database
This was a retrospective cohort study using a database composed of EMRs of patients from the outpatient Endocrinology Unit of a tertiary public hospital from Southern Brazil. The complete dataset consists of EMRs of patients who had their first medical appointment in the period between January 1, 2015, and December 31, 2017, totaling 2,973 patients. Data within the 2-year period after the first medical appointment were used. Only patients diagnosed with diabetes were selected. They were identified via information from the International Classification of Diseases (ICD) of the first medical appointment and/or result of the first plasma glucose (≥ 126 mg/ dL) and/or the first glycated hemoglobin (HbA1c) (≥ 6.5%) measurements. The final dataset contained 617 patients, 512 (82.9%) were not hospitalized within the 2-year period and 105 (17.0%) were hospitalized at least once during that period. The EMR contained relevant information such as age, skin color, gender, number of outpatient visits, and laboratory tests (creatinine, plasma glucose, HbA1c, and urinary albumin concentration (UAC). The presence of diabetic kidney disease (DKD) was assessed with UAC and estimated glomerular filtration rate (eGFR) calculation using the CKD-EPI equation (19). Diabetic kidney disease was defined as an eGFR < 60 mL/min/1.73 m2 and/or UAC from a single urinary sample ≥ 14 mg/L (20,21,22). All textual records were written in Brazilian Portuguese.
This study was approved by the hospital’s Ethical Committee under number 43431521.0.0000.5327. The data were obtained via anonymized query and the consent form was waived.
Intelligent system protocol
We adopted a five-step method (Figure 1) to predict the occurrence of hospitalization in patients with diabetes from their EMRs. In the first step, we gathered, pre-processed the data, and discarded repeated observations. Missing data were imputed using the k-Nearest Neighbor method (23,24). After pre-processing, the dataset was rescaled using the max-min scaling in Eqn. 1, the values of all continuous features were in the range [0, 1]. In the equation, X represents the feature’s values, and Xmax and Xmin are the largest and smallest feature values in the dataset.
In step two, we divided the complete dataset to obtain train and validation portions using the κ-fold cross-validation technique (25). The complete dataset is partitioned into five subsets (κ = 5), and the model is trained on four of them while being validated on the remaining subset. This process was iterated five times, ensuring each fold acted as validation set at least once. At each fold, we used a stratified randomized sampling approach to ensure proportional representation of each class (hospitalized and not hospitalized) and reflect the class proportions of the complete sample (26). To obtain a better generalization of the model, we repeated the five-fold cross-validation process 20 times. In each repetition, we randomly shuffled the dataset before dividing it into five folds for cross-validation. This ensured the model was trained and validated on multiple different combinations of data sets, allowing it to capture more general and robust patterns in the data. At the end of these 20 repetitions, we obtained 100 validation results (five folds × 20 repetitions).
Our goal was to correctly identify patients at risk of hospitalization, i.e., the minority class. Therefore, in step three of the method we applied resampling techniques to the training portion. These techniques are recommended when dealing with highly unbalanced class problems, in which results can be influenced by the majority class. We tested six resampling techniques: Instance Hardness Threshold (IHT), Random Under Sampler (RUS), Synthetic Minority Oversampling Technique (SMOTE), adaptive synthetic sampling (ADASYN), synthetic minority oversampling technique and edited nearest neighbor (SMOTEENN) and synthetic minority oversampling technique with Tomek links technique (SMOTETomek) (Table S1). Resampling was not applied to the validation set, as we wanted to evaluate classification results in a reallife situation.
In step four, we performed feature selection using a wrapper method (27). In step five, we used the best subset of features in step four in the validation portion of the dataset for each machine learning algorithm tested; they are logistic regression (LR), K-nearest neighbors (KNN), support vector machine (SVM), Extreme Gradient Boosting (XGBoost), and Bagging Classifier using Decision Trees (Table S2). We averaged the one hundred validation results for the following performance metrics: accuracy, positive and negative predictive values (PPV and NPV, respectively), sensitivity, specificity, F1-Score, and are under the Receiver Operating Characteristic (ROC) curve.
We finally selected the model with the best predictive performance and used the SHapley Additive exPlanations (SHAP) method to analyze the results. This involves calculating SHAP values, which are obtained using a game-theoretic approach (28). The calculation of SHAP values involves evaluating the contribution of each feature to the model’s prediction by comparing the prediction with and without the feature, while considering all possible feature combinations. This process is applied to every feature and observation in the dataset, producing a matrix of SHAP values that reveals the relative importance of each feature to each observation. That enables a deeper understanding of the importance of each feature in the prediction and allows identifying features with the greatest impact on the model’s output (29,30).
Statistical analysis
A convenience sample was used. The algorithm performance was measured by the area under the ROC curve generated by plotting sensitivity versus one minus specificity. Based on the two operating points, 2×2 tables were developed to characterize the sensitivity and specificity of the algorithm. All statistical analyses, methods, techniques, and machine learning algorithms were implemented via Python (version 3.9.12).
RESULTS
Figure 2 shows the features used as predictors and the frequency in which they were selected by the best performing classification algorithm. Table 1 reports the validation results, with the best performers for each metric highlighted in bold.
Features used as inputs and frequency with which they were retained in the one hundred validations of the best predictive model.
Average predictive performance and standard deviations obtained from one hundred replicates of the dataset validation portion for different combinations of resampling technique and classification algorithm
Combining the XGBoost and IHT models yielded the highest sensitivity (value of 0.93) in correctly classifying hospitalization events and an acceptable AUC of 0.72. In addition, it yielded the lowest standard deviation for sensitivity and AUC values, indicating more generalizable results compared to other models. We used SHAP to interpret the features most frequently selected by the XGBoost-IHT combination and provide insights about the importance of each feature (Figure 3) and its effect on the classification result (Figure 4).
SHapley Additive exPlanations values (features impacts on model outputs). SHapley Additive exPlanations values were computed across the entire dataset using the XGBoost-IHT combination model. Each point on the graph corresponds to a single data observation. The distribution of SHapley Additive exPlanations values is displayed along the horizontal axis through violin plots. Positive SHapley Additive exPlanations values indicate features that contribute to the accurate classification of hospitalization cases, with larger values meaning greater impact. Conversely, negative SHapley Additive exPlanations values are associated with features influencing the prediction of non-hospitalization cases. The color spectrum in the graph represents the actual values of data observations, transitioning from blue to red as the value increases.
Relationship between the features' real values (displayed along the horizontal axes) and their respective SHapley Additive explanations values (displayed along the vertical axes).
Figure 3 allows a visual assessment of the features’ importance for the classification of hospitalization cases. Outpatient visits in the 2-year period and amplitude of eGFR were important hospitalization predictors, as demonstrated by their high importance values in Figure 3. Regarding the outpatient visits in the 2-year period, it is clear that many visits (turning red) are associated with larger SHAP values, i.e., it increases the probability of hospitalization. The amplitude behavior of eGFR feature is similar, suggesting that the greater the difference in exam results, the higher the probability of hospitalization. Although in Figure 3 some features show relatively low importance, it is important to assess their clinical relevance in conjunction with other features and make informed decisions. A case in point is the DKD, which, despite its relatively modest SHAP contribution, is informative due to a shift in SHAP values from negative to positive when transitioning from the presence to absence of DKD.
The graphs of Figure 4 show in a clearer way the relationship between features’ values and the probability of hospitalization. This is because the horizontal axes represent the actual values of each feature, while the vertical axes represent their corresponding SHAP values. Age is the third most important feature in predicting hospitalizations. As shown in Figure 4, SHAP values for age vary according to the patient's age group. For patients under 24 years old, positive SHAP values indicate a higher probability of hospitalization for this age group. In opposition, mostly negative SHAP values associated with the age group between 25 and 65 years old suggest lower probability of hospitalization for this age group. Finally, for patients over 65 years old the relation between feature and prediction is not clear, varying according to other features present in the model.
DISCUSSION
Considering the limited number of characteristics, we were able to analyze with the available data, the area under the ROC curve reported for the constructed classifier is indeed a considerable achievement for predicting the hospitalization of patients with diabetes, showing that with few easy-to-obtain EMR data it is possible to predict the probability of the patient’s hospitalization, which enables to identify the patient at greatest risk, allowing the allocation of available resources before the outcome occurs.
The results of our study showed that the top three most important features for predicting hospitalization are: number of outpatient visits, amplitude of eGFR, and age (patients under 24 years old and 65 to 70 years old have a higher risk).
The classifier showed that patients who had the highest number of medical consultations during the last 2 years were those with the highest risk of hospitalization. The number of outpatients visits can correlate with the complexity of the patient’s health status, since patients who have more comorbidities or more serious illnesses are those who have a higher number of consultations. A study that evaluated the magnitude and predictors of hospital admission in patients with type 2 diabetes at public hospitals of Eastern Ethiopia showed that medical conditions, including the number of comorbidities and the presence of chronic diabetes complications, are determinant for hospital admission (15). The number of inpatient healthcare visits can also predict the chance of readmission within 30 days after hospital discharge (31).
The contribution of eGFR amplitude to predict hospitalization is because patients with varying eGFR may have presented a worsening of their renal function in the last 2 years. A study that evaluated the variability in eGFR in patients with diabetes showed they were at greater risk of major clinical outcomes (major macrovascular events, new or worsening nephropathy, and all-cause mortality) (32). The study assessed the association between 20-month eGFR variability and the risk of major clinical outcomes in type 2 diabetes among 8,241 patients. Variability in eGFR was calculated from three serum creatinine measurements over 20 months. Compared with low variability, greater 20-month eGFR variability was independently associated with higher risk of the primary outcome with evidence of a positive linear trend (p = 0.015).
As our study showed, age is a known predictor for hospitalization in patients with diabetes. The study revealed that patients who are younger than 24 years old were more likely to be hospitalized, as well as patients between 65 and 70 years old. The fact that younger people are more likely to be hospitalized may represent patients with type 1 diabetes who have diabetic ketoacidosis (DKA). The precipitating factors of DKA were evaluated in a public Brazilian hospital, showing that the mean age of patients with DKA from January 2005 to March 2010 was 26 ± 13 years (33), remaining at 26.2 ± 14.5 years in the period from April 2010 to January 2017 (34), treatment noncompliance being the leading precipitating factor in both periods (33,34). A study that evaluated care indicators for patients with diabetes in our country showed that in 2019, worse indicators were observed for younger individuals (35).
Although patients between 65 and 70 years old had a higher risk of hospitalization than those over 70 years old, we were unable to identify differences in this subgroup. Dennis et al. showed in their predictive model that hospitalized patients with diabetes were older. In addition, age can also predict length of stay (36) and is an important feature for predicting 1-year mortality (37).
Our study has limitations. First, from a clinical perspective, our data do not include patients’ medical history such as insulin use, type of diabetes, and duration of diabetes. Second, the database lacks information on important comorbidities and anthropometry. Third, the database had limited sociodemographic information; previous studies showed that low education, low socioeconomic status, high alcohol use, longer diabetes mellitus duration can predict hospitalization (15,16,17). Moreover, patients attending our hospital’s outpatient Endocrinology clinic could have been hospitalized in another hospital and this would not have been identified because the search for outcomes was done exclusively in our hospital database.
Despite these limitations, it is essential to consider the financial implications of implementing predictive models in healthcare settings. Implementation costs can vary significantly; according to Al Meslamani (38), these considerations encompass various expenses, including the initial model development, its real-world implementation, user training, and ongoing maintenance. Cost estimates can vary widely; for instance, simple models may incur operational costs ranging from $60,750 to $94,500, whereas training more complex models can cost tens of millions. The sustainability of funding, coupled with potential secondary savings—such as increased efficiency for providers, reduced hospitalizations, and shorter lengths of stay—is vital for justifying investments in these models. For example, Brisimi et al. revealed that in 2012 the average hospitalization cost in the United States was $9,500. A predictive model with an 81% detection rate could reduce preventive measure costs to just $320 per patient, potentially saving up to $1 billion in avoidable hospitalizations across the Unites States (39).
CONCLUSION
The proposed model demonstrated predictive capability and may help identify patients with diabetes who are at higher risk of hospitalization. That allows paying special attention to these patients during outpatient follow-up and identify patients with greatest risk upon arrival at the emergency room, optimizing resources allocation. The factors that most contribute to the prediction are the number of outpatient visits, amplitude of estimated glomerular filtration rate and age (patients under 24 years old and 65 to 70 years old present higher probability).
Appendix
-
Funding:
this study was funded by the Research Incentive Fund of the Hospital de Clínicas de Porto Alegre and the Graduate Program in Medical Sciences: Endocrinology of the Faculty of Medical Sciences of the Universidade Federal do Rio Grande do Sul. The research was partially funded by the Brazilian Federal Agency for Support and Evaluation of Graduate Education (Capes) – Brazil – Financial Code 001; the Brazilian National Council for Scientific and Technological Development (CNPq); the Instituto de Avaliação de Tecnologia em Saúde (IATS); and the Fundação de Amparo à Pesquisa do Estado do Rio Grande do Sul (FAPERGS).
-
Consent for publication:
all authors revised the final version of the manuscript and agree with the publication of the results presented.
Availability of data:
dataset that supports the findings of this study are available from the corresponding author, upon reasonable request.
REFERENCES
- 1 Magliano DJ, Boyko EJ. IDF Diabetes Atlas 10th Edition. Bruxelas: International Diabetes Federation; 2022.
-
2 Brazilian Society of Diabetes. Available from: https://diabetes.org.br Accessed on: 2024 Oct 20. doi:10.1080/13696998.2023.2285186
» https://doi.org/10.1080/13696998.2023.2285186» https://diabetes.org.br -
3 Sun H, Saeedi P, Karuranga S, Pinkepank M, Ogurtsova K, Duncan BB, et al. IDF Diabetes Atlas: Global, regional and country-level diabetes prevalence estimates for 2021 and projections for 2045. Diabetes Res Clin Pract. 2022;183:109119. doi: 10.1016/j.diabres.2021.109119. Erratum in: Diabetes Res Clin Pract. 2023 Oct;204:110945. doi: 10.1016/j.diabres.2023.110945
» https://doi.org/10.1016/j.diabres.2021.109119 -
4 Telo GH, Cureau FV, De Souza MS, Andrade TS, Copês F, Schaan BD. Prevalence of diabetes in Brazil over time: A systematic review with meta-analysis. Diabetol Metab Syndr. 2016;8(1) doi:10.1186/s13098-016-0181-1
» https://doi.org/10.1186/s13098-016-0181-1 -
5 Alonso-Morán E, Orueta JF, Esteban JI, Axpe JM, González ML, Polanco NT, et al. Multimorbidity in people with type 2 diabetes in the Basque Country (Spain): Prevalence, comorbidity clusters and comparison with other chronic patients. Eur J Intern Med. 2015 Apr;26(3):197-202. doi: 10.1016/j.ejim.2015.02.005
» https://doi.org/10.1016/j.ejim.2015.02.005 - 6 Struijs JN, Baan CA, Schellevis FG, Westert GP, van den Bos GA. Comorbidity in patients with diabetes mellitus: impact on medical health care utilization. BMC Health Serv Res. 2006 Jul 4;6:84
-
7 Pereda P, Boarati V, Guidetti B, Duran AC. Direct and Indirect Costs of Diabetes in Brazil in 2016. Ann Glob Health. 2022 Mar 3;88(1):14. doi: 10.5334/aogh.3000
» https://doi.org/10.5334/aogh.3000 -
8 Ellahham S. Artificial Intelligence: The Future for Diabetes Care. Am J Med. 2020 Aug;133(8):895-900. doi: 10.1016/j.amjmed.2020.03.033
» https://doi.org/10.1016/j.amjmed.2020.03.033 -
9 Contreras I, Vehi J. Artificial Intelligence for Diabetes Management and Decision Support: Literature Review. J Med Internet Res. 2018 May 30;20(5):e10775. doi: 10.2196/10775
» https://doi.org/10.2196/10775 -
10 Kavakiotis I, Tsave O, Salifoglou A, Maglaveras N, Vlahavas I, Chouvarda I. Machine Learning and Data Mining Methods in Diabetes Research. Comput Struct Biotechnol J. 2017 Jan 8;15:104-16. doi: 10.1016/j. csbj.2016.12.005
» https://doi.org/10.1016/j. csbj.2016.12.005 - 11 Vijiyakumar K, Lavanya B, Nirmala I, Caroline SS. Random Forest Algorithm for the Prediction of Diabetes. In: Proceeding of International Conference on Systems Computation Automation and Networking . ; 2019.
-
12 Mujumdar A, Vaidehi V. Diabetes Prediction using Machine Learning Algorithms. In: Procedia Computer Science. Vol 165. Elsevier B.V.; 2019:292-9. doi:10.1016/j.procs.2020.01.047
» https://doi.org/10.1016/j.procs.2020.01.047 -
13 Singla R, Singla A, Gupta Y, Kalra S. Artificial intelligence/machine learning in diabetes care. Indian J Endocrinol Metab. 2019;23(4):495-7. doi:10.4103/ijem.IJEM_228_19
» https://doi.org/10.4103/ijem.IJEM_228_19 -
14 Lucini FR, Fogliatto FS, da Silveira GJ, Neyeloff JL, Anzanello MJ, Kuchenbecker RS, et al. Text mining approach to predict hospital admissions using early medical records from the emergency department. Int J Med Inform. 2017 Apr;100:1-8. doi: 10.1016/j. ijmedinf.2017.01.001
» https://doi.org/10.1016/j. ijmedinf.2017.01.001 -
15 Regassa LD, Tola A. Magnitude and predictors of hospital admission, readmission, and length of stay among patients with type 2 diabetes at public hospitals of Eastern Ethiopia: a retrospective cohort study. BMC Endocr Disord. 2021;21(1) doi:10.1186/s12902-021-00744-3
» https://doi.org/10.1186/s12902-021-00744-3 -
16 Sajjad MA, Holloway KL, de Abreu LLF, Mohebbi M, Kotowicz MA, Pedler D, et al. Comparison of incidence, rate and length of all-cause hospital admissions between adults with normoglycaemia, impaired fasting glucose and diabetes: a retrospective cohort study in Geelong, Australia. BMJ Open. 2018 Mar 23;8(3):e020346. doi: 10.1136/ bmjopen-2017-020346
» https://doi.org/10.1136/ bmjopen-2017-020346 -
17 Begum N, Donald M, Ozolins IZ, Dower J. Hospital admissions, emergency department utilisation and patient activation for self-management among people with diabetes. Diabetes Res Clin Pract. 2011;93(2):260-7. doi: 10.1016/j.diabres.2011.05.031
» https://doi.org/10.1016/j.diabres.2011.05.031 -
18 Dennis S, Taggart J, Yu H, Jalaludin B, Harris MF, Liaw ST. Linking observational data from general practice, hospital admissions and diabetes clinic databases: can it be used to predict hospital admission? BMC Health Serv Res. 2019;19(1):526. doi: 10.1186/s12913-019-4337-1
» https://doi.org/10.1186/s12913-019-4337-1 -
19 Stevens PE, Levin A; Kidney Disease: Improving Global Outcomes Chronic Kidney Disease Guideline Development Work Group Members. Evaluation and management of chronic kidney disease: synopsis of the kidney disease: improving global outcomes 2012 clinical practice guideline. Ann Intern Med. 2013;158(11):825-30. doi: 10.7326/0003-4819-158-11-201306040-00007
» https://doi.org/10.7326/0003-4819-158-11-201306040-00007 -
20 Incerti J, Zelmanovitz T, Camargo JL, Gross JL, de Azevedo MJ. Evaluation of tests for microalbuminuria screening in patients with diabetes. Nephrol Dial Transplant. 2005;20(11):2402-7. doi: 10.1093/ ndt/gfi074
» https://doi.org/10.1093/ ndt/gfi074 -
21 Zelmanovitz T, Gross JL, Oliveira JR, Paggi A, Tatsch M, Azevedo MJ. The receiver operating characteristics curve in the evaluation of a random urine specimen as a screening test for diabetic nephropathy. Diabetes Care. 1997 Apr;20(4):516-9. doi: 10.2337/diacare.20.4.516
» https://doi.org/10.2337/diacare.20.4.516 -
22 ElSayed NA, Aleppo G, Aroda VR, Bannuru RR, Brown FM, Bruemmer D, et al., on behalf of the American Diabetes Association. 11. Chronic Kidney Disease and Risk Management: Standards of Care in Diabetes-2023. Diabetes Care. 2023;46(Suppl 1):S191-S202. doi: 10.2337/dc23-S011
» https://doi.org/10.2337/dc23-S011 -
23 Alwan JK, Jaafar DS, Ali IR. Diabetes diagnosis system using modified Naive Bayes classifier. Indonesian Journal of Electrical Engineering and Computer Science. 2022;28(3):1766-74. doi:10.11591/ijeecs.v28. i3.pp1766-1774
» https://doi.org/10.11591/ijeecs.v28. i3.pp1766-1774 -
24 Blume CA, Brust-Renck PG, Rocha MK, Leivas G, Neyeloff JL, Anzanello MJ, et al. Development and Validation of a Predictive Model of Success in Bariatric Surgery. Obes Surg. 2021;31(3):1030-7. doi: 10.1007/ s11695-020-05103-0.
» https://doi.org/10.1007/ s11695-020-05103-0 -
25 Wu SL, Christian Hospital C, Batur Çolak A. Risk factor identification and prediction models for prolonged length of stay in hospital after acute ischemic stroke using artificial neural networks. Front Neurol. 2023;14(1085178.). http://taiwanstrokeregistry.org/
» http://taiwanstrokeregistry.org/ -
26 Kumar M, Ang LT, Ho C, Soh SE, Tan KH, Chan JKY, et al. Machine Learning-Derived Prenatal Predictive Risk Model to Guide Intervention and Prevent the Progression of Gestational Diabetes Mellitus to Type 2 Diabetes: Prediction Model Development Study. JMIR Diabetes. 2022 Jul 5;7(3):e32366. doi: 10.2196/32366
» https://doi.org/10.2196/32366 -
27 Bolón-Canedo V, Rego-Fernández D, Peteiro-Barral D, Alonso-Betanzos A, Guijarro-Berdiñas B, Sánchez-Maroño N. On the scalability of feature selection methods on high-dimensional data. Knowl Inf Syst. 2018;56(2):395-442. doi:10.1007/s10115-017-1140-3
» https://doi.org/10.1007/s10115-017-1140-3 -
28 Mangalathu S, Hwang SH, Jeon JS. Failure mode and effects analysis of RC members based on machine-learning-based SHapley Additive exPlanations (SHAP) approach. Eng Struct. 2020;219. doi:10.1016/j. engstruct.2020.110927
» https://doi.org/10.1016/j. engstruct.2020.110927 -
29 Ullah I, Liu K, Yamamoto T, Zahid M, Jamal A. Modeling of machine learning with SHAP approach for electric vehicle charging station choice behavior prediction. Travel Behav Soc. 2023;31:78-92. doi:10.1016/j.tbs.2022.11.006
» https://doi.org/10.1016/j.tbs.2022.11.006 -
30 Yang C, Chen M, Yuan Q. The application of XGBoost and SHAP to examining the factors in freight truck-related crashes: An exploratory analysis. Accid Anal Prev. 2021;158:106153. doi: 10.1016/j. aap.2021.106153
» https://doi.org/10.1016/j. aap.2021.106153 -
31 Eby E, Hardwick C, Yu M, Gelwicks S, Deschamps K, Xie J, et al. Predictors of 30 day hospital readmission in patients with type 2 diabetes: a retrospective, case-control, database study. Curr Med Res Opin. 2015;31(1):107-14. doi: 10.1185/03007995.2014.981632
» https://doi.org/10.1185/03007995.2014.981632 -
32 Jun M, Harris K, Heerspink HJ, Badve SV, Jardine MJ, Harrap S, et al; ADVANCE Collaborative Group. Variability in estimated glomerular filtration rate and the risk of major clinical outcomes in diabetes: Post hoc analysis from the ADVANCE trial. Diabetes Obes Metab. 2021;23(6):1420-5. doi: 10.1111/dom.14351
» https://doi.org/10.1111/dom.14351 -
33 Weinert LS, Scheffel RS, Severo MD, Cioffi AP, Teló GH, Boschi A, et al. Precipitating factors of diabetic ketoacidosis at a public hospital in a middle-income country. Diabetes Res Clin Pract. 2012 Apr;96(1):29-34. doi: 10.1016/j.diabres.2011.11.006
» https://doi.org/10.1016/j.diabres.2011.11.006 -
34 da Rosa Carlos Monteiro LE, Garcia SP, Bottino LG, Custodio JL, Telo GH, Schaan BD. Precipitating factors of diabetic ketoacidosis in type 1 diabetes patients at a tertiary hospital: a cross-sectional study with a two-time-period comparison. Arch Endocrinol Metab. 2022;66(3):355-61. doi: 10.20945/2359-3997000000480
» https://doi.org/10.20945/2359-3997000000480 -
35 Malta DC, Ribeiro EG, Gomes CS, Alves FT, Stopa SR, Sardinha LM, et al. Indicators of the line of care for people with diabetes in Brazil: National Health Survey 2013 and 2019. Epidemiol Serv Saude. 2022;31(spe1):e2021382. doi: 10.1590/SS2237-9622202200011. especial
» https://doi.org/10.1590/SS2237-9622202200011 -
36 Barsasella D, Bah K, Mishra P, Uddin M, Dhar E, Suryani DL, et al. A Machine Learning Model to Predict Length of Stay and Mortality among Diabetes and Hypertension Inpatients. Medicina (Kaunas). 2022;58(11):1568. doi: 10.3390/medicina58111568
» https://doi.org/10.3390/medicina58111568 -
37 Alimbayev A, Zhakhina G, Gusmanov A, Sakko Y, Yerdessov S, Arupzhanov I, et al. Predicting 1-year mortality of patients with diabetes mellitus in Kazakhstan based on administrative health data using machine learning. Sci Rep. 2023;13(1):8412. doi: 10.1038/ s41598-023-35551-4
» https://doi.org/10.1038/ s41598-023-35551-4 -
38 Al Meslamani AZ. Beyond implementation: the long-term economic impact of AI in healthcare. J Med Econ. 2023;26(1):1566-9. doi: 10.1080/13696998.2023.2285186
» https://doi.org/10.1080/13696998.2023.2285186 -
39 Brisimi TS, Xu T, Wang T, Dai W, Paschalidis IC. Predicting diabetes-related hospitalizations based on electronic health records. Stat Methods Med Res. 2019;28(12):3667-82. doi: 10.1177/0962280218810911
» https://doi.org/10.1177/0962280218810911 -
40 Smith MR, Martinez T, Giraud-Carrier C. An instance level analysis of data complexity. Mach Learn. 2014;95(2):225-256. doi:10.1007/ s10994-013-5422-z
» https://doi.org/10.1007/ s10994-013-5422-z -
41 Hanafy M, Ming R. USING MACHINE LEARNING MODELS TO COMPARE VARIOUS RESAMPLING METHODS IN PREDICTING INSURANCE FRAUD. J Theor Appl Inf Technol. 2021;30(12). http://www.jatit.org
» http://www.jatit.org -
42 Wei Z, Zhang L, Zhao L. Minority-prediction-probability-based oversampling technique for imbalanced learning. Inf Sci (N Y). 2023;622:1273-95. doi:10.1016/j.ins.2022.11.148
» https://doi.org/10.1016/j.ins.2022.11.148 - 43 Chawla N V, Bowyer KW, Hall LO, Kegelmeyer WP. SMOTE: Synthetic Minority Over-Sampling Technique. Vol 16.; 2002.
- 44 Chawla N V, Bowyer KW, Hall LO, Kegelmeyer WP. SMOTE: Synthetic Minority Over-Sampling Technique. Vol 16.; 2002.
-
45 Batista GE, Prati RC, Monard MC. A Study of the Behavior of Several Methods for Balancing Machine Learning Training Data. Sigkdd Explorations. 2004;6(1):20-9. doi: https://doi.org/10.1145/1007730.1007735
» https://doi.org/10.1145/1007730.1007735 -
46 Godoy LC, Farkouh ME, Austin PC, Shah BR, Qiu F, Sud M, et al. Predicting left main stenosis in stable ischemic heart disease using logistic regression and boosted trees. Am Heart J. 2023 Feb;256:117-127. doi: 10.1016/j.ahj.2022.11.004
» https://doi.org/10.1016/j.ahj.2022.11.004 -
47 Wu X, Kumar V, Ross QJ, et al. Top 10 algorithms in data mining. Knowl Inf Syst. 2008;14(1):1-37. doi:10.1007/s10115-007-0114-2
» https://doi.org/10.1007/s10115-007-0114-2 -
48 Álvarez-Alvarado JM, Ríos-Moreno JG, Obregón-Biosca SA, Ronquillo-Lomelí G, Ventura-Ramos E, Trejo-Perea M. Hybrid techniques to predict solar radiation using support vector machine and search optimization algorithms: A review. Applied Sciences (Switzerland). 2021;11(3):1-17. doi:10.3390/app11031044
» https://doi.org/10.3390/app11031044 - 49 Grollmuss O, Zhu B, Demetrio P, Zhang H. The predictive value of pressure recording analytical method for the duration of mechanical ventilation in children undergoing cardiac surgery with an XGBoost-based machine learning model. Front Cardiovasc Med. 2022;9.
-
50 Samih A, Ghadi A, Fennan A. Enhanced sentiment analysis based on improved word embeddings and XGboost. International Journal of Electrical and Computer Engineering. 2023;13(2):1827-1836. doi:10.11591/ijece.v13i2.pp1827-1836
» https://doi.org/10.11591/ijece.v13i2.pp1827-1836 -
51 Bbeiman L. Bagging Predictors. 1996 [cited 2025 Feb 4];24:123-140. Available from: https://link.springer.com/article/10.1007/BF00058655
» https://link.springer.com/article/10.1007/BF00058655





SHAP: SHapley Additive exPlanations.
Standard deviation measures the dispersion of data, amplitude represents the range of values, and bias indicates a consistent error or distortion in data. eGFR: estimated glomerular filtration rate; HbA1c: glycated hemoglobin.
eGFR: estimated glomerular filtration rate; HbA1c: glycated hemoglobin.
SHAP: SHapley Additive explanations