Open-access Machine learning and cancer survival: analysis of the Registro Hospitalar de Câncer do Estado de São Paulo

ABSTRACT

OBJECTIVE  To compare the performance of different Survival Machine Learning (SML) algorithms in predicting the survival of cancer patients.

METHODS  Data from the Registro Hospitalar de Câncer do Estado de São Paulo (São Paulo State Cancer Registry Hospital) were used, covering the five most incident types of cancer (breast, prostate, lung, colorectal and cervix). Six algorithms were evaluated: Gradient Boosting Survival (GBS), Random Survival Forest (RSF), Support Vector Machine Survival (SVM-Survival), XGBoost Cox, XGBoost Accelerated Failure Time (AFT), and LightGBM. Performance was measured by the Concordance Index (C-Index), C-Index IPCW and Integrated Brier Score (IBS) metrics.

RESULTS  The XGBoost AFT model showed the best C-Index results for breast (0.7845), lung (0.7368), colorectal (0.7618), and cervix (0.7726), while GBS was superior for prostate (0.7574). Clinical staging was consistently the most important variable, according to the explainability analysis.

CONCLUSION  The SML algorithms showed good predictive performance, regardless of cancer type, sample size and censoring proportion. These models show potential for subsidizing cancer planning and supporting strategic decisions in the organization of cancer care networks.

DESCRIPTORS:
Neoplasms; Survival Analysis; Machine Learning

RESUMO

OBJETIVO  Comparar o desempenho de diferentes algoritmos de Survival Machine Learning (SML) na predição da sobrevida de pacientes com câncer.

MÉTODOS  Utilizaram-se dados do Registro Hospitalar de Câncer do Estado de São Paulo, contemplando os cinco tipos de câncer mais incidentes (mama, próstata, pulmão, colorretal e colo do útero). Foram avaliados seis algoritmos: Gradient Boosting Survival (GBS), Random Survival Forest (RSF), Support Vector Machine Survival (SVM-Survival), XGBoost Cox, XGBoost Accelerated Failure Time (AFT) e LightGBM. O desempenho foi medido pelas métricas Concordance Index (C-Index), C-Index IPCW e Integrated Brier Score (IBS).

RESULTADOS  O modelo XGBoost AFT apresentou os melhores resultados de C-Index para mama (0,7845), pulmão (0,7368), colorretal (0,7618) e colo do útero (0,7726), enquanto o GBS foi superior para próstata (0,7574). O estadiamento clínico foi consistentemente a variável mais importante, segundo a análise de explicabilidade.

CONCLUSÃO  Os algoritmos de SML demonstraram bom desempenho preditivo, independentemente do tipo de câncer, do tamanho amostral e da proporção de censura. Esses modelos mostram potencial para subsidiar o planejamento oncológico e apoiar decisões estratégicas na organização das redes de atenção ao câncer.

DESCRITORES:
Neoplasias; Análise de Sobrevida; Aprendizado de Máquina

INTRODUCTION

The incidence of cancer has risen sharply in recent years due to an ageing population and changes in lifestyle habits, with an estimated 20 million new cases and 9.7 million deaths per year1. It is a major public health problem, with an impact on health systems around the world.

Survival analysis is a type of analysis in which the dependent variable of interest is the time until a certain event occurs2. In cancer, these studies aim to analyze the time between diagnosis and/or admission to a specialized service to treat the disease and the patient’s death3. An important aspect of this type of analysis is the inclusion of the observation time of censored individuals. Censoring occurs when a patient is lost to follow-up or when the study ends and they continue without having presented the outcome of interest. In contrast, analyses of dichotomous outcomes, such as five-year survival, need to use only those individuals for whom there is certainty about survival or death during the follow-up period, and it is necessary to exclude censored individuals from the analysis.

Survival analysis models based on traditional statistics depend on assumptions about the distribution of the data. The Cox proportional hazards model (Cox-PH) is the most widely used and is a semi-parametric model that assumes the proportionality of risks for different individuals is constant over time3. As such, its use requires tests to confirm that the assumptions are met, requiring adaptations and adjustments when this is not the case2. This model is often used to estimate the effect of explanatory variables on the risk of an event over time but can be adapted for individual risk prediction using estimators such as Breslow.

The use of Survival Machine Learning (SML) algorithms for survival analysis has been growing in recent years4,5. Their main advantage is the ability to analyze large data sets with variables of different formats, without the need to meet certain assumptions. This is especially useful in the case of attributes with non-linear associations with the variable of interest.

In Brazil, some studies have looked at cancer mortality and survival6,7, using data from the Registro Hospitalar de Câncer do Estado de São Paulo (RHC/SP – São Paulo State Cancer Registry Hospital) of the Fundação Oncocentro do Estado de São Paulo (FOSP – São Paulo State Oncocentro Foundation). However, these studies did not consider survival time as the dependent variable of interest, but rather the occurrence or non-occurrence of the event (death) after certain time intervals. Thus, in these analyses, it was necessary to exclude individuals due to loss to follow-up.

In a previous study, SML models were compared with machine learning models for dichotomous survival outcomes at one, three and five years in patients with colorectal cancer, showing that the dichotomous models underestimated the survival of these patients8. This article included patients with the five most frequent types of cancer in the state of São Paulo, which show significant differences in their incidence and in the proportion of censored cases. The main SML models were tested to identify the algorithms with the best prediction performance and possible differences in performance related to sample size and the proportion of censored cases.

METHODS

The RHC/SP, managed by FOSP, gathers socioeconomic and clinical information on patients diagnosed with cancer since 2000, provided by 81 institutions in the state of São Paulo. This database, with anonymized information, is public and can be consulted by completing a registration form on the institution’s website9. For this study, cases of the five most common types of cancer in the state of São Paulo were selected: breast, prostate, lung, colorectal and cervix.

Before carrying out the study, the authors identified the variables that would be included in the model and the inclusion and exclusion criteria, considering the characteristics of the tumors and the available data set. The decision was made to exclude variables that showed high collinearity with the chosen variables, a low proportion of completeness or that were related to the type of treatment the individual had undergone. The intention was to work with data that is available before treatment begins, so that these models can be used in later studies to build scenarios and/or optimize the organization of the health services network.

The data sets for the five types of cancer under study were prepared in two stages: general selections and column adjustments – applied to all the topographies analyzed – and specific selections, made according to the particularities of each type. For all types of cancer, patients were excluded if they were under 20 years of age, did not live in the state of São Paulo, had an undefined clinical stage or had not been diagnosed, had carcinoma in situ, had no microscopic confirmation of the diagnosis, had uncertain morphology, or had undergone a bone marrow transplant.

In the specific selections, additional criteria were applied for each type of cancer. Patients with lung cancer who had received hormone therapy were excluded, since this treatment is generally intended for patients with other types of tumors with lung metastases. In the case of colorectal cancer, only patients with adenocarcinoma morphology (code 81403) were considered, since the other morphologies are rarer and show different behavior. For breast cancer, only women were included, given that this type of cancer in men is rare and behaves very differently.

The set of variables analyzed included individual patient information, such as age (IDADE), sex (SEXO), and schooling (ESCOLARI); variables related to place of residence, such as the city’s IBGE code (IBGE) and Regional Health Department (DRS); and variables relating to the care institution, including institution code (INSTITU), category of care (CATEATEND), IBGE code of the institution’s city (IBGEATEN), DRS of the institution (DRS_INST), and category of qualification in high complexity oncology (HABILIT2). Clinical variables were also considered, such as diagnosis prior to admission (DIAGPREV), topography (TOPO), morphology (MORFO), clinical staging (EC), and year of diagnosis (ANODIAG). Finally, the categorized time between the first consultation and the start of treatment (TRATCONS_CAT) and between diagnosis and the start of treatment (DIAGTRAT_CAT) were evaluated. The gender variable was only used for lung and colorectal cancer, the topography variable was not considered for prostate, and the morphology variable was not considered for colorectal. The data was divided into two parts, with 80% of the patients being used to train the survival algorithms and the other 20% to validate these models.

The outcome was the time between diagnosis and death from all causes, including censored patients. Six machine learning models were used for survival analysis: Gradient Boosting Survival (GBS), Random Survival Forest (RSF), Survival Support Vector Machine (SSVM), XGBoost Cox (XGB-Cox), XGBoost Accelerated Failure Time (XGB-AFT), and LightGBM (LGBM). Each adopts a different strategy for modeling the time until an event occurs and was chosen for their ability to deal with censorship in the data, capture non-linear patterns and provide more flexible predictions compared to traditional statistical approaches. The methodological aspects relating to the treatment of the database, as well as the specificities of each type of SML model, are described in more depth in another paper8.

Three metrics were used to evaluate the survival models: Concordance Index (C-Index), Inverse Probability of Censoring Weighted Concordance Index (C-Index IPCW), and Integrated Brier Score (IBS). The C-Index is a metric used to measure the discriminatory capacity of the model, i.e. its ability to correctly order the survival times of individuals10. Values close to 1 indicate perfect agreement, while a value of 0.5 suggests performance equivalent to chance. The C-Index IPCW is a variation of the C-Index that corrects for biases introduced by non-informative or covariate-dependent censoring11. Finally, the IBS measures the accuracy of predictions of survival probability over time and is calculated from the survival curves estimated by the models. Lower values indicate better performance, with 0 being a perfect prediction12. As IBS depends on the survival curves, it was only calculated for the models that had adequate survival curves and were consistent with the empirical observations.

The hyperparameters of the machine learning models were selected using Optuna13, using three different samplers. RandomSampler combines the parameters randomly; TPESampler14 optimizes the choice using probabilistic models; and CmaEsSampler15 employs an evolutionary strategy that adjusts a Gaussian distribution over the space of hyperparameters. Each sampler was evaluated with 150 different combinations, applying cross-validation in 10 folds for each set of parameters. For each type of cancer, the model with the best performance among the three samplers was selected.

The SHapley Additive exPlanations (SHAP) and the Permutation Importance (PI)17 were used to assess the impact of the different variables on the predictions of the survival models. SHAP estimates the average contribution of each variable to the predictions, considering all possible combinations of values, while PI evaluates the loss of model performance when the values of a variable are randomly shuffled.

Ethical Aspects

This study was carried out using a public database, anonymized and without sensitive variables that could identify individuals, preserving confidentiality. As such, it was not submitted to a Research Ethics Committee, in accordance with Resolution No. 510/2016 of the National Health Council.

RESULTS

The database was obtained from the RHC/SP in September 2024 with data on 1,223,973 individuals. After general and specific selections by type of cancer, 141,726 patients with breast cancer, 111,406 with prostate cancer, 45,719 with lung cancer, 44,856 with colorectal cancer, and 27,850 with cervical cancer were included, as described in Chart 1.

Chart 1
Adjustments and selections made to the RHC/SP datasets.

The profile of the patients included in the study is shown in Table 1. The number of patients included for each type of cancer ranged from 27,850 cervical cases to 141,726 breast cases, with 44,856 colorectal, 45,719 lung, and 111,406 prostate cases. In lung cancer and colorectal cancer, there was a predominance of male patients, corresponding to 60.2% and 52.0% of cases, respectively.

Table 1
Profile of patients included in the study.

The proportion of deaths from all causes varied considerably between the types of cancer included. Lung cancer cases had the highest proportion of deaths (85.1%), while breast and prostate had the lowest proportion of deaths, with values close to each other (32.9% and 32.7%, respectively). Colorectal (48.8%) and cervix (44.9%) were at intermediate levels. The distribution according to topography, as well as the division between training and testing can be found in the Supplementary Materiala. The Kaplan-Meier curves stratified by clinical stage can also be found in the Supplementary Material.

Table 2 shows the performance of the SML models for predicting survival after optimizing the hyperparameters for the five types of cancer. In breast cancer, the XGB-AFT model showed the best performance in terms of C-Index (0.7845) and C-Index IPCW (0.7570), while GBS obtained the lowest IBS (0.1325). As for prostate cancer, GBS stood out with the best C-Index (0.7574) and C-Index IPCW (0.7332) values, as well as the lowest IBS (0.1298). In the lung cancer scenario, the XGB-AFT model obtained the highest C-Index (0.7368) and the highest IPCW C-Index (0.7325), although the lowest IBS was recorded by the GBS model (0.1165). For colorectal cancer, the XGB-AFT again showed the highest C-Index (0.7618) and C-Index IPCW (0.7532) values, while the lowest IBS was recorded by the GBS (0.1552). Finally, in the case of cervical cancer, the XGB-Cox model achieved the highest C-Index (0.7726), while the GBS obtained the highest IPCW C-Index (0.7643) and the lowest IBS (0.1473). The survival curves are presented in the Supplementary Material.

Table 2
Summary of the results obtained by the best models after searching for hyperparameters.

Chart 2 shows the most relevant variables for predicting survival, identified using the SHAP and PI methods, for each type of cancer analyzed. The models selected were those with the best performance in Table 2 according to the IPCW C-Index: XGB-AFT for breast, lung and colorectal; GBS for prostate and cervix. The variable CE (clinical staging) was consistently classified as the most important in the breast, prostate, cervical and colorectal cancer models, by both the SHAP and PI methods, while in lung cancer it was the most important only according to the SHAP method, and in the PI method the categorized time between consultation and treatment was the most important variable. The Figure shows the distribution of SHAP values.

Chart 2
Importance of features according to the SHAP and Permutation Importance methods.

Figure
SHAP values for interpretability of the best performing models, according to type of cancer.

DISCUSSION

This study stands out by demonstrating the applicability of SML models for predicting survival in patients with five types of cancer. The models performed as assessed by the C-Index, with values between 0.6686 and 0.7845, showing the potential of SML algorithms to predict survival from data collected by cancer registries. The GBS and XGB-AFT models had the best metrics for the various types of cancer. GBS was superior in prostate cancer, while XGB-AFT was the best among the other types of tumors. RSF obtained the third best result for all topographies. It should also be noted that the best algorithms had good predictive performance, above 0.73, regardless of the proportion of censored cases in the sample and the total number of patients included. GBS showed the best performance in terms of IBS across all tumor categories..

In a previous study, our group had already identified good predictive capacity using data from the cancer registry6. Considering five-year survival, an accuracy of 77.9% and an AUC of 0.858 were obtained with XGBoost. However, in that study it was necessary to exclude individuals who did not have a complete follow-up period. Thus, out of a total of 31,916 eligible patients, only 23,338 (73%) could be used in the analysis. In another publication, when comparing these results with the SML algorithms, the results showed that the algorithms for dichotomous outcomes underestimated the survival of patients with colorectal cancer by up to 18% in the first year, demonstrating that the SML models are more suitable for this type of analysis.

Considering the weight of the inputs in the model results, the CE variable was the most important in all the analyses, except in the PI for lung cancer, where it came second. This finding is consistent with that of other authors18, and demonstrates the importance of staging at diagnosis as the main prognostic factor for different types of cancer.

In breast cancer, the XGB-AFT model performed best, with a C-Index of 0.7845, surpassing the 0.73 obtained in a study of 36,958 patients diagnosed in the Netherlands between 2005 and 200819. In this study, the XGB outperformed the Cox-PH, RSF and SSVM models, whose C-Indexes ranged from 0.63 to 0.64. In the analysis of the importance of variables by SHAP, age and staging stood out as the main predictors, just as we found in our study.

For prostate cancer, the best result was achieved with the GBS model (C-Index = 0.7574), higher than that found in a study using the Surveillance, Epidemiology, and End Results (SEER) database, which brings together data from 18 population-based cancer registries in the United States20. Data from patients diagnosed with prostate cancer between 2000 and 2019, with a positive lymph node and no metastases, was used, with a total of 3,280 patients. In this investigation, a C-Index value of 0.745 was obtained with the Gradient Boosting Survival Analysis (GBSA) technique and with the Extra Survival Trees (EST), higher than the other models analyzed, RSF and Cox-PH, which ranged between 0.734 and 0.743.

In the case of lung cancer, the XGB-AFT also stood out (C-Index = 0.7368), surpassing the results obtained in a study using the population-based cancer registry of a state in Germany21. Data from patients diagnosed with lung cancer between 2016 and 2021 was used, with a total of 10,383 patients. In this literature study, the RSF model with variable imputation obtained a C-Index of 0.703, while the other approaches (Cox-PH, DeepSurv, and TabNet) performed less well (0.556 to 0.701).

For colorectal cancer, XGB-AFT obtained the highest C-Index in the analysis (0.7618), although this was lower than that found in a study using a hospital database in China22. Data from patients diagnosed with colorectal cancer between 2012 and 2019 was used, with a total of 2,157 patients. In the aforementioned study, a C-Index value of 0.789 was obtained for DeepHit, while the other techniques analyzed (Cox-PH, RSF, GB, DeepSurv, Cox-Time, and Neural Multitask Logistic Regression) varied between 0.781 and 0.787. The most important variable in the SHAP analysis of the DeepHit model was staging. It is worth noting that a key variable in the study, the surgical resection margin, is not available in the RHC/SP.

Finally, in cervical cancer, the best algorithm was XGB-AFT, which had a C-Index of 0.7726, lower than the value found in a study using the SEER database23. Data from patients diagnosed with cervical cancer between 2013 and 2015 was used, with a total of 3,810 patients. In this literature study, a C-Index value of 0.95 was obtained for the RSF, while the Cox-PH and Weibull obtained 0.81 and 0.80, respectively. The most important variable was T stage, followed by tumor size and staging. Variables related to treatment were used, such as chemotherapy and radiotherapy, in addition to the measurement of tumor size, which are variables we chose not to use in our study.

The use of variables restricted to individual information, such as place of residence, institution of care, clinical characteristics, and time until treatment, without including type of treatment or tumor recurrence, is a strength of the study. This choice broadens the applicability of the models in real planning contexts, as this data is available early in the records. Thus, the algorithms developed can support strategic health management decisions, even in scenarios with limited clinical data. However, the absence of more specific clinical information, such as tumor markers and recurrence data, restricts its use for individual clinical prediction, placing its main contribution in supporting health management and scenario building.

This study has some limitations. Firstly, it is important to note that the study used information from the RHC/SP database. This data is collected in hospitals from medical records by professionals from the institutions themselves, with heterogeneous teams. Although all the registrars receive specific training and use the SISRHC software, which has internal checks and consistency rules capable of standardizing the information, such as restrictions on the selection of morphologies compatible with each topography, inaccuracies can still occur, especially in more complex variables. One attempt to mitigate this risk was to exclude patients without staging or with undefined staging, which suggests poorer quality of the record.

In addition, the database currently does not capture some factors relevant to survival, such as race/skin color, risk factors such as smoking and clinical data such as tumor markers. The latter are particularly important, as they have an impact on more precise definitions of staging and the probability of survival. Therefore, updating registration systems to include new variables could increase the predictive potential of the models.

Additionally, it should be borne in mind that no direct comparison of performance was made with Cox models. The application of this type of model requires the proportionality of risks over time, which may make it unfeasible to include variables that do not meet this premise2. Thus, its use could restrict the set of predictor variables considered in the analysis.

Another relevant limitation is that it was not possible to calculate the IBS for the XGB-AFT, XGB-COX, and LGBM models, due to the lack of survival curve estimates with adequate adherence to the observed data, even after adjustment attempts. Future research could explore variations in parameterization, pre-processing or specific methodological adaptations for these algorithms, to make it possible to obtain reliable curves and, consequently, calculate the IBS. It is also worth noting that it was not possible to use SurvSHAP24, a specific method for assessing the impact of different variables on survival model predictions. This was because its application requires the generation of individual explanations over multiple survival time points, making processing particularly onerous on extensive bases and in more complex models. We therefore opted to use the traditional SHAP16 and PI17 for the global assessment of the importance of the variables.

Finally, although the models performed well on RHC/SP data, their application to records from other regions of the country has not been tested. It should be noted that one variable of great importance to the models, clinical staging, has high non-completion values in some regions of the country, which may limit its application. This lack of external validation limits the extrapolation of the results to other states, which could be the subject of future studies exploring the robustness of these models in more heterogeneous national cancer registry databases.

CONCLUSIONS

The SML algorithms proved to be applicable to the RHC/SP data for survival analysis. GBS, XGB-AFT, and RSF performed best, regardless of sample size or censoring proportion. The results reinforce the potential of these models to support evidence-based public policies, identifying profiles of patients or services at greater risk and contributing to better targeting of resources and organization of care networks.

Supplementary material

available from: https://doi.org/10.5281/zenodo.19250128

REFERENCES

  • 1 Bray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229-63. https://doi.org/10.3322/caac.21834
    » https://doi.org/10.3322/caac.21834
  • 2 Kleinbaum DG, Klein M. Survival analysis: a self-learning text. 3a ed. New York: Springer; 2012. (Statistics for biology and health).
  • 3 Bustamante-Teixeira MT, Faerstein E, Latorre MR. Técnicas de análise de sobrevida. Cad Saude Publica. 2024;18:579-94. https://doi.org/10.1590/S0102-311X2002000300003
    » https://doi.org/10.1590/S0102-311X2002000300003
  • 4 Kourou K, Exarchos TP, Exarchos TP, Exarchos K, Exarchos KP, Karamouzis MV, et al. Machine learning applications in cancer prognosis and prediction. Comput Struct Biotechnol J. 2015;13:8-17. https://doi.org/10.1016/j.csbj.2014.11.005
    » https://doi.org/10.1016/j.csbj.2014.11.005
  • 5 Tizi W, Berrado A. Machine learning for survival analysis in cancer research: a comparative study. Sci Afr. 2023;21:e01880. https://doi.org/10.1016/j.sciaf.2023.e01880
    » https://doi.org/10.1016/j.sciaf.2023.e01880
  • 6 Cardoso LB, Parro VC, Peres SV, Curado MP, Fernandes GA, Wünsch Filho V, et al. Machine learning for predicting survival of colorectal cancer patients. Sci Rep. 2023 Jun 1;13(1):8916. https://doi.org/10.1038/s41598-023-35649-9
    » https://doi.org/10.1038/s41598-023-35649-9
  • 7 Silva GFS, Duarte LS, Shirassu MM, Peres SV, Moraes MA, Chiavegatto Filho A. Machine learning for longitudinal mortality risk prediction in patients with malignant neoplasm in São Paulo, Brazil. Artif Intell Life Sci. 2023;3:100061. https://doi.org/10.1016/j.ailsci.2023.100061
    » https://doi.org/10.1016/j.ailsci.2023.100061
  • 8 Cardoso LB, Angelo SA, Bonilha YPG, Maia F, Ribeiro AG, Curado MP et al. Methodology for comparing machine learning algorithms for survival analysis. arXiv; 2025. [cited 2025 Oct 29]. Available from: https://arxiv.org/abs/2510.24473 doi: 10.48550/ARXIV.2510.24473.
    » https://doi.org/10.48550/ARXIV.2510.24473» https://arxiv.org/abs/2510.24473
  • 9 Fundação Oncocentro de São Paulo. Banco de dados do RHC. São Paulo: FOSP; 2024 [cited 2024 Jul 20]. Available from: https://fosp.saude.sp.gov.br/fosp/diretoria-adjunta-de-informacao-e-epidemiologia/rhc-registro-hospitalar-de-cancer/banco-de-dados-do-rhc/
    » https://fosp.saude.sp.gov.br/fosp/diretoria-adjunta-de-informacao-e-epidemiologia/rhc-registro-hospitalar-de-cancer/banco-de-dados-do-rhc/
  • 10 Harrell FE, Califf RM, Pryor DB, Lee KL, Rosati RA. Evaluating the yield of medical tests. JAMA. 1982;247(18):2543-6.
  • 11 Robins JM, Rotnitzky A, Zhao LP. Estimation of regression coefficients when some regressors are not always observed. J Am Stat Assoc. 1994;89(427):846-66. https://doi.org/10.1080/01621459.1994.10476818
    » https://doi.org/10.1080/01621459.1994.10476818
  • 12 Graf E, Schmoor C, Sauerbrei W, Schumacher M. Assessment and comparison of prognostic classification schemes for survival data. Stat Med. 1999;18(17-18):2529–45. https://doi.org/10.1002/(sici)1097-0258(19990915/30)18:17/18<2529::aid-sim274>3.0.co;2-5
    » https://doi.org/10.1002/(sici)1097-0258(19990915/30)18:17/18<2529::aid-sim274>3.0.co;2-5
  • 13 Akiba T, Sano S, Yanase T, Ohta T, Koyama M. Optuna: a next-generation hyperparameter optimization framework. In: Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. Anchorage: ACM; 2019 [cited 2026 Oct 22]. P2623-31. Available from: https://dl.acm.org/doi/10.1145/3292500.3330701
    » https://dl.acm.org/doi/10.1145/3292500.3330701
  • 14 Bergstra J, Bardenet R, Bengio Y, Kégl B. Algorithms for hyper-parameter optimization. In: Shawe-Taylor J, Zemel R, Bartlett P, Pereira F, Weinberger KQ, editors. Advances in neural information processing systems. New York: Curran Associates, Inc.; 2011 [cited 2026 Apr 22]. Available from: https://proceedings.neurips.cc/paper_files/paper/2011/file/86e8f7ab32cfd12577bc2619bc635690-Paper.pdf
    » https://proceedings.neurips.cc/paper_files/paper/2011/file/86e8f7ab32cfd12577bc2619bc635690-Paper.pdf
  • 15 Hansen N. The CMA evolution strategy: a tutorial. arXiv; 2016. https://doi.org/10.48550/ARXIV.1604.00772
    » https://doi.org/10.48550/ARXIV.1604.00772
  • 16 Lundberg S, Lee SI. A unified approach to interpreting model predictions. arXiv; 2017. https://doi.org/10.48550/ARXIV.1705.07874
    » https://doi.org/10.48550/ARXIV.1705.07874
  • 17 Breiman L. Random forests. Mach Learn. 2001;45(1):5-32. doi: 10.1023/A:1010933404324.
    » https://doi.org/10.1023/A:1010933404324
  • 18 Yu W, Lu Y, Shou H, Xu H, Shi L, Geng X, et al. A 5-year survival status prognosis of nonmetastatic cervical cancer patients through machine learning algorithms. Cancer Med. 2023;12(6):6867-76. https://doi.org/10.1002/cam4.5477
    » https://doi.org/10.1002/cam4.5477
  • 19 Moncada-Torres A, van Maaren MC, Hendriks MP, Siesling S, Geleijnse G. Explainable machine learning can outperform Cox regression predictions and provide insights in breast cancer survival. Sci Rep. 2021;11(1):6968. https://doi.org/10.1038/s41598-021-86327-7
    » https://doi.org/10.1038/s41598-021-86327-7
  • 20 Peng ZH, Tian J, Chen B, Zhou H, Bi H, He M et al. Development of machine learning prognostic models for overall survival of prostate cancer patients with lymph node-positive. Sci Rep. 2023;13(1):18449. https://doi.org/10.1038/s41598-023-45804-x
    » https://doi.org/10.1038/s41598-023-45804-x
  • 21 Germer S, Rudolph C, Labohm L, Katalinic A, Rath N, Rausch K et al. Survival analysis for lung cancer patients: a comparison of Cox regression and machine learning models. Int J Med Inform. 2024;191:105607. https://doi.org/10.1016/j.ijmedinf.2024.105607
    » https://doi.org/10.1016/j.ijmedinf.2024.105607
  • 22 Yang XJ, Qiu H, Wang LY, Wang X. Predicting colorectal cancer survival using time-to-event machine learning: retrospective cohort study. J Med Internet Res. 2023;25:e44417. https://doi.org/10.2196/44417
    » https://doi.org/10.2196/44417
  • 23 Kolasseri AE, Venkataramana B. Comparative study of machine learning and statistical survival models for enhancing cervical cancer prognosis and risk factor assessment using SEER data. Sci Rep. 2024;14(1):22203. https://doi.org/10.1038/s41598-024-72790-5
    » https://doi.org/10.1038/s41598-024-72790-5
  • 24 Krzyzinski M, Spytek M, Baniecki H, Biecek P. SurvSHAP(t): time-dependent explanations of machine learning survival models. Knowl Based Syst. 2023;262:110234. https://doi.org/10.1016/j.knosys.2022.110234
    » https://doi.org/10.1016/j.knosys.2022.110234

Edited by

Data availability

The data used in this study come from the Registro Hospitalar de Câncer do Estado de São Paulo (São Paulo State Cancer Registry Hospital) and can be accessed at by registering at: https://fosp.saude.sp.gov.br/fosp/diretoria-adjunta-de-informacao-e-epidemiologia/rhc-registro-hospitalar-de-cancer/banco-de-dados-do-rhc/.

Publication Dates

  • Publication in this collection
    24 July 2026
  • Date of issue
    2026

History

  • Received
    25 Nov 2025
  • Accepted
    1 Apr 2026
location_on
Faculdade de Saúde Pública da Universidade de São Paulo Avenida Dr. Arnaldo, 715, 01246-904 São Paulo SP Brazil, Tel./Fax: +55 11 3061-7985 - São Paulo - SP - Brazil
E-mail: revsp@usp.br
rss_feed Acompanhe os números deste periódico no seu leitor de RSS
Ir para o topo Reportar erro