INTRODUCTION
For more than 40 years, clinicians and researchers have used intensive care unit (ICU) scoring systems. These tools have been extensively tested and applied to assess the severity of acute illnesses, estimate the mean mortality, and evaluate ICU performance. Although the Acute Physiology and Chronic Health Evaluation (APACHE IV) and Simplified Acute Physiology Score (SAPS 3) are essential tools for ICU management, region-specific models have recently been developed and implemented to provide more precise predictions.1-3In the present study, we focused on how intensivists should use general scoring systems to assess ICU performance and how this information can guide management.
What are intensive care unit predictive scores?
Intensive care unit scoring systems, such as the APACHE and SAPS, are not disease specific; they reflect the overall severity of disease in critically ill patients at ICU admission and predict the risk of in-hospital mortality.4These scores consider chronic health conditions, physiological and laboratory variables, reasons for ICU admission and the use of supportive therapies. These variables, when applied in an equation, provide a number that represents the severity of critical illness at baseline. Subsequent equations estimate the risk of hospital mortality and the ICU length of stay (LOS) (Figure 1A).5
How can intensive care unit score results be interpreted?
Intensive care unit scores do not discriminate well between patients who live and those who die on an individual basis.6However, they do provide a mortality risk estimate. Thus, ICU managers can calculate the standardized mortality rate (SMR) by dividing the “mean observed mortality” by the “mean estimated mortality” for all patients admitted to an ICU during a given period.
When the actual hospital mortality rate is higher than the predicted mortality rate, the SMR is > 1. When the actual rate is lower than the predicted rate, the SMR is < 1. Suppose that during a specific year, 35% of all admitted patients died before hospital discharge, but the predicted mortality rate was 30%. In this case, the ICU SMR was 1.17 (i.e., 0.35/0.30). Thus, an SMR of 1.17 indicates an excess mortality rate of 17%, reflecting a potential imbalance between the quality of care and patients’ predicted conditions and outcomes.
Similarly, ICU managers can assess the standardized LOS according to severity, dividing the average observed ICU LOS by the expected ICU LOS.7The SAPS 3 group called this metric standardized resource use (SRU) and considered only the ICU LOS of survivors.8The same rationale described for the SMR can be applied here, where an SRU > 1 implies that more resources are used than are predicted by patient severity. The SRU estimates the average amount of resources (ICU days) used per surviving patient and is a well-validated proxy for hospital cost and efficiency (Figure 1B).9
Variations in organizational features across different ICUs can explain discrepancies between the predicted and obtained outcomes. These risk-adjusted metrics provide a data-driven approach to identify areas for improvement, implement changes, and measure their effectiveness. This continuous quality improvement cycle ensures optimal patient care within the ICU (Figure 1C).10
Pitfalls in intensive care unit prognostic score use (and how to overcome them):
-
Individual mortality prediction: ICU scoring systems should be used to predict mortality in a population, not in individuals. We should interpret an “x% predicted mortality” as follows: “for every 100 patients with the same characteristics, x% will probably die”.
-
Mortality prediction to guide patient care: the ICU score can inform managers but is not an adequate tool to guide goals of care or ICU interventions.
-
ICU allocation: scores are not tools for triage. ICU admission should always be guided by a clinical judgment and a prioritization framework.11
-
Timing: a premature analysis may not encompass all the outcomes for admitted patients. For example, if results from the first trimester are reported on April 1st, there are likely patients admitted in the last week of the first quarter who are still in the hospital. Therefore, we recommend allowing longer periods before definitive analysis and conclusions can be drawn.
-
“Transfer and discharge bias”: variation in the SMR could be partially explained by ICU and hospital discharge patterns. Every transferred patient is considered alive at discharge. Intensive care units that deal with frequent out-of-hospital transfers can artificially decrease the SMR, leading to biased comprehension of ICU quality.12 Considering these transfer rates allows a critical appraisal of this point.
-
“Garbage in, garbage out”: data collection is a crucial step of ICU management. Missing data and inconsistencies associated with data collection compromise calculated scores and their predictions.13
-
An extremely low SMR does not mean that the quality of care is good for all patients: the SMR represents an average value and fails to capture the potential variability in care and outcomes across different groups.
-
The distribution of illness severity can impact performance metrics: in ICUs with greater proportion of lower-risk patients, the scoring system may (globally) overestimate the mortality rate, a situation known as Simpson’s paradox.14
How can benchmarking be used for intensive care unit management?
-
The ICU performance metrics of a recent period (month, trimester, year) can be presented. Context should always be provided if appropriate, such as changes in structure, process, and case mix that could be associated with the outcomes (Table 1S - Supplementary Material)
-
Trends should be analyzed over time. Performance evaluation is a dynamic process. Observing how the performance metrics of ICUs change over time (instead of checking an isolated time interval) is strongly recommended.
-
Comparisons of ICU performance to that of recognized benchmarks and general ICUs should be performed. In addition, comparisons with similar profiles, especially for specialized ICUs (e.g., neurological, surgical), may be relevant. Finally, the case mix should always be considered to allow for an adequate comparison.15
Benchmarking SMRs and lengths of stay provides ICU and hospital managers with a broader view of quality improvement. Other domains suitable for benchmarking include adherence to care processes, patient safety, economic outcomes, and patient or family satisfaction. Currently, benchmarking in real time by using electronic multinational platforms such as Epimed,1ANZICS,2and NICE is possible.3By leveraging real-time data, ICU managers can identify areas for improvement, implement targeted interventions, and continuously monitor intervention effectiveness.
Conclusions
A pragmatic and rigorous interpretation of SMRs and risk-adjusted lengths of stay allows data-driven approaches to assess and improve ICU performance. These are essential principles for intensivists, as prognostic scores are valuable tools for continuous ICU evaluation and management. Understanding the benefits and limitations of these tools may ensure broader implementation and drive better organizational interventions to improve ICU quality and efficiency.
REFERENCES
- 1 Soares M, Borges LP, Bastos LD, Zampieri FG, Miranda GA, Kurtz P, et al. Update on the Epimed Monitor Adult ICU Database: 15 years of its use in national registries, quality improvement initiatives and clinical research. Crit Care Sci. 2024;36:e20240150en.
- 2 Burrell AJ, Udy A, Straney L, Huckson S, Chavan S, Saethern J, et al. "The ICU efficiency plot": a novel graphical measure of ICU performance in Australia and New Zealand. Crit Care Resusc. 2023;23(2):128-31.
- 3 van de Klundert N, Holman R, Dongelmans DA, de Keizer NF. Data Resource Profile: the Dutch National Intensive Care Evaluation (NICE) Registry of Admissions to Adult Intensive Care Units. Int J Epidemiol. 2015;44(6):1850-1850h.
- 4 Zampieri FG, Granholm A, Møller MH, Scotti AV, Alves A, Cabral MM, et al. Customization and external validation of the Simplified Mortality Score for the Intensive Care Unit (SMS-ICU) in Brazilian critically ill patients. J Crit Care. 2020;59:94-100.
- 5 Keegan MT, Soares M. What every intensivist should know about prognostic scoring systems and risk-adjusted mortality. Rev Bras Ter Intensiva. 2016;28(3):264-9.
- 6 Sinuff T, Adhikari NK, Cook DJ, Schünemann HJ, Griffith LE, Rocker G, et al. Mortality predictions in the intensive care unit: comparing physicians with scoring systems. Crit Care Med. 2006;34(3):878-85.
- 7 Peres IT, Hamacher S, Oliveira FL, Bozza FA, Salluh JI. Prediction of intensive care units length of stay: a concise review. Rev Bras Ter Intensiva. 2021;33(2):183-7.
- 8 Rothen HU, Stricker K, Einfalt J, Bauer P, Metnitz PG, Moreno RP, et al. Variability in outcome and resource use in intensive care units. Intensive Care Med. 2007;33(8):1329-36.
- 9 Rapoport J, Teres D, Zhao Y, Lemeshow S. Length of stay data as a guide to hospital economic performance for ICU patients. Med Care. 2003;41(3):386-97.
- 10 Quintairos A, Pilcher D, Salluh JI. ICU scoring systems. Intensive Care Med. 2023;49(2):223-5.
- 11 Nates JL, Nunnally M, Kleinpell R, Blosser S, Goldner J, Birriel B, et al. ICU Admission, Discharge, and Triage Guidelines: A Framework to Enhance Clinical Operations, Development of Institutional Policies, and Further Research. Crit Care Med. 2016;44(8):1553-602.
- 12 Reineck LA, Pike F, Le TQ, Cicero BD, Iwashyna TJ, Kahn JM. Hospital factors associated with discharge bias in ICU performance measurement. Crit Care Med. 2014;42(5):1055-64.
- 13 Afessa B, Keegan MT, Gajic O, Hubmayr RD, Peters SG. The influence of missing components of the Acute Physiology Score of APACHE III on the measurement of ICU performance. Intensive Care Med. 2005;31(11):1537-43.
- 14 Marang-van de Mheen PJ, Shojania KG. Simpson's paradox: how performance measurement can fail even with perfect risk adjustment. BMJ Qual Saf. 2014;23(9):701-5.
- 15 Soares M, Salluh JIF, Zampieri FG, Bozza FA, Kurtz PMP. A decade of the ORCHESTRA study: organizational characteristics, patient outcomes, performance and efficiency in critical care. Crit Care Sci. 2024;36:e20240118en.
-
Funding:
Dr. Jorge Ibrain Figueira Salluh is supported in part by individual research grants from Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq) and Fundação Carlos Chagas Filho de Amparo à Pesquisa do Estado do Rio de Janeiro (FAPERJ).
Edited by
-
Responsible editor:
Pedro Henrique Rigotti Soares


(A) Quality metrics: from patient data to intensive care unit data; (B) Benchmarking. Rothen et al.