Abstract
Those interested in measuring violent crime use the number of homicides, based on the never-verified assumption that the seriousness of the crime implies negligible underreporting. In Brazil, the Mortality Information System (SIM) is the source usually used to measure the number of homicides. Between 1996 and 2021, among deaths from external causes, 8.2% were classified as deaths from external causes of undetermined intent (MVCI), indicating the possibility of homicide underreporting. From the microdata on external cause deaths recorded in the SIM, a supervised learning algorithm learns the characteristics of homicides, accidents and suicides and estimates the probability that MVCIs are homicides incorrectly recorded as MVCI, allowing us to estimate the number of shadow homicides in the SIM. The result suggests that between 1996 and 2021, 128,567 shadow homicides were not accounted for, or 43.6% of MVCI s are actually homicides. Evidence contrary to the hypothesis of reduced underreporting, implying significant impacts on public policy evaluations, econometric procedures and the need to account for these hidden homicides in crime measurements.
Key words:
Machine Learning; Homicide; Crime; Underreporting; DATASUS
Resumo
Interessados em mensurar a criminalidade violenta utilizam o número de homicídios, a partir de hipótese, nunca verificada, de que a gravidade do crime implica subnotificação desprezível. No Brasil, o Sistema de Informações sobre Mortalidade (SIM) é fonte habitualmente utilizada na aferição do número de homicídio. Entre 1996 e 2021, dentre as mortes por causa externa, 8,2% são classificadas como mortes por causa externa de intenção indeterminada (MVCI), indicando possibilidade de subnotificação de homicídios. A partir dos microdados sobre morte por causa externa registradas no SIM, algoritmo de aprendizado supervisionado aprende as características de homicídios, acidentes e suicídios e estima probabilidade de as MVCIs serem na realidade homicídio incorretamente registrados como MVCI, permitindo-nos estimar o número de homicídios ocultos no SIM. O resultado encontrado sugere que, entre 1996 e 2021, não foram contabilizados 128.567 homicídios ocultos ou 43,6% das MVCI são na realidade homicídios. Evidência contrária à hipótese de reduzida subnotificação, implicando impactos significativos em avaliações de políticas públicas, procedimentos econométricos e necessidade de contabilização destes homicídios ocultos em medições da criminalidade.
Palavras-chave:
Aprendizado de Maquinas; Homicídio; Crime; Subnotificação; DATASUS
Resumen
Quienes se interesan en medir los delitos violentos utilizan el número de homicidios, basándose en la hipótesis, nunca verificada, de que la gravedad del delito implica un subregistro insignificante. En Brasil, el Sistema de Información de Mortalidad (SIM) es la fuente comúnmente utilizada para medir el número de homicidios. Entre 1996 y 2021, el 8,2% de las muertes por causas externas se clasificaron como muertes por causas externas de intención indeterminada (MVCI), lo que indica la posibilidad de subregistro. Utilizando microdatos sobre muertes por causas externas registrados en el SIM, un algoritmo de aprendizaje supervisado aprende las características de homicidios, accidentes y suicidios y estima la probabilidad de que las MVCI sean en realidad homicidios registrados incorrectamente como MVCI, lo que permite estimar el número de homicidios ocultos en el SIM. Los resultados sugieren que, entre 1996 y 2021, 128.567 homicidios ocultos no se reportaron, o que el 43,6% de las MVCI fueron en realidad homicidios. Esta evidencia contradice la hipótesis de reducción del subregistro, lo que implica impactos significativos en las evaluaciones de políticas públicas, los procedimientos econométricos y la necesidad de contabilizar estos homicidios ocultos en las mediciones de la delincuencia.
Palabras clave:
Aprendizaje automático; Homicidio; Delincuencia; Subdeclaración; DATASUS
Introduction
The use of the homicide rate per 100,000 inhabitants as a proxy for criminal dynamics appears justified and recommended in the literature in the hypothesis of alleged reduced underreporting1. In Brazil, the Ministry of Health’s Mortality Information System (SIM), a database used in several criminology studies (e.g., the Atlas of Violence2), uses the International Classification of Diseases (ICD-10) to classify causes of death. Chapter XX of the ICD-10 covers deaths from external causes, i.e., accidents, suicides, homicides, medical sequelae, and deaths from external causes of undetermined intent (ECUI). Deaths classified as ECUI are events in which, in principle, it is impossible to distinguish between homicides, accidents, and suicides3. This study investigates the possibility that deaths from ECUIs are, in fact, unidentified homicides and, therefore, hidden from researchers. This possibility has not been considered in crime analyses and could have significant impacts on public policy assessments.
From 1996 to 2021, 8.2% of deaths from external causes were classified as deaths from external causes of undetermined intent (ECUI). That is, on average, the health system was unable to identify the intentionality of 11,336 deaths per year, a figure higher than the annual homicide average in São Paulo. In certain states, the number of ECUIs exceeded the number of homicides, hindering an understanding of the proper level of violence. For example, in São Paulo, from 2018 to 2021, the number of ECUIs was, on average, 27.9% higher than the number of homicides. The national average of 6.05 ECUIs per 100,000 inhabitants brings Brazil closer to the group of European countries with the most significant uncertainty about the intentionality of deaths from external causes. Therefore, there is significant uncertainty regarding the intentionality of deaths from external causes in Brazil.
Analyzing the death records of Eastern European countries as a whole, the evidence presented in the Russian case4 and the Baltic countries5 indicates that a significant portion of ECUIs are, in fact, unreported homicides. In Brazil, inconsistencies in SIM records and proposals for rectification are well-known6. There is evidence of under-enumeration of deaths in the SIM, indicating, in general, that there are erroneously classified deaths, including homicide cases that were recorded as ECUIs. The proposals to use fieldwork techniques and, therefore, the need for a local focus, make it impossible to generalize these proposals to national records. An alternative to official records, victimization surveys aim to circumvent underreporting by questioning the population and are developed to overcome the limitations of police data. However, in addition to being subject to revealed preference bias7, in the Brazilian case, there are only three nationwide victimization surveys8, yet these are incomparable in time and focused on property crimes.
Daniel Cerqueira9 developed a method that can assign the probability that each ECUI is actually a homicide. The method identifies differentiated death patterns based on the situational characteristics of the incident and the individual characteristics of the victims. In this work, we improved the adopted methodology and expanded the period analyzed. Thus, based on microdata on deaths from external causes recorded in the SIM, a supervised learning algorithm learns the characteristics of homicides, accidents, and suicides and estimates the probability that deaths from ECUIs recorded from 1996 to 2021 are actually homicides incorrectly recorded as ECUIs, allowing us to estimate the number of hidden homicides in the SIM.
The results suggest that, from 1996 to 2021, 128,567 hidden homicides went uncounted, or that 43.6% of deaths from ECUIs are in fact hidden homicides. The 9.6% increase in the total number of homicides, when considering hidden homicides, allows for a more accurate portrayal of violent crime, for example, by indicating a fictitious reduction in homicides in some states. It also suggests, at least in the Brazilian case, violations of commonly adopted assumptions in the treatment of variables with measurement error. This highlights the relevance of the results and indicates a non-negligible impact on diagnoses and public policy analyses, inconsistent with the expected reduced underreporting advocated in the literature.
Following this introduction, section two presents the basis used and a summary of the literature on the quality of SIM records. Section three discusses ECUIs and presents the methodology used. Hidden and projected homicides are discussed in section four. Finally, the final section summarizes the study’s conclusions.
Methods
By law, no burial should take place in the absence of a death certificate10. In the case of unnatural deaths, regulations11,12 make it mandatory for the forensic medical service to provide a Death Certificate (DC) stating the underlying cause of death.
The death certificate used in SIM records of deaths due to external causes must be issued by a forensic doctor, following a forensic autopsy report. Based on this forensic examination and information provided by family members, individuals who helped the victim, or the police, the forensic doctor attempts to establish the underlying cause of death. That is, to determine whether the death due to external causes was caused by: i) accidents (e.g., car accident); ii) intentional self-harm (i.e., suicide); iii) assaults (e.g., assault with a firearm); iv) legal interventions and war operations; v) deaths by undetermined intent (ECUI); vi) complications from medical care, medical sequelae, or natural accidents. Based on the information gathered by the medical examiner, coders at the municipal health secretariats will enter the ICD-10 code underlying the death. If it is impossible to distinguish between homicide, accident, and suicide, the underlying cause is classified as death by external causes of undetermined intent (ECUI).
Thus, considering deaths belonging to chapter XX of ICD-10 (e.g., external causes of morbidity and mortality), the records were aggregated into four death groups, presented in Chart 1, according to the intentionality of the basic cause of death and the instrument that initiated the morbid process: i) death of determined intention resulting from an unnatural accident or suicide are classified in the accidents/suicides group; ii) deaths of determined intent resulting from aggression or legal interventions are classified in the homicide group; iii) deaths due to external causes of undetermined intent (ECUI) are classified in the undetermined group; iv) Deaths of determined intent resulting from medical sequelae, food deprivation and natural accidents, such as crocodile bite, volcanic eruption and lightning strike, are excluded from further analysis, given the impossibility of this type of death being represented by ECUI.
The instrument variable was created based on the ICD-10 code recorded in the underlying cause of death. This variable identifies the instrument responsible for triggering the morbid process, capturing the dynamics associated with deaths. Ten instrument categories were thus established (i.e., poisoning, hanging, drowning, gunshot wound (GWP), blunt force, fire, piercing, blunt force, unknown, and vehicle). Deaths caused by impact result from a variety of events, including falls, falling objects, crushing injuries caused by tools and utensils, and boiler and other material explosions. The piercing category includes deaths caused by piercing or cutting objects. Blunt force includes a variety of actions, such as striking, beating, kicking, and biting. Hanging also includes cases of strangulation. Fire includes deaths caused by smoke inhalation resulting from fire and arson. Poisoning results from the ingestion of a wide variety of substances, such as alcohol, psychoactive drugs, medications, and solvents.
Chart 1 shows the groups of deaths classified by intentionality, the instrument responsible for the underlying cause of death, and the ICD-10 code considered. Thus, for example, accidental death of a pedestrian in a collision with a car (ICD-10:V03) is an unnatural accident of determined intent and, therefore, is included in the accident/suicide group, caused by the vehicle instrument.
The relevance of deaths categorized under the ICD-10 in mortality research has prompted analyses that seek to assess the quality of these records and present proposals to overcome existing limitations. In a group of former Soviet republics5, this suggests the use of ECUIs in an attempt to conceal the number of suicides and homicides. In Russia4, from 2000 to 2011, 33% of ECUIs were, in fact, incorrectly recorded homicides. Meanwhile, in Finland13, 10% of ECUIs were suicides. In Argentina14, from 1997 to 2018, 28.5% of ECUIs were identified as hidden homicides.
In the Brazilian case, from 2000 to 2009, using active search, Viçosa15 identified that 57.7% of the 391 deaths from undetermined external causes were in fact homicides. In the State of Rio de Janeiro16, there is a fragility in the flow of information between those responsible for registering deaths and those responsible for determining the intentions of deaths. Finally, using a Bayesian approach and 2001 data, Luciana Tricai Cavalini17 suggests an underreporting of 5.9% of deaths. Therefore, in several contexts, the literature suggests underreporting of deaths from external causes, that is, the failure to register deaths from external causes or the registration of these deaths in poorly defined categories, deaths from external causes of undetermined intent (ECUI).
In an ecologically designed study, the investigation of the number of hidden homicides uses information recorded in death certificates due to external causes, accessed through SIM microdata from 1996 to 2021, creating a data panel with all the Federation Units (UFs) plus the Federal District.
As described in Chart 1, the investigation considers three types of death by external causes: the total homicide (H*) group includes assaults and legal interventions; the total accident/suicide (AS*) group is defined as the union of deaths resulting from unnatural accidents or suicide; and, finally, the ECUI group. Total accident/suicide (AS*) and total homicide (H*) are only partially observed; that is, the unobserved portion is misclassified as death by external causes of undetermined intent (ECUI). Equations 1, 2, and 3 summarize the problem:
Where HO and ASO are, respectively, hidden homicides and hidden accidents/suicides incorrectly recorded as deaths due to external causes of undetermined intent (ECUI), and H and AS are, respectively, the homicides and accidents/suicides recorded in the SIM, whose intent was determined by the health system. This study aims to identify hidden homicides (HO) registered in the SIM as deaths due to external causes of undetermined intent (ECUI). Failure to register a homicide in the SIM or recording it under the heading of natural deaths implies that hidden homicides found are considered a lower limit for incorrect homicide registration. In turn, as long as the cause of death due to external causes has been determined, the classification is assumed to be correct.
The identification of hidden homicides follows a methodology used in several classification problems18 and occurs in three stages: hyperparameter streamlining (model selection), generalization testing (model assessment), and final model adjustment.
In the model selection stage, the database composed of the binary dependent variable of the cause of death (homicides (H) or accidents/suicides (AS)) and predictor variables was divided in a stratified random manner into a training base (70%) and a test base (30%), in order to maintain the proportion of homicides and accidents/suicides of the original database (stratified random sampling) in each partition. Due to the imbalance in the training base, composed of 42.9% of homicides, the training base was resampled using the Synthetic Minority Oversampling Technique (SMOTE)19. Next, hyperparameter streamlining occurs using 10-fold cross-validation, adjusting seven algorithms based on training (logistic regression, penalized logistic regression (ridge, lasso, and elastic net), decision tree, bagging decision trees, and random forest). There are several possible hyperparameter combinations, and this step seeks to identify the combination with the best out-of-sample predictive ability for each algorithm. Therefore, the hyperparameter combination with the highest average area under the ROC curve is selected.
The model assessment stage investigates the possibility of overfitting and the generalizability of the combination of hyperparameters selected in model selection when classifying the test base, in other words, the ability to accurately classify new observations and, in this work, the algorithm’s ability to identify hidden homicides. The algorithm and its respective combination of hyperparameters with the most significant area under the ROC curve in the test base will be used to identify hidden homicides. In the hidden homicide identification stage, the algorithm selected in the model assessment stage is adjusted across the entire dataset (i.e., training and testing), establishing the predictive model used20.
Predictive variables are those with information throughout the analyzed period. Thus, the predictors are mostly categorical variables, i.e., gender, ethnicity/skin color, marital status, victim’s schooling, incident location, underlying cause of death instrument, year, month, day of the week of death, and state of occurrence. The categorical variables are transformed into binary variables of the variable’s components (indicator variable). The victim’s age is the only continuous variable. Missing values for the victim’s age are imputed using the k-nearest neighbors method (k=5) from the information on known deaths and using the other variables as imputation predictors. Finally, the model interpretation uses SHAP Value to verify the contribution of the predictor variables to the estimated probability21,22.
The Shapley Additive Explanations (SHAP) method explains the estimated probability of each observation by assigning to each predictor variable the percentage contribution of this variable to the estimated probability, indicating the direction and intensity of the association between the values of the predictor variables and the estimated probability. Inspired by the Shapley value23, the method was initially designed to distribute the reward of a cooperative game among coalition participants. In machine learning, the estimated probability of each observation is distributed among the P predictor variables, such that the SHAP of the j-th variable is the weighted average of the marginal contribution of variable j to the estimated probability across all possible coalitions of variables p!. Formally, the SHAP of the j-th variable is:
That is, SHAP for the “j” variable performs the sum of the S coalitions of variables that do not contain the j-th variable. The marginal contribution of the j-th variable in each coalition will be the difference between the estimated probability including the j-th variable and without the j-th variable , weighted by the probability of variable j contributing to coalition S. In summary, SHAP assigns importance to the predictor variable by comparing the model’s prediction with and without the variable. However, because the order in which a model adjusts variables can affect its predictions, comparisons are performed in all possible orders so that variables are compared fairly.
Model estimation and interpretation
Table 1a presents the average predictive performance metrics of the seven algorithms after applying 10-fold cross-validation to the training dataset. At this stage, the algorithms are streamlined; therefore, for each algorithm, we present the performance of the hyperparameter combination with the most significant area under the ROC curve. In the training dataset cross-validation process, the algorithms achieved excellent predictive performance, and random forest (RF) was the best performer, outperforming the other algorithms in the metrics considered.
According to the evidence presented in Table 1b, similar to the model selection process in the training dataset, the algorithms used presented excellent predictive performance in predicting new observations from the test dataset. The good performance in classifying new observations suggests adequate generalization capability24-26 and the lack of overfitting. In the model assessment stage, the random forest (RF), by a small margin, displayed the best performance in the considered metrics and, therefore, was selected as the model to be used in the identification of hidden homicides.
Figure 1 presents the SHAP Summary Plot of the deaths from the ECUI database, that is, the mean of the absolute SHAP values of each variable across all estimated probabilities of the deaths from ECUIs being concealed homicide. The y-axis shows, in decreasing order, the influence of the predictive variables, that is, the ranking of the relative importance of each variable in the calculation of the estimated probabilities. The result indicates the importance of the instrument variable responsible for the underlying cause of death in the calculation of the classifications, with the highest absolute mean SHAP value (0.225), followed by incident location (0.052) and age (0.046). The contribution of the other variables appears to be negligible.
Results and discussion
From 1996 to 2021, the health system failed to identify 128,567 homicides. Therefore, the findings suggest that 43.6% of the deaths from ECUIs recorded in Brazil are actually hidden homicides. This figure is 10 percentage points higher than a similar study investigating deaths from ECUIs in Russia4. On average, the state failed to identify 4,944 homicides per year, a figure close to the median homicide rate in Bahia. In other words, in each year of the analyzed period, an average number of homicides similar to that observed in the state with the third-highest absolute number of homicides was not recorded. This finding suggests that, at least in the Brazilian case, contrary to what is advocated in the literature1, the extent of the undercounting of homicides is not negligible.
The trend in the proportion of hidden homicides incorrectly classified in the deaths from ECUIs category reveals three patterns of undercounting. Between 1996 and 2009, the mean proportion of 47.6% was the highest in the entire period. In this context of high uncertainty, 1999 recorded the highest proportion in the series, at 53.4%. From 2010 and during the subsequent seven years, we observed a decline in the absolute number of deaths from ECUIs, while the mean proportion of hidden homicides in the deaths from ECUIs reached 37.1%, the lowest average in the series, with a minimum value of 35.4%. Even so, in these years of lower undercounting, at least one-third of deaths from ECUIs were hidden homicides. In the final four years, uncertainty about the intentionality of deaths increased again, and 2019 recorded the highest absolute number of deaths from ECUIs. However, there was no similar increase in hidden homicides, although the annual average increased to 40.9% in those same four years.
To avoid the statistical bias resulting from underreporting of crime indicators, some studies assume that the proportion of underreporting remains constant over time in the areas under investigation27. The evidence presented points to substantial variability in the proportion of underreported homicides, contradicting, at least in the case of Brazilian homicides, the hypothesis generally adopted in the literature.
At the subnational level, the five states with the highest absolute number of recorded homicides - São Paulo, Rio de Janeiro, Bahia, Minas Gerais, and Pernambuco - are also those with the highest incidence of hidden homicides, accounting for 78.2% of the total number of hidden homicides in the country. Notably, these states remained among those with the five highest absolute numbers of hidden homicides throughout the period, with São Paulo having the highest total in absolute terms, having been the state with the highest absolute number of hidden homicides for twenty years.
Figure 2 presents the impact of hidden homicides on violent crime assessments by comparing the rates per 100,000 inhabitants of registered and projected homicides, that is, the sum of hidden homicides and registered homicides, among the ten most violent states of each year, in addition to the registered and projected rates for Brazil.
In the case of federal states, considering hidden homicides indicates a fictitious reduction in violent crime in some states. For example, between 1998 and 2009, Rio de Janeiro, the state with the highest median hidden homicide rate, rose nine times. In 2019, only after accounting for hidden homicides did Rio de Janeiro appear among the ten most violent states. This year, 2,480 hidden homicides were identified in Rio de Janeiro, representing 70.0% of the homicides recorded in the state. On four occasions, Roraima also appears among the most violent states, only after including hidden homicides. Furthermore, including hidden homicides in the calculation of state homicide rates results in a shift in the rankings of the most violent states every year, and in eight years, the most violent state has changed.
A similarity exercise between recorded and projected rates finds high values for the root mean square error (RMSE) metric and indicates the states most affected by accounting for hidden homicides. Thus, six states (Rio de Janeiro, Bahia, Sergipe, Rio Grande do Norte, Roraima, and São Paulo) appear with RMSEs greater than 4, a significant result given the state averages for recorded homicides. For example, in São Paulo, the projected rate is on average 17.7% higher than the recorded rate over the entire period, or 69.1% higher considering only the final four years.
This new picture of violent crime alters diagnoses and assessments of security policies. For example, in Rio de Janeiro, the reduction in homicides during the period of accelerated implementation of the Pacification Police Unit28, that is, between 2009 and 2013, of -6.9% in the registered homicide rate, becomes a 26.3% reduction in the projected homicide rate, an effect caused by high underreporting in 2009. Also, in Rio de Janeiro, in 2019, a period of high uncertainty about the intentionality of deaths in the State, the annual growth rate of registered and projected homicides presents a singular distance among the other States and is relevant in public policy assessments. The registered rate indicates a drop of 45.6%, while the projected rate suggests a retraction of 12.1%.
In the national aggregate, the comparison between the recorded and projected homicide rates using predictive metrics reveals a significant difference. For example, 2.76 units in the case of RMSE (Root Mean Square Error). This high error value, relative to the mean number of recorded homicides, suggests that accounting for hidden homicides brings new meaning to analyses of Brazilian homicide dynamics. For example, while the 2019 annual growth rate indicates a 22.1% reduction in the recorded homicide rate, in the case of the projected homicide rate, this indicator indicates a 17.0% reduction, a difference of five percentage points, a relevant result in public policy evaluations.
The difference between the recorded and projected annual homicide growth rate summarizes the impact of including hidden homicides in the Brazilian homicide time series. According to Figure 3, the most significant differences between the annual growth rates occur at the beginning and end of the analyzed series, periods of high hidden homicide numbers. This finding, in itself, suggests a new interpretation of criminal dynamics. Additionally, in intermediate periods, a reversal in the annual variation was observed on five occasions, a finding that alters public safety diagnoses and highlights the importance of questioning the hypothesis of reduced underreporting in homicides. As a result of the different annual variations, on average, the projected homicide rate exceeds the recorded rate by 10.0%, and, in the period analyzed, the cumulative number of projected homicides exceeded the recorded rate by 9.6%, totaling 1,454,544 homicides. These findings show the imperative of including hidden homicides in violent crime investigations. The lack of these findings precludes a diagnosis capable of guiding intervention targeting the most critical states or characterizing the most victimized populations, hindering the development of policies aimed at reducing violent crime.
Conclusion
Adopting homicide as an indicator of criminal dynamics is recommended by the literature and justified by researchers based on the unverified hypothesis of lower underreporting1,7. In Brazil, the SIM is a reference for information on homicides and is a source of several studies on crime2,29. However, high underreporting of deaths in the SIM6, that is, the failure to register deaths or registration in poorly defined categories, and the evidence found in the international4,14 and national15 literature, suggest that a portion of deaths from ECUIs are, in fact, hidden homicides.
From the microdata produced when filling out death certificates for deaths due to external causes registered in the SIM, between 1996 and 2021, using the individual and situational characteristics of these deaths, a set of supervised learning models learned the characteristics of homicides and accidents/suicides and, then, the model with the best generalization capacity classified deaths due to external causes of undetermined intent as homicides or accidents/suicides, identifying the lower limit of hidden homicides in the SIM.
Thus, 128,567 hidden homicides were identified, meaning 43.6% of deaths from ECUIs in indeed hidden homicides. The number of hidden homicides identified can alter diagnoses of criminal dynamics, public policy assessments, and invalidate econometric procedures typically adopted to correct underreporting. For example, while variations in the time distribution of hidden homicides among states and the frequency of undercounting invalidate strategies traditionally used to correct underreporting in panel data27 or to redistribute homicides30, the identification of hidden homicides indicates a fictitious reduction in homicides in some states, a shift in positions among the most violent states, and a reversal in the annual variation rate of Brazilian homicides, altering the understanding of criminal dynamics, impacts revealed in the underestimated results of the UPPs on the trajectory of homicides in Rio de Janeiro.
The evidence of undercounting presented by Soares Filho6 should serve as a warning about the problems in the SIM’s death accounting system. Of these problems, the high number of deaths from external causes of undetermined intent is, per se, the main objection to using homicides recorded in the SIM under the assumption of reduced underreporting. The identification of 128,567 hidden homicides suggests that the SIM fails to record, on average, in each year of the analyzed period, a number of homicides similar to that of Bahia, the state with the third-highest absolute number of homicides, figures that cannot be overlooked. Therefore, the significant number of hidden homicides contradicts the commonly held hypothesis of reduced underreporting in homicide records1. Given the potential negative impacts of all kinds, using the number of homicides recorded in the SIM is a mandatory strategy to reduce the bias introduced by hidden homicides. The proposed method, however, does not exhaust the possibilities for identifying hidden homicides. Its generalizability is positively correlated with the quality of information provided on deaths with known intent. Therefore, high uncertainty about death characteristics can affect the model’s predictive capacity and should not be understood as a substitute for the need to improve the quality of SIM information.
References
- 1 Pinotti P. The Credibility Revolution in the Empirical Analysis of Crime. Ital Econ J 2020; 6(2):207-220.
-
2 Cerqueira D, Bueno S, Lima RS, Alves PP, Marques D, Lins GOA. Atlas da Violencia 2023 [Internet]. 2023 [acessado 2021 jun 10]. Disponível em: https://www.ipea.gov.br/atlasviolencia/arquivos/artigos/9350-223443riatlasdaviolencia2023-final.pdf
» https://www.ipea.gov.br/atlasviolencia/arquivos/artigos/9350-223443riatlasdaviolencia2023-final.pdf -
3 World Health Organization (WHO). International statistical classification of diseases and related health problems, 10th revision [Internet]. 5ª ed. 2016 [acessado 2021 jun 10]. Available: https://apps.who.int/iris/handle/10665/246208
» https://apps.who.int/iris/handle/10665/246208 - 4 Andreev E, Shkolnikov VM, Pridemore WA, Nikitina SY. A method for reclassifying cause of death in cases categorized as 'event of undetermined intent'. Popul Health Metr 2015; 13(1):23.
- 5 Värnik P, Sisask M, Värnik A, Yur'yev A, Kõlves K, Leppik L, Nemtsov A, Wasserman D. Massive increase in injury deaths of undetermined intent in ex-USSR Baltic and Slavic countries: Hidden suicides? Scand J Public Health 2010; 38(4):395-403.
- 6 Soares Filho AM, Cortez-Escalante JJ, França E. Review of deaths correction methods and quality dimensions of the underlying cause for accidents and violence in Brazil. Cien Saude Colet 2016; 21(12):3803-3818.
- 7 Tabarrok A, Heaton P, Helland E. The Measure of Vice and Sin: A Review of the Uses, Limitations and Implications of Crime Data. In: Benson BL, Zimmerman PR, editors. Handbook on the Economics of Crime. Cheltenham: Edward Elgar Publishing; 2010. p. 53-81.
- 8 Monteiro J, Caballero B. Crime e Violência. In: Guia brasileiro de análise de dados: armadilhas & soluções. 1ª ed. Brasília: Enap; 2021. p. 126-169.
- 9 Cerqueira D. Mortes violentas não esclarecidas e impunidade no Rio de Janeiro. Econ Apl 2012; 16(2):201-235.
- 10 Brasil. Lei nº 6.015, de 31 de dezembro de 1973. Dispõe sobre os registros públicos, e dá outras providências. Diário Oficial da União; 1973.
- 11 Brasil. Ministério da Saúde (MS). Conselho Federal de Medicina. Centro Brasileiro de Classificação de Doenças. A declaração de óbito: documento necessário e importante. 3ª ed. Brasília: MS; 2009.
- 12 Conselho Federal de Medicina (CFM). Resolução CFM n° 1.779/2005. Regulamenta a responsabilidade médica no fornecimento da Declaração de Óbito. Diário Oficial da União; 2005.
- 13 Ohberg A, Lonnqvist J. Suicides hidden among undetermined deaths. Acta Psychiatr Scand 1998; 98(3):214-218.
- 14 Santoro A. Recálculo de las tendencias de mortalidad por accidentes, suicidios y homicidios en Argentina, 1997-2018. Rev Panam Salud Publica 2021; 44:1.
- 15 Melo CM, Bevilacqua PD, Barletto M, França EB. Qualidade da informação sobre óbitos por causas externas em município de médio porte em Minas Gerais, Brasil. Cad Saude Publica 2014; 30(9):1999-2004.
- 16 Lopes AS, Passos VMA, Souza MFM, Cascão AM. Melhoria da qualidade do registro da causa básica de morte por causas externas a partir do relacionamento de dados dos setores Saúde, Segurança Pública e imprensa, no estado do Rio de Janeiro, 2014. Epidemiol Serv Saude 2018; 27(4):e2018058.
- 17 Cavalini LT, Ponce de Leon ACM. Correção de sub-registros de óbitos e proporção de internações por causas mal definidas. Rev Saude Publica 2007; 41(1):85-93.
- 18 Nascimento CF, Santos HG, Batista AFM, Lay AAR, Duarte YAO, Chiavegatto Filho ADP. Cause-specific mortality prediction in older residents of São Paulo, Brazil: a machine learning approach. Age Ageing 2021; 50(5):1692-1698.
- 19 Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP. SMOTE: Synthetic Minority Over-sampling Technique. J Artif Intell Res 2002; 16:321-357.
- 20 Murphy KP. Probabilistic Machine Learning: An introduction. 1ª ed. Cambridge: Press, MIT; 2020.
- 21 Lundberg SM, Lee SI. A unified approach to interpreting model predictions. In: Guyon I, Von Luxburg U, Bengio S, Wallach H, Fergus R, Vishwanathan S, Garnett R, editors. Advances in Neural Information Processing Systems 30 (NIPS 2017). Long Beach: NIPS 2017; 2017.
- 22 Aas K, Jullum M, Løland A. Explaining individual predictions when features are dependent: More accurate approximations to Shapley values. Artif Intell 2021; 298:103502.
- 23 James G, Witten D, Hastie T, Tibshirani R. Resampling Methods. In: An Introduction to Statistical Learning. New York: Springer; 2021. p. 197-223.
- 24 Hastie T, Tibshirani R, Friedman J. Model Assessment and Selection. In: The Elements of Statistical Learning. New York: Springer; 2009. p. 219-259.
- 25 Kuhn M, Johnson K. Over-Fitting and Model Tuning. In: Applied Predictive Modeling. 1ª ed. New York: Springer; 2013. p. 61-92.
- 26 Shapley LS. A Value for n-Person Games. In: Contributions to the Theory of Games (AM-28), Volume II. Princeton: Princeton University Press; 1953. p. 307-318.
- 27 Bianchi M, Buonanno P, Pinotti P. Do Immigrants Cause Crime? J Eur Econ Assoc 2012; 10(6):1318-1347.
- 28 Montes GC, Lins GO. Deterrence effects, socio-economic development, police revenge and homicides in Rio de Janeiro. Int J Soc Econ 2018; 45(10):1406-1423.
- 29 Cerqueira D, Soares RR. The Welfare Cost of Homicides in Brazil: Accounting for Heterogeneity in the Willingness to Pay for Mortality Reductions. Health Econ 2016; 25(3):259-276.
- 30 Garcia LP, Freitas LRS, Silva GDM, Höfelmann DA. Estimativas corrigidas de feminicídios no Brasil, 2009 a 2011. Rev Panam Salud Publica 2015; 37(4/5):251-257.
The data sources used in the research are indicated in the body of the article.




Source: Authors.


Source: MS/SVS/Data/SIM.
Source: MS/SVS/Data/SIM.