ABSTRACT
Objective: This study evaluated the use of synthetic populations, generated by conditional generative adversarial networks (cGAN), to estimate abusive alcohol consumption in small areas of Belo Horizonte (MG).
Methods: Data from the Vigitel survey from 2006 to 2018 were used. Sociodemographic and lifestyle variables were considered, and the synthetic records were assigned to nine vulnerability clusters.
Results: The inclusion of synthetic data reduced standard errors and narrowed confidence intervals, especially in more vulnerable areas. Higher prevalence was observed in less vulnerable areas.
Conclusions: The approach proved promising for overcoming representativeness limitations and enhancing risk factor monitoring, but it still depends on the quality of the original data, may introduce biases, and requires external validation and cautious use.
Keywords:
Abusive alcohol consumption; Generative adversarial networks; Synthetic population; Population surveys
RESUMO
Objetivos: Este estudo avaliou o uso de populações sintéticas, geradas por redes generativas adversariais condicionais (cGAN), para estimar o consumo abusivo de bebidas alcoólicas em pequenas áreas de Belo Horizonte (MG).
Métodos: Usou-se dados do Sistema de Vigilância de Fatores de Risco e Proteção para Doenças Crônicas por Inquérito Telefônico (Vigitel) de 2006 a 2018. Foram consideradas variáveis sociodemográficas e de estilo de vida, e os registros sintéticos foram atribuídos a nove clusters de vulnerabilidade.
Resultados: A inclusão de dados sintéticos reduziu erros padrão e estreitou intervalos de confiança, sobretudo em áreas mais vulneráveis. Observou-se maior prevalência em áreas menos vulneráveis.
Conclusões: A abordagem mostrou-se promissora para superar limitações de representatividade e ampliar o monitoramento de fatores de risco, mas ainda depende da qualidade dos dados originais, pode introduzir vieses e requer validação externa e uso cauteloso.
Palavras-chaves:
Consumo abusivo de bebidas alcóolicas; Redes generativas adversariais; População sintética; Inquéritos populacionais
INTRODUCTION
The analysis of risk factors in small areas is essential for understanding social inequalities in health and for supporting the planning of local public policies. Brazil, despite advances in surveillance, still faces difficulties in producing valid estimates in small areas because of the high costs and challenges in data collection, which make the implementation of representative surveys unfeasible. The Surveillance System of Risk and Protective Factors for Chronic Diseases by Telephone Survey (Vigitel), an annual telephone survey covering the state capitals, is one of the main sources of monitoring, but its dependence on landlines limits representativeness in more vulnerable groups, such as low-income populations and residents of peripheral areas1.
These limitations have stimulated the development of alternative methodologies to increase accuracy in small areas. The international literature suggests the use of approaches such as Bayesian models, multilevel models, and small area estimation techniques2.
More recently, innovations in artificial intelligence, such as conditional generative adversarial networks (cGANs), have been applied to generate realistic synthetic populations, enabling simulations that expand the analytical capacity in contexts where real data are scarce, confidential, or difficult to obtain3. This study aimed to evaluate the use of synthetic populations, generated by cGANs, to assess alcohol abuse in small areas of Belo Horizonte (MG), exploring its methodological potential and discussing its implications.
METHODS
This was an ecological study conducted with Vigitel data for Belo Horizonte (MG), between 2006 and 2018. The addresses of the interviewees, provided by the operators, were georeferenced and linked to the National Register of Addresses for Statistical Purposes, allowing association with census tracts2.
Abusive alcohol consumption was defined as consuming four or more drinks on one occasion for women and five or more for men4. Nine vulnerability clusters (LO-0, LO-1, MD-0, MD-1, HI-0, HI-1, HI-2, VH-0, VH-1) were considered “small areas,” derived from the four categories of the Health Vulnerability Index (HVI) (Low, Medium, High, and Very High). This regrouping strategy was proposed in a previous study5 to reduce the variability of HVI. Synthetic populations were generated by cGAN networks implemented in PyTorch, following the classic GAN architecture, where the generator and discriminator are jointly optimized in an adversarial scenario. Training occurred over 1,024 epochs, with a batch size of 500, a learning rate of 2×10⁻⁴, and the Adam optimizer (default parameters). The stopping criterion was based solely on the number of epochs, without early stopping, and the adopted loss function was the canonical adversarial loss (minimax log loss). The variables used were sociodemographic and lifestyle variables (e.g., sex, age, race/skin color, education level, and daily fruit consumption), selected for their epidemiological relevance and availability in the database. Between 5,000 and 100,000 synthetic records were generated per cluster, from which subsamples of 700 individuals were selected (defined on the basis of the stabilization of the standard error at ~1%) and incorporated, using supervised models such as XGBoost, into the Vigitel database in the high and very high vulnerability clusters, totaling 3,500 interviews per period (Table 1). The similarity between real and synthetic distributions was determined using the KSComplement index. The prevalence and 95% confidence interval (95%CI) were estimated by cluster and period (2006-2009; 2010-2013; 2014-2018), considering post-stratification weights. Differences between clusters within the same period were assessed by the non-overlapping of the 95%CIs.
The study was approved by the Research Ethics Committee of the Federal University of Minas Gerais (Opinion No. 6.538.883).
Data Availability Statement:
Data are available upon request from the corresponding author.
RESULTS
The prevalence of alcohol abuse was consistently higher in the less vulnerable clusters across all periods analyzed, both in the real data (T1 28.2%; 95%CI 25.1-31.3; T2 26.2%; 95%CI 23.0-29.4; T3 27.0%; 95%CI 22.9-31.0) and in the synthetic data (T1 27.4%; 95%CI 24.7-30.0; T2 27.7%; 95%CI 24.7-30.7; T3 26.5%; 95%CI 22.6-30.5) (Table 2).
The inclusion of synthetic populations in the database reduced standard errors and narrowed confidence intervals, especially in the most vulnerable areas, where the original samples were smaller. For example, in cluster VH-0 and period T1, the 95%CI ranged 7.3-24.4% (SE=4.3) with real data and 12.4-18.2% (SE=1.5) with synthetic data (Table 2). The KSComplement index of 0.96 confirmed high similarity between the observed and synthetic distributions.
DISCUSSION
The findings reinforce that the use of synthetic populations generated by cGAN can mitigate limitations of population-based surveys by producing more accurate estimates in small areas. The higher consumption of alcoholic beverages in less vulnerable areas is consistent with the national and international literature4,6 and can be explained by the fact that, in higher-income contexts, there is not only greater financial availability for the purchase of alcoholic beverages, but also a greater supply of commercial establishments and social spaces that encourage drinking. In addition, cultural and behavioral aspects, such as the valuing of social practices linked to alcohol consumption, contribute to reinforcing this pattern. On the other hand, more vulnerable areas may show a lower reported prevalence, but are more exposed to harms resulting from alcohol because of the lower availability of health services and social support.
The use of synthetic data contributed to greater stability of the estimates, reducing variability between periods and proving to be a promising approach to mitigate the scarcity of collected data and reduce distribution and representation biases in small geographical areas.
Despite the methodological advances, the technique should be considered exploratory. Its application depends on the quality of the real data, may introduce biases related to modeling, and requires external validation in different contexts7. Ethical issues emerge from the use of synthetic data, including the need for transparency and safeguards against inappropriate uses. When compared to traditional small-area approaches, such as Bayesian and multilevel models, cGANs show competitive and complementary performance2,7. However, its adoption requires additional studies that evaluate sensitivity, generalization, and replication in other populations.
Thus, it is concluded that synthetic populations represent an innovative resource for health surveillance, but their use must be accompanied by critical assessments, continuous validation, and ethical reflection. This innovation can support the formulation of more precise and equitable public policies, provided it is used responsibly.
ACKNOWLEDGMENTS:
Deborah Carvalho Malta thanks the National Council for Scientific and Technological Development (CNPQ) for the scholarship received (CNPQ 0-0202/771013). Crizian Saar Gomes thanks the São Paulo State Research Support Foundation (FAPESP) for the research scholarship (Process No. 2024/07524-0).
REFERENCES
-
1. Malta DC, Silva MMA, Moura L, Morais Neto OL. A implantação do Sistema de Vigilância de Doenças Crônicas Não Transmissíveis no Brasil, 2003 a 2015: alcances e desafios. Rev Bras Epidemiol 2017; 20(4): 661-75. https://doi.org/10.1590/1980-5497201700040009
» https://doi.org/https://doi.org/10.1590/1980-5497201700040009 -
2. Bernal RTI, Carvalho QH, Pell JP, Leyland AH, Dundas R, Barreto ML, et al. A methodology for small area prevalence estimation based on survey data. Int J Equity Health 2020; 19(200): 124. https://doi.org/10.1186/s12939-020-01220-5
» https://doi.org/https://doi.org/10.1186/s12939-020-01220-5 -
3. Murtaza H, Ahmed M, Khan NF, Murtaza G, Zafar S, Bano A. Synthetic data generation: state of the art in health care domain. Comput Sci Rev 2023; 48: 100546. https://doi.org/10.1016/j.cosrev.2023.100546
» https://doi.org/https://doi.org/10.1016/j.cosrev.2023.100546 - 4. World Health Organization. Global status report on alcohol and health 2018. Geneva: WHO; 2018.
-
5. Faria TMTR, Vasconcelos MA, Bernal RTI, Mielke GI, Souza JB, Gomes CS, et al. Improving the mapping of leisure-time physical activity inequities: the use of artificial intelligence to advance estimates of small-areas in Brazil. Public Health 2025; 243: 105727. https://doi.org/10.1016/j.puhe.2025.105727
» https://doi.org/https://doi.org/10.1016/j.puhe.2025.105727 -
6. Malta DC, Silva AG, Prates EJS, Alves FTA, Cristo EB, Machado IE. Convergência no consumo abusivo de álcool nas capitais brasileiras entre sexos, 2006 a 2019: o que dizem os inquéritos populacionais. Rev Bras Epidemiol 2021;24: e210022.SUPL1. https://doi.org/10.1590/1980-549720210022.supl.1
» https://doi.org/https://doi.org/10.1590/1980-549720210022.supl.1 -
7. Kang J, Kim Y, Imran MM, Jung GS, Kim YB. Generating population synthesis using a diffusion model. WSC ‘23: Proceedings of the Winter Simulation Conference. 2024;2944-55. https://dl.acm.org/doi/10.5555/3643142.3643387
» https://dl.acm.org/doi/10.5555/3643142.3643387
-
HOW TO CITE THIS ARTICLE:
Gomes CS, Bernal TRI, Araújo LF, França Júnior CR, Alves SN, Souza JB, et al. Use of synthetic population to assess alcohol abuse in small areas. Rev Bras Epidemiol. 2026; 29: e260013. https://doi.org/10.1590/1980-549720260013
-
FUNDING:
This study was funded by the following sources: Center for Innovation and Artificial Intelligence for Health (CI-IA Saúde), partly with funds from the São Paulo Resarch Foundation (FAPESP), Process No. 2020/09866-4, from the Minas Gerais Research Support Foundation (FAPEMIG), Process No. PPE-00030-21, and from UNIMED Belo Horizonte; CNPq and Decit/SECTICS/MS (call 21/2023); Fapemig call B01/2025 (APQ 02385-25); Health Ministry/National Health Fund (TED No. 114/2024 - “Artificial Intelligence Center for Health - NIAR-Saúde - UFMG”).
Edited by
-
ASSOCIATE EDITOR:
Vera Maria Vieira Paniz https://orcid.org/0000-0003-3186-9991
-
SCIENTIFIC EDITOR:
Francisco Chiaravalloti Neto https://orcid.org/0000-0003-2686-8740
