Open-access Calibration of SB Brasil 2023 examiners: use of technologies associated with the In-Lux method

Abstract

This methodological study presents the development and implementation of technological tools for the online training and calibration of dentists participating in the SB Brasil 2023 survey, using the Moodle® platform. The training and in-lux calibration process employed 10 and 25 sets of photographs, respectively. Conditions including dental crown status (dmft and DMFT indexes), caries consequences (pufa/PUFA index), malocclusion (canine relation, overjet, overbite, posterior crossbite, and Dental Aesthetic Index), and dental trauma were evaluated, based on WHO or SB Brasil 2010 criteria. Software was developed to record codes, automated calculations of agreement coefficients (overall percentage, score-specific, and simple and weighted kappa), and generate reports. Examiners were allowed to repeat calibration attempts until they achieved a minimum agreement (Kappa ≥ 0.61). In total, 1,513 examiners used the software, and 728 successfully completed the calibration with substantial or higher agreement across all conditions/indexes. Individual reports described discrepancies and monitored attempts. Occlusal conditions had the lowest percentages of almost perfect agreement and required more attempts to achieve calibration. The technological tools implemented in SB Brasil 2023 enabled online training and calibration, promoting consistency among examiners for fieldwork. These findings demonstrate the feasibility of remote strategies for calibration in epidemiological surveys, particularly in scenarios involving multiple geographically distributed examiners, with potential applications in other contexts and health areas.

Health Surveys; Observer Variation; Education, Distance; Calibration; Epidemiology

Introduction

Population-based oral health surveys estimate the prevalence of major oral health conditions, providing data to guide public policies and plan dental services effectively. These surveys are complex, costly, and require rigorous methodological standards to ensure unbiased and reliable results. A key challenge in such surveys is the involvement of multiple examiners, requiring standardized diagnostic criteria to ensure findings’ reliability, comparability, and validity.1,2

Since the 1990s, the World Health Organization (WHO) has recommended training and calibration protocols to standardize the interpretation and application of diagnostic criteria for obtaining dental index.2,3 The training familiarizes examiners with research protocols, diagnostic criteria, and assessment codes.4 At the same time, calibration compares their measurements to reference standards (validity) and evaluates consistency in and among examiners (intra- and inter-examiner reliability).3,4 The goal of calibration is to ensure uniform interpretation and application of diagnostic criteria, ensure examiners consistently adhere to a standard, and minimize variability among examiners.

Traditionally, calibration has been conducted through in vivo clinical examinations, which simulate actual data collection conditions. However, this method has logistical and ethical challenges,4-6 including participant recruitment, repeated examinations, addressing low prevalence conditions, and the extensive time and infrastructure required.4-7 Furthermore, in vivo calibration may not always be feasible due to resource limitations or specific population challenges, such as young children, individuals with disabilities, or older adults.4

The use of photographic records for calibration—known as the in-lux method—has been proposed to overcome these limitations. This approach uses high-quality photographs representing various dental conditions. First implemented in the 2010 Brazilian Oral Health Survey (SB Brasil 2010) to assess examiner reliability for dental trauma and fluorosis,8 the in-lux method has since been successfully applied in other surveys, including a regional survey in Minas Gerais9 and a schoolchildren oral health survey in Pelotas, Brazil,10 mainly to evaluate low-prevalence conditions. Also, previous studies have demonstrated the validity of photographs to evaluate dental caries,4,7,11 and plaque scoring system.12

A comparative analysis of the in vivo and in-lux methods showed that both can effectively identify examiners’ reliability to assess dental caries (DMFT index) and malocclusion (DAI), with the in-lux method offering greater feasibility in resource-constrained settings or when involving many examiners.7 The in-lux method provides several advantages: it enables the calibration of many examiners in a shorter timeframe, reduces logistical costs, minimizes ethical concerns for participants, and ensures consistent assessments across geographically dispersed examiners. Additionally, it facilitates shorter intervals between calibration and data collection and enhances reliability assessment for low-prevalence conditions.7,9

Reporting agreement metrics among examiners is essential for validating survey results. The WHO recommends global agreement percentages and simple kappa statistics for consistency evaluation. The weighted kappa statistic is preferred for ordinal measures, as it accounts for the degree of disagreement, penalizing more severe misclassifications.13 Traditionally, spreadsheet software calculates agreement coefficients manually during calibration exercises. This approach can be time-consuming and may contribute to fatigue among participants, especially when repeated attempts are necessary due to unsatisfactory results.

The integration of new technologies could streamline this process by enabling examiners to record results faster and calculate agreement coefficients more efficiently. Furthermore, these technologies could facilitate repeated calibration exercises with greater ease. During the COVID-19 pandemic, the 2020 Brazilian National Epidemiological Survey (here called SB Brasil 2023) was suspended, highlighting the need for innovative solutions to ensure continuity in research. In this context, alternative strategies were developed, adapting training and calibration processes to online formats mediated by digital technologies. Given prior evidence and experiences with the in-lux method, the SB Brasil 2023 survey adopted this approach, incorporating electronic technologies to optimize data collection, storage, and analysis.7,9 This article describes the development and implementation of calibration software and a data management system that supported the in-lux calibration method used in the SB Brasil 2023 survey.

Methods

Methodological study on the development and implementation of technological tools for the training and calibration of examiners in the SB Brasil 2023, coordinated by the Brazilian Ministry of Health, with data collection taking place in 2023 and 2024. The development and testing of the technological tool occurred from June 2020 to August 2021.

Training and calibration of the field team

The examiners were dentists working in Primary Health Care within the Unified Health System (SUS), who performed oral exams at participants’ homes in a probabilistic sample of participants at the index-ages of 5 and 12 years, and age groups 15–19, 35–44, and 65–74 years, as defined by the WHO, to evaluate oral health conditions among children, adolescents, adults, and elderly individuals, respectively. The exams were conducted following codes and criteria defined by WHO or those adopted in SB Brasil 2010, aiming to maintain the historical data series for oral health surveillance, as detailed in the technical project.14 This process was monitored by field supervisor researchers in each region of the country, defined as Local References.

The theoretical-practical training was conducted online, with synchronous and asynchronous activities, using a virtual learning environment built on the Moodle® platform. The theoretical training on the research methodology, codes, and criteria for assessing oral health conditions was conducted for a minimum of 16 hours. The codes and criteria for the assessed conditions were described and illustrated with photographs in video lessons, in addition to being detailed in the Examiner’s Manual.

The practical training involved one exercise and calibration. In these two activities, the in-lux method was adopted, which uses photographs of the mouth and teeth to replicate the situation the examiner would encounter in the field.7 Ten sets of photographs were used for the training and 25 for the calibration of the examiners, aiming to encompass the diversity of conditions encountered in fieldwork. The photographs were organized into separate PDF files. Each photograph or set of photographs was labeled with the term “volunteer,” numbered sequentially (e.g., volunteer 1, volunteer 2). A Nikon camera model D750, Macro Lens 100 mm, and Nikon circular and twin flashes were used to capture the photographs by a professional photographer who had completed dental photography courses totaling 200 hours of training. Most of the volunteers were patients undergoing treatment at the dental clinics of the Faculty of Dentistry at UFMG. However, due to the COVID-19 pandemic and the interruption of clinical services, photographs of children were also obtained in other locations, as they had not yet been collected in sufficient numbers. The photographs for training and calibration of trauma and clinical consequences of dental caries were obtained from the archives of professionals and professors from the dental trauma extension project at the Faculty of Dentistry at UFMG.

In the photographs, the examiners assessed crown caries in deciduous and permanent teeth for the evaluation of the dmft/DMFT (Decayed, Missing, and Filled Teeth) index, clinical consequences of untreated caries using the pufa/PUFA (pulpal involvement, ulceration, fistula, and dentoalveolar abscess) index, occlusal condition in deciduous dentition (canine relation, overbite, overjet, and posterior crossbite), conditions for obtaining the Dental Aesthetic Index (DAI) and dental trauma.

Each set of photographs included 12 images of the entire mouth, with frontal, occlusal, and lateral shots (in Maximum Habitual Intercuspation) to assess dental crown caries in all deciduous and permanent teeth and occlusal conditions. For the crown condition of each of the 20 deciduous teeth, a total of 200 records were made during training and 500 records during calibration. For the 28 permanent teeth, these numbers were 560 and 700 records, respectively. Some of these photographs were taken with a WHO-type periodontal probe positioned to measure the occlusal conditions, with overbite in the deciduous dentition (defined as increased when the overbite measurement in millimeters was > 2 mm), and the following conditions assessed for the DAI: incisal diastema (space, in millimeters, between the two upper central incisors), maxillary anterior misalignment (largest irregularity, in millimeters, between upper incisors), mandibular anterior misalignment (largest irregularity, in millimeters, between lower incisors), maxillary anterior overjet (distance, in millimeters, from the incisal edge of the most prominent upper incisor to the vestibular surface of the corresponding lower incisor), mandibular anterior overjet (measurement, in millimeters, of mandibular protrusion), and anterior vertical open bite (measurement, in millimeters, of the distance between the upper and lower incisal edges). For the assessment of clinical consequences of dental caries and dental trauma, individual photographs with marked teeth to be evaluated were used. For occlusal conditions, clinical consequences of dental caries and dental trauma, only one record was made per volunteer, meaning 10 records during training and 25 records during calibration.

Standard Examination

The sets of photographs were evaluated by a team of five researchers to define the codes for each assessed condition and volunteer/tooth, just as was done during the training and calibration by the examiners. The researchers were graduates in Dentistry, with prior experience in epidemiological surveys, and underwent theoretical training prior to evaluating the photographs using the same material provided to the examiners: the Examiner’s Manual and the video lessons about the oral health conditions addressed in the training and calibration. Each researcher independently assigned the codes, recording them in a form on Google Forms®. Meetings were held to review the codifications. In case of disagreements, the codes and criteria were discussed based on the material provided until a consensus was reached, assigning a final code to each tooth or volunteer. This consensus outcome was referred to as the standard examination, which was considered for determining the agreement among the examiners in the study. In the context of in-lux calibration, the reference standard served as the benchmark examiner, an experienced assessor assumed to be error-free or nearly so.15 Data from training and calibration exercises were used to evaluate the level of agreement between the examiners and the benchmark. This process assessed inter-examiner agreement, requiring all examiners to achieve reproducible results compared to the reference standard. The consensus outcome was recorded in the data management system as the “reference standard” and used for calculating agreement coefficients.

After the consensus, the results of the standard examination for calibration showed that the volunteers had an average dmft of 3.36, with 56% having dmft > 1. Thirteen children had at least one decayed tooth, with an average of 2.24 decayed teeth. Filled teeth were identified in 10 children, with an average of 1.08 per child. In the permanent dentition, the volunteers had at least one decayed tooth, with an average of 2.24 decayed teeth. Additionally, 80% had one or more filled teeth (average of 6.68), and 52% had missing teeth due to caries (average of 2.76). Regarding the pufa/PUFA index, 32% of the teeth showed no clinical consequences of caries, while 4% had dentoalveolar abscesses, 24% had fistulas, 32% had pulpal involvement, and 8% had ulcers. As for dental trauma, the distribution was: no trauma (8%), treated fracture (16%), enamel fracture (20%), enamel and dentin fracture (24%), fracture with pulpal involvement (12%), tooth loss due to trauma (8%), other damage (8%), and one record of an examination not performed due to the absence of incisors. In the DAI index, the distribution was: no malocclusion (44%), defined malocclusion (32%), severe malocclusion (12%), and very severe malocclusion (12%). In the deciduous dentition, based on the canine relation evaluation, 76% of the children were class I, 16% were class II, and 4% were class III. For four children, code 9 was assigned due to inability to evaluate. Regarding overbite, 36% had a normal relationship between the incisors, 8% had reduced bites, 28% had deep bites, and for 28% of the children it was not possible to evaluate. As for overjet, 40% had a normal relationship, 28% had an increased bite, and 92% did not have posterior crossbite.

Data recording, data management system recording, and report generation

For the recording of the codes and criteria for the conditions assessed by the examiners in the research, a software integrated into a data management system, named SC Brasil, was developed. This system enabled the research team to issue reports for monitoring the calibration process. The software development considered functionalities designed to meet the specific needs of SB Brasil 2023, as described in Table 1.

Table 1
Functionalities of the data recording software for training and calibration of examiners in SB Brasil 2023, and the data management and reporting system.

The data recording software was developed using JavaScript programming language, with the implementation of all the functionalities required by SB Brasil 2023. Encryption and user authentication mechanisms were incorporated to ensure the security and integrity of the data. After data collection, the information was exported to the data management system using an integrated export module that converted the data into a format compatible with statistical processing. The data management software was developed in PHP, due to its robustness and compatibility with web systems. It was designed to process the collected data, and algorithms were implemented for data validation, cleaning, and safe storage, as well as for the calculation of the agreement coefficients.

Obtaining the agreement coefficients

The agreement coefficients were calculated both in the training and calibration phases for each condition assessed in an automated manner, using algorithms implemented in the data management system.

The overall and score-specific agreement percentages were calculated for all conditions. For posterior crossbite, which allowed binary response (presence/absence), the simple Kappa was estimated. The agreement among examiners for conditions with ordinal scale responses was evaluated using the weighted Kappa coefficient. These conditions included crown conditions, clinical consequences of untreated dental caries, canine relation, overbite, overjet in the deciduous dentition, DAI classification, and dental trauma. Specifically, regarding crown condition, the weighted Kappa coefficient was estimated, considering all codes or only those included in the calculation of the dmft/DMFT index (carious, filled with caries, filled without caries, and missing). The DAI index evaluates malocclusion according to three dimensions of dentition, spacing and occlusion based on 11 occlusal characteristics. The scores of DAI were calculated by adding the item scores, which were multiplied by their coefficients (weights). A constant is then added to the summated score. Higher values indicated worse malocclusion and greater orthodontic treatment need. DAI scores were grouped into four categories: no abnormality or minor malocclusion (DAI ≤ 25); definite malocclusion (DAI = 26–30); severe malocclusion (DAI = 31–35) and very severe malocclusion (DAI ≥ 36) according to Jenny and Cons.16 This classification was used for analyzing the agreement coefficients among the examiners.

The global agreement or agreement proportion observed, which is also known as the crude or raw agreement, is the simplest method for summarizing an agreement for categorical variables. It reflects the percentage of the total number of units inspected where there is agreement between the examiner and the standard examination. The agreement percentage was calculated by dividing the total number of concordant records by the total number of records.2 Score-specific agreement is a complement to the global agreement evaluation and is obtained by the ratio between the total number of agreements for a given score (e.g., code A for the crown condition) and the total number of examinations evaluated with that score by both examiners.

The Cohen’s kappa statistic (κ) is used to measure agreement of binary values. It is a relative measure that determines the excess of observed agreement to chance agreement. A negative value is assumed if there is complete disagreement; kappa is zero if there is no more agreement that can be expected due to chance, and one if there is perfect agreement.2,3,17 Weighted kappa is recommended for determining examiner reliability for ordinal data.13 This statistic incorporates the factor of agreement by chance alone and also weights proportional agreement. The weighted Kappa was calculated by applying linear weights, according to Cicchetti and Allison.18 The formulas used to calculate the weighted Kappa are described in Fleiss, Levin, and Paik.19

The agreement coefficients were obtained by creating cross tables between the examiner’s records and the standard examination, with row i of the table corresponding to the examiner’s records and column j to the standard examination. The formulas implemented in the data management system to obtain the weighted Kappa are presented in Table 2.

Table 2
2x2 table and formulas for calculating agreement coefficients between examiners and the standard examination

The interpretation of the Kappa coefficient was performed according to Landis and Koch, with an examiner being considered fit for fieldwork if they achieved substantial agreement, i.e., they presented a Kappa coefficient equal to or greater than 0.61 for all the assessed conditions.20

Software functionality tests

The first testing phase consisted of evaluating the software’s compatibility with the Mobile Data Collection Device (DMC) provided by the Brazilian Institute of Geography and Statistics for fieldwork, as well as checking the data entry fields, screen layout, registration process, restricted access, and data export. For this, data entry simulations were performed by five researchers, and reports were created on the necessary adjustments.

Subsequently, typing tests were conducted and data spreadsheet were generated, as well as data export to the management system, and validation of the method for calculating the agreement coefficients and demonstration in the reports. For this, data from five volunteers were entered, and the calculations were performed based on the standard examination for all the considered conditions. To validate the calculations, all kappa estimates, standard errors, and confidence intervals were performed simultaneously by a researcher in Microsoft Excel® spreadsheets and Stata® version 17 (StataCorp LLC) using the command “kap condition examiner condition_standard, wgt(w) tab”. The checks were carried out until identical values for the kappa coefficient and confidence intervals were obtained.

In-lux calibration result

The calibration results were presented descriptively, considering the total number of examiners with records for all the assessed conditions. The minimum and maximum values for the kappa coefficient found were obtained, as well as the number of attempts required to achieve a kappa coefficient ≥ 0.61. Additionally, the distribution of examiners with substantial or excellent agreement was presented, along with the number of examiners who achieved substantial agreement in a single attempt, 2 to 4 attempts, and 5 or more attempts.

Results

Data recording software, demonstration of calculations, and reports

The software enabled the recording of codes for each condition and volunteer evaluated in pre-coded fields or in whole number entries, in the case of the DAI, through restricted access for the examiner. A dental chart was implemented for recording the crown conditions, including caries, as well as specific fields to record other conditions in the volunteers. Figure 1 illustrates the fields for recording the crown conditions of each deciduous tooth, with pre-registered options. Each set of photographs per volunteer generated 20 records, as shown in the example. The figure also illustrates the field for recording dental trauma observed in the evaluated tooth. In this case, each photograph per volunteer generated a single record (Figure 1). Examiners could complete their records in up to 10 attempts, both in the practical training and calibration phases. If the examiner, using their login and password, entered data more than once, each new entry was automatically recorded as a new attempt.

Figure 1
Illustrative images of the fields for recording the codes for the condition of the crown of each deciduous tooth (dmft) and dental trauma.

The data management system generated reports for monitoring calibration by the Local References. The individual report from each examiner was issued separately for each index, presenting the cross-table between the records made by the examiner and the results of the standard examination. In the example shown in Figure 2, the individual examiner report displays the calibration results for evaluating the crown condition of a deciduous tooth (a total of 500 records). The blue diagonal shows the number of records in which there was agreement between the examiner's result and the standard examination records for the codes that constitute the dmft index. The results outside the diagonal indicate disagreement. The horizontal rows represent the examiner’s records, while the vertical rows show the results of the standard examination. For this examiner, 401 teeth were considered healthy, with 393 records agreeing with the standard examination. There were eight discordant records, of which four teeth, considered healthy by the examiner, were deemed decayed according to the standard examination. The overall agreement was 93.40%. The lowest agreement was observed for the condition “filled, but with caries” (26.31%). For this condition, of the 18 records made by the examiner, five agreed with the standard examiner. The report presents the values for the observed weighted proportion of agreement (Pow) and the chance-expected weighted proportion of agreement (Pew). The Kw was 0.8700, considering all conditions in the crown evaluation. This value was calculated based on Pow (0.9825) and Pew (0.8654), with the calculations shown in Table 3. The same calculations were repeated considering the components of the dmft index, resulting in a Kw of 0.8341.

Figure 2
Individual report demonstrating the agreement between the examiner’s records and the standard examination for the deciduous tooth crown conditions, with the results of the overall agreement coefficients, score-specific agreement, weighted Kappa considering all categories, and weighted Kappa considering only the components of the dmft index.

Table 3
Demonstration of obtaining the Kw coefficient for agreement between the examiner and the standard examination for the condition of the deciduous tooth crown.

In addition to the reports separated by evaluated condition, the system generated consolidated individual reports, showing the coefficients of agreement generated in all calibration attempts made by the same examiner for all evaluated conditions/indices in a single document. In Figure 3, for example, the Kw for clinical consequences of dental caries was 0.739 in the 2nd attempt. In the 1st attempt, the value was 0.584 (Figure 3). For the DAI, Kappa ≥ 0.61 was obtained in the 4th attempt, with the value highlighted in green.

Figure 3
Model of an individual consolidated report with results from calibration attempts by the same examiner for all assessed conditions.

Calibration Results

A total of 1,513 examiners registered in the system, of which 728 (48.12%) completed calibration for all conditions and indices, with K or Kw values ≥ 0.61. The remaining registered examiners either did not have records for all indices or did not achieve substantial or higher agreement. The minimum Kappa values were 0.61 for almost all conditions, except for Posterior Crossbite, where the lowest Kappa value was 0.638. Kappa = 1 was also the maximum observed value for almost all conditions, except for the condition of the crown of deciduous teeth (dmft index). Higher percentages of examiners with almost-perfect agreement were observed for the crown condition in permanent dentition, untreated dental caries clinical consequence, and dental trauma. The highest percentage of examiners with Kappa values between 0.61 and 0.80 was observed for the DAI (Table 4). Regarding the number of attempts, occlusion conditions of deciduous and permanent dentition were those that most frequently required 2 to 4 attempts to achieve Kappa ≥ 0.61 (Figure 4).

Table 4
Minimum and maximum values of the Kappa coefficients obtained by the examiners for each assessed condition and the distribution of examiners according to substantial or almost-perfect agreement levels (n = 728 examiners).

Figure 4
Distribution of examiners according to the number of calibration attempts until achieving substantial agreement with the standard examination (kappa ≥ 0.61).

Discussion

This methodological study described the development and implementation of technological tools for the training and calibration of examiners for SB Brasil 2023. The use of the in-lux calibration method and a virtual learning environment (Moodle®) enabled remote training, aiming for uniformity in the application of criteria and codes for assessing oral conditions among examiners from different regions of the country. This approach represented an advance and may be particularly useful in large-scale surveys involving multiple examiners in different locations, as in this national survey.

The adoption of the in-lux method proved to be suitable for calibrating a large number of examiners, meeting the need to carry out the process remotely. The observed levels of agreement contributed to demonstrating the feasibility of the process, which was based on the use of a standard examination obtained through consensus and demonstrated consistency among examiners. Consensus was the strategy employed to define the standard examination, as it minimizes errors among examiners.21 The selection of a standard examiner is one of the main objectives of calibration, aiming to measure how far each examiner is from the standard, which is assumed to be the true value.3 When a standard examiner is not established, it is possible for all values observed by different examiners to be close to each other (high Kappa values) but distant from the presumed true value (that of the standard examiner). Still, in this in-lux calibration process, the number of photographs was defined based on WHO recommendations, which suggest that the examiner should assess at least 25 subjects to measure variations among examiners (inter-examiner reproducibility) after practicing the examination on a group of 10 subjects.2

The selection of volunteers for obtaining photographs was planned to anticipate the conditions that would be encountered in the field, according to the criteria of Pine, Pitts and Nugent,22 which is crucial in assessing consistency among examiners.1 This approach aims to represent the conditions to be evaluated during field research. Another important aspect is the prevalence of the condition being evaluated, which can influence the Kappa coefficient,15,23 the method chosen for this study as it is one of the most commonly used and robust methods to assess the reliability of measurements24. In this study, the average DMFT and dmft values were high, and more than 50% of the volunteers had all conditions of the index. For dental trauma and clinical consequence of dental caries, all conditions were represented in the photographs. However, for occlusal conditions, there was an absence or low prevalence of certain conditions in the primary dentition (open bite, edge-to-edge bite, anterior crossbite, and posterior crossbite). For these conditions, the most frequent was normal occlusion. Therefore, the higher number of sound volunteers (fewer diagnostic errors) compared to volunteers with occlusal alterations (more diagnostic errors) may dilute the errors attributed to occlusal alterations in the primary dentition, leading to a positive perception of the results achieved in the examiners’ calibration. However, occlusal conditions in the primary dentition and DAI required the highest number of attempts to achieve substantial agreement. This result may also reveal the difficulties in assessing these conditions using photographs. Thus, the decision to use the same set of photographs for the evaluation of crowns in both primary and permanent teeth, as well as for occlusal conditions in the primary dentition, complicated the process due to the low occurrence of certain conditions.25 Future studies should delve deeper into understanding the challenges associated with in-lux calibration using photographs specifically selected to better represent occlusal conditions in the primary dentition. Strategies that consider the use of three-dimensional images or more advanced technologies could be explored to overcome the observed limitations and enhance consistency in the evaluation of these conditions.22

According to WHO, where there are major discrepancies, the examiners should be re-examined so that inter-examiner differences can be reviewed and resolved by group discussion.3 If the variability is large, the examiner should review the interpretation of the criteria and conduct additional examinations until acceptable consistency is achieved. Otherwise, the examiner should not proceed with data collection. This aspect of calibration was implemented through the demonstration of individual reports, which were used by research teams to discuss calibration results with the examiners, reinforcing diagnostic criteria in situations with more disagreements compared to the standard examination. Thus, for some examiners it was necessary to repeat the calibration exercise two or more times. The calibration software allowed for multiple attempts to be recorded, allowing for a situation similar to “repetition” of examinations.

The functionalities developed in the software met the technical and logistical requirements of the survey, and the automation of calculations facilitated the monitoring of the calibration process by the research team. Some challenges related to the configuration of mobile data collection devices (DMCs) and internet access for data export were encountered. The desktop version aimed to overcome the app’s limitation of compatibility with only Android® operating systems. The data recorded throughout the process also contributed to generating detailed information about the calibration process, considering the number of registered examiners, the number of examiners with incomplete records, examiners who did not achieve the minimum agreement required for fieldwork, agreement coefficient values, among others. This study focused on presenting the Kappa index, as it was employed to determine whether or not the examiner could continue fieldwork. Future publications should analyze overall agreement indices and score-specific agreement, along with their relationship with the Kappa index, aiming to further detail the main discrepancies.

By integrating photographic records and specialized software, this innovative calibration process offers a scalable solution for diverse and challenging contexts, ensuring methodological consistency and technological efficiency. It is believed that this study contributes to strengthening oral health surveillance in Brazil by promoting standardized methodologies for data collection. It is expected that the experiences shared in this study may support future epidemiological surveys, with enhancements to address the challenges encountered.

A limitation of the calibration process in SB Brasil 2023 was the absence of duplicate evaluations to assess intra-examiner consistency over time. Without these evaluations, it was not possible to verify whether examiners maintained consistent application of diagnostic criteria throughout the data collection period. This step was not carried out as it proved to be unfeasible during data collection, given the difficulty of obtaining subjects’ authorization to conduct a second examination. Additionally, for occlusal conditions, the high prevalence of volunteers without alterations may have diluted diagnostic errors, creating an optimistic perception of inter-examiner agreement for these conditions. Despite the limitations, the technological tools implemented in SB Brasil 2023 enabled online training and calibration, allowing for the attainment of examiner consistency for fieldwork. These results highlight the feasibility of remote strategies for calibration in epidemiological surveys, particularly in scenarios involving multiple geographically distributed examiners, with potential applications in other contexts and health fields.

Acknowledgments

The National Oral Health Survey was conducted with funds from the Ministry of Health. RCF receives a research productivity grant from Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq - Process: 310938/2022-8)

References

  • 1 Assaf AV, Tagliaferro EP, Meneghim MC, Tengan C, Pereira AC, Ambrosano GM, et al. A new approach for interexaminer reliability data analysis on dental caries calibration. J Appl Oral Sci. 2007 Dec;15(6):480-5. https://doi.org/10.1590/S1678-77572007000600005
    » https://doi.org/10.1590/S1678-77572007000600005
  • 2 World Health Organization. Oral health surveys: basic methods. 5th ed. Geneva: World Health Organization; 2013.
  • 3 World Health Organization. Calibration of examiners for oral health epidemiology surveys. Geneva: World Health Organization; 1993.
  • 4 Christian B, Amezdroz E, Calache H, Gussy M, Sore R, Waters E. Examiner calibration in caries detection for populations and settings where in vivo calibration is not practical. Community Dent Health. 2017 Dec;34(4):248-53. https://doi.org/10.1922/CDH_4102Christian06
    » https://doi.org/10.1922/CDH_4102Christian06
  • 5 Andrade FR, Narvai PC, Montagner MA. The ethics of in vivo calibrations in oral health surveys. Rev Bras Epidemiol. 2016;19(4):812-21. Portuguese. https://doi.org/10.1590/1980-5497201600040011
    » https://doi.org/10.1590/1980-5497201600040011
  • 6 Martins AM, Silveira MF, Freitas CV, Eleutério NB, Oliveira PH, Ferreira RC. Desafios de um exercício de calibração para estudo epidemiológico envolvendo variáveis quantitativas e categóricas ordinais. Arq Odontol. 2011;47(4):196-207. https://doi.org/10.7308/aodontol/2012.48.1.03
    » https://doi.org/10.7308/aodontol/2012.48.1.03
  • 7 Pinto RD, Vettore MV, Abreu MH, Palmier AC, Moura RN, Roncalli AG. Reliability analysis using the in-lux examination method for dental indices in adolescents for use in epidemiological studies. Community Dent Oral Epidemiol. 2023 Oct;51(5):847-53. https://doi.org/10.1111/cdoe.12775
    » https://doi.org/10.1111/cdoe.12775
  • 8 Ministério da Saúde (BR). Secretaria de Vigilância em Saúde. Secretaria de Atenção à Saúde. Departamento de Atenção Básica. Coordenação Nacional de Saúde Bucal. SB Brasil 2010: Pesquisa Nacional de Saúde Bucal: resultados principais. Brasília, DF: Ministério da Saúde; 2011.
  • 9 Pinto RS, Lopes DL, Santos JS, Roncalli AG, Projeto SB. Minas Gerais 2012: pesquisa das condições de saúde bucal da população mineira: métodos e resultados principais. Arq Odonto. 2018;54 e14: https://doi.org/10.7308/aodontol/2018.54.e14
    » https://doi.org/10.7308/aodontol/2018.54.e14
  • 10 Goettems ML, Correa MB, Vargas-Ferreira F, Torriani DD, Marques M, Domingues MR, et al. Methods and logistics of a multidisciplinary survey of schoolchildren from Pelotas, in the Southern Region of Brazil. Cad Saude Publica. 2013 May;29(5):867-78. https://doi.org/10.1590/S0102-311X2013000500005
    » https://doi.org/10.1590/S0102-311X2013000500005
  • 11 Felsch M, Meyer O, Schlickenrieder A, Engels P, Schönewolf J, Zöllner F, et al. Detection and localization of caries and hypomineralization on dental photographs with a vision transformer model. NPJ Digit Med. 2023 Oct;6(1):198. https://doi.org/10.1038/s41746-023-00944-2
    » https://doi.org/10.1038/s41746-023-00944-2
  • 12 Avenetti DM, Martin MA, Gansky SA, Ramos-Gomez FJ, Hyde S, Van Horn R, et al. Calibration and reliability testing of a novel asynchronous photographic plaque scoring system in young children. J Public Health Dent. 2023 Mar;83(1):108-15. https://doi.org/10.1111/jphd.12557
    » https://doi.org/10.1111/jphd.12557
  • 13 Cohen J. Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit. Psychol Bull. 1968 Oct;70(4):213-20. https://doi.org/10.1037/h0026256
    » https://doi.org/10.1037/h0026256
  • 14 Ministério da Saúde (BR). SB Brasil 2020: Pesquisa Nacional de Saúde Bucal: projeto técnico. Brasília, DF: Ministério da Saúde.; 2022.
  • 15 Agbaje JO, Mutsvari T, Lesaffre E, Declerck D. Measurement, analysis and interpretation of examiner reliability in caries experience surveys: some methodological thoughts. Clin Oral Investig. 2012 Feb;16(1):117-27. https://doi.org/10.1007/s00784-010-0475-x
    » https://doi.org/10.1007/s00784-010-0475-x
  • 16 Jenny J, Cons NC. Establishing malocclusion severity levels on the Dental Aesthetic Index (DAI) scale. Aust Dent J. 1996 Feb;41(1):43-6. https://doi.org/10.1111/j.1834-7819.1996.tb05654.x
    » https://doi.org/10.1111/j.1834-7819.1996.tb05654.x
  • 17 Cohen J. A coefficient of agreement for nominal scales. Educ Psychol Meas. 1960;20(1):10. https://doi.org/10.1177/001316446002000104
    » https://doi.org/10.1177/001316446002000104
  • 18 Cicchetti DV, Allison T. A new procedure for assessing reliability of scoring EEG sleep recordings. Am J EEG Technol. 1971;11(3):101-10. https://doi.org/10.1080/00029238.1971.11080840
    » https://doi.org/10.1080/00029238.1971.11080840
  • 19 Fleiss JL, Levin B, Paik MC. Statistical methods for rates and proportions. 3rd ed. New York: Wiley; 2003. https://doi.org/10.1002/0471445428
    » https://doi.org/10.1002/0471445428
  • 20 Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977 Mar;33(1):159-74. https://doi.org/10.2307/2529310
    » https://doi.org/10.2307/2529310
  • 21 Frias AC, Antunes JL, Narvai PC. Precisão e validade de levantamentos epidemiológicos em saúde bucal: cárie dentária na Cidade de São Paulo, 2002. Rev Bras Epidemiol. 2004;7(2):10. https://doi.org/10.1590/S1415-790X2004000200004
    » https://doi.org/10.1590/S1415-790X2004000200004
  • 22 Pine CM, Pitts NB, Nugent ZJ. British Association for the Study of Community Dentistry (BASCD) guidance on the statistical aspects of training and calibration of examiners for surveys of child dental health: a BASCD coordinated dental epidemiology programme quality standard. Community Dent Health. 1997 Mar;14 Suppl 1:18-29.
  • 23 Shankar V, Bangdiwala SI. Observer agreement paradoxes in 2x2 tables: comparison of agreement measures. BMC Med Res Methodol. 2014 Aug;14(1):100. https://doi.org/10.1186/1471-2288-14-100
    » https://doi.org/10.1186/1471-2288-14-100
  • 24 McHugh ML. Interrater reliability: the kappa statistic. Biochem Med (Zagreb). 2012;22(3):276-82. https://doi.org/10.11613/BM.2012.031
    » https://doi.org/10.11613/BM.2012.031
  • 25 Agbaje JO, Mutsvari T, Lesaffre E, Declerck D. Examiner performance in calibration exercises compared with field conditions when scoring caries experience. Clin Oral Investig. 2012 Apr;16(2):481-8. https://doi.org/10.1007/s00784-011-0523-1
    » https://doi.org/10.1007/s00784-011-0523-1

Publication Dates

  • Publication in this collection
    19 May 2025
  • Date of issue
    Apr 2025

History

  • Received
    27 Jan 2025
  • Accepted
    29 Jan 2025
  • Reviewed
    3 Feb 2025
location_on
Sociedade Brasileira de Pesquisa Odontológica - SBPqO Av. Prof. Lineu Prestes, 2227, 05508-000 São Paulo SP - Brazil, Tel. (55 11) 3044-2393/(55 11) 9-7557-1244 - São Paulo - SP - Brazil
E-mail: office.bor@ingroup.srv.br
rss_feed Acompanhe os números deste periódico no seu leitor de RSS
Ir para o topo Reportar erro