Open-access Prediction of healthcare costs on consumer direct health plan in the Brazilian context

Resumo

The rise in healthcare costs has led to the adoption of cost-sharing devices in health plans. This article explores this discussion by simulating Health Savings Accounts (HSAs) to cover medical and hospital expenses, supported by catastrophic insurance. Simulating 10 million lives, we evaluate the utilization of catastrophic insurance and the balances of HSAs at the end of working life. To estimate annual expenditures, a Markov Chain approach - distinct from the usual ones - was used based on recent past expenditures, age range, and gender. The results suggest that HSAs do not create inequalities, offering a viable method to sustain private healthcare financing for the elderly.

JEL classification. I13, C02, C15, C53.

1. Introduction

Consumer directed health plans (CDHP) are products intended to fund healthcare expenses, which have been consolidated in the United States since the end of the nineties as a possible approach to contain the rise in healthcare costs. In CDHP, the insured has a personal account to pay for his or her medical procedures, what leads to the increase in awareness about healthcare costs (Gabel et al. (2002) and Bundorf (2016)).

According to Bundorf (2016), there are three features often associated to CDHP: relatively high deductibles, a personal account to accumulate resources, and availability of information about healthcare costs. On the one hand, these features diverge from models based purely on mutualism, as they seek to accomplish a higher perception of well-being and economy efficiency, while mitigating the disruptive impacts due to moral hazard. On the other hand, the attachment of individual expenses to a personal account raises questions about justice, as individuals could persistently experience health shocks during their work life, making it unlikely for them to save enough to cover their expenses after retirement. This would cause, from a population point of view, a great variation on the balance of these accounts, causing a part of the population to have a great amount of resources at retirement (savings), while other part would not have accumulated anything (the product would have features of a self-insurance) (Eichner, 1998).

In recent years, a great rise in healthcare costs has been observed around the world. In the Brazilian supplementary health system, for example, the mean annual healthcare expense, including dental benefits, increased 73% between 2013 and 2018 according to the National Supplementary Health Agency (ANS, 2019-2021a), what has led to a public debate about the sustainability of the sector (ANS, 2019-2021b). Following the example of other countries, it has been proposed an alternative to the system in an attempt to establish norms and legal certainty for products with co-payment and deductible devices, as it is a reasonable assumption that Brazil faces similar challenges as other countries in what concerns the essential cause of healthcare costs increase, that is, spending more than needed to treat a specific health problem. This over-consumption and oversupply pattern is understood as the result of two factors. On the one hand, consumers may partake in more risky activities when protected by a health plan, increasing the probability of needing more care (ex-ante moral hazard) or simply consuming more care than necessary (ex-post moral hazard). On the other hand, providers may supply more care than needed, inducing the demand (Ehrlich and Becker (1972) and Zweifel and Manning (2000)).

In order to mitigate the impact of these factors, in 2018 the Brazilian regulatory agency (ANS) proposed a resolution to regulate co-payment and deductibles in health plans (RN nº 433, June 27 2018), which had a negative public repercussion, that led ANS to suspend the norms, leaving the market uncovered in what concerns the understanding of these devices. The negative repercussion was caused by a misunderstanding about the real use of these devices in health plans, what evidences the need to further the discussion about them, including the possibility of considering them as part of a new market segment.

This paper aims to subsidize the discussion with the analysis of empirical data by simulating health savings accounts (HSA). A product derived from CDHP, the HSA are individual accounts set to exclusively cover expenses with healthcare goods and services, which is composed by a resources accumulation phase, typically subsidized by the employer with participation of the employee, followed by a decumulation phase, generally after retirement. The objective of this form of funding is the implementation of a risk pooling over the individual cycle of life, mobilizing resources from various sources to fund the decumulation phase. Resource mobilization to a savings account for the stages of the life cycle when the increase in healthcare expenses exceeds income is a response to policies aimed at maintaining private funding for healthcare for the elderly population.

Since the end of the nineties, countries as the United States, China, Singapore and South Africa have taken steps to incorporate these alternatives in their health systems. In the US experience, the HSA products combine a high deductible health plan with tax breaks to form a savings account. These products have a mixed framework, which seeks to reduce the individual exposition to the risk of extreme events, by combining two devices: a catastrophic health insurance associated with a personal savings account.

In this paper, we simulate a HSA product combined with a catastrophic health insurance during a labor period of forty years (from 25 to 65 years-old) in the Brazilian context. The main objective of the simulation is to study, from a population point of view, the variation of HSA balances at the end of the accumulation phase. A great discrepancy in these balances may be interpreted as the result of persistent health shocks, understood here as the cause of the financial need for healthcare goods and services, and measured through the costs covered by a health plan under the individual perspective. The persistence of shocks is a fundamental aspect in the debate about CDHP since one could argue that individuals with poor health would not be able to accumulate savings overtime, so these accounts would reflect inequalities between individuals with distinct overall health conditions.

In order to simulate the annual healthcare expenses of an individual, we propose an approach based on Markov Chains (Taylor and Karlin, 1998), which differs from the usual methods of predicting healthcare expenses based on regression models (Jones, 2000). In this approach, we define levels of annual expense, which are ranges of expense, and apply a prediction technique which estimates the probability of an individual to have annual expenses in each level, based on his or her levels in the previous two years, age range and sex. This model does not try to predict the exact value of the annual expenses by assuming some kind of functional relation between the expenses and the independent variables, but rather predicts the probability of each level of expense for each combination of categories of the independent variables sex, age range and previous expense levels.

Freed from assumptions about regression error distributions and the restricted form of the functional relation between dependent and independent variables, which are often not satisfied in real datasets, a Markov Chain approach makes a mild assumption about the dependence between the expenses in a year with that in previous years, and tries to achieve a more realistic goal, which is to predict an expense level rather than the exact value of the expense. This more realistic approach may lead to improved prediction models, as it can be applied in the context of any country and naturally considers their specific behaviors.

In Section 2 we present the dataset used in the simulation, and discuss the considered CDHP product, the modelling of healthcare expenses and the proposed Markov Chain approach. In Section 3 we present the results of the simulation studies, and in Section 4 we comment on how they may be interpreted to subsidize the debate about CDHP products in Brazil.

2. Materials and methods

2.1 Dataset

The dataset contains claims of a Brazilian self-management health plan. There are six types of health plans operators in Brazil, which are characterized by the juridical nature of their operations. The self-management health plans are similar to the self-insured group health plans in the US, which offer plans for employees and dependents, assuming the financial risk, which can be partially funded by the employees. This portfolio was followed longitudinally for five years (2005 to 2009) and all claims (expenses) are corrected by the inflation index IPCA (Extended National Consumer Price Index) to the December 2009 value.

As the objective of this paper is to study the persistence of health shocks during the work life, we will consider only working age individuals, that is, those with 25 years completed at January 2005 and with at most 65 years completed at December 2009, which amounts to approximately 39,000 lives. During the five years of study, the portfolio size changed because of the entrance of new individuals, death and other motives, so from all these lives, around 11,000 were not followed during all the period, so there will be considered in the analysis only the 27,780 individuals between 25 and 65 years-old which were in the portfolio for the whole period.

Table 1 presents some descriptive statistics of the annual expenses of the individuals within the considered age range which stayed in the portfolio during the five-year period. The percentage of individuals with zero annual expenses is between 5 and 6% in all years. As Table 1 portrays the behavior over time of a closed cohort, it is expected that a shift in the mean expense occurs due to the aging of the group. However, even though the variation in the mean expense from 2005 to 2009 was around 36% and the expenses are corrected by a general inflation index, there is no guarantee that the real variation due to aging can still be observed in the costs of the healthcare goods and services of this portfolio. Therefore, in reality, both the aging effect and the real variation in the costs are reflected in these values.

Table 1
Descriptive statistics of the annual expenses of individuals between 25 and 65 years-old which stayed on the portfolio during the whole five-year period. The percentiles, mean and standard deviation are calculated considering only the individuals with positive expenses.

As expected (Seshamani and Gray (2004), Zweifel et al. (2004) and Werblow et al. (2007)), large expenses are concentrated in a small parcel of individuals: in 2008, lesser than 1% of the individuals were responsible for expenses ranging from R$ 31,158 to R$ 1,044,525 where we observe that the maximum expense is 385 times the mean one. Also, in 2005, around 5% of the individuals were responsible for annual expenses ranging from R$ 7,083 to R$ 426,772, and the maximum expense was 206 times the mean one.

In orderto compare men and women according to their annual expenses, we present in Tables A.1 and A.2 in Appendix A some descriptive statistics for the annual expenses for each sex. From the total of individuals, there are 13,539 (49%) women and 14,241 (51%) men. Over time, the percentage of women without expenses ranged between 5.1 and 5.5%, while the percentage of men ranged from 5.8 to 6.8%. When considering only the individuals with positive expense, we see that the mean annual expenses of women is greater than that of men, as it varied between R$ 2,341 and R$ 3,062 for the female sex, and between R$ 1,795 and R$ 2,576 for the male sex. In the same manner, the median annual expenses of women ranged from R$ 998 to R$ 1,104 and that of men from R$ 580 to R$ 721.

In Tables A.3, A.4, A.5, A.6, A.7, A.8, A.9, A.10, A.11, A.12, A.13, A.14, A.15, A.16, A.17, A.18, A.19, A.20 in Appendix A we present the descriptive statistics of the annual expenses by sex and age range, where it can be seen in all years that the mean annual expenses increase with age, highlighting the effect of age range on healthcare expenses. When we assess the factors which influence individual healthcare expenses, age is always presumed to have a positive effect since, as age increases, so does the probability of occurrence of chronicle diseases and loss of functional capabilities. Therefore, it is expected that high expenses are related to advanced age (Duncan et al. (2016) and Frees et al. (2014)).

In Figure A.1 in Appendix A we present the mean annual expenses for each combination of sex and age range, with an error bar representing one standard error. On top of each error bar, we present the size of each group. When we compare men and women, we see a distinct pattern, as women tend to spend more in the early age ranges, a tendency which inverts itself in later ranges. This feature is found in other references in the literature (Yamamoto, 2013). Over time, we observe a greater increment in the expenses of the later age ranges.

While the mean annual expenses of men increase constantly with age, the mean annual expenses of women is stable until the age of 50, when we see an increase in the mean of the age ranges 51-55 and 56-60, followed by a small reduction in age range 61-65. Thus, we see that women spend more than men in practically all age ranges, although this difference decreases with age. In 2006 and 2008, the mean annual expenses of men in the age range 61-65 surpassed that of women.

2.2 Persistence of costs

Persistence of costs is defined as the continuity (permanence) of elevated health costs of an individual over time. In order to analyze the persistence of costs, we divide the individual annual expenses of each year i ∈ {2005,2006,2007,2008,2009} into four levels, namely:

  • F1,i: annual expenses lesser or equal to R$ 300 at year i;

  • F2,i: annual expenses greater than R$ 300 and lesser or equal to R$ 1,000 at year i;

  • F3,i: annual expenses greater than R$ 1,000 and lesser or equal to R$ 5,000 at year i; and

  • F4,i: annual expenses greater than R$ 5,000 at year i.

Dividing the expenses this way we have, for each individual, a sequence of five levels describing his or her expenses throughout the years. For example, an individual with costs {0;200;500; 1,500;200} has sequence F1,2005, F1,2006, F2,2007, F3,2008, F1,2009 as expense levels.

In Figure 1 we present the sample proportion of individuals which transitioned between each pair (Fk,i, Fl,j), j > 1 of states, in which rows represent the origin state Fk,i and columns the destination state Fl,j. The color of the matrices entries is related to the estimated transition probability: red means great probability (around 0.5), gray means medium probability (around 0.2) and white means low probability (around zero). In this figure, there are ten 4 by 4 matrices, one for each pair of distinct years, which present the transition probabilities from the ranges of the early year to that of the later. For example, the matrix in the lower-left corner presents the transition probabilities between the expense’s levels of 2005 and 2006, and the matrix in the upper between the expense’s levels of 2008 and 2009. The leading diagonal of each matrix is the proportion of individuals which stayed at the same expense level at both years.

Figure 1
Transition matrix of expenses for each two years combination.

In Figure 1 we see, for example, that within individuals with expenses greater than R$ 5,000 in 2008, 46% had expenses between R$ 1,000 and R$ 5,000, and 27% maintained the expenses at level greater than R$ 5,000, in 2009; within individuals with annual expenses lesser or equal to R$ 300 in 2007, 52% had expenses lesser than R$ 300, and only 3% had expenses greater than R$ 5,000, in 2008. We also note that high values of probabilities are concentrated around the leading diagonals, except at the greatest expense level. However, these high values decrease with the distance between the years, highlighting the fact, also observed in Eichner et al. (1996), that persistence of costs decays over time. It is important to note that the matrices were constructed considering only the individuals between 25 and 65 years-old that were alive in 2009. If the individuals which died in the considered period were also considered, then there would be greater probabilities in the diagonal point of the greatest expense level, as it is known that individuals tend to have consistently high expenses in the period leading up to their death.

In order to evaluate the effect of age on the persistence of costs, we present in Fig. A.2 in Appendix A the transition matrices calculated considering only the younger (21 to 40 years-old) and older (41 to 65 years-old) individuals, respectively. When comparing the diagonal of the matrices of both figures, we see that, for younger individuals, the probability of those in expense level 3 remaining in the same level or moving up to expense level 4 is approximately 56%. However, for the population aged 40 and above, this proportion is around 66%, i.e., 10 percentage points higher, in average.

2.3 Health Savings Accounts (HSA)

The health savings accounts (HSA) are devices in which a savings account attached to each individual receives annual contributions with the exclusive purpose of covering healthcare expenses. The annual expenses covered by HSA are limited by a predefined value while its balance is positive. In case expenses exceed the limit or the account balance, then a catastrophic insurance is activated to cover them. Although there is a limitation on the expenses covered by funds from the savings account, in this device there is no limitation on individual annual healthcare expenses, as the insurance covers any expenses beyond the threshold.

The CDHP product considered in this paper is a HSA in which each individual has an account, started at 25 years-old, from which are deducted his or her healthcare expenses. The account dynamic is as follows:

  • At the beginning of each year, the employer deposits R$ 2,500 on the individual account.

  • In case the annual healthcare expenses of an individual do not exceed the account balance or R$ 5,000, they are fully paid by funds from the account. If the annual expenses surpass R$ 5,000 or the balance of the account, then R$ 5,000, or the balance, is withdrawn to partially cover them, and the remaining value is covered by a catastrophic insurance. Therefore, all healthcare expenses up to R$ 5,000 or the account balance, if lesser, are covered by the individual, and the remaining value is covered by the insurance.

In what follows, we present an approach based on Markov Chain to simulate the product described above, which seeks to predict the expense level of each individual at each year of his or her work life in order to assess the balance of the accounts at age 65, and the frequency and severity of catastrophic insurance use, from a population point of view. In the next section, we discuss the approaches to healthcare expenses prediction at the individual level used in the literature, presenting their qualities and shortcomings.

2.4 Prediction of future healthcare expenses at individual level

Modelling techniques to assess the risk associated with events which incur in healthcare expenses are relatively recent (Jones, 2000), for their development is dependent on the volume and quality of available information, which have only improved in the last couple of decades. Also, the sector has been demanding new frameworks for covering and management of health risks, in contrast to its historical foundation, based purely on refunding, increasing the need for improved quantitative models.

In order to model healthcare expenses patterns in the US, Frees et al. (2011) considered more than thirty factors related to an individual’s demography, socioeconomic level, health condition, employment and availability of health insurance. There were considered two dependent variables, namely, frequency and severity of expenses, and the significant factors found for each of them differed. Also, Duncan et al. (2016) explored the effects on healthcare expenses of approximately one hundred independent variables related to demography, age and comorbidities history, among others.

Pope et al. (2004) and Duncan et al. (2016) pointed out the limitations of models which consider only demographic factors and health condition history, affirming the importance of taking into account past expenses while developing prediction models. Indeed, according to Duncan et al. (2016), one of the best sources of information for predicting healthcare expenses are past expenses.

About issues encountered when modelling individual healthcare expenses, Duncan et al. (2016) evidenced that: (a) annual expenses follow approximately a log-normal distribution and are characterized by high variance; (b) the error distribution is clearly heteroscedastic: the conditional variance for individuals with high expenses is greater than that of individuals with low expenses; (c) there exists interaction between the independent variables, and the relationship between expenses and health condition is expected to be non-linear, what implies non-linear relationships between dependent and independent variables; (d) many of the independent variables are highly correlated and some of them have very rare categories. Due to the great variability of annual expenses in a population and the great number of independent variables, the quality of fitted models is quite low.

Other problems encountered when modelling annual healthcare expenses are the excess of zeros and the presence of extreme values. The first problem is due to individuals which do not use their health insurance in the period of a year. The second is caused by the fact that a very small portion of a population has annual expenses hundreds of times the mean one. In order to attack these problems, Marcondes et al. (2018) proposed a hurdle model for heavy-tailed data, which is specifically useful for annual healthcare expenses data. A hurdle model takes special care of the zero values, while a heavy-tail model takes into account values far from the mean. The mixture of models treating these two features is a better fit for annual healthcare expenses than the usual regression models, as evidenced by the application in Marcondes et al. (2018), which considered a dataset of annual healthcare expenses.

Based on the literature, we see that there is a need to develop better techniques to model healthcare expenses in a dynamic scenario characterized by the rapid increase in the amount of available information and data about individuals. With this purpose, we consider a Markov Chain approach to model healthcare expenses in order to simulate an HSA product.

2.5 Methodology for prediction of future healthcare expenses at individual level

In this section, we present the methodology proposed in this paper to predict future healthcare expenses at the individual level, which is based on Markov Chains of order 2 that are defined in Appendix B. The prediction technique will be employed to estimate annual individual healthcare expenses in order to simulate the expenses which will be covered by a HSA during the labor life of simulated individuals.

2.5.1 Simulation of individual annual healthcare expenses over time

Our simulation study is based on the assumption that the pattern of annual healthcare expenses of an individual is similar to that of those of the same sex, age range and expenses history. The study consists in the simulation of 10,000 lives, starting at 25 years-old, which are followed annually until 65 years-old, what corresponds to the forty-one years of a work life. The annual expenses of these individuals in their first year (25 years-old) are sampled from the 2009 annual expenses empirical distribution of age range 25-30. There are 1,683 individuals in the age range 25-30 in 2009, so in order to obtain 10,000 lives from them, we will sample 10,000 lives with repetition from the 1,683 individuals, setting the age of all new sampled individuals to exactly 25 years-old.

From the sex, age range, expenses history and expenses in the first year of the 10,000 lives, we simulate their personal HSA over time, a process fulfilled by iterating two consecutive procedures. In the first step, we predict to which expense level each individual will go, based on probabilities depending on his or her sex, age range at the present year and expense level on the current and previous year. On the second step, we sample the exact value to be deducted from each individual account to cover the expenses from the empirical distribution of the annual expenses of the individuals of the same sex, age range and predicted expense level. This empirical distribution is the one in Fig. B.1 which is related to the individual’s age range and sex. The distribution is built by combining individual’s medical expenses from years 2005 to 2009, stratified by age and sex, obtaining the possible points that can be sampled within the expense level sampled in the first step, i.e., between the respective dotted lines in Fig. B.1. The simulation was replicated i,000 times. In Appendix B, we further explain the simulation dynamics.

3. Results

In order to assess the effectiveness of the HSA, we present in this section the results of the simulation study regarding the balance of the individual accounts over time, and the frequency and severity of the catastrophic insurance use.

3.1 Balance of the individual health savings accounts

In Table 2 we present descriptive statistics of the individual HSA of the simulated 10,000 lives at every five years between 25 and 65 years-old.

Table 2
Descriptive statistics of the simulated individual mean HSA balance at every five years between 25 and 65 years-old. We present the mean (µ) and standard deviation (σ) over the 1,000 simulations.

According to the simulation, at 30 years-old 26% of the lives had a zero balance in their HSA and, at 40 years-old, this percentage was lesser than 3%, in average. As the annual health expenses are smaller at young age, the mean increase in the individual HSA balance is greater in early ages, and in later age ranges this balance is more spread around the mean. At 40 years-old, only 5% of the lives had a balance greater than R$ 30,000 in average and, at 45 years-old, between 50% and 75% of the lives had this kind of balance. At 55 years-old only 15% of the lives had a balance lesser than R$ 20,000, while 15% had a balance greater than R$ 50,000 in average and, at 60 years-old, only 25% of the lives had a balance lesser than R$ 30,000 in average and 50% had a balance greater than R$39,000 in average.

At 65 years-old, the end of the work life, only 1% of the lives had zero balance and 5% a balance lesser than R$11,000 in average, while half the lives had a balance greater than R$41,000 in average, with 25% with a balance greater than R$53,000 in average. The mean balance at 65 years-old is around R$40,000 in average, with a maximum balance around R$89,000 in average. In Fig. 2 we present the empirical distribution of the individual HSA balance in the 10,000,000 simulated individuals at 65 years-old. This distribution is approximately symmetrical (Skewness coefficient = -0.1202) and 01 outlier was detected (R$93,142.83). According to the simulation, the HSA is a good alternative to fund healthcare costs at old age, as a great parcel of the simulated population had a considerable balance at retirement. Note that the simulation is based on the expenses of a portfolio in which there is no individual HSA, and that we considered a savings account without interest. Therefore, the balance of the individual HSA could in reality be greater than that simulated.

Figure 2
Empirical distribution of the simulated HSA balances at 65 years-old of the 10,000,000 simulated individuals.

3.2 Frequency and severity of the catastrophic insurance use

In this section we analyze the use of the catastrophic insurance, more specifically its frequency and severity. In Table 3 we present the percentage (frequency), and the total value covered (severity), of the lives which used the catastrophic insurance each number of times.

Table 3.
Mean Frequency and severity of the catastrophic insurance use by the simulated lives during the 40 years period. We present the mean (µ) and standard deviation (σ) over the 1,000 simulations.

In average, only 5.6% of the simulated lives did not use the insurance at the 40 years period, approximately 50% used at most three times, and 80% used at most six times. For 0.5% of the lives, it was necessary to use the insurance 13 times or more, and for 13 lives the insurance was triggered between 20 and 30 times among the 40 possibles. Furthermore, we note that around 50% of the expenses covered by the catastrophic insurance were with lives that used it at most five times, which corresponds to 72% of the lives, while 10% of the expenses covered were with lives that used it 10 times or more, which corresponds to 3.5% of lives in average. Also, around 0.5% of the expenses covered by the insurance were with only 11 lives in average , which used it at least 21 times.

From a population point of view, the simulation of 1,000 repetitions of the portfolio during 40-years each, revealed a mean healthcare expense of R$1,082,350,725 with standard deviation of R$7,688,268, from which R$489,367,594 (45.2%) with standard deviation of R$6,862,835 was covered by the catastrophic insurance (in average), and the remaining R$592,983,131 (54.8%) with standard deviation of R$1,784,781 was covered by the individual HSA (in average).

In Figure 3 we present the distributions of percentage of coverage for healthcare expenses, either by HSA or catastrophic insurance.

Figure 3
Distribution of percentage of coverage for healthcare expenses from HSA and catastrophic insurance in the 10,000,000 simulated individuals.

3.3 Savings account balance, coverage by HSA and coverage by catastrophic insurance

The proposed CDHP product has three main features: the HSA balance at 65 years-old, the expenses covered by the HSA and the expenses covered by catastrophic insurance. The first two features are more susceptible to the effects of a HSA on an individual behavior, as the decision to consume healthcare goods and services may be inhibited by this CDHP product. Table 4 displays descriptive statistics of these three features.

Table 4
Descriptive statistics of the HSA balance at 65 years-old, the expenses covered by the HSA and the expenses covered by catastrophic insurance over the work life. The percentiles, mean and standard deviation are calculated considering only the values greater than zero (n=10,000). We also present the mean (µ) and standard deviation (a) over the 1,000 simulations.

If a life did not have any expenses in the 40 years period, its account balance would be R$ 100,000, although the simulated mean and median balance at 65 years-old were between R$41,000 and R$41,535, in average with s.d. of R$248, and the maximum was R$ 89,251 in average, with s.d. of R$ 1,355. Therefore, the mean account balance at retirement is in average 42% the value which has been deposited. The mean expense covered during the work life by the insurance was around R$51,800, in average with s.d. of R$713, and the maximum was around R$1,176,000, in average with s.d. of R$ 119,167, which is 22 times the mean value, evidencing the presence of outliers. For 25% of the lives, the insurance covered lesser than R$ 12,000, and for half the lives it covered at most R$29,700 in average with s.d. of R$459. On the other hand, the 5% of lives with the greatest expenses covered by the insurance had more than R$ 175,000 covered in average. When considering the percentage of expenses covered by the insurance, half the lives had at most 40% of them covered by the insurance in average, while 25% of the lives had almost half of their expenses covered by it in average, and 5% of the lives had more than 75% of their expenses covered by the insurance in average.

Figures 4 and 5 display the empirical distribution of the severity of catastrophic insurance use, and of its logarithm transformation, where it can be noted the presence of outliers, evidencing the coverage of great healthcare expenses by the insurance. Coverages under R$1.00 were ignored in the logarithmic transformation.

Figure 4
Severity of catastrophic insurance use in the 10,000,000 simulated individuals.

Figure 5
Logarithm transform of severity of catastrophic insurance usage in the 10,000,000 simulated individuals.

In Figure 6 we present the dispersion plot between the total expenses at the 40 year period of each life and the percentage of these expenses which were covered by catastrophic insurance. The color of the points refers to the number of times in which the insurance was used overthe work life.

Figure 6
Dispersion plot between the total expenses at the 40 year period of each life in the 10,000,000 simulated individuals and the percentage of these expenses which were covered by catastrophic insurance. The color refers to the number of times in which the insurance was used over the work life.

4. Discussion and Conclusions

An increase in healthcare costs has been observed recently in several countries, especially in Brazil, where it has led to a discussion about co-payment and deductible devices for health plans, which culminated in a resolution by the regulatory agency (ANS) to incorporate these devices to the Brazilian supplementary health system. A negative public repercussion based on misunderstandings followed the resolution, forcing ANS to revoke it, leaving the market uncovered regarding the understanding of these devices. In this context, this paper aimed to subsidize the discussion about co-payment and deductibles by simulating a consumer directed health plan product which has never been commercialized in Brazil, the health savings accounts, which combines a high deductible health plan with tax breaks to form a savings account aiming to reduce the individual exposition to the risk of extremal events by combining a catastrophic insurance with an individual savings account.

This mixed product benefits both the employer and employee. From the employee point of view, the accredited network of healthcare providers is superior than the standard, and the individual is not limited to the accredited network. Moreover, the balance may be invested to increase income, but due to the purpose of this work, this issue of investment has not been considered. From the employer point of view, it is expected lower individual healthcare expenses, since the employees, when paying for their own healthcare goods and services, tend to be more selective in choosing their providers and more thrifty when using the available services.

In order to carry out our simulation, we had to first establish a prediction technique for annual healthcare expenses. To address this challenge, we proposed an approach to healthcare expenses prediction based on Markov Chains, which seeks to predict the annual expenses of an individual based on his or her sex, age range and expenses in the last two years. In this approach, first the expense level is estimated and then, the exact expense is chosen from an empirical distribution.

Although classical in Statistics and Probability Theory, and applied in health economics to predict health conditions and costs related to it (Komorowski and Raffa (2016), Sato and Zouain (2010), Jeong et al. (2014), Graves et al. (2016), Bala and Mauskopf (2006), Dobi and Zempléni (2019) and Garg et al. (2010)), it seems that a Markov Chain approach may also be useful to predict annual healthcare expenses, as it is known that expenses history is an important factor in predicting healthcare expenses (Duncan et al., 2016). A Markov Chain approach based on previous expenses, sex, and age range may be more viable for pricing health plans than a method based on health condition due to ethical issues.

Based on a dataset containing annual expenses of a portfolio of a Brazilian self-management health plan, we simulated the proposed HSA product for 10,000 lives from 25 to 65 years-old in order to study the balance of the individual savings accounts over time and at retirement, and the frequency and severity of catastrophic insurance use. In this simulation, we took into account the persistence of costs phenomenon, widely known to be part of individual healthcare expenses. The results evidenced a low prevalence of persistence of costs at the individual level, as at 65 years-old the balance of the individual savings accounts is symmetrically spread around the average of R$ 44,000, supporting that the adoption of this model is not a mechanism of inequalities generation. During the 40 years period in the 10,000,000 simulated individuals, the total healthcare expenses of the simulated portfolio was R$1,082,350,725 in average with s.d. of R$7,688,268, of which R$489,367,594 in average (45.2%) with s.d. R$6,862,835 was covered by catastrophic insurance. From the individual perspective, half the lives had 40% of their expenses covered by catastrophic insurance and only 5% of them had more than 75% of their expenses covered by catastrophic insurance. At retirement, the mean balance of the individual accounts was 45.2% of the value deposited over the work life, so the mean expense covered by the HSA during the work life was 54.8% that deposited, in average.

Therefore, we may conclude that a mixed CDHP product combining a HSA and a catastrophic insurance may be viable in Brazil from a population point of view as it does not create inequalities caused by persistence of costs. In order to implement such a product, one needs to adapt the value deposited by the employer every year and the limit covered by the HSA to suit the health system costs.

We leave some topics for future research. An interesting topic would be to compare a Markov Chain approach to healthcare expenses prediction with the usual methods based on regression. Also, the development of a method which incorporates the best features of both approaches seems promising. About the proposed CDHP, there are some hyperparameters which can be optimized in order to obtain a more realistic simulation: the break points of the expense levels, the value deposited by the employer every year and the limit of expenses covered by the HSA. In order to optimize these quantities, it would be needed data with better quality and greater quantity than that of this paper. Finally, to further the discussion in Brazil about co-payment and deductible devices, it would be interesting to study other kinds of CDHP products in the Brazilian context.

Appendix A Descriptive Statistics and Transition Matrix

A.1 Descriptive Statistics of Annual Expenses, by sex

In this section, it is presented descriptive statistics by sex of the annual expenses individuals by sex, between 25 and 65 years-old which stayed on the portfolio during the whole five years period. The percentiles, mean and standard deviation are calculated considering only the individuals with positive expenses.

Table A.1
Descriptive Statistics Annual Expenses (Female)
Table A.2
Descriptive Statistics Annual Expenses (Male)
A.2 Descriptive statistics of annual expenses, by sex and age category

In this section, it is presented descriptive statistics by age range of the annual expenses of individuals between 25 and 65 years-old which stayed on the portfolio during the whole five years period. The percentiles, mean and standard deviation are calculated considering only the individuals with positive expenses.

Table A.3
Descriptive Statistics Annual Expenses (Female) 21a24
Table A.4
Descriptive Statistics Annual Expenses (Male) 21a24
Table A.5
Descriptive Statistics Annual Expenses (Female) 25a30
Table A.6
Descriptive Statistics Annual Expenses (Male) 25a30
Table A.7
Descriptive Statistics Annual Expenses (Female) 31a35
Table A.8
Descriptive Statistics Annual Expenses (Male) 31a35
Table A.9
Descriptive Statistics Annual Expenses (Female) 36a40
Table A.10
Descriptive Statistics Annual Expenses (Male) 36a40
Table A.11
Descriptive Statistics Annual Expenses (Female) 41a45
Table A.12
Descriptive Statistics Annual Expenses (Male) 41a45
Table A.13
Descriptive Statistics Annual Expenses (Female) 46a50
Table A.14
Descriptive Statistics Annual Expenses (Male) 46a50
Table A.15
Descriptive Statistics Annual Expenses (Female) 51a55
Table A.16
Descriptive Statistics Annual Expenses (Male) 51a55
Table A.17
Descriptive Statistics Annual Expenses (Female) 56a60
Table A.18
Descriptive Statistics Annual Expenses (Male) 56a60
Table A.19
Descriptive Statistics Annual Expenses (Female) 61a65
Table A.20
Descriptive Statistics Annual Expenses (Male) 61a65

Figure A.1.
Profile plot of the mean annual expenses by year, for each sex and age range considering only individuals between 25 and 65 years-old which stayed on the portfolio during the whole five years period.

Figure A.2
Transition matrices of expenses for each two years combination, i.e. from 2005 to 2009, considering individuals with ages between (a) 21 to 40 years-old, and (b) 41 to 65 years-old.

Appendix B Markov Chains and HSA simulation

The Markov Chain approach proposed in this paper aims to predict the expense level of an individual in a given year by estimating the probability of his or her annual expenses being at every level given the expense level of previous years. Formally, considering the definition of Fj,i presented in Section 2.5.1, the probability of an individual with annual expenses in Fk,i·at year i having annual expenses in Fl,j at year j is estimated as:

(1) P ^ ( F k , i , F l , j ) = Number of individuals in  F k , i which are also in  F l , j Number of individuals in  F k , i , with j > i .

In Markov Chain Theory, P^(Fk,i,Fl,j) is an estimator of the transition probabilities between states Fk,i· and Fl,j. Indeed, the ratio above is the proportion of individuals which were in Fk,i· at year i and transitioned to Fl,j at year j. Therefore, if we were to estimate the probability of an individual in Fk,i· to be in Fl,j at year j we could use the proportion P^(Fk,i,Fl,j), which is a consistent estimator for such probability (Anderson and Goodman, 1957).

The approach above may also be applied to estimate the transition probabilities for individuals of a given sex and/or age range. This probability would allow to study the persistence of costs in distinct groups of individuals, as it is believed to behaviour differently for each age range and sex. The estimation of these probabilities would be done in the same manner as above, but the number of individuals considered in the ratio of P^(Fk,i,Fl,j) would be that within the group of interest. Proceeding this way, we have a prediction for the probability of an individual, with given sex, age range and expense level in a previous year, to be in each expense level in the current year.

A transition matrix characterizes a Markov Chain, that is a sequence of random variables in which the distribution of the current random variable depends only on the value of the previous k and is given by the probabilities of the transition matrix, in which k is called the order of the Markov Chain. If k = 1 then the transition matrix is calculated as the ratio given in (1) when j = i+1. If k = 2, then the probability of being in a given state after visiting some in the past depends only on the last two states visited. In this case, the Markov Chain is generated by the transition probabilities estimated as which refers to the transition to Fl,i after being in Fk,i−2, Fm,i−1 at the previous two years. These probabilities may also be estimated for a given group of individuals, considering the numbers in the ratio to be that within the group.

(2) P ^ ( F k , i 2 , F m , i 1 ; F l , i ) = Number of individuals in  F k , i 2 which are also in  F m , i 1 and F l , i Number of individuals in  F k , i 2 which are also in F m , i 1

With a Markov Chain of order 2 we may estimate the probability of an individual with given sex and age range to be in each expense level as a function of the levels he or she was in the last two years. This is the method we use to predict the expense level of each life in the simulated portfolio. Observe that it diverges from the usual methods based on regression as we do not try to predict the exact value of the annual expense, but rather the expense level, so we do not have to assume a distribution for the expenses, nor a functional relation between the current year expenses and the independent variables sex, age range and previous two years expenses.

Transition Matrices: In order to perform the first step of the simulation, we need to estimate a transition matrix between the expense levels. For this purpose, we suppose that the expense level follows a homogeneous Markov Chain of order 2 (Taylor and Karlin, 1998). This means the transition from one state to another does not depend on the year, but only on the states. Formally, this means that

(3) P ( F k , i 2 , F m , i 1 ; F l , i ) = P ( F k , j 2 , F m , j 1 ; F l , j )

for any i, j, in which P is the population transition probability, in contrast to the estimated transition probability . With this assumption, we may estimate the transition probabilities by that of the time period 2007 to 2009. Therefore, we estimate the transition from expense levels Fk, Fm to Fi as

(4) P ^ ( F k , F m ; F l ) = P ^ ( F k , 2007 , F m , 2008 ; F l , 2009 ) ,

in which we consider only the individuals within the respective sex and age range group in the ratios that define these transitions. This assumption is supported by the matrices in Figure 1, where we see that those which compare years with the same distance are similar (these are the matrices in a same diagonal of Figure 1).

Under this approach, we have 16 transition matrices, one for each combination of sex and age range (26-30, 31-35, 36-40, 41-45, 46-50, 51-55, 56-60, 61-65). Each matrix has 16 rows, one for each pair of expense levels in the current and last year, and 4 columns, one for each possible expense level in the next year. The entries of these matrices are transition probabilities from the state given by the pair to each one of the expense levels. To estimate the transitions, we consider that the age range of each individual in the dataset is the one which he or she was part of for the most number of months in the triennial 2007-2009.

First step The expense level of an individual next year is predicted based on his or her sex, age range and expense level last year and in the current year. This prediction is sampled from the estimated conditional distribution of the expense levels of given sex, age range and two last expense levels, which is the row associated to the last two expense levels of the transition matrix of the given sex and age range. We start this process at the initial values for 24 and 25 years-old to estimate the level at 26 years-old. We then iterate this process to estimate the level for the following ages: use the estimated values for 25 and 26 to predict 27 years-old; that of 26 and 27 to predict 28 years-old, and so on until the age of 65. Therefore, for each one of the 10,000 lives, we have a sequence of 41 levels corresponding to its predicted expense level for each work life year. Observe that, from age 27 on, the expenses history is predicted in the previous two years.

Second step After we simulate the expense level of a life in a given year, we need to simulate the value to be withdrawn from its account to cover a healthcare expense in such level. This is done by sampling a point from the empirical distribution of the aggregated annual expenses of all years which contain only the individuals with the same sex and age range of the life. Within this empirical distribution, we sample one of the points inside the expense level predicted on step one, i.e., inside the respective dotted lines in Figure B.1, which shows the aggregated empirical distributions of the logarithm of the annual healthcare expenses between 2005 and 2009 for each sex and age range.

In Table B.1 we see the number of expenses in the dataset at each expense level by sex and age range. These percentages aggregate all expenses of the period, so they refer to the percentage of expenses, rather than the percentage of individuals, as each individual has five expenses, one for each year. In the dataset, 25% of women’s expenses and 45% of men’s in the age range 25-30 are in the first level (up to R$ 300), percentages which decrease with age, attaining the values of 15% and 23% for women’s and men’s, respectively, at the age range 61-65. As the age increases, the same reduction is observed in the percentage of expenses in the second expense level (R$ 300 - R$ 1,000) for both sexes. On the other hand, the percentage ofexpenses in the third expense level (R$ 1,000 - R$ 5,000) increases with age and the difference between these percentages in age ranges 25-30 and 61-65 is 14 and 17 percentage points for women’s and men’s, respectively. The same is observed for the fourth expense level (greater than R$ 5,000): 10% of women’s expenses and 4% of men’s in the age range 25-30 are greater than R$ 5,000 in opposition to 16% of women’s and 12% of men’s in the age range 61-65. We see in the last three age ranges (after 51 years-old) that around 58% to 62% of women’s expenses, and 40% to 48% of men’s, are greater than R$ 1,000.

Table B.1.
Percentage of expenses in each expense level by sex and age range. The expenses of all years are aggregated, so these percentages refer to the percentage of expenses, rather than the percentage of individuals, as each individual has five expenses, one for each year.

Figure B.1
Empirical distributions of the logarithm of the annual healthcare expenses between 2005 and 2009. The distributions aggregate points from the five years. The dotted lines are respectively log(300), log(1000) and log(5000), which are the break points of the expense levels Fi,j.

Bibliography

  • Anderson, T. W. and L. A. Goodman (1957): “Statistical Inference about Markov Chains”, Ann. Math. Statist, 1 (28), 89–110. [34]
  • ANS (2019-2021a): “Dados Gerais - ANS,” Website, Agência Nacional De Saúde Suplementar, acessed on October 9th, 2019. [2]
  • ANS (2019-2021b): “Mapa Estratégico 2019-2021 - ANS,” Website, Agência Nacional De Saúde Suplementar, acessed on October 9th, 2019. [2]
  • Bala, M. V. and J. A. Mauskopf (2006): “Optimal assignment of treatments to health states using a Markov decision model”, PharmacoEconomics, 4 (24), 345–354. [17]
  • Bundorf, M. K. (2016): “Consumer-directed health plans: A review of the evidence”, The Journal Risk and Insurance, 83 (1), 9–41. [1, 2]
  • Dobi, B. and A. Zempléni (2019): “Markov chain-based cost-optimal control charts for healthcare data,” Preprint 1903.06675, arXiv, acessed on October 9th, 2019. [17]
  • Duncan, I., M. Loginov, and M. Ludkovski (2016): “Testing alternative regression frameworks for predictive modeling of health care costs”, North American Actuarial Journal, 1 (20), 65–87. [5, 8, 17]
  • Ehrlich, I. and G. S. Becker (1972): “Market insurance”, American Economic Review, 4 (80), 623–648. [2]
  • Eichner, M., M. McClellan, and D. Wise (1996): “Insurance or self-insurance?: Variation, persistence, and individual health accounts,” Working Paper Series 5640, National Bureau of Economic Research. [7]
  • Eichner, M. J. (1998): “The demand for medical care: What people pay does matter”, American Economic Review, 2 (88), 117–121. [2]
  • Frees, E. W, R. A Derrig, and G. Meyers (2014): Predictive modeling applications in actuarial science, vol. 1, Cambridge University Press. [5]
  • Frees, E. W, J. Gao, and M. A. Rosenberg (2011): “Predicting the frequency and amount of health care expenditures”, North American Actuarial Journal, 3 (15), 377–392. [8]
  • Gabel, J. R., A. T. Lo Sasso, and T. Rice (2002): “Consumer-driven health plans: Are they more than talk now?” Health Affairs, 21 ((Suppl1)), W395–W407. [1]
  • Garg, L., S. McClean, B. Meenan, and P. Millard (2010): “A non-homogeneous discrete time Markov model for admission scheduling and resource planning in a cost or capacity constrained healthcare system”, Health Care Manag Sci, 2 (13), 155–169. [17]
  • Graves, N, C Wloch, J Wilson, A Barnett, A Sutton, N Cooper, K Merollini, V McCreanor, Q Cheng, E Burn, T Lamagni, and A. Charlett (2016): “A cost-effectiveness modelling study of strategies to reduce risk of infection following primary hip replacement based on a systematic review”, Health Technology Assessment, 54 (20), 1–144. [17]
  • Jeong, S., C-H. Youn, and Y-W. Kim (2014): “Predicted cost model for integrated healthcare systems using Markov process,” in Ubiquitous Information Technologies and Applications, Springer, Berlin, Heidelberg, vol. 280 of Lecture Notes in Electrical Engineering, 181–187. [17]
  • Jones, A. (2000): “Health Econometrics”, Handbook of Health Economics, 265–344. [3, 8]
  • Komorowski, M. and J. Raffa (2016): “Markov Models and Cost Effectiveness Analysis: Applications in Medical Research”, Secondary Analysis of Electronic Health Records, 351–367. [17]
  • Marcondes, D., C. P. Peixoto, and A. C. Maia (2018): “A survey of a hurdle model for heavy-tailed data based on the generalized lambda distribution”, Communications in Statistics - Theory and Methods. [9]
  • Pope, G. C, J. Kautter, R. P Ellis, A. S Ash, J. Z Ayanian, L. I Lezzoni, M. J Ingber, J. M Levy, and J. Robst (2004): “Risk adjustment of Medicare capitation payments using the CMS-HCC model”, Health Care Financing Review, 4 (25), 119–141. [8]
  • Sato, R. C. and D. M. Zouain (2010): “Markov models in health care”, Einstein (São Paulo), 3 (8), 376–379. [17]
  • Seshamani, M. and A. Gray (2004): “Time to death and health expenditure: An improved model for the impact of demographic change on health care costs”, Age and Ageing, 6 (33), 556–561. [4]
  • Taylor, H. M and S. Karlin (1998): An Introduction to Stochastic Modeling, Academic Press Limited., 3 ed. [3, 35]
  • Werblow, A., S. Felder, and P. Zweifel (2007): “Population ageing and health care expenditure: A school of “red herrings”?” Health Economics, 10 (16), 1109–1126. [4]
  • Yamamoto, D. H. (2013): “Health care costs - from birth to death,” Health Care Cost Institute’s Independent Report Series 2013-1, Society of Actuaries, acessed on October 1st, 2019. [5]
  • Zweifel, P., S. Felder, and A. Werblow (2004): “Population ageing and health care expenditure: New evidence on the “red herring””, The Geneva Papers on Risk and Insurance, 4 (29), 652–666. [4]
  • Zweifel, P. and W. G. Manning (2000): Moral hazard and consumer incentives in health care, Elsevier. [2]

Datas de Publicação

  • Publicação nesta coleção
    19 Set 2025
  • Data do Fascículo
    2025
location_on
Fundação Getúlio Vargas Praia de Botafogo, 190 11º andar, 22253-900 Rio de Janeiro RJ Brazil, Tel.: +55 21 3799-5831 , Fax: +55 21 2553-8821 - Rio de Janeiro - RJ - Brazil
E-mail: rbe@fgv.br
rss_feed Acompanhe os números deste periódico no seu leitor de RSS
Ir para o topo Reportar erro