10 Causal Inference: Basic Concepts
In previous chapters we have discussed two primary kinds of formal analyses. Regression model analyses allow us to describe correlations between key predictors and outcomes while adjusting for other confounding factors. Prediction analyses allow us to build models that can estimate outcome values for future observations; there is typically no distinction between key predictors or confounders with prediction models.
There is another type of analysis that we might want to do that we will call a causal analysis. In these analyses we want to make inference about what might be the effect on an outcome of changing a key predictor and setting it to be a different value. These causal inferences are often desired because they can lead us to identify treatments or interventions that can help us or improve our lives in different ways. Leaders of organizations are often interested in knowing what are the various “levers” at their disposal that they can adjust in order to make an outcome change in a particular direction. For example, a product manager at a company might want to know if changing the advertisement for a product will increase its sales or if it might be better to lower the price instead. Causal analyses can sometimes be useful for identifying such targets (or levers) of intervention.
Conceptually, the simplest scenario where we might draw a causal inference is when we have data from an experiment where we actively controlled something. Some examples of controlled experiments are
Laboratory experiments where we want to know how a certain type of cell grows under different environmental conditions that we directly control;
Agricultural field experiments where we specifically plant crops with different fertilizer mixes in order to determine which fertilizer combination gives the largest yield;
Controlled clinical trials where we study whether a new drug is effective by randomizing patients to receive either a placebo or an experimental drug.
In each of these cases, we as scientists or experimenters manipulated something and compared the outcome to some alternative. If you want to know what is the effect of manipulating something, the ideal data come from an experiment where you manipulated something. Drawing causal inferences from such an experiment will then require making some assumptions about how the data were generated, which we will detail later in this chapter.
While we might ideally design a study where we carefully manipulated some factor and watched out the outcome changes, this kind of approach is unfortunately not always possible. Sometimes it is simply too expensive to conduct a designed study. Conducting a large controlled clinical trial for a new drug can cost millions of dollars and may not be worth the expense for some pharmaceutical companies. If we want to know whether a certain chemical causes cancer, it would be fundamentally unethical to deliberately expose someone to that chemical and observe the effects. Political leaders often want to know the effect of large policy changes, but it can be impractical (or take too much time) to test those changes on the population. As a result, if we want to know what the effect might be of changing certain factors, we often have to resort to looking at a different kind of data.
One key challenge with making causal inferences is that we often have data from a study where nothing was manipulated. Such a study is often referred to as an observational study, because we merely observe the predictors and outcomes, as opposed to directly manipulating them as in the examples above. With observational studies, it can be a challenge to make causal inferences from the data without making unrealistic assumptions. Although some statistical techniques have been developed to help with mitigating the need for unrealistic assumptions, no statistical technique is available that can guarantee the validity of a causal inference.
Two key questions we want to address in this Chapter are:
What does it mean for something to cause something else? Most people have some intuitive sense of what this means in the real world. But for the sake of doing causal analyses, it is useful to have a precise definition of what this means.
When can we conclude that something causes something else? This is a difficult question and it’s the key one we want to answer when we’re doing causal analyses. Making conclusions about causal effects requires important assumptions and requires an understanding of the broader scientific process.
10.1 Accumulation of Evidence
In this chapter, the phrase “causal inference” is a shorthand for a specific set of statistical methods that are grounded in a framework that supports drawing conclusions about causal relationships between two factors, rather than just correlations (or associations) between two factors. However, “causal inference” has a broader, more general meaning, in that it refers to inferring a causal relationship between two factors. And the basis on which we might infer a causal relationship extends well beyond the specific sub-field of statistics referred to as “causal inference.” This is because drawing a conclusion about the nature of the relationship between two factors is not based on the use of any single technique or method or any single study. Instead, it is the overall body of research and more specifically, the evidence produced by that body of research, that supports (or not) a causal relationship between two factors.
The reason we need multiple independent studies to draw causal conclusions is that in any given study there may be hidden factors that can confound the causal relationship that we are interested in. As a result, we may draw an inappropriate conclusion from a single study. However, across multiple studies, the hope is that with different study designs, different methodologies, different datasets, different investigators, and different statistical analyses, it becomes much less likely that the same hidden factor would be present (and not accounted for) in every single study. If we can draw a similar conclusion about a relationship across multiple independent studies, then it suggests that the relation may be causal. That said, even with multiple independent studies, it is still possible to miss important factors, and so the presence of a relationship in multiple studies is not a guarantee of causality.
In general, we infer causality by considering the totality of the evidence that has been generated about the nature of the relationship between the two factors, and assess the strength of the evidence to inform the degree of confidence we have in whether the relationship is causal. To the extent that employing causal inference techniques improves the quality or strength of the evidence available, this is a good thing, and that is why we have included these chapters in this book.
10.2 Using Causal Language
Causal language is used in everyday writing and speech, typically without much thought. One might say one of the following statements:
“The light turned on because I flipped the switch”
“The ibuprofen made my headache go away”
“I got a good job because I graduated from UT Austin”
While such statements might seem reasonable, it is worth examining them more closely. One thing to note about each statement is the lack of comparison in each of them. Each of the statements seems to have an implied comparison that is obvious.
For example, the statement “The light turned on because I flipped the switch” might be implicitly being compared to whether the light would turn on if I hadn’t flipped the switch. So the alternative scenario is one in which I stood there not flipping the light switch. The question then is what would have happened with the light in that scenario where the switch is not flipped? It seems clear, based on a basic knowledge of electricity and physics, that the light would not turn on.
We can revise each of the causal statements above to include an alternative scenario that serves as a comparison:
“The light turned on because I flipped the switch…compared to if I had not flipped the switch”
“The ibuprofen made my headache go away…compared to if I had taken no medicine”
“I got a good job because I graduated from UT Austin…compared to if I had not attended UT Austin”
In each of these scenarios, the conclusion seems fairly straightforward. We know ibuprofen can help with headaches and so compared to taking nothing, it seems reasonable to say that “the ibuprofen caused my headache to go away”. But what if we changed that statement to say
- “The ibuprofen made my headache go away…compared to if I had taken Tylenol”?
Well, now we might reasonably assume that Tylenol and ibuprofen are equally likely to make a headache go away. Therefore, we might conclude that the ibuprofen did not have an effect compared to what might have happened if Tylenol had been taken because although ibuprofen is better than nothing, it’s not necessarily better than Tylenol. In this case, we might conclude that taking medicine is better than taking nothing but whether we take ibuprofen or Tylenol is not quite so important.
How about the last statement about getting a job after graduating from college? Graduating from college seems like it should lead to getting a good job. But what does it mean to not attend UT Austin? That “alternative scenario” includes a variety of possible activities that are not specified here. For example, we could change that statement to
- “I got a good job because I graduated from UT Austin…compared to if I had graduated from the University of Oklahoma”.
Now, UT Austin and University of Oklahoma are both good colleges. Without any specific knowledge, it’s not clear that my chances of getting a good job are any better after graduating from one college or the other.
In everyday conversation we often make use of causal language without explicit reference to a comparison or alternate scenario. In many cases, that alternate scenario is obvious and does not need to be said. However, even in everyday situations, different people may not agree on what is the implied alternate scenario. In scientific settings, it is critical to specify the alternate scenario to which a given intervention or action is being compared.
10.3 Potential Outcomes Framework
The primary framework that statisticians use to describe what it means for one thing to cause something else is known as the potential outcomes framework. This framework provides a systematic notation for describing outcomes of experiments before we observe them and for defining what is a causal effect. In addition, it provides a set of assumptions that need to hold in order to conclude that one thing causes another thing, i.e. to infer a causal effect.
The simplest version of the potential outcomes framework considers an intervention that has two levels. Think of the example with ibuprofen versus taking nothing. That intervention (i.e. “taking medicine”) has two possibilities:
Take the ibuprofen and observe the state of the headache
Take no medicine and similarly observe the state of the headache
When thinking about interventions, we think of one level of the intervention as being the “active” or “treatment” level, which is usually the thing that we are interested in studying. Meanwhile the other level is the “inactive” or “control” level, which is the reference to which we are comparing. For the purposes of this book, we will always refer to either the “treatment” or the “control”. In this example, we might think of taking ibuprofen as the treatment and taking nothing as the control.
In addition to specifying the nature of the intervention, we also need to specify an outcome. In the simplest case, where the intervention only has two levels, we have to imagine what the outcome would be under each of the intervention scenarios. These are the potential outcomes. Of course, we do not yet know what will happen under either scenario but we can imagine abstractly what they might be. In the headache example we might consider using a “headache score”, where a score of 10 indicates “severe pain” and a score of 1 indicates “no pain”. The potential outcomes questions then are what would the headache score be under the ibuprofen scenario and what would the headache score be under the “no medicine” scenario? Ultimately, we are interested in the difference of the outcomes between these two scenarios. That difference is the causal effect of taking ibuprofen versus taking no medicine.
We can generalize the concept of potential outcomes and causal effects by introducing a little mathematical notation. The first concept we will introduce is the concept of units on which we do interventions and we observe outcomes. The units are very often going to be individual people, but they may be other things like communities in the United States or products sold by a company.
We will use \(X\) to refer to the intervention and \(Y\) to refer to the outcome. With the intervention, we will say that
\(X=0\) indicates the control level
\(X=1\) indicates the treatment level
The outcome is represented by \(Y\) and we say that
\(Y(0)\) is the outcome observed with \(X=0\) (usually this is the control)
\(Y(1)\) is the outcome observed with \(X=1\) on the same unit (this is the treatment)
Here, \(Y(0)\) and \(Y(1)\) are the potential outcomes. Finally, a causal effect is defined as
- \(D = Y(1)-Y(0)\), which is the difference in potential outcomes between the two intervention levels on the same unit.
Continuing the example from above, we would say that
\(X=0\) indicates the control level of taking “no medicine”
\(X=1\) indicates the treatment level of taking ibuprofen
For the potential outcomes, we have
\(Y(0)\) the headache score when taking no medicine
\(Y(1)\) the headache score when taking ibuprofen
The causal effect is then \(D = Y(1)-Y(0)\), which is the difference in headache score between taking ibuprofen and taking no medicine.
One point to emphasize so far is that when we describe the intervention and the potential outcomes in this section, we imagine them all happening to the same unit. So \(Y(0)\) is the headache score when a person takes no medicine and \(Y(1)\) is the headache score when the same person takes ibuprofen. We will highlight the implications of this framework in the next section.
10.3.1 An Impossible Experiment
Let’s consider a different example based on the following causal statement:
“Robert got a good job because he graduated from UT Austin…compared to if he had graduated from the University of Oklahoma”
How would we describe this using the potential outcomes framework? First we need to define an outcome, which is a bit vague in the statement above. What does it mean to have a “good job”? One way we can make this more specific is by focusing on the salary that the job offers. Therefore, we can refine our statement to be
“Robert got a job with a higher salary because he graduated from UT Austin…compared to if he had graduated from the University of Oklahoma”
Our outcome is then
- \(Y =\) the salary that Robert makes at his first job after college.
We have two different intervention levels:
\(X=1\) indicates that Robert graduated from UT Austin
\(X=0\) indicates that Robert graduated from University of Oklahoma
Our potential outcomes are now
\(Y(1)\) is the salary that Robert makes at his first job after graduating from UT Austin
\(Y(0)\) is the salary that Robert makes at his first job after graduating from University of Oklahoma
The causal effect on Robert’s salary of attending UT Austin versus University of Oklahoma is then \(D=Y(1)-Y(0)\), i.e. the difference in salaries.
One problem with this entire narrative so far is that it assumes that Robert can attend UT Austin and University of Oklahoma at the same time. If we suspend reality for just a second, we can assume that two identical versions of Robert attended different colleges and obtained jobs with different salaries. But of course, this is not possible in reality, so as a result it is not possible to compute the causal effect \(D\) for Robert as an individual. This is known in the literature as the fundamental problem of causal inference.
So what are we supposed to do now? The next section presents a possible (but problematic) alternative.
10.3.2 A Possible Experiment
Let’s consider the same example as in the previous section but with a few modifications. First, instead of just focusing on Robert’s outcomes, let’s include Robert and his friend Linda, who is also going to college at the same time as he is. The outcome that we will look at is still their salaries after college, so the outcome will be modified to be
\(Y_R =\) the salary that Robert makes at his first job after college.
\(Y_L =\) the salary that Linda makes at her first job after college.
We still have two different intervention levels:
\(X=1\) indicates graduating from UT Austin
\(X=0\) indicates graduating from University of Oklahoma
Our potential outcomes are now (slightly revised)
\(Y_R(1)\) is the salary that Robert makes at his first job after graduating from UT Austin
\(Y_L(0)\) is the salary that Linda makes at her first job after graduating from University of Oklahoma
Let’s define the following difference: \(D^\star=Y_R(1)-Y_L(0)\) is the difference between Robert’s salary and Linda’s salary after college.
Is \(D^\star\) a causal effect? In general, the answer is no, because it is a difference that is defined on two different units, i.e. Robert and Linda. It is not a difference that is defined on the same unit or person as in the previous section. Therefore, it does not conform to the definition of a causal effect.
If we wanted to define a genuine causal effect we could define either
\[ D_R = Y_R(1)-Y_R(0) \]
which would be the causal effect for Robert, or we could define
\[ D_L = Y_L(1)-Y_L(0) \]
which would be the causal effect for Linda. However, in the first case we do not actually observe \(Y_R(0)\), which is Robert’s salary when he graduates from University of Oklahoma, and in the second case we do not observe \(Y_L(1)\), which is Linda’s salary when she graduates from UT Austin. Therefore, once again we cannot calculate either \(D_R\) or \(D_L\) because of the fundamental problem of causal inference.
\(D^\star\) is sometimes referred to as an association. That is because \(D^\star\) provides the difference in salary associated with Robert attending UT Austin compared to Linda attending University of Oklahoma. This is different from say \(D_R\), which provides the difference in salary caused by Robert attending UT Austin compared to Robert attending University of Oklahoma. We will dive deeper into this distinction between \(D^\star\) and \(D_R\) in the next section.
If we had concluded that \(D^\star\) was a causal effect, then in general we would have been making an inappropriate conclusion. However, at this point it’s not immediately clear why, other than it violates the definition of a causal effect. The general reason is because Robert and Linda are different and the ways in which they are different can cause problems when drawing causal conclusions. In the following sections, we will talk more about why comparing different units can cause problems in making causal inferences.
10.4 Confounding
When considering the effect of a treatment on an outcome, confounding is a general concept that describes the effect of other factors on the relationship between the treatment and outcome. In the previous section, we described \(D_R\), the causal effect of attending UT Austin (relative to University of Oklahoma) on Robert’s salary, and \(D^\star\), the association between Robert’s salary (after attending UT Austin) and Linda’s salary (after attending University of Oklahoma). Simply put, if \(D_R\ne D^\star\), then we have confounding. The catch, of course, is that we can never check if those two quantities are equal because we cannot compute \(D_R\).
In general, there is no reason to expect that \(D_R\) will equal \(D^\star\), so it is usually safe to assume that there is some confounding. Otherwise, we could always just compute \(D^\star\) and have our estimate of the causal effect. The question then is why is it problematic to compare an outcome across different people? In other words, why is confounding a problem for determining causal effects?
Let’s continue the hypothetical example from above with Robert and Linda and consider what are the differences between them. To begin with, we know that Robert attended UT Austin and Linda attended University of Oklahoma. Indeed, this is the comparison we are most interested in with respect to their starting salaries. Now let’s suppose Robert comes from a wealthy family that owns a successful large business in Texas. The cost of college is not a major concern for his family and they are willing to pay for him to attend either college. Meanwhile, Linda’s family has a middle class background and lives in Oklahoma. Because the in-state tuition for University of Oklahoma would be much less than the out-of-state tuition for UT Austin, her parents would prefer that she attend the University of Oklahoma.
At this point it would be fair to say that one difference between Robert and Linda is that they have different family wealth. Two questions that are important to ask here are:
Given what you know about Robert and Linda, can you make a prediction about who is more likely to attend UT Austin (or University of Oklahoma)?
Given what you know about Robert and Linda, can you make a prediction about who is more likely to have a higher (or lower) salary after college?
If you think the answer is “Yes” to both questions above, then “family wealth” is a potentially confounding factor when considering the causal effect of attending UT Austin on salary. Let’s play out one hypothetical version of the story below.
Robert decided early on that he wanted to attend UT Austin. After graduation, Robert’s father offers him a job at the family business which can afford to pay him a starting salary of $150,000 per year. Linda was accepted to both UT Austin and University of Oklahoma, but because of financial reasons, she decided to attend University of Oklahoma. After graduation, Linda takes a job with a salary of $50,000 per year.
In this example we can calculate \(D^\star = 150,000 - 50,000 = 100,000\). But would it be fair to say that this $100,000 difference in salary was caused by attending one school over another? In this case, another plausible explanation is that the difference in salary is caused by the difference in family wealth between Robert and Linda and that difference in family wealth caused them to go to different schools. Given the information we currently have, there is no way to distinguish between the potential effect of attending UT Austin vs. University of Oklahoma and the potential effect of coming from a wealthy vs. middle class family.
This example illustrates the problem of confounding. Confounding makes it so that the effect of the treatment we are interested in is hard to distinguish from the effect of some other factor (the confounder). Accounting for confounders, or confounding factors, is a major area of study and we will introduce a specific method for doing so in Chapter 11. In the next section we will just give a basic idea of what can be done to address this problem.
One important aspect to note about confounders is that variables or characteristics are not inherently confounders. Rather a variable can only be a confounder when considered in the context of another causal relationship. So “family wealth” can be a confounder in the context of the university–salary relationship. But “family wealth” is not a confounder all by itself. Indeed, it doesn’t make sense to call a variable a confounder without reference to a specific causal relationship of interest.
10.4.1 Adjusting for Confounding
Once we have identified that confounding is a problem, or that we have a potential confounding factor, what can we do about it? In some situations we can attempt to “adjust” for confounding and here we will give an example of how that might work.
Continuing the example from the previous section, suppose Robert has a friend Steven from childhood who has a similarly wealthy family. Robert and Steven are arguably comparable, given their similar family wealth and having grown up in the same area. Steven decided to go to University of Oklahoma for college and after graduation he was offered a job that paid a salary of $120,000 per year.
Comparing Robert to Steven seems a bit more reasonable in the sense that even though they are not literally the same people, they are quite similar. In particular, they both seem to have similar family wealth, which was the confounding factor that we were concerned about previously. If we look at the difference in salary between Robert and Steven we find that it is \(150,000-120,000 = 30,000\).
Linda also has a childhood friend Sharon who grew up in the same area as Linda and comes from a middle class background. However, unlike Linda, Sharon attended UT Austin for college and after graduation took a job that paid $70,000 per year. Comparing Linda to Sharon also seems like a reasonable comparison given their similar family wealth. When we look at the difference in their salaries we get \(70,000-50,000 = 20,000\).
What we have done here is taken each of our original subjects of comparison—Robert and Linda—and matched each of them to someone who is similar to them with respect to family wealth but attended a different college. That way, we can compare each person to their match and hopefully remove the effect of family wealth on a person’s salary after college. Note that when we looked at the differences in salary between the matched pairs of people ($30,000 and $20,000), the differences were smaller than they were when we just compared Robert and Linda ($100,000). The hope is that any difference that we observe between the matched pairs is solely caused by their attendance at different colleges and not by their family wealths. We have summarized the four individuals introduced here in Table 10.1.
| Attended UT Austin | Attended U of Oklahoma | ||
|---|---|---|---|
| Robert | Wealthy family | Linda | Middle class family |
| Salary: $150,000 | Salary: $50,000 | ||
| Sharon | Middle class family | Steven | Wealthy family |
| Salary: $70,000 | Salary: $120,000 | ||
| Average Salary | $110,000 | $85,000 |
We can take a look at the two groups of people that either went to UT Austin (left two columns) or University of Oklahoma (right two columns). Both groups have one person from a wealthy family and one person from a middle class family. In that sense, both groups of people are similar because each group has the same mix of family wealth At the bottom of Table 10.1 we have put the average of the salaries for the people who attended UT Austin and the people who attended University of Oklahoma. If we take the difference of the two averages, we get \(110,000 - 85,000 = 25,000\). Thus, we could say that the average difference in salary between people who attended UT Austin and University of Oklahoma is $25,000. This quantity is known as an average treatment effect.
What can we say about whether this average treatment effect is caused by attending UT Austin vs. University of Oklahoma? In this case, it would be more difficult to argue than before that the difference in salary is caused by differing family wealth because even though individual people have different family wealth, both groups of people have the same mix of family wealth. Therefore, we could argue that the two groups of people are comparable—the only difference between the groups is where they attended college.
Now, there could be other factors that are different between the two groups that we haven’t recorded here, so it’s always good to be on the lookout for other possible differences. But at least for now, we can argue that we have adjusted for differences in family wealth. When we say that we have “adjusted” for a certain confounding factor, this is the basic idea: We have constructed a dataset where the two groups that we are comparing are similar with respect to that confounding factor. Therefore, if we see a difference in the outcome between the two treatment groups, it cannot be explained by differences in this confounding factor.
10.5 Constructing Comparable Groups
The example in the previous section was intentonally small so that we could show the details of how people were similar or different from each other with respect to different factors. In general, we will be studying large groups of people and will want to know the effect of a treatment relative to some control in those groups. How can we ensure that the groups of people that we are comparing are similar to each other with respect to potential confounding factors? How can we construct groups of people that are comparable to each other except for their assignment to the treatment or control groups?
10.5.1 Randomized Assignment
If we are designing a study and we are in control of assigning who will receive the treatment and who will receive the control, then the simplest way to ensure that the two groups will be comparable is to randomize them to one group or the other. For each person, we can flip a fair coin and assign them the treatment if it comes up heads and assign them the control if it comes up tails. The reason this approach works is because the coin is not affected by any of the characteristics of the people in the study. The coin is not going to assign the treatment to all the tall people and the control to all the short people. It’s not going to be influenced by the wealthy backgrounds of some of the people or the hair color of other people. When people are randomized to receive either the treatment or control we often say that we are conducting a randomized controlled trial or RCT.
Randomized controlled trials are considered the gold standard for estimating causal effects because they statistically guarantee that both groups will be similar on all factors, regardless of whether we measure them. Therefore, the only way that the groups will differ (statistically) is in which treatment group they have been assigned. However, in order to do an RCT we need to be in a position where we can assign the treatment level to a given individual and there are some important scenarios where this is simply not possible.
In some cases, randomly assigning people to a treatment is unethical. For example, if you wanted to study the health effects of a certain chemical, it would likely be unethical to deliberately expose a group of people to that chemical, especially if there was already some evidence that it was harmful. Therefore, in the study of potentially harmful chemicals, it is not common to see randomized controlled trials. In other situations it may not be practical to conduct an RCT. If a city wanted to know whether increasing the sales tax would affect economic development, it would not be practical to randomize some citizens to receive the increased tax and some citizens to receive the current tax. Finally, RCTs are expensive to conduct and require careful execution. Very often, the resources required to run an RCT are simply not available.
When we cannot conducted a controlled trial where we randomize people to receive different treatments, we have to resort to data that were collected where nothing was controlled and the mechanism by which people received treatments is unknown. Such studies are known as observational studies and we have to employ analytical techniques that allow us to construct comparable groups for studying treatments. We briefly discuss two of those approaches in the next sections.
10.5.2 Direct Matching
In the college comparison example discussed in Section 10.4 the approach that we used to make comparisons of people with similar family wealth is known as direct matching. Direct matching pairs a person who received the treatment with a person who received the control and who is an exact match on a potential confounding factor. So in the college comparison example, we matched a person who attended UT Austin with a person who attended University of Oklahoma and had the same family wealth. That way, when we compared the salaries between the two people, we knew that any difference could not have been caused by a difference in family wealth.
Direct matching is intuitive and effective at eliminating any differences between people with respect to potential confounders. If a dataset is large enough, we can often take someone who received the treatment and match them with a person with the exact same confounder profile but who received the control. As long as we can find matches for everyone in the treatment group, we should be able to create a dataset that effectively adjusts for the confounder on which we based the matching.
The problem with direct matching is that we often do not want to adjust for a single confounding variable. Often, there are multiple confounding variables that we want to control. For example, we might want to adjust for family wealth and height. Now, for each person who received the treatment, we need to find someone who received the control and has the same family wealth and has the same height. This might still be possible with two factors, but as we increase the number of factors that we want to match on, it becomes increasingly difficult to find exact matches.
Ultimately, direct matching is very effective at controlling for confounding factors, but only if there are just a handful of factors. With the large size of datasets these days, we can often match people on a larger number of factors than was previously possible. But because the matching problem increases in complexity exponentially with each additional confounder, we can quickly run out of data for matching.
10.5.3 Propensity Score Matching
Propensity score matching is a technique for creating comparable groups when there is a large number of potential confounding variables and direct matching is not feasible. The technique involves building a model that predicts who receives the treatment or control based on the collection of potential confounders. We will not get into the details of how this is done right now as propensity score matching is discussed in detail in Chapter 11.
10.6 Estimating Average Treatment Effects
As we saw in Section 10.4.1, even though we cannot directly compute the causal effect of a treatment on an individual unit or person, we can compute average treatment effects by comparing relatively similar groups of people (where similarity is with respect to potential confounding factors). That is, we have substituted the problem of estimating the effect of a treatment on one person with the problem of estimating the effect of a treatment on two groups of similar people.
From a statistical perspective, two groups of units are comparable if on average, they are the same on every characteristic that you might measure or observe. Because we are talking about groups of units and not individual units, we will be making comparisons of averages instead of making direct comparisons. Using this concept, we can define the average treatment effect as a difference in average outcomes between two groups that received different treatments but are otherwise comparable. In practice, we often estimate the average treatment effect and use that as our best estimate of the individual causal effect.
Statistical Definition
We can define the average treatment effect a bit more precisely using statistical notation. Given the random potential outcomes \(Y(1)\) and \(Y(0)\) corresponding to the treatment and control for person \(i\) in a population, then the average treatment effect is
\[ ATE = \mathbb{E}[Y(1)]-\mathbb{E}[Y(0)] \] where the expectation \(\mathbb{E}\) is the average value taken over the entire population of people. Again, this quantity is not directly estimable because it requires observing two outcomes on the same person.
Now suppose we have an observational study where a group of people happen to receive the treatment (\(X=1\)) and another group of people happen to receive the control (\(X=0\)). In this scenario we didn’t specifically assign people to different treatments. Then a different quantity that we can estimate is \[ \mathbb{E}[Y\mid X=1]-\mathbb{E}[Y\mid X=0], \] where again the expectation \(\mathbb{E}\) is the average taken over the entire population of people. This difference is the difference in means between the group of people who received the treatment and the group of people who received the control. If these two groups are comparable, then this quantity is equal to the average treatment effect. If the groups are not comparable, than this difference is not equal to the average treatment effect and is interpreted as an association between \(X\) and \(Y\).
With a set of data on \(n\) individuals where \(Y_1,Y_2,\dots,Y_n\) represent the outcomes for each of the \(n\) individuals and \(X_1,X_2,\dots,X_n\) represent the treatment each individual received (either 0 or 1), we can estimate this quantity above as
\[ \widehat{ATE} = \frac{1}{n}\sum_{i=1}^n Y_i X_i - \frac{1}{n}\sum_{i=1}^n Y_i(1-X_i). \] Again, we are assuming that the group for which \(X_i=1\) and the group for which \(X_i=0\) are comparable.
The quantity \(\widehat{ATE}\) is something we can estimate from data and leverages the fact that we have information on multiple people and not just one person. If the two groups that we are comparing are similar with respect to potential confounding factors, then we can interpret \(\widehat{ATE}\) as an average of the individual causal effects of each individual.
10.7 Causal Inference Assumptions
After all this, we can state the key assumptions that are needed in order make causal inferences or conclusions from the data we analyze (Table 10.2). These assumptions are in regards to the data-generating process and the manner in which the treatment and control are assigned to individual units in the sample.
| Name of Assumption | Definition |
|---|---|
| Stable unit treatment value assumption (SUTVA) (no interference and no hidden treatments) | This assumption requires that the treatment level received by one unit does not affect the potential outcomes for another unit. So if I receive the treatment, that doesn’t have any influence on the potential outcomes of another subject in the study, regardless of whether that subject receives the treatment or control. In addition, it should be clear that the treatment everyone receives is the same, so that there aren’t any hidden variations of the treatment that people actually receive. |
| Positivity | This assumption says that every unit in the study sample has some non-zero probability of receiving either the treatment or control. |
| Ignorability | This assumption states that there are no hidden confounders; in other words, the groups receiving the treatment and control are comparable. |
The first two assumptions are often satisfied in fairly common analysis scenarios. However, the third assumption is the most difficult one to verify, if it is even possible. In Table 10.3 below we show some examples of how each of these assumptions could be violated.
| Assumption | Example of Violation |
|---|---|
| SUTVA | A study where the units have the potential to communicate with each other could result in a SUTVA violation. For example, if my friend and I are enrolled in a study about a new medication and I experience bad side effects while taking the medication, I might discourage my friend from taking the same medication, thereby affecting the potential outcomes for my friend. |
| Positivity | Any study where a sub-group of people simply cannot be assigned to a treatment. For example, in a study of two drugs (call them drug A and drug B), you might have some people who are allergic to drug A. Therefore, there is no chance that those people would be assigned to receive drug A. |
| Ignorability | Any observational study where the assignment of the treatment or control was not directly controlled has the potential to have hidden confounders, hence violating ignorability. It is the job of the analyst to argue ignorability holds in such settings based on a thorough understanding of how the data are generated. |
10.8 Causal Inference in Real-World Studies
At this point it should be clear that when attempting to draw causal inferences about treatments and outcomes, we want to be as sure as possible that we are comparing groups that are similar to each other with respect to potential confounding factors. We also want to be as sure as possible that the data we are analyzing satisfy the key assumptions described in the previous section. In this section we provide three real-world examples from our own work that each raise a causal question. We discuss whether the groups that are being compared in each study are likely to be similar to each other and the extent to which they satisfy the assumptions needed to make causal inferences.
10.8.1 Example: PREACH Study of Air Cleaners and Asthma
The Particulate Reduction Education in City Homes (PREACH)1 study was a randomized controlled trial that examined the effectiveness of using indoor air cleaners to improve asthma morbidity in children who lived with a smoker. The study had one group that was randomized to receive an educational module about indoor air pollution and its relationship with asthma (the “control” intervention) and another group that received an air cleaner in their home in addition to the educational module (the “treatment” intervention). One of the outcomes that was examined was the number of “symptom-free days”, which is the number of days in the past two weeks that the study participant did not experience any symptoms (higher is better). Note that the study actually had a third intervention level that involved a health coach, but we will ignore that for this example.
The study found that after six months, on average, the group that received the air cleaners plus the educational module experienced about 1.3 more symptom-free days than the group that only received the educational module. Based on this information, one might be inclined to conclude that “The use of an air cleaner in the home, relative to not using an air cleaner, caused symptom-free days to increase, on average, across the study participants.” Is this a reasonable conclusion? The answer depends on whether the two groups of people we are studying are comparable.
The comparison being made is between a group that received an educational module and a group that received an air cleaner and an educational module. In this example, where the individual people were randomized to receive either intervention level using a coin flip, we can be reasonably sure that the two groups are comparable, so that on average, we expect them to be fairly close to each other with respect to potential confounding factors. The only way in which they differ is by whether they got the treatment or the control.
According to Table 1 from the paper, before the study started (but after people were randomized to groups),
- 50% of the control group was male while 59% of the air cleaner group was male;
- The average age of the control group was 9.2 years old while the averate age of the air cleaner group was 9 years old;
- 89% of the control group used Medicaid or some other public insurance program while 88% of the air cleaner group used public insurance
- 32% of the control group had severe persistent asthma while 29% of the air cleaner group ahd severe persistent asthma.
As you can see from these four characteristics, the percentages between the two groups are not identical, but they are relatively close.
Now, our definition of “comparable” did not say that the groups had to be similar on average for just four characteristics, but for all characteristics. How can we know if they are similar for all characteristics if we cannot possibly measure everything? This is where we can rely on the randomization in the assignment of the treatment levels. Because the randomization scheme is indifferent to any study participant characteristics, we can be reasonably sure that both groups, on average, will have similar characteristics, even for things that we do not measure.
For a randomized controlled trial like PREACH, where people are randomly assigned to receive a specific treatment (air cleaner or education module), the argument that the two groups are comparable is perhaps the strongest. This is because the manner with which people are assigned to receive a treatment (the coin flip mechanism) ignored any and all characteristics about the people themselves. Therefore, it is not likely that either group will be biased with respect to one characteristic or another. For example, it’s not likely that one group will be much taller than the other group because roughly half of the tall people will be assigned to one group and the other half of the tall people will be assigned to the other group. So both groups will have a similar proportion of tall people. The same argument can be used for any other characteristic of the people in the study.
The randomized assignment of treatment in this study suggests that the ignorability assumption is satisfied and that there are unlikely to be any hidden confounders. We can also consider the other two key assumptions needed to make causal inferences from the data in the study.
SUTVA: Given that the participants were recruited independently from a wide swath of areas across Baltimore City, it’s reasonable to assume that the participants acted separately from each other and didn’t actively influence the decisions, behaviors, or outcomes of the other participants.
Positivity: In this study, anyone who qualified to be in the study in the first place had some chance of receiving either intervention level. This is not the same as saying that anyone could participate in the study, because there were a number of eligibility requirements for the study itself (for example, participants had to live with someone who was a smoker). But once the eligibility requirements were satisified, there were no other barriers to receiving either intervention level.
Source: Butz et al. (2011)
10.8.2 Example: Coarse Particulate Matter and Asthma
Decades of research have provided strong evidence that particulate matter is related to a variety of health problems, including premature mortality. Much of that work is focused on the fine fraction of particulate matter, also known as fine PM or PM2.5. Substantially less research has focused on the coarse fraction of particulate matter, or coarse PM. While the larger size particles are not thought to be as harmful as the smaller particles which can travel deep into the lungs and airways, there is comparatively little research characterizing the health risks of coarse PM.
We conducted a national study of coarse PM and asthma outcomes in the United States to see if there was any relationship between long-term average concentrations of outdoor coarse PM and asthma hospitalizations amongst children enrolled in the Medicaid system, which is a public insurance program in the U.S. Full details of how the study was conducted and how the data were analyzed can be found in the journal article2. We only provide a brief synopsis here.
One key finding from the paper was that communities where there were higher concentrations of coarse PM tended to also have higher rates of asthma hospitalizations. If a community had a long-term average coarse PM concentration that was 1 \(\mu\)g/m\(^3\) higher than another community, then on average, the community with the higher coarse PM had a 2.3% higher asthma hospitalization rate. Given the findings of the study, it might be natural to ask whether higher levels of coarse PM, relative to lower levels of coarse PM, cause asthma hospitalizations, thereby leading to higher hospitalization rates.
This study differs from the PREACH study described in the previous section in a variety of ways. First, this was an observational study, where data were collected and analyzed without any active role played on the part of the investigators. Second, the “treatment” being examined is the concentration of coarse PM, which is a continuously varying quantity with more than just two levels. Third, the units of analysis were not individual people, but rather communities of people. Despite the differences in the two studies, we can still ask the basic causal question of whether high levels of coarse PM cause asthma hospitalizations relative to low levels. But the question of whether the groups of people that we are studying (i.e. the different communities) are comparable to each other is much more complicated. The study itself used a variety of statistical approaches to adjust for potential confounding factors to ensure that the communities being compared were comparable.
Given the description of the study, we can assess whether the three key assumptions are likely to be satisfied in this study.
SUTVA: The units in this analysis were individual communities and because of the spatially dynamic nature of air pollution, if we knew that pollution was high in one community, we could be reasonably sure that it would be high in neighboring communities. However, this doesn’t necessarily imply a SUTVA violation because the question is whether high levels of coarse PM in one location affect asthma hospitalizations in another community under high or low coarse PM levels (i.e. the other location’s potential outcomes). One way to imagine how SUTVA might be violated here is if high levels of coarse PM in one community caused hospitals to be overwhelmed and forcing patients to go elsewhere, thereby affecting medical services in neighboring communities. That said, there is no evidence of such a phenomenon occurring in response to typical air pollution levels in the United States.
Positivity: Because all of the communities in this study were susceptible to either high or low coarse PM levels, and all communities had “access” to the “treatment”, the positivity assumption was satisfied in this study.
Ignorability: The question to answer here is whether communities with high coarse PM are comparable to communities with low coarse PM. One could quickly come up with reasons why the answer is no. Communities with high air pollution in general tend to have lots of sources of air pollution, such as cars, trucks, ports, and power stations. Bigger communities tend to have more of all those things so we might expect communities with high coarse PM to be bigger and with larger populations. That said, factors like population are easily measured and can be directly adjusted for in the analysis. However, one could think of other factors that were not measured that might differ between the two groups. For example, economic activity or conditions that might affect both coarse PM levels and asthma outcomes were not measured and could not be directly controlled. For this study, it would be difficult to argue that there were no hidden confounders. Rather one would have to argue that any unmeasured or hidden confounders did not have a significant impact on the causal relationship of interest.
Source: Keet et al. (2018)
10.8.3 Example: Mobility Asthma Project
The last real-world example we will discuss comes from a study that was conducted in Baltimore, Maryland in the United States that followed families who moved from the city of Baltimore to surrounding communities. As part of the resolution of a lawsuit that was filed against the Baltimore City housing department regarding the racial segregation of public housing within the city, the Baltimore Regional Housing Partnership was created in 2015 to provide “fair housing opportunities for African-American public housing residents throughout the Baltimore region.” As part of their efforts, BRHP provides housing mobility vouchers to low-income families that help those families move to higher opportunity areas in the Baltimore metropolitan region.
We designed the Mobility Asthma Project (MAP) study to follow families with children that moved to the higher opportunity areas to see if there were any changes to their home environment and to the children’s asthma morbidity. Assessments of the child and the home were taken both before and after moving in order to track any changes. The unit of analysis was the child in the family and the primary outcome was the number of asthma exacerbations in the previous 3 months.
The study found that moving was associated with a 70% decrease in the rate of exacerbations3. The paper noted that the “magnitude of reduction of exacerbations associated with moving was greater than that observed for individual- and household-level interventions for asthma in racialized populations, larger than the effect of inhaled corticosteroids, and similar to that observed for the effect of biologic agents.” In other words, the decline in asthma exacerbations observed in the study was larger than that normally seen with expensive and powerful medications. The causal question here is whether moving to a higher opportunity area (relative to staying in the same location) causes asthma exacerbations in children to decrease.
This study shares some features with the previous study on coarse PM. The MAP study was an observational study in that the investigators did not actively intervene on the study subjects. Rather the study leveraged an event that was known to be happening in the future for each family (moving to a new neighborhood). The intervention in this case had two levels: Living in the original home was the “control” level and moving to the new home was the “treatment” level. What makes this study different compared to the others is that the units (i.e. the children) in the treatment group and the units in the control group are the same units. In other words, the comparison being made is between the children in their original homes and the same children in their new homes.
It’s important to note that we have not circumvented the fundamental problem of causal inference, because while we are comparing the same children, we are comparing their health status at different times. Therefore, it’s not necessarily true that the children in the treatment group and the children in the control group will be comparable, even though they are the same children. The reason is that the children might have changed over time (e.g. grown) in ways that makes them different. Therefore, we still have a question of whether we can make comparisons between the two treatment groups to draw causal conclusions.
Regarding the three key causal inference assumptions, we can make the following assessment.
SUTVA: Much like the PREACH study, the participants in this study are reasonably assumed to have acted independently and had little impact, if any, on the potential outcomes of other participants. It’s difficult to think of how there might be a violation of this assumption in this study.
Positivity: Because of the study design, essentially every participant received both intervention levels at some point in the study. Therefore, by definition, everyone had a positive probability of receiving either treatment.
Ignorability: Although this study has a special design where each participant serves as their own “control”, there is still the possibility of confounding by other factors. Given the structure of the study, the active intervention level (i.e. moving to the new house) always occurred later in time relative to the control level (i.e. living in the original house). Therefore, there is a possibility that the participant (which in this case was a child) could have “changed” in a manner that made the later version of that child not comparable to their former selves. At a minimum, the child will be older when the move occurs, although age was a factor that was directly controlled for in the analysis. They key issue here is whether there were any other factors that had the potential to change over the course a few months that could affect asthma exacerbations.
Source: Pollack et al. (2023)
10.9 Summary
Very often in data analyses we are interested in answering causal questions. But making causal conclusions from data analyses require that some critical assumptions be made about the data generating process and the underlying mechanisms being studied. The potential outcomes framework provides a formal statistical representation of a causal effect and shows why it is fundamentally not possible to compute causal effects directly. However, there are ways to estimate other quantities, such as average treatment effects, from observed data, provided the necessary assumptions are met. Randomized controlled trials provide the clearest path to making causal inferences from data. In situations where we cannot control who receives the treatment, matching techniques can be useful. We will see in the next chapter how observational studies can also serve as an important source of evidence for causal conclusions.
10.10 Exercises
A researcher rides her bike to a meeting at a university and gets there 20 minutes early. She notes to a fellow meeting attendee, “I got here early because I rode my bike instead of driving my car.” What is the outcome in this scenario? What is the “treatment” and “control” in this scenario? What comparison is the researcher making?
A professor wants to test two versions of a question on an exam to see if one of the versions might be too difficult. He divides the class into two groups based on their current standing in the course. All of the students who have a “B” grade or above in the class get Version 1 of the question and all of the students who have a “C” grade or below get Version 2 of the question. After the exam, the professor finds that the students who got Version 2 of the question generally performed much worse than the students who got Version 1. He then concludes Version 2 is too hard relative to Version 1. Is a causal conclusion justified in this case? In particular, are the two groups of students comparable?
A published paper about the PREACH study can be found at https://pmc.ncbi.nlm.nih.gov/articles/PMC6413330/↩︎
The full published study of coarse PM and asthma outcomes can be found at https://pmc.ncbi.nlm.nih.gov/articles/PMC9135134/↩︎
Full details about the study can be found at https://pmc.ncbi.nlm.nih.gov/articles/PMC10189571/↩︎