4 Causation and Scientific Inquiry

Learn how causal questions differ from associations, how experiments and observational studies support causal conclusions, and what assumptions and limitations shape the evidence.

From association to a causal question

A causal question asks what would happen to an outcome if a factor changed. It is different from asking whether the factor and outcome occur together: an observed association alone does not establish that changing the factor would change the outcome.

A useful causal question specifies the intervention, the alternative condition, the population, and the period. For example, among students eligible for a tutoring program, compare average end-of-term scores if they were offered the program with scores if they were not offered it.

and the missing comparison

describe the outcome a unit—such as a person, classroom, or farm—would have under each possible condition. A causal effect is a contrast between those outcomes. For any one unit, however, only one alternative is observed at a time: a student either received the tutoring offer or did not. The unobserved alternative is the counterfactual, so researchers use comparison groups to estimate it.

The causal comparison depends on how the groups were formed. A claim is more informative when it states which intervention is being compared with which alternative, for whom, and over what period.

What experiments can establish

In a , researchers assign units to conditions by chance. For example, to test whether a new fertilizer increases crop yield, a researcher could randomly assign comparable plots to receive the fertilizer or a control treatment, then measure yields using the same procedure. Random assignment tends to balance known and unknown background causes between groups, making a difference in average yield more credibly attributable to the assigned treatment than a simple comparison of plots that happened to receive different treatments.

Randomization does not guarantee that groups will be perfectly identical in a particular experiment; chance imbalances can occur, especially with small samples. Causal conclusions also depend on sound measurement, adequate follow-up, and attention to noncompliance, missing data, and unintended differences between groups. A result may apply to the treatment, participants, and setting studied without automatically applying to other populations or versions of the intervention.

When experiments are not feasible

Experiments may be impractical, unethical, or impossible. Researchers cannot ethically assign people to many harmful exposures, and some causes, including past events and large-scale policies, cannot be controlled directly. In these cases, scientists use observational evidence and seek designs that approximate a fair comparison. A natural experiment is one approach: an outside event or rule creates plausibly independent differences in exposure.

Challenges in observational studies

In an , researchers record exposure and subsequent outcomes without assigning the exposure. Its central challenge is : a third factor influences both exposure and outcome, creating or distorting their association. For instance, students with lower prior scores might be more likely to receive tutoring and still earn lower final scores than students who did not receive tutoring. A raw comparison could therefore make tutoring appear ineffective or harmful even if it helped those students.

Other challenges can also distort a comparison:

  • : the outcome, or an early stage of it, may influence the supposed cause. An association between illness and a behavior, for example, may occur because illness changed the behavior.

  • : people included in a study may differ systematically from those left out, or inclusion may depend on both exposure and outcome-related factors.

  • : exposures, outcomes, or confounders may be recorded inaccurately.

  • Ambiguous treatments: labels such as “exercise” or “high-quality instruction” may cover different versions, making the causal comparison unclear.

Researchers can measure likely confounders and use stratification, matching, or statistical adjustment to make groups more comparable. These methods do not automatically remove bias: an unmeasured or poorly measured confounder may remain, and adjusting for an unsuitable variable can itself distort an estimate. Observational conclusions therefore depend on assumptions, not only on the size or precision of an association.

Assumptions for observational comparisons

Three common assumptions used to support observational causal conclusions are:

  • : groups are comparable in their , perhaps after adjustment.

  • : each relevant type of unit has some chance of receiving each condition.

  • : the treatment being compared is defined clearly enough that its versions represent the same intervention.

These assumptions can be difficult to satisfy. In particular, poorly measured or unmeasured can undermine comparability, while an ambiguous treatment definition can make the intended comparison unclear.

Building and evaluating causal arguments

Causal diagrams and mechanistic knowledge can help researchers state what they think causes what, identify plausible confounders, and decide what to measure. A mechanism—a sequence of processes linking a cause to an outcome—can make a claim more credible, but a plausible story alone does not establish that the cause made a difference. Conversely, well-designed evidence can support a even when every detail of the mechanism is not yet understood.

A careful analysis defines the intervention, outcome, population, and time period, then examines how the comparison was formed, what alternative explanations remain, whether measures are trustworthy, and whether results hold across methods or settings. Agreement among randomized experiments, observational studies, mechanistic evidence, and natural experiments can strengthen a conclusion because these approaches have different weaknesses. Disagreement is informative too: it may reveal differences in populations, treatment definitions, measurements, or assumptions.

Causal conclusions are evidence-based judgments about how outcomes would change under specified conditions, not mere correlations or absolute certainties. Their strength is limited by the quality of the design and the plausibility of its assumptions.