5. Planning Samples and Experiments

A practical guide to selecting samples, identifying bias, distinguishing observational studies from experiments, and judging when statistical conclusions support generalization or causation.

1. Identify the and

A statistical result is only as trustworthy as the design that produced the data. Before calculating a mean, percentage, confidence interval, or test statistic, identify who was studied, how the units were selected, what was measured, whether treatments were imposed, and whether another variable could explain the result.

The is the complete group about which a conclusion is intended. The is the group actually observed. A case, or study unit, is one individual, object, organization, or other unit from which data are collected. A variable is a characteristic measured for each study unit.

For example, if a school wants to estimate the proportion of all enrolled students who regularly exercise, all enrolled students form the , the selected students form the , and exercise frequency is the variable measured. The sampling frame is the list or procedure used to identify members of the .

A useful should represent the target as well as possible. Sampling variability is unavoidable, but systematic problems in selection or measurement can create . A large does not automatically fix a poor design: a large biased can give a precise estimate of the wrong quantity.

Takeaway: Define the , , study units, sampling frame, and variables before interpreting numerical results.

2. Compare Sampling Methods

Different sampling methods use different structures and have different strengths.

  • Simple random sampling: Every possible of the chosen size has an equal chance of being selected. A random-number generator could select students from a complete enrollment list.

  • Stratified random sampling: The is divided into meaningful, nonoverlapping groups called strata, and a random is selected from each stratum. This can ensure representation of important subgroups.

  • Cluster sampling: The is divided into natural groups called clusters. The researcher randomly selects clusters and studies every member of selected clusters or samples members within them. This can reduce cost, but selected clusters may differ from one another.

  • Systematic sampling: The researcher selects every kth case from an ordered list, usually after choosing a random starting point. A repeating pattern in the list can create if it aligns with the sampling interval.

  • Convenience : Cases are chosen because they are easiest to reach. The results may reflect accessibility rather than the .

  • Voluntary-response : People choose whether to participate, as in an online poll. People with especially strong opinions or unusual experiences may be more likely to respond.

Convenience and voluntary-response methods are easy to conduct but often represent respondents better than they represent the intended . The quality of a sampling method depends on the research question, the sampling frame, and the possible differences between those selected and those not selected.

Takeaway: Probability-based sampling methods are designed to improve representativeness, while convenience and voluntary-response methods are especially vulnerable to .

3. Diagnose and Reduce

is a systematic tendency for a study's method to produce results that differ from the truth. It is different from random error: a larger can reduce random sampling variability but generally does not remove a flawed design.

Important sources of include:

  • Undercoverage: Some members have little or no chance of selection.

  • Nonresponse : Selected individuals do not participate, and those who respond differ systematically from those who do not.

  • Voluntary-response : Participants select themselves, often because they have unusually strong opinions.

  • Convenience-selection : Cases are chosen because they are easy to reach rather than because they represent the .

  • Response : Participants give inaccurate answers because of privacy concerns, faulty memory, social pressure, or question wording.

  • Interviewer or researcher : An interviewer's behavior, expectations, or wording influences responses or measurements.

  • Measurement : An instrument or procedure consistently overstates or understates the quantity being measured.

  • Attrition : Participants who leave a study differ systematically from those who remain.

A stronger survey design uses a clearly defined target , an appropriate sampling frame, probability-based selection when possible, neutral wording, confidentiality protections, consistent procedures, and follow-up with nonrespondents. Questions should avoid leading language.

Takeaway: Ask not only how many observations were collected, but also who could be selected, who participated, and how the variables were measured.

4. Separate Selection from Assignment

The difference between and determines what kind of conclusion a study can support.

determines who enters a . When the is appropriately selected and other sources of are limited, it supports generalizing from the to the .

determines which treatment a study participant receives. It helps create comparable treatment groups and supports cause-and-effect conclusions by reducing systematic differences between groups.

A study can use one without the other. For example, an with volunteer participants can support a causal conclusion for those participants, but its results may not generalize broadly to the entire . Conversely, a representative can support -level descriptions or associations but usually cannot establish causation.

A useful rule is:

  • helps answer, “To whom do the results apply?”

  • helps answer, “Did the treatment cause the observed difference?”

Takeaway: Generalization depends mainly on how units enter the study; causal inference depends mainly on how treatments are assigned.

5. Distinguish Observational Studies and Experiments

In an , researchers measure variables without assigning treatments or deliberately changing the explanatory variable. For example, researchers might compare health outcomes for people who choose to take a vitamin supplement with outcomes for people who do not. The researchers observe existing behavior.

An can reveal associations and suggest hypotheses. However, an observed association by itself does not establish that one variable causes another, because groups may differ in other important ways.

In an , researchers impose one or more treatments, assign study units to treatment conditions, and measure the responses. For example, researchers could assign participants to receive Method A or Method B and then compare changes in blood pressure.

Experiments are especially useful for studying causation because the researcher controls the treatment and can use to balance other differences between groups. Ethical limits still apply: participants must be protected, and unreasonable risks must not be imposed.

The design determines the strength of the conclusion:

  • A representative generally supports association and, if is limited, generalization.

  • A randomized with volunteers supports cause-and-effect conclusions for the experimental units, but broad generalization may be limited.

  • A randomized using a random provides stronger support for both causation and generalization when feasible and ethical.

  • A convenience usually describes only the observed cases.

Takeaway: Observational studies primarily show association; well-designed randomized experiments can provide evidence of cause and effect.

6. Recognize Confounding

A is a variable whose effects cannot be separated from the effects of the explanatory variable. A related idea is a lurking variable: an unmeasured variable that helps explain an observed relationship.

Suppose a study finds that people who carry lighters have higher rates of lung disease. Carrying a lighter does not cause lung disease. Smoking is a because smokers are more likely to carry lighters and are also more likely to develop lung disease.

Tutoring provides another example. If students who attend tutoring have higher test scores than students who do not, the difference may reflect tutoring, but it may also reflect prior achievement, motivation, available time, or parental support. When students choose whether to attend, these variables may be confounded with tutoring participation.

Confounding threatens a causal interpretation because several explanations remain possible. It is not the same as random variability. More observations may reduce random error, but a larger study with the same confounding structure can continue to produce a misleading association.

Takeaway: When a relationship is observed, identify other variables related to both the explanatory variable and the response before claiming causation.

7. Strengthen Experimental Design

A strong commonly combines comparison, randomization, control of other conditions, replication, and—when appropriate— and .

  • Comparison: Compare outcomes under different treatments. A control group may receive no active treatment, a standard treatment, or a placebo.

  • Randomization: Use chance to place study units into treatment groups. This tends to balance known and unknown characteristics, although it does not guarantee perfectly identical groups in one .

  • Control of other conditions: Keep procedures, timing, measurement instruments, and instructions as similar as possible except for the treatment being studied.

  • : Group units by a known characteristic expected to affect the response, then randomly assign treatments within each group. For example, participants could be grouped by age range before assignment.

  • Replication: Apply each treatment to many study units or repeat the study. Replication reveals natural variation and makes conclusions less dependent on an unusual individual or chance result.

  • : Prevent participants, researchers, or both from knowing treatment assignments. A double-blind study keeps both participants and the researchers interacting with them unaware of assignments when feasible.

These principles work together. Comparison provides a reference, randomization reduces systematic group differences, control limits alternative explanations, replication assesses consistency, addresses important known variation, and reduces expectation effects.

Takeaway: A persuasive is not defined by size alone; it is strengthened by a design that makes treatment groups comparable and alternative explanations less plausible.

8. Apply a Design Checklist

Before collecting data, use this checklist:

  1. What is the target ?

  2. What cases or study units will be observed?

  3. Is the sampling frame complete enough?

  4. How will the be selected?

  5. Could undercoverage, nonresponse, or voluntary participation create ?

  6. What are the explanatory and response variables?

  7. Is the study observational or experimental?

  8. If it is an , who receives each treatment, and how will assignment be randomized?

  9. What variables could confound the relationship?

  10. Should the design use a control group, , replication, or ?

  11. Will the design support description, generalization, association, or causation?

The final conclusion should match the design. A representative may justify generalization, while may justify a causal interpretation. Neither conclusion should be claimed when , confounding, weak measurement, or ethical limitations undermine the design.

Final takeaway: Begin with the target and selection process, inspect possible sources of and confounding, distinguish association from causation, and state only the conclusions that the design supports.