8. Sampling Distributions and the Central Limit Theorem
A guided explanation of sampling distributions, the Central Limit Theorem, standard error, sample size, and their role in statistical inference.
The Purpose of Sampling Distributions
Statistical inference uses information from a sample to learn about a larger population. A statistic calculated from one sample can change when a different sample is selected, even when the same sampling method is used.
A describes the values of a statistic across many possible random samples of the same size from the same population. It helps answer three questions:
What value does the statistic tend to have?
How much does it vary from sample to sample?
How can that variation be used to quantify uncertainty?
A distribution of individual observations describes population or sample data values. A instead describes a statistic, such as a or .
Takeaway: Sampling distributions provide the probability model behind estimates and measures of uncertainty.
The and Its
For a random sample from a population with mean and standard deviation , the is
Its has center and spread given by
and
Thus, the is centered at the population mean and is an of . Its sampling variability becomes smaller as increases.
For example, if delivery times have population standard deviation minutes and random samples contain deliveries, then
This means that the collection of sample means typically has a spread of about minutes. It does not mean that every individual will be exactly minutes from .
When is unknown, the estimated is commonly
where is the sample standard deviation.
Takeaway: The is centered at the population mean, and its decreases in proportion to .
The
The explains why normal distributions are central to many statistical procedures. If observations are independent and come from a population with finite mean and finite variance , then, as the sample size becomes sufficiently large, the of the is approximately normal:
Equivalently, the standardized is approximately standard normal:
The theorem does not say that individual observations become normal. It also does not require the original population to be normal, although the sample size needed for a good approximation depends on the population's shape. A normal population gives a normal of the mean for every sample size when the observations are independent.
A moderately skewed population may need a sample size around to for a useful approximation in some settings, while a strongly skewed or heavy-tailed population may require a much larger sample. There is no universal cutoff that works for every population.
The theorem also does not repair a biased sampling process. A large sample selected in a systematically biased way can produce a precise estimate of the wrong quantity.
Takeaway: The concerns the shape of a statistic's , not the quality of the sampling method or the distribution of individual observations.
Sampling Distributions for Proportions
When observations are classified as successes or failures, the is
If the population proportion of successes is , then
and
When the expected numbers of successes and failures are sufficiently large, the of is approximately normal. A commonly used check is
If is unknown, an estimated for a often replaces with :
The normal approximation should be used cautiously when the sample is small or when the observed proportion is near or . Specific procedures may use different conditions.
Takeaway: Sample proportions have their own and require adequate expected counts before a normal approximation is used.
Sampling Design and What It Supports
supplies the probability model needed to describe sampling variation. In a simple random sample, every possible sample of the specified size has the same chance of selection. More generally, a probability-based design gives population units known, nonzero chances of selection.
makes a sample representative on average, but no single sample is guaranteed to match the population perfectly. It supports generalization because the resulting sampling variation can be quantified.
Do not confuse with :
concerns how units are selected from a population and supports generalizing results to that population.
concerns how participants are allocated to treatments and supports conclusions about cause and effect in an experiment.
Increasing the sample size reduces variation, but it does not remove systematic problems such as undercoverage, nonresponse bias, measurement error, or selection bias. For example, a very large group of volunteers responding to an online advertisement may still differ systematically from the target population.
Takeaway: Precision from a large sample cannot compensate for a systematically biased sampling process.
How Sample Size Changes Inference
Sample size affects both the spread and the shape of a .
For a ,
For a ,
Both standard errors decrease at the rate . Therefore:
Multiplying the sample size by divides the by .
Multiplying the sample size by divides the by .
Doubling the sample size multiplies the by , approximately , rather than cutting it in half.
Larger samples also generally make the of a mean more nearly normal, especially when the population is skewed. However, the sample size needed depends on the population distribution, dependence among observations, outliers, and the statistic being studied. The rule is not universal.
Sample size is not the same as population size. When sampling without replacement from a finite population and the sample is not negligible relative to the population, a may be appropriate:
Here, is the population size. The correction reflects the reduced uncertainty that results from sampling a substantial fraction of a finite population.
Takeaway: More observations generally improve precision and normal approximation, but they do not eliminate bias or justify a universal sample-size rule.
From Sampling Distributions to Statistical Inference
Sampling distributions connect probability models to confidence intervals and significance tests.
A generally has the form
A larger produces a wider interval. Because increasing the sample size generally decreases the , larger samples often produce narrower intervals when other conditions remain the same.
A is designed so that, over many repetitions of the sampling procedure, approximately of intervals produced by that procedure contain the true population parameter when the assumptions are satisfied. It is not correct to say that one completed interval has a probability of containing a fixed parameter.
A compares an observed statistic with the expected under a null hypothesis. For a mean, a standardized test statistic can be written as
where is the null-hypothesis mean. A result farther from the null value in standard-error units is less compatible with the null model, subject to the assumptions of the procedure.
The supports approximately normal reference distributions for many statistics when the sampling design and other assumptions are appropriate. Nevertheless, statistical significance does not by itself establish practical importance, causation, or freedom from bias.
Takeaway: Sampling distributions quantify uncertainty, while confidence intervals and significance tests use that uncertainty to support carefully qualified conclusions.
Worked Example: Mean Waiting Time
Consider a population of waiting times with mean minutes and standard deviation minutes. A random sample of waiting times is selected.
First, find the center and spread of the :
and
Because the sample size is reasonably large, the suggests the approximation
where the second value represents the standard deviation of the .
Suppose the observed is minutes. Its standardized distance from the population mean is
Thus, the observed mean is standard errors above the population mean. This is noticeably higher than expected from ordinary sampling variation, but a complete probability calculation or formal test is needed to assess how unusual it is.
The result does not prove that the population mean changed. It could be produced by variation. A responsible conclusion considers the sampling design, the model assumptions, the , the research question, and plausible alternative explanations.
Takeaway: Standardization expresses an observed statistic in units of its sampling variability, making it possible to compare the observation with a reference distribution.
Key Checks and Final Synthesis
Keep the following distinctions in view:
A sample distribution displays the observed data values in one sample; a displays a statistic across many possible samples.
The standard deviation of individual observations is , whereas the standard deviation of sample means is .
The does not make biased data valid.
Doubling does not halve the ; a fourfold increase is needed to halve it.
The condition is not a universal guarantee of a normal approximation.
The usual formulas rely on independence or an appropriate dependence structure. Clustered, repeated, or time-series observations may require different methods.
Statistical significance is not the same as practical importance, causation, or freedom from bias.
The main chain of reasoning is:
Select observations using an appropriate sampling design.
Identify the statistic and its .
Determine the center and .
Check whether a normal or other approximation is justified.
Use the to construct an interval or evaluate a test statistic.
Interpret the result in light of bias, dependence, assumptions, and practical importance.
Final takeaway: Sampling distributions turn sample-to-sample variation into a formal basis for estimating population quantities and evaluating statistical evidence.