8. Sampling Distributions and the Central Limit Theorem

A guided explanation of sampling distributions, the Central Limit Theorem, standard error, sample size, and their role in statistical inference.

The Purpose of Sampling Distributions

Statistical inference uses information from a sample to learn about a larger population. A statistic calculated from one sample can change when a different sample is selected, even when the same sampling method is used.

A describes the values of a statistic across many possible random samples of the same size from the same population. It helps answer three questions:

  • What value does the statistic tend to have?

  • How much does it vary from sample to sample?

  • How can that variation be used to quantify uncertainty?

A distribution of individual observations describes population or sample data values. A instead describes a statistic, such as a or .

Takeaway: Sampling distributions provide the probability model behind estimates and measures of uncertainty.

The and Its

For a random sample X1,X2,…,XnX_1,X_2,\ldots,X_n from a population with mean μ\mu and standard deviation σ\sigma, the is

Xˉ=X1+X2+⋯+Xnn.\bar{X}=\frac{X_1+X_2+\cdots+X_n}{n}.

Its has center and spread given by

E(Xˉ)=μE(\bar{X})=\mu

and

SD(Xˉ)=SE(Xˉ)=σn.SD(\bar{X})=SE(\bar{X})=\frac{\sigma}{\sqrt{n}}.

Thus, the is centered at the population mean and is an of μ\mu. Its sampling variability becomes smaller as nn increases.

For example, if delivery times have population standard deviation 1212 minutes and random samples contain 3636 deliveries, then

SE(Xˉ)=1236=2 minutes.SE(\bar{X})=\frac{12}{\sqrt{36}}=2\text{ minutes}.

This means that the collection of sample means typically has a spread of about 22 minutes. It does not mean that every individual will be exactly 22 minutes from μ\mu.

When σ\sigma is unknown, the estimated is commonly

SE^(xˉ)=sn,\widehat{SE}(\bar{x})=\frac{s}{\sqrt{n}},

where ss is the sample standard deviation.

Takeaway: The is centered at the population mean, and its decreases in proportion to 1/n1/\sqrt{n}.

The

The explains why normal distributions are central to many statistical procedures. If observations are independent and come from a population with finite mean μ\mu and finite variance σ2\sigma^2, then, as the sample size becomes sufficiently large, the of the is approximately normal:

Xˉ≈N(μ,σn).\bar{X}\approx N\left(\mu,\frac{\sigma}{\sqrt{n}}\right).

Equivalently, the standardized is approximately standard normal:

Z=Xˉ−μσ/n.Z=\frac{\bar{X}-\mu}{\sigma/\sqrt{n}}.

The theorem does not say that individual observations become normal. It also does not require the original population to be normal, although the sample size needed for a good approximation depends on the population's shape. A normal population gives a normal of the mean for every sample size when the observations are independent.

A moderately skewed population may need a sample size around 2525 to 3030 for a useful approximation in some settings, while a strongly skewed or heavy-tailed population may require a much larger sample. There is no universal cutoff that works for every population.

The theorem also does not repair a biased sampling process. A large sample selected in a systematically biased way can produce a precise estimate of the wrong quantity.

Takeaway: The concerns the shape of a statistic's , not the quality of the sampling method or the distribution of individual observations.

Sampling Distributions for Proportions

When observations are classified as successes or failures, the is

p^=number of successesn.\hat{p}=\frac{\text{number of successes}}{n}.

If the population proportion of successes is pp, then

E(p^)=pE(\hat{p})=p

and

SE(p^)=p(1−p)n.SE(\hat{p})=\sqrt{\frac{p(1-p)}{n}}.

When the expected numbers of successes and failures are sufficiently large, the of p^\hat{p} is approximately normal. A commonly used check is

np≥10andn(1−p)≥10.np\ge 10\quad\text{and}\quad n(1-p)\ge 10.

If pp is unknown, an estimated for a often replaces pp with p^\hat{p}:

SE^(p^)=p^(1−p^)n.\widehat{SE}(\hat{p})=\sqrt{\frac{\hat{p}(1-\hat{p})}{n}}.

The normal approximation should be used cautiously when the sample is small or when the observed proportion is near 00 or 11. Specific procedures may use different conditions.

Takeaway: Sample proportions have their own and require adequate expected counts before a normal approximation is used.

Sampling Design and What It Supports

supplies the probability model needed to describe sampling variation. In a simple random sample, every possible sample of the specified size has the same chance of selection. More generally, a probability-based design gives population units known, nonzero chances of selection.

makes a sample representative on average, but no single sample is guaranteed to match the population perfectly. It supports generalization because the resulting sampling variation can be quantified.

Do not confuse with :

  • concerns how units are selected from a population and supports generalizing results to that population.

  • concerns how participants are allocated to treatments and supports conclusions about cause and effect in an experiment.

Increasing the sample size reduces variation, but it does not remove systematic problems such as undercoverage, nonresponse bias, measurement error, or selection bias. For example, a very large group of volunteers responding to an online advertisement may still differ systematically from the target population.

Takeaway: Precision from a large sample cannot compensate for a systematically biased sampling process.

How Sample Size Changes Inference

Sample size affects both the spread and the shape of a .

For a ,

SE(Xˉ)=σn.SE(\bar{X})=\frac{\sigma}{\sqrt{n}}.

For a ,

SE(p^)=p(1−p)n.SE(\hat{p})=\sqrt{\frac{p(1-p)}{n}}.

Both standard errors decrease at the rate 1/n1/\sqrt{n}. Therefore:

  • Multiplying the sample size by 44 divides the by 22.

  • Multiplying the sample size by 99 divides the by 33.

  • Doubling the sample size multiplies the by 1/21/\sqrt{2}, approximately 0.7070.707, rather than cutting it in half.

Larger samples also generally make the of a mean more nearly normal, especially when the population is skewed. However, the sample size needed depends on the population distribution, dependence among observations, outliers, and the statistic being studied. The rule n≥30n\ge 30 is not universal.

Sample size is not the same as population size. When sampling without replacement from a finite population and the sample is not negligible relative to the population, a may be appropriate:

SE(Xˉ)=σnN−nN−1.SE(\bar{X})=\frac{\sigma}{\sqrt{n}}\sqrt{\frac{N-n}{N-1}}.

Here, NN is the population size. The correction reflects the reduced uncertainty that results from sampling a substantial fraction of a finite population.

Takeaway: More observations generally improve precision and normal approximation, but they do not eliminate bias or justify a universal sample-size rule.

From Sampling Distributions to Statistical Inference

Sampling distributions connect probability models to confidence intervals and significance tests.

A generally has the form

estimate ± (critical value)(standard error).\text{estimate}\ \pm\ (\text{critical value})(\text{standard error}).

A larger produces a wider interval. Because increasing the sample size generally decreases the , larger samples often produce narrower intervals when other conditions remain the same.

A 95%95\% is designed so that, over many repetitions of the sampling procedure, approximately 95%95\% of intervals produced by that procedure contain the true population parameter when the assumptions are satisfied. It is not correct to say that one completed interval has a 95%95\% probability of containing a fixed parameter.

A compares an observed statistic with the expected under a null hypothesis. For a mean, a standardized test statistic can be written as

t=xˉ−μ0SE(xˉ),t=\frac{\bar{x}-\mu_0}{SE(\bar{x})},

where μ0\mu_0 is the null-hypothesis mean. A result farther from the null value in standard-error units is less compatible with the null model, subject to the assumptions of the procedure.

The supports approximately normal reference distributions for many statistics when the sampling design and other assumptions are appropriate. Nevertheless, statistical significance does not by itself establish practical importance, causation, or freedom from bias.

Takeaway: Sampling distributions quantify uncertainty, while confidence intervals and significance tests use that uncertainty to support carefully qualified conclusions.

Worked Example: Mean Waiting Time

Consider a population of waiting times with mean μ=8\mu=8 minutes and standard deviation σ=6\sigma=6 minutes. A random sample of n=36n=36 waiting times is selected.

First, find the center and spread of the :

E(Xˉ)=8E(\bar{X})=8

and

SE(Xˉ)=636=1 minute.SE(\bar{X})=\frac{6}{\sqrt{36}}=1\text{ minute}.

Because the sample size is reasonably large, the suggests the approximation

Xˉ≈N(8,1),\bar{X}\approx N(8,1),

where the second value represents the standard deviation of the .

Suppose the observed is xˉ=10\bar{x}=10 minutes. Its standardized distance from the population mean is

z=10−81=2.z=\frac{10-8}{1}=2.

Thus, the observed mean is 22 standard errors above the population mean. This is noticeably higher than expected from ordinary sampling variation, but a complete probability calculation or formal test is needed to assess how unusual it is.

The result does not prove that the population mean changed. It could be produced by variation. A responsible conclusion considers the sampling design, the model assumptions, the , the research question, and plausible alternative explanations.

Takeaway: Standardization expresses an observed statistic in units of its sampling variability, making it possible to compare the observation with a reference distribution.

Key Checks and Final Synthesis

Keep the following distinctions in view:

  1. A sample distribution displays the observed data values in one sample; a displays a statistic across many possible samples.

  2. The standard deviation of individual observations is σ\sigma, whereas the standard deviation of sample means is σ/n\sigma/\sqrt{n}.

  3. The does not make biased data valid.

  4. Doubling nn does not halve the ; a fourfold increase is needed to halve it.

  5. The condition n≥30n\ge 30 is not a universal guarantee of a normal approximation.

  6. The usual formulas rely on independence or an appropriate dependence structure. Clustered, repeated, or time-series observations may require different methods.

  7. Statistical significance is not the same as practical importance, causation, or freedom from bias.

The main chain of reasoning is:

  1. Select observations using an appropriate sampling design.

  2. Identify the statistic and its .

  3. Determine the center and .

  4. Check whether a normal or other approximation is justified.

  5. Use the to construct an interval or evaluate a test statistic.

  6. Interpret the result in light of bias, dependence, assumptions, and practical importance.

Final takeaway: Sampling distributions turn sample-to-sample variation into a formal basis for estimating population quantities and evaluating statistical evidence.