4 Sampling and Sampling Distributions
Learn how sampling designs affect the information a sample provides, how statistics vary across repeated samples, and when the central limit theorem supports normal approximations.
Populations, samples, and statistics
A is the full group of interest, while a is the subset observed. A characteristic, such as its mean or proportion, is a ; a number calculated from a is a . Statistics use data to learn about populations, so how a is selected matters.
Sampling methods
Sampling methods determine which members can enter a study. Probability-based methods use random selection and allow sampling uncertainty to be quantified.
Simple random sampling: Select individuals at random so each possible of the chosen size has an equal chance of selection. For example, randomly draw customer IDs from a complete customer list.
Systematic sampling: Choose a random starting point in an ordered list, then select every th individual. For example, survey every th order after choosing a random starting order. A repeating pattern in the list can bias results.
Stratified sampling: Divide the into meaningful subgroups, or strata, then randomly from each. For example, a company might customers from every region to ensure each region is represented.
Cluster sampling: Divide the into groups, or clusters, randomly select some clusters, then survey everyone—or a of people—in those clusters. For example, a retailer might randomly select several stores and survey their customers.
Convenience sampling: Select people who are easiest to reach, such as customers who happen to visit one store. It is quick, but may systematically miss parts of the .
Bias and coverage
Nonrandom methods, including convenience sampling, can produce biased results, and a larger alone will not fix that bias. A may also be affected by nonresponse or incomplete coverage of the . Better sampling design limits bias; increasing size generally reduces random sampling variability but cannot by itself repair a biased design.
Sampling distributions
A varies from to . Its describes the probability distribution of that across all possible samples of a fixed size drawn using the same method. It is not the distribution of individual observations in one .
A from one is only one possible result. Its describes its expected value and variability, providing a basis for estimating quantities and judging how much confidence to place in a business decision.
Standard errors and size
The standard deviation of a ’s is its . For independent random observations from a with mean and standard deviation , the mean has mean and standard error :
If is unknown, the standard error for the mean is commonly estimated by , where is the standard deviation. Increasing size reduces the standard error at a rate proportional to ; for example, quadrupling the size halves it.
For a proportion , based on independent observations from a with proportion ,
When sampling without replacement from a finite , observations are not fully independent. If the is a substantial fraction of the , a reduces the standard error. When the is a small fraction, this adjustment is often negligible.
The central limit theorem
The explains why normal distributions are useful in statistical inference. Under suitable conditions—most commonly, independent observations from the same with finite variance—the of the mean becomes approximately normal as size grows, even if the itself is not normally distributed. If the is normal, the mean is normally distributed for any size.
When the approximation is appropriate, the mean has an approximately normal with mean and standard deviation :
The approximation generally improves with larger samples, but there is no universal -size cutoff. Strongly skewed populations, extreme outliers, dependence, or other unusual features can require more care. The CLT concerns the distribution of the across repeated samples; it does not say that the individual values in a large become normally distributed.
Interpreting a -mean example
Suppose order values have mean and standard deviation . For random independent samples of size , the mean has mean and standard error
If the central limit theorem approximation is reasonable, the distribution of -average order values is approximately normal with mean and standard deviation . This describes how averages would vary across repeated samples, not how individual order values are distributed.