06 Sampling and Sampling Distributions
A structured guide to populations, sampling methods, sampling distributions, standard errors, Normal approximations, and the quality of sample-based estimates.
Populations, Samples, and Numerical Summaries
A population is the complete set of individuals or measurements of interest. A sample is a subset selected from that population.
A numerical characteristic of a population is a , while a numerical summary calculated from a sample is a . For example, the population mean is , while the sample mean is . Because different samples contain different observations, a generally changes from sample to sample and is treated as a random variable before the sample is observed.
Why the distinction matters
Statistical inference uses the behavior of sample statistics to learn about unknown population parameters. A sample does not usually reproduce the population perfectly, so conclusions must account for sampling variability.
Takeaway: Parameters describe populations; statistics describe samples and help estimate population parameters.
Choosing a
The goal of a well-designed is to give population members a known and appropriate chance of selection while reducing systematic . The design should match the research question, the structure of the population, and available resources.
Simple random sampling: Every possible sample of the specified size has an equal chance of selection. For example, a random-number generator could select 100 distinct students from a numbered list of 2,000 students.
Stratified sampling: The population is divided into meaningful, nonoverlapping strata, and a random sample is selected from each stratum. This helps ensure representation of important subgroups.
Cluster sampling: The population is divided into naturally occurring clusters. A random sample of clusters is selected, and researchers survey every member or take a further sample within selected clusters. This can reduce cost, but members of the same cluster may be similar.
Systematic sampling: After a random starting point, researchers select every th member of an ordered list. With population size and sample size , a common interval is approximately .
Convenience samples include people who are easiest to reach. Voluntary-response samples include people who choose whether to participate. Both may be affected by selection because participants can differ systematically from nonparticipants. Periodic patterns in an ordered list can also create problems for systematic sampling.
A larger sample reduces random variation but does not, by itself, correct a biased sampling design.
Takeaway: The sampling design affects whether an estimate represents the target population; sample size alone cannot eliminate systematic .
Understanding Sampling Distributions
A describes the probability distribution of a over all possible random samples of a specified size from a population. It is different from the distribution of individual observations in one sample.
Three features are especially important:
Center: The typical value of the .
Spread: How much the varies from sample to sample.
Shape: Whether the distribution is symmetric, skewed, bell-shaped, or otherwise structured.
For an unbiased estimator, the center of the equals the being estimated. For the sample mean,
Thus, repeated sample means may fall above or below , but they are centered at over repeated sampling.
For example, if a population has mean , repeatedly taking random samples of size and calculating produces a distribution of sample means centered near 50. Some means will be close to 50 and others will be farther away.
A population distribution describes individual population values, a sample distribution describes individual values within one sample, and a describes a computed across many samples.
Takeaway: A describes the long-run behavior of a , including its center, spread, and shape.
The Sample Mean and Its
Let be the mean of a random sample of size from a population with mean and standard deviation . When observations are independent or approximately independent,
and
The standard deviation of this is the :
When is unknown, it is commonly estimated with the sample standard deviation :
The square-root relationship has an important consequence: multiplying the sample size by four cuts the in half. If sampling without replacement takes a substantial fraction of a finite population, a finite population correction may be appropriate:
Here, is the population size. The correction reflects the reduced uncertainty that results when a large fraction of a finite population has already been sampled.
Takeaway: Larger samples make sample means less variable, but precision improves according to a square-root relationship.
Normal Approximations for Sample Means
The explains why sample means often have an approximately Normal . Under suitable conditions, as increases,
The corresponding standardized is
If the population itself is Normal, the of is Normal for every sample size when the usual independence conditions hold. For a non-Normal population, the approximation depends on sample size and population shape.
Before applying the theorem, consider:
Randomness: The sample should come from a probability-based method.
Independence: Observations should be independent, or the sampling fraction should be small enough for dependence from sampling without replacement to be negligible.
Population shape and sample size: Mild skewness may be acceptable with moderate sample sizes, while heavy skewness, outliers, or heavy tails may require much larger samples.
The rule of thumb is only a rough guideline. It does not apply equally to every population or .
Takeaway: Normal approximations require conditions; sample size must be considered together with randomness, independence, and population shape.
Sampling Distributions for Proportions
For a population proportion , let represent the proportion of successes in a random sample of size . Its center and are
and
When the expected numbers of successes and failures are sufficiently large, commonly checked by
the is approximately Normal:
Because is often unknown, an estimated can use the observed :
The same three ideas—center, spread, and shape—organize the interpretation of a 's .
Takeaway: Sample proportions are centered at the population proportion, become less variable with larger samples, and are approximately Normal when both expected counts are sufficiently large.
, Precision, and Estimator Quality
The quality of an estimator involves more than its center. is the difference between an estimator's expected value and the it is intended to estimate:
A is unbiased when its is centered at the target . Under random sampling, the sample mean is unbiased for the population mean, and the is unbiased for the population proportion.
Sampling variability describes the spread of an estimator's . A smaller means repeated samples tend to produce more similar estimates. Larger samples generally reduce this variability because averaging smooths some variation in individual observations.
An estimator is if it tends to get closer to its target as the sample size increases. For the sample mean,
Random sampling reduces random error, but it does not automatically fix undercoverage, nonresponse, leading questions, measurement problems, or a flawed sampling frame.
Worked calculation
Suppose , , and . Then
If the approximation conditions are satisfied, . A sample mean of 84 has standardized value
Therefore, 84 is two standard errors above the population mean in the . This statement concerns the sample mean, not every individual observation.
Takeaway: Reliable estimation requires both low sampling variability and a design that limits systematic .