09 Inference for Means and Proportions
A structured guide to selecting, calculating, and interpreting confidence intervals and hypothesis tests for proportions, means, independent groups, and paired observations.
The common framework
Statistical inference uses sample data to estimate population parameters and evaluate claims about those parameters. The response variable determines the type of parameter:
A binary response leads to a population proportion, such as , or a difference in proportions, .
A quantitative response leads to a population mean, such as , a difference in means, such as , or a mean paired difference, .
A gives one best numerical estimate. A gives a range of plausible parameter values. A Hypothesis test evaluates evidence against a specific null claim. These methods share the following structures.
For an interval:
For a test:
A statistically significant result is not automatically practically important. Interpretation should consider the estimated effect, its uncertainty, and the context of the problem.
Takeaway: Identify the parameter and decide whether the goal is estimation or evaluation of a claim before selecting a formula.
Conditions for reliable inference
Before calculating an interval or test statistic, verify that the design and data support the procedure.
Randomness: Data should come from a random sample or randomized experiment when the goal is to generalize or make causal conclusions.
Independence: Observations should be independent unless the design intentionally creates pairs or repeated measurements.
10% condition: When sampling without replacement from a finite population, the sample should generally be no more than of the population.
Normality for means: A one-sample or two-sample procedure is most reliable when the population is approximately normal or the sample is large enough for the sampling distribution of the mean to be approximately normal. Strong skewness and outliers are especially concerning for small samples.
Success-failure conditions: A normal approximation for a sample proportion requires sufficiently many expected successes and failures. A common check is
For a hypothesis test, use the null proportion when checking these counts; for an interval, use the sample proportion when appropriate. If the conditions fail, consider an exact binomial method or a small-sample interval such as the Wilson interval.
Takeaway: A correct formula cannot repair a poor design, dependence, severe outliers, or inadequate sample-size conditions.
One-sample inference for a proportion
For one binary population, let be the number of successes in a sample of size . The sample proportion is
The parameter of interest is the population proportion .
Estimating one proportion
For a sufficiently large sample, an approximate two-sided is
For a interval, . The correct interpretation is procedural: over many repetitions using the same method, approximately of the intervals would contain the true value of . It is not correct to say that the fixed parameter has a probability of being in one already-calculated interval.
For example, if of sampled voters support a proposal, then . The approximate interval is
or approximately .
Testing one proportion
To test
against the two-sided alternative , use
The null value , rather than , appears in the because the test assumes that the null hypothesis is true. Report the test statistic and its , then compare the with the chosen significance level . Reject when the is less than or equal to .
Takeaway: Use in the usual interval , but use in the for a one-proportion hypothesis test.
One-sample inference for a mean
For one quantitative population, let be the sample mean, the sample standard deviation, and the sample size. The parameter is the population mean .
Estimating one mean
When the population standard deviation is unknown, use a . A is
where the critical value comes from a distribution with degrees of freedom. The distribution accounts for the additional uncertainty caused by estimating the population standard deviation with .
Testing one mean
To test
against , use
with degrees of freedom. A one-sided alternative uses the corresponding one-sided tail.
For example, suppose a random sample has , mL, and mL, while the claimed mean is mL. Then
Whether this is statistically significant depends on the alternative hypothesis and the resulting .
Takeaway: For a mean with unknown population standard deviation, use the sample standard deviation and a distribution rather than treating the standard deviation as known.
Comparing two independent proportions
Suppose two independent groups have sample proportions and , based on sample sizes and . The parameter is , and the is .
for a difference in proportions
An approximate interval is
This is generally unpooled: each sample proportion is estimated separately because the interval allows the true difference to be any plausible value.
Hypothesis test for equal proportions
For
the null hypothesis says that the two population proportions are equal. Combine the samples to estimate their common proportion:
The is
and the test statistic is
For example, if of customers in Group 1 renew and of customers in Group 2 renew, then , , and the observed difference is approximately . A test of equal renewal rates uses the pooled estimate, while an interval for the difference generally uses the separate sample proportions.
Takeaway: For two proportions, confidence intervals are generally unpooled, whereas tests of an equal-proportion null use a .
Comparing two independent means
For two independent quantitative groups, summarize the samples with and . The parameter is usually , estimated by .
Welch's unpooled method
The does not assume equal population variances. Its is
A is
with degrees of freedom estimated by
For a null difference , the test statistic is
Pooled method
A pooled procedure assumes a common population variance:
The pooled variance estimate is
and the is
The corresponding test statistic is
with degrees of freedom. Use the pooled method when equal variances are scientifically justified; otherwise, Welch's method is generally the safer default.
Takeaway: Independent means require attention to variance assumptions. Welch's unpooled method is usually preferred when those assumptions are uncertain.
Paired observations
When observations are naturally matched, calculate one difference for each pair. Examples include before-and-after measurements on the same person, two measurements from the same experimental unit, or matched subjects.
Define the difference consistently, for example:
Then analyze the differences as a one-sample mean problem. Let and be the mean and standard deviation of the differences.
A for the population mean difference is
To test whether the mean difference is zero, use
and
with degrees of freedom. The normality condition concerns the differences, not necessarily the separate before-and-after measurements.
Pairing can reduce variation because each subject or unit serves as its own control. For example, blood pressure measured before and after treatment for each patient should usually be analyzed with a rather than an independent two-sample procedure.
Takeaway: Preserve the matching, compute within-pair differences, and perform one-sample inference on those differences.
Choosing the procedure
Use the following decision sequence to select a procedure.
Identify the response type. A binary response produces a proportion; a quantitative response produces a mean.
Count the groups or conditions. Decide whether there is one group, two independent groups, or one paired group.
Check dependence. Measurements from the same subjects or matched units are paired rather than independent.
Name the parameter. Common choices are , , , , and .
Check assumptions. Examine randomness, independence, normality, outliers, and success-failure counts.
Choose the objective. Use a to estimate a parameter or a hypothesis test to evaluate a specific claim.
Interpret in context. State the estimated effect or interval, its direction, the or confidence level, and its practical meaning.
Useful matches include:
One binary sample and a claimed proportion: one-proportion test or exact binomial test.
One binary sample and estimation: one-proportion .
Two independent binary samples: two-proportion test or interval.
One quantitative sample: .
Two independent quantitative samples: Welch procedure by default; pooled procedure only with justified equal variances.
Paired quantitative observations: on the differences.
Takeaway: The correct method is determined mainly by the response type, number of groups, and dependence structure—not merely by the number of columns in a data table.
Interpreting intervals and tests
For a two-sided test at significance level , a corresponding two-sided at level often gives the same decision:
If the null value lies outside the interval, reject the null hypothesis.
If the null value lies inside the interval, do not reject the null hypothesis.
For example, a for of excludes zero. At the corresponding significance level, the data support a difference in population means. The interval also indicates that the estimated difference is positive and gives plausible values for its size.
Do not say that a large proves the null hypothesis. The appropriate conclusion is that the data do not provide sufficient evidence against the null hypothesis. Likewise, statistical significance does not establish practical importance; consider the magnitude and consequences of the effect.
Final takeaway: Confidence intervals quantify plausible effect sizes, while hypothesis tests assess evidence against a null claim. Both conclusions depend on the study design, assumptions, and context.