09 Inference for Means and Proportions

A structured guide to selecting, calculating, and interpreting confidence intervals and hypothesis tests for proportions, means, independent groups, and paired observations.

The common framework

Statistical inference uses sample data to estimate population parameters and evaluate claims about those parameters. The response variable determines the type of parameter:

  • A binary response leads to a population proportion, such as pp, or a difference in proportions, p1−p2p_1-p_2.

  • A quantitative response leads to a population mean, such as μ\mu, a difference in means, such as μ1−μ2\mu_1-\mu_2, or a mean paired difference, μd\mu_d.

A gives one best numerical estimate. A gives a range of plausible parameter values. A Hypothesis test evaluates evidence against a specific null claim. These methods share the following structures.

For an interval:

estimate±(critical value)(standard error).\text{estimate} \pm (\text{critical value})(\text{standard error}).

For a test:

test statistic=observed statistic−null valuestandard error under H0.\text{test statistic}=\frac{\text{observed statistic}-\text{null value}}{\text{standard error under }H_0}.

A statistically significant result is not automatically practically important. Interpretation should consider the estimated effect, its uncertainty, and the context of the problem.

Takeaway: Identify the parameter and decide whether the goal is estimation or evaluation of a claim before selecting a formula.

Conditions for reliable inference

Before calculating an interval or test statistic, verify that the design and data support the procedure.

  • Randomness: Data should come from a random sample or randomized experiment when the goal is to generalize or make causal conclusions.

  • Independence: Observations should be independent unless the design intentionally creates pairs or repeated measurements.

  • 10% condition: When sampling without replacement from a finite population, the sample should generally be no more than 10%10\% of the population.

  • Normality for means: A one-sample or two-sample tt procedure is most reliable when the population is approximately normal or the sample is large enough for the sampling distribution of the mean to be approximately normal. Strong skewness and outliers are especially concerning for small samples.

  • Success-failure conditions: A normal approximation for a sample proportion requires sufficiently many expected successes and failures. A common check is

np≥10andn(1−p)≥10.np\ge 10\quad\text{and}\quad n(1-p)\ge 10.

For a hypothesis test, use the null proportion when checking these counts; for an interval, use the sample proportion when appropriate. If the conditions fail, consider an exact binomial method or a small-sample interval such as the Wilson interval.

Takeaway: A correct formula cannot repair a poor design, dependence, severe outliers, or inadequate sample-size conditions.

One-sample inference for a proportion

For one binary population, let XX be the number of successes in a sample of size nn. The sample proportion is

p^=Xn.\hat p=\frac{X}{n}.

The parameter of interest is the population proportion pp.

Estimating one proportion

For a sufficiently large sample, an approximate two-sided is

p^±z∗p^(1−p^)n.\hat p\pm z^*\sqrt{\frac{\hat p(1-\hat p)}{n}}.

For a 95%95\% interval, z∗≈1.96z^*\approx 1.96. The correct interpretation is procedural: over many repetitions using the same method, approximately 95%95\% of the intervals would contain the true value of pp. It is not correct to say that the fixed parameter has a 95%95\% probability of being in one already-calculated interval.

For example, if 220220 of 400400 sampled voters support a proposal, then p^=0.55\hat p=0.55. The approximate 95%95\% interval is

0.55±1.960.55(0.45)400≈0.55±0.0487,0.55\pm1.96\sqrt{\frac{0.55(0.45)}{400}}\approx 0.55\pm0.0487,

or approximately (0.501,0.599)(0.501,0.599).

Testing one proportion

To test

H0:p=p0H_0:p=p_0

against the two-sided alternative Ha:p≠p0H_a:p\ne p_0, use

z=p^−p0p0(1−p0)n.z=\frac{\hat p-p_0}{\sqrt{\frac{p_0(1-p_0)}{n}}}.

The null value p0p_0, rather than p^\hat p, appears in the because the test assumes that the null hypothesis is true. Report the test statistic and its , then compare the with the chosen significance level α\alpha. Reject H0H_0 when the is less than or equal to α\alpha.

Takeaway: Use p^\hat p in the usual interval , but use p0p_0 in the for a one-proportion hypothesis test.

One-sample inference for a mean

For one quantitative population, let xˉ\bar x be the sample mean, ss the sample standard deviation, and nn the sample size. The parameter is the population mean μ\mu.

Estimating one mean

When the population standard deviation is unknown, use a . A is

xˉ±tα/2, n−1∗sn,\bar x\pm t^*_{\alpha/2,\,n-1}\frac{s}{\sqrt n},

where the critical value comes from a tt distribution with n−1n-1 degrees of freedom. The tt distribution accounts for the additional uncertainty caused by estimating the population standard deviation with ss.

Testing one mean

To test

H0:μ=μ0H_0:\mu=\mu_0

against Ha:μ≠μ0H_a:\mu\ne\mu_0, use

t=xˉ−μ0s/n,t=\frac{\bar x-\mu_0}{s/\sqrt n},

with n−1n-1 degrees of freedom. A one-sided alternative uses the corresponding one-sided tail.

For example, suppose a random sample has n=25n=25, xˉ=503\bar x=503 mL, and s=8s=8 mL, while the claimed mean is μ0=500\mu_0=500 mL. Then

t=503−5008/25=1.875.t=\frac{503-500}{8/\sqrt{25}}=1.875.

Whether this is statistically significant depends on the alternative hypothesis and the resulting .

Takeaway: For a mean with unknown population standard deviation, use the sample standard deviation and a tt distribution rather than treating the standard deviation as known.

Comparing two independent proportions

Suppose two independent groups have sample proportions p^1\hat p_1 and p^2\hat p_2, based on sample sizes n1n_1 and n2n_2. The parameter is p1−p2p_1-p_2, and the is p^1−p^2\hat p_1-\hat p_2.

for a difference in proportions

An approximate interval is

(p^1−p^2)±z∗p^1(1−p^1)n1+p^2(1−p^2)n2.(\hat p_1-\hat p_2)\pm z^*\sqrt{\frac{\hat p_1(1-\hat p_1)}{n_1}+\frac{\hat p_2(1-\hat p_2)}{n_2}}.

This is generally unpooled: each sample proportion is estimated separately because the interval allows the true difference to be any plausible value.

Hypothesis test for equal proportions

For

H0:p1−p2=0,H_0:p_1-p_2=0,

the null hypothesis says that the two population proportions are equal. Combine the samples to estimate their common proportion:

p^pool=x1+x2n1+n2.\hat p_{\text{pool}}=\frac{x_1+x_2}{n_1+n_2}.

The is

SE0=p^pool(1−p^pool)(1n1+1n2),SE_0=\sqrt{\hat p_{\text{pool}}(1-\hat p_{\text{pool}})\left(\frac{1}{n_1}+\frac{1}{n_2}\right)},

and the test statistic is

z=p^1−p^2SE0.z=\frac{\hat p_1-\hat p_2}{SE_0}.

For example, if 7272 of 120120 customers in Group 1 renew and 5454 of 110110 customers in Group 2 renew, then p^1=0.600\hat p_1=0.600, p^2≈0.491\hat p_2\approx0.491, and the observed difference is approximately 0.1090.109. A test of equal renewal rates uses the pooled estimate, while an interval for the difference generally uses the separate sample proportions.

Takeaway: For two proportions, confidence intervals are generally unpooled, whereas tests of an equal-proportion null use a .

Comparing two independent means

For two independent quantitative groups, summarize the samples with xˉ1,s1,n1\bar x_1,s_1,n_1 and xˉ2,s2,n2\bar x_2,s_2,n_2. The parameter is usually μ1−μ2\mu_1-\mu_2, estimated by xˉ1−xˉ2\bar x_1-\bar x_2.

Welch's unpooled method

The does not assume equal population variances. Its is

SEWelch=s12n1+s22n2.SE_{\text{Welch}}=\sqrt{\frac{s_1^2}{n_1}+\frac{s_2^2}{n_2}}.

A is

(xˉ1−xˉ2)±t∗SEWelch,(\bar x_1-\bar x_2)\pm t^*SE_{\text{Welch}},

with degrees of freedom estimated by

ν≈(s12/n1+s22/n2)2(s12/n1)2n1−1+(s22/n2)2n2−1.\nu\approx\frac{\left(s_1^2/n_1+s_2^2/n_2\right)^2}{\dfrac{(s_1^2/n_1)^2}{n_1-1}+\dfrac{(s_2^2/n_2)^2}{n_2-1}}.

For a null difference Δ0\Delta_0, the test statistic is

t=(xˉ1−xˉ2)−Δ0s12/n1+s22/n2.t=\frac{(\bar x_1-\bar x_2)-\Delta_0}{\sqrt{s_1^2/n_1+s_2^2/n_2}}.

Pooled method

A pooled procedure assumes a common population variance:

σ12=σ22=σ2.\sigma_1^2=\sigma_2^2=\sigma^2.

The pooled variance estimate is

sp2=(n1−1)s12+(n2−1)s22n1+n2−2,s_p^2=\frac{(n_1-1)s_1^2+(n_2-1)s_2^2}{n_1+n_2-2},

and the is

SEpool=sp1n1+1n2.SE_{\text{pool}}=s_p\sqrt{\frac{1}{n_1}+\frac{1}{n_2}}.

The corresponding test statistic is

t=(xˉ1−xˉ2)−Δ0sp1/n1+1/n2,t=\frac{(\bar x_1-\bar x_2)-\Delta_0}{s_p\sqrt{1/n_1+1/n_2}},

with n1+n2−2n_1+n_2-2 degrees of freedom. Use the pooled method when equal variances are scientifically justified; otherwise, Welch's method is generally the safer default.

Takeaway: Independent means require attention to variance assumptions. Welch's unpooled method is usually preferred when those assumptions are uncertain.

Paired observations

When observations are naturally matched, calculate one difference for each pair. Examples include before-and-after measurements on the same person, two measurements from the same experimental unit, or matched subjects.

Define the difference consistently, for example:

di=afteri−beforei.d_i=\text{after}_i-\text{before}_i.

Then analyze the nn differences as a one-sample mean problem. Let dˉ\bar d and sds_d be the mean and standard deviation of the differences.

A for the population mean difference μd\mu_d is

dˉ±tα/2, n−1∗sdn.\bar d\pm t^*_{\alpha/2,\,n-1}\frac{s_d}{\sqrt n}.

To test whether the mean difference is zero, use

H0:μd=0H_0:\mu_d=0

and

t=dˉsd/n,t=\frac{\bar d}{s_d/\sqrt n},

with n−1n-1 degrees of freedom. The normality condition concerns the differences, not necessarily the separate before-and-after measurements.

Pairing can reduce variation because each subject or unit serves as its own control. For example, blood pressure measured before and after treatment for each patient should usually be analyzed with a rather than an independent two-sample procedure.

Takeaway: Preserve the matching, compute within-pair differences, and perform one-sample tt inference on those differences.

Choosing the procedure

Use the following decision sequence to select a procedure.

  1. Identify the response type. A binary response produces a proportion; a quantitative response produces a mean.

  2. Count the groups or conditions. Decide whether there is one group, two independent groups, or one paired group.

  3. Check dependence. Measurements from the same subjects or matched units are paired rather than independent.

  4. Name the parameter. Common choices are pp, p1−p2p_1-p_2, μ\mu, μ1−μ2\mu_1-\mu_2, and μd\mu_d.

  5. Check assumptions. Examine randomness, independence, normality, outliers, and success-failure counts.

  6. Choose the objective. Use a to estimate a parameter or a hypothesis test to evaluate a specific claim.

  7. Interpret in context. State the estimated effect or interval, its direction, the or confidence level, and its practical meaning.

Useful matches include:

  • One binary sample and a claimed proportion: one-proportion zz test or exact binomial test.

  • One binary sample and estimation: one-proportion .

  • Two independent binary samples: two-proportion test or interval.

  • One quantitative sample: .

  • Two independent quantitative samples: Welch procedure by default; pooled procedure only with justified equal variances.

  • Paired quantitative observations: on the differences.

Takeaway: The correct method is determined mainly by the response type, number of groups, and dependence structure—not merely by the number of columns in a data table.

Interpreting intervals and tests

For a two-sided test at significance level α\alpha, a corresponding two-sided at level 100(1−α)%100(1-\alpha)\% often gives the same decision:

  • If the null value lies outside the interval, reject the null hypothesis.

  • If the null value lies inside the interval, do not reject the null hypothesis.

For example, a 95%95\% for μ1−μ2\mu_1-\mu_2 of (2.1,7.4)(2.1,7.4) excludes zero. At the corresponding 5%5\% significance level, the data support a difference in population means. The interval also indicates that the estimated difference is positive and gives plausible values for its size.

Do not say that a large proves the null hypothesis. The appropriate conclusion is that the data do not provide sufficient evidence against the null hypothesis. Likewise, statistical significance does not establish practical importance; consider the magnitude and consequences of the effect.

Final takeaway: Confidence intervals quantify plausible effect sizes, while hypothesis tests assess evidence against a null claim. Both conclusions depend on the study design, assumptions, and context.