10. Significance Tests and Statistical Inference

A structured guide to formulating hypotheses, carrying out significance tests, interpreting p-values, evaluating decision errors and power, and distinguishing statistical evidence from practical importance.

Formulating Hypotheses and Test Direction

Significance testing uses sample data to evaluate a claim about a population. The central question is whether the observed result would be unusual if a specified baseline claim were true. The procedure assesses the strength of evidence against a null model; it does not prove a hypothesis.

A hypothesis must concern a population parameter rather than only a sample statistic. Common parameters include:

  • A population proportion, pp.

  • A population mean, μ\mu.

  • A difference between two population proportions, p1−p2p_1-p_2.

  • A difference between two population means, μ1−μ2\mu_1-\mu_2.

The contains equality and gives the reference value. The states the direction or form of departure that would support the research claim.

For example, to ask whether a population proportion differs from 0.500.50, use H0:p=0.50H_0:p=0.50 and Ha:p≠0.50H_a:p\ne0.50. To ask whether it exceeds 0.500.50, use Ha:p>0.50H_a:p>0.50. To ask whether a population mean is less than 7070, use H0:μ=70H_0:\mu=70 and Ha:μ<70H_a:\mu<70.

The alternative must be chosen before examining the results. A two-sided alternative is appropriate when departures in either direction matter; a one-sided alternative is appropriate when only one direction is relevant.

Takeaway: Define the population parameter, include equality in the , and choose the alternative direction in advance.

The Significance-Test Process

A significance test follows a connected sequence of decisions:

  1. State the null and alternative hypotheses in terms of the population parameter.

  2. Calculate a that measures the distance between the sample result and the null value in standard-error units.

  3. Find and interpret the using the appropriate reference distribution and the direction of the .

  4. Compare the with the selected α\alpha.

  5. State a conclusion about the population parameter in the context of the problem.

The general structure is:

test statistic=observed statistic−null valuestandard error under H0.\text{test statistic}=\frac{\text{observed statistic}-\text{null value}}{\text{standard error under }H_0}.

A large in magnitude indicates that the sample result is far from the null value relative to the expected sampling variability. The exact meaning of “large” depends on the test and whether the alternative is left-tailed, right-tailed, or two-sided.

The decision rule is:

  • If the is less than or equal to α\alpha, reject H0H_0.

  • If the is greater than α\alpha, fail to reject H0H_0.

“Fail to reject” is more accurate than “accept.” A nonsignificant result may occur because the null is reasonable, but it may also reflect a small sample, high variability, or low-quality measurements.

Takeaway: A decision based on a must be followed by a contextual conclusion, not just the words “significant” or “not significant.”

Testing a Population Proportion

For a one-sample test of a population proportion, let p^\hat p be the sample proportion, let p0p_0 be the null value, and let nn be the sample size. The standard normal statistic is:

z=p^−p0p0(1−p0)n.z=\frac{\hat p-p_0}{\sqrt{\frac{p_0(1-p_0)}{n}}}.

The standard error uses p0p_0, not p^\hat p, because the test evaluates the sampling distribution that would result if the were true.

For example, suppose 120120 of 200200 randomly sampled people support a policy. Then p^=120/200=0.60\hat p=120/200=0.60. To test whether more than half of the population supports it, use H0:p=0.50H_0:p=0.50 and Ha:p>0.50H_a:p>0.50. The statistic is:

z=0.60−0.500.50(0.50)200≈2.83.z=\frac{0.60-0.50}{\sqrt{\frac{0.50(0.50)}{200}}}\approx2.83.

For a right-tailed test, this gives a of about 0.0020.002. With α=0.05\alpha=0.05, reject H0H_0. The conclusion is that the sample provides statistically significant evidence that more than 50%50\% of the population supports the policy.

For the usual normal approximation, the sample should be random or representative, and the expected counts under the null should commonly satisfy:

np0≥10andn(1−p0)≥10.np_0\ge10\quad\text{and}\quad n(1-p_0)\ge10.

If these conditions fail, an exact or simulation-based method may be more appropriate.

Takeaway: In a one-proportion test, standardize the sample proportion using the null proportion and check the expected-count conditions.

Testing a Population Mean

For a one-sample test of a population mean when the population standard deviation is unknown, let xˉ\bar x be the sample mean, ss the sample standard deviation, and μ0\mu_0 the null value. The statistic is:

t=xˉ−μ0s/n,t=\frac{\bar x-\mu_0}{s/\sqrt n},

with n−1n-1 degrees of freedom. The tt-distribution accounts for the extra uncertainty from estimating the population standard deviation with ss.

Suppose a random sample of 2525 customer-support calls has mean 7272 seconds and standard deviation 1010 seconds. To test whether the population mean differs from 7070 seconds, use H0:μ=70H_0:\mu=70 and Ha:μ≠70H_a:\mu\ne70. Then:

t=72−7010/25=1.00.t=\frac{72-70}{10/\sqrt{25}}=1.00.

The test has 2424 degrees of freedom and a two-sided of approximately 0.330.33. At α=0.05\alpha=0.05, fail to reject H0H_0. The sample does not provide sufficient evidence that the population mean call time differs from 7070 seconds. This does not prove that the true mean is exactly 7070; it means the observed difference is not unusual enough, given the variability and sample size, to provide strong evidence against the .

Observations should be independent, and the sampling method should be random or reasonably representative. For small samples, the population distribution should be approximately normal and free of extreme outliers. Larger samples generally provide more robustness, although severe skewness or outliers can still cause problems.

Takeaway: A one-mean test uses a tt-statistic when the population standard deviation is unknown and requires attention to independence, sampling, and distributional conditions.

P-Values, Significance Levels, and Errors

A describes how compatible the observed data are with the null model. A of 0.030.03 means that, if the were true and the sampling process were repeated under the stated assumptions, results at least as extreme as the observed result would occur about 3%3\% of the time.

A does not represent:

  • The probability that the is true.

  • The probability that the is true.

  • The probability that the result occurred through “random chance alone.”

  • The size or practical importance of the effect.

A small provides evidence against the null model, but it does not establish that an effect is large, useful, causal, or scientifically important. Conversely, a large does not prove the ; it indicates that the data do not provide sufficiently strong evidence against it under the chosen procedure.

The α\alpha is selected before examining the data. It is the maximum tolerated probability of a , which occurs when a true is rejected. Common levels include 0.100.10, 0.050.05, and 0.010.01, but 0.050.05 is a convention rather than a universal requirement.

A occurs when a false is not rejected. Its probability is denoted by β\beta. is the probability of correctly rejecting a false , so equals 1−β1-\beta. generally increases with a larger sample, a larger true effect, lower variability, or a less stringent . Reducing α\alpha generally makes Type I errors less likely but can make Type II errors more likely, all else equal.

Takeaway: Interpret the under the null model, and consider both kinds of decision error when planning or evaluating a test.

From Statistical Evidence to Practical Meaning

and answer different questions. asks whether the data provide evidence that an effect differs from the null value relative to sampling variability. asks whether the size of the effect matters in the real application.

A very large sample can make a tiny difference statistically significant. For example, a reduction of 0.20.2 seconds in processing time among 50,00050{,}000 employees might produce a very small , while having little effect on cost, safety, or productivity. Conversely, a meaningful effect in a small or highly variable sample may fail to reach .

A complements a significance test by showing a range of plausible parameter values. It provides information about the direction, possible size, and precision of an effect that a alone does not provide. For a two-sided test at α=0.05\alpha=0.05, a corresponding 95% generally excludes the null value when the test rejects H0H_0, and includes the null value when the test fails to reject it.

When assessing practical importance, ask whether the entire is above a meaningful benefit threshold, below a meaningful harm threshold, or includes effects that are too small to matter. Report the estimated effect, an uncertainty measure such as a , the when appropriate, and the real-world context.

Takeaway: Statistical evidence should be interpreted alongside effect size, uncertainty, study design, sample size, and the consequences of the decision.

Writing a Complete Conclusion

A complete conclusion connects the calculation to the population and the real question. Include these four elements:

  1. The decision: reject or fail to reject H0H_0.

  2. The evidence: report or characterize the relative to α\alpha.

  3. The population parameter and the direction of the conclusion.

  4. The practical meaning and any important limitations.

For the policy example, a complete conclusion is:

Because the of 0.0020.002 is less than α=0.05\alpha=0.05, we reject H0H_0. The sample provides statistically significant evidence that the population proportion supporting the policy exceeds 0.500.50. The estimated proportion is 0.600.60, so the effect may also be practically meaningful, although its importance depends on the policy decision and the quality of the sampling process.

Avoid statements such as “the is proven false,” “the is true,” or “there is a 2%2\% chance that the is true.” These statements misinterpret the logic of hypothesis tests and p-values.

Use this final check:

  • Did the hypotheses concern a population parameter?

  • Was the alternative direction chosen before viewing the results?

  • Were the assumptions and conditions considered?

  • Was the correct and reference distribution used?

  • Was the interpreted as a probability under H0H_0?

  • Were statistical and distinguished?

  • Was the conclusion stated in context?

Takeaway: A strong conclusion reports both what the data support statistically and what the estimated effect means in practice.