10. Significance Tests and Statistical Inference
A structured guide to formulating hypotheses, carrying out significance tests, interpreting p-values, evaluating decision errors and power, and distinguishing statistical evidence from practical importance.
Formulating Hypotheses and Test Direction
Significance testing uses sample data to evaluate a claim about a population. The central question is whether the observed result would be unusual if a specified baseline claim were true. The procedure assesses the strength of evidence against a null model; it does not prove a hypothesis.
A hypothesis must concern a population parameter rather than only a sample statistic. Common parameters include:
A population proportion, .
A population mean, .
A difference between two population proportions, .
A difference between two population means, .
The contains equality and gives the reference value. The states the direction or form of departure that would support the research claim.
For example, to ask whether a population proportion differs from , use and . To ask whether it exceeds , use . To ask whether a population mean is less than , use and .
The alternative must be chosen before examining the results. A two-sided alternative is appropriate when departures in either direction matter; a one-sided alternative is appropriate when only one direction is relevant.
Takeaway: Define the population parameter, include equality in the , and choose the alternative direction in advance.
The Significance-Test Process
A significance test follows a connected sequence of decisions:
State the null and alternative hypotheses in terms of the population parameter.
Calculate a that measures the distance between the sample result and the null value in standard-error units.
Find and interpret the using the appropriate reference distribution and the direction of the .
Compare the with the selected .
State a conclusion about the population parameter in the context of the problem.
The general structure is:
A large in magnitude indicates that the sample result is far from the null value relative to the expected sampling variability. The exact meaning of “large” depends on the test and whether the alternative is left-tailed, right-tailed, or two-sided.
The decision rule is:
If the is less than or equal to , reject .
If the is greater than , fail to reject .
“Fail to reject” is more accurate than “accept.” A nonsignificant result may occur because the null is reasonable, but it may also reflect a small sample, high variability, or low-quality measurements.
Takeaway: A decision based on a must be followed by a contextual conclusion, not just the words “significant” or “not significant.”
Testing a Population Proportion
For a one-sample test of a population proportion, let be the sample proportion, let be the null value, and let be the sample size. The standard normal statistic is:
The standard error uses , not , because the test evaluates the sampling distribution that would result if the were true.
For example, suppose of randomly sampled people support a policy. Then . To test whether more than half of the population supports it, use and . The statistic is:
For a right-tailed test, this gives a of about . With , reject . The conclusion is that the sample provides statistically significant evidence that more than of the population supports the policy.
For the usual normal approximation, the sample should be random or representative, and the expected counts under the null should commonly satisfy:
If these conditions fail, an exact or simulation-based method may be more appropriate.
Takeaway: In a one-proportion test, standardize the sample proportion using the null proportion and check the expected-count conditions.
Testing a Population Mean
For a one-sample test of a population mean when the population standard deviation is unknown, let be the sample mean, the sample standard deviation, and the null value. The statistic is:
with degrees of freedom. The -distribution accounts for the extra uncertainty from estimating the population standard deviation with .
Suppose a random sample of customer-support calls has mean seconds and standard deviation seconds. To test whether the population mean differs from seconds, use and . Then:
The test has degrees of freedom and a two-sided of approximately . At , fail to reject . The sample does not provide sufficient evidence that the population mean call time differs from seconds. This does not prove that the true mean is exactly ; it means the observed difference is not unusual enough, given the variability and sample size, to provide strong evidence against the .
Observations should be independent, and the sampling method should be random or reasonably representative. For small samples, the population distribution should be approximately normal and free of extreme outliers. Larger samples generally provide more robustness, although severe skewness or outliers can still cause problems.
Takeaway: A one-mean test uses a -statistic when the population standard deviation is unknown and requires attention to independence, sampling, and distributional conditions.
P-Values, Significance Levels, and Errors
A describes how compatible the observed data are with the null model. A of means that, if the were true and the sampling process were repeated under the stated assumptions, results at least as extreme as the observed result would occur about of the time.
A does not represent:
The probability that the is true.
The probability that the is true.
The probability that the result occurred through “random chance alone.”
The size or practical importance of the effect.
A small provides evidence against the null model, but it does not establish that an effect is large, useful, causal, or scientifically important. Conversely, a large does not prove the ; it indicates that the data do not provide sufficiently strong evidence against it under the chosen procedure.
The is selected before examining the data. It is the maximum tolerated probability of a , which occurs when a true is rejected. Common levels include , , and , but is a convention rather than a universal requirement.
A occurs when a false is not rejected. Its probability is denoted by . is the probability of correctly rejecting a false , so equals . generally increases with a larger sample, a larger true effect, lower variability, or a less stringent . Reducing generally makes Type I errors less likely but can make Type II errors more likely, all else equal.
Takeaway: Interpret the under the null model, and consider both kinds of decision error when planning or evaluating a test.
From Statistical Evidence to Practical Meaning
and answer different questions. asks whether the data provide evidence that an effect differs from the null value relative to sampling variability. asks whether the size of the effect matters in the real application.
A very large sample can make a tiny difference statistically significant. For example, a reduction of seconds in processing time among employees might produce a very small , while having little effect on cost, safety, or productivity. Conversely, a meaningful effect in a small or highly variable sample may fail to reach .
A complements a significance test by showing a range of plausible parameter values. It provides information about the direction, possible size, and precision of an effect that a alone does not provide. For a two-sided test at , a corresponding 95% generally excludes the null value when the test rejects , and includes the null value when the test fails to reject it.
When assessing practical importance, ask whether the entire is above a meaningful benefit threshold, below a meaningful harm threshold, or includes effects that are too small to matter. Report the estimated effect, an uncertainty measure such as a , the when appropriate, and the real-world context.
Takeaway: Statistical evidence should be interpreted alongside effect size, uncertainty, study design, sample size, and the consequences of the decision.
Writing a Complete Conclusion
A complete conclusion connects the calculation to the population and the real question. Include these four elements:
The decision: reject or fail to reject .
The evidence: report or characterize the relative to .
The population parameter and the direction of the conclusion.
The practical meaning and any important limitations.
For the policy example, a complete conclusion is:
Because the of is less than , we reject . The sample provides statistically significant evidence that the population proportion supporting the policy exceeds . The estimated proportion is , so the effect may also be practically meaningful, although its importance depends on the policy decision and the quality of the sampling process.
Avoid statements such as “the is proven false,” “the is true,” or “there is a chance that the is true.” These statements misinterpret the logic of hypothesis tests and p-values.
Use this final check:
Did the hypotheses concern a population parameter?
Was the alternative direction chosen before viewing the results?
Were the assumptions and conditions considered?
Was the correct and reference distribution used?
Was the interpreted as a probability under ?
Were statistical and distinguished?
Was the conclusion stated in context?
Takeaway: A strong conclusion reports both what the data support statistically and what the estimated effect means in practice.