08 Hypothesis Testing
A practical guide to formulating hypotheses, evaluating sample evidence with test statistics and p-values, making decisions, and interpreting errors and statistical power.
The purpose of hypothesis testing
Hypothesis testing uses sample data to evaluate a claim about a population parameter, such as a population mean , proportion , or difference between means. Because random samples vary, an observed difference may reflect a real population effect or ordinary sampling variation.
The goal is not to prove a claim with certainty. Instead, the procedure measures how compatible the observed data are with a specified null claim and then reports a decision at a preselected level of evidence.
Takeaway: A hypothesis test connects sample evidence to a carefully stated claim about a population.
Formulating the hypotheses
Begin by identifying the population parameter of interest. The is the default claim being tested and is written as . It commonly states that there is no difference, no association, or no effect, and it includes equality:
For a bottle claimed to contain an average of 500 milliliters, a suitable is .
The is the claim for which evidence is sought. It must be chosen before examining the results:
Right-tailed:
Left-tailed:
Two-tailed:
Use a two-tailed alternative when departures in either direction matter. Use a one-tailed alternative only when the research question and consequences justify focusing on one direction. The hypotheses concern population parameters, not individual observations or sample statistics, and they should be mutually exclusive and exhaustive.
Takeaway: Define the parameter and choose the direction of the alternative before looking at the sample results.
Setting the standard and measuring departure
The , denoted by , is selected before conducting the test. It is the maximum probability of rejecting when is actually true. Common choices include , , and .
A smaller requires stronger evidence before rejecting the . The is not the probability that is true, and it is not the probability that the eventual conclusion is correct.
Next, select an appropriate . A general standardized form is shown below.
For a sample mean with known population standard deviation , a common statistic is
When is unknown and the sample standard deviation is used, the statistic is often
with degrees of freedom under the usual one-sample conditions. The appropriate statistic and reference distribution depend on the parameter, study design, assumptions, and whether population variability is known.
Takeaway: Set in advance and match the to the parameter and study conditions.
Using p-values and rejection regions
The measures how unusual the observed would be if the were true. It is calculated in the direction specified by the .
For a right-tailed test, use the area to the right of the observed statistic.
For a left-tailed test, use the area to the left.
For a two-tailed test, include results at least as far from the null value in either direction.
Use the approach by comparing it with :
The critical-value approach gives the same decision by comparing the with a rejection boundary. For a two-tailed standard normal test with , the rule is approximately
For example, if and , reject . If and , fail to reject . A small indicates evidence against , but it does not give the probability that is true.
Takeaway: Compare the with the preselected , or use the equivalent critical-value rule.
Interpreting the decision
Rejecting the means that the data provide statistically significant evidence in favor of the at the chosen . Failing to reject the means that the data do not provide sufficient evidence against it.
Failing to reject does not mean accepting or proving the . A nonsignificant result may occur because the null claim is reasonable, the true effect is small, the data are variable, or the sample is too small to detect the effect.
A complete conclusion should identify the population, state the direction of the result, and address practical context. For example: “At the 5% , the sample provides sufficient evidence that the population mean battery life exceeds 10 hours.” Statistical significance does not necessarily imply practical importance; a very large sample can produce a small for a trivial effect.
Takeaway: State what the data support, avoid claiming proof, and distinguish statistical significance from practical importance.
Errors, power, and study design
There are two possible errors when making a hypothesis-testing decision. If is true and the decision is to reject it, the result is a . If is false and the decision is to fail to reject it, the result is a . The other two combinations are correct decisions.
A occurs when a true is rejected. Its probability is controlled by :
A occurs when a false is not rejected. Its probability is denoted by . Unlike , generally depends on the particular alternative value, effect size, sample size, variability, and test design.
is the probability of rejecting when a specified real effect exists:
Power generally increases when the sample size increases, the true effect is larger, measurement variability decreases, the increases, or a justified one-tailed test is used instead of a two-tailed test. With sample size fixed, lowering usually lowers power because the rejection region becomes smaller. A power analysis before data collection can identify the sample size needed to detect a scientifically meaningful effect with a target power, often 80% or 90%.
Takeaway: Every decision involves possible errors, and power describes the ability to detect a specified real effect.
A complete testing workflow
A complete hypothesis test follows a consistent sequence:
Define the population parameter and state and .
Choose before analyzing the data.
Select the and verify the relevant assumptions.
Calculate the from the sample.
Find the or determine whether the statistic lies in the rejection region.
Reject or fail to reject .
Interpret the result in the context of the research question, including direction and practical size.
Example: battery lifetime
A company claims that its batteries last an average of 10 hours. An independent sample gives , hours, and known population standard deviation hours. To test whether the true mean lifetime is greater than 10 hours, use
At , calculate
For a right-tailed standard normal test, gives a of about . Since , reject . The data provide statistically significant evidence that the population mean battery life exceeds 10 hours. This does not establish that every battery lasts more than 10 hours or that the additional hours is practically important.
Final checklist: State the parameter, specify the hypotheses and direction, set , verify assumptions, calculate the statistic, determine the or rejection region, make the decision, and interpret both statistical and practical meaning.