08 Hypothesis Testing

A practical guide to formulating hypotheses, evaluating sample evidence with test statistics and p-values, making decisions, and interpreting errors and statistical power.

The purpose of hypothesis testing

Hypothesis testing uses sample data to evaluate a claim about a population parameter, such as a population mean μ\mu, proportion pp, or difference between means. Because random samples vary, an observed difference may reflect a real population effect or ordinary sampling variation.

The goal is not to prove a claim with certainty. Instead, the procedure measures how compatible the observed data are with a specified null claim and then reports a decision at a preselected level of evidence.

Takeaway: A hypothesis test connects sample evidence to a carefully stated claim about a population.

Formulating the hypotheses

Begin by identifying the population parameter of interest. The is the default claim being tested and is written as H0H_0. It commonly states that there is no difference, no association, or no effect, and it includes equality:

H0:μ=μ0H_0:\mu=\mu_0

For a bottle claimed to contain an average of 500 milliliters, a suitable is H0:μ=500H_0:\mu=500.

The is the claim for which evidence is sought. It must be chosen before examining the results:

  • Right-tailed: Ha:μ>μ0H_a:\mu>\mu_0

  • Left-tailed: Ha:μ<μ0H_a:\mu<\mu_0

  • Two-tailed: Ha:μ≠μ0H_a:\mu\ne\mu_0

Use a two-tailed alternative when departures in either direction matter. Use a one-tailed alternative only when the research question and consequences justify focusing on one direction. The hypotheses concern population parameters, not individual observations or sample statistics, and they should be mutually exclusive and exhaustive.

Takeaway: Define the parameter and choose the direction of the alternative before looking at the sample results.

Setting the standard and measuring departure

The , denoted by α\alpha, is selected before conducting the test. It is the maximum probability of rejecting H0H_0 when H0H_0 is actually true. Common choices include α=0.10\alpha=0.10, α=0.05\alpha=0.05, and α=0.01\alpha=0.01.

A smaller α\alpha requires stronger evidence before rejecting the . The is not the probability that H0H_0 is true, and it is not the probability that the eventual conclusion is correct.

Next, select an appropriate . A general standardized form is shown below.

test statistic=observed estimate−null valuestandard error of the estimate\text{test statistic}=\frac{\text{observed estimate}-\text{null value}}{\text{standard error of the estimate}}

For a sample mean with known population standard deviation σ\sigma, a common statistic is

z=xˉ−μ0σ/nz=\frac{\bar{x}-\mu_0}{\sigma/\sqrt{n}}

When σ\sigma is unknown and the sample standard deviation ss is used, the statistic is often

t=xˉ−μ0s/nt=\frac{\bar{x}-\mu_0}{s/\sqrt{n}}

with n−1n-1 degrees of freedom under the usual one-sample conditions. The appropriate statistic and reference distribution depend on the parameter, study design, assumptions, and whether population variability is known.

Takeaway: Set α\alpha in advance and match the to the parameter and study conditions.

Using p-values and rejection regions

The measures how unusual the observed would be if the were true. It is calculated in the direction specified by the .

  • For a right-tailed test, use the area to the right of the observed statistic.

  • For a left-tailed test, use the area to the left.

  • For a two-tailed test, include results at least as far from the null value in either direction.

Use the approach by comparing it with α\alpha:

{Reject H0,p≤α,Fail to reject H0,p>α.\begin{cases} \text{Reject }H_0, & p\le\alpha,\\ \text{Fail to reject }H_0, & p>\alpha. \end{cases}

The critical-value approach gives the same decision by comparing the with a rejection boundary. For a two-tailed standard normal test with α=0.05\alpha=0.05, the rule is approximately

∣z∣≥1.96|z|\ge 1.96

For example, if p=0.032p=0.032 and α=0.05\alpha=0.05, reject H0H_0. If p=0.081p=0.081 and α=0.05\alpha=0.05, fail to reject H0H_0. A small indicates evidence against H0H_0, but it does not give the probability that H0H_0 is true.

Takeaway: Compare the with the preselected α\alpha, or use the equivalent critical-value rule.

Interpreting the decision

Rejecting the means that the data provide statistically significant evidence in favor of the at the chosen . Failing to reject the means that the data do not provide sufficient evidence against it.

Failing to reject does not mean accepting or proving the . A nonsignificant result may occur because the null claim is reasonable, the true effect is small, the data are variable, or the sample is too small to detect the effect.

A complete conclusion should identify the population, state the direction of the result, and address practical context. For example: “At the 5% , the sample provides sufficient evidence that the population mean battery life exceeds 10 hours.” Statistical significance does not necessarily imply practical importance; a very large sample can produce a small for a trivial effect.

Takeaway: State what the data support, avoid claiming proof, and distinguish statistical significance from practical importance.

Errors, power, and study design

There are two possible errors when making a hypothesis-testing decision. If H0H_0 is true and the decision is to reject it, the result is a . If H0H_0 is false and the decision is to fail to reject it, the result is a . The other two combinations are correct decisions.

A occurs when a true is rejected. Its probability is controlled by α\alpha:

P(Type I error)=αP(\text{Type I error})=\alpha

A occurs when a false is not rejected. Its probability is denoted by β\beta. Unlike α\alpha, β\beta generally depends on the particular alternative value, effect size, sample size, variability, and test design.

is the probability of rejecting H0H_0 when a specified real effect exists:

Power=P(reject H0∣H0 is false)=1−β\text{Power}=P(\text{reject }H_0\mid H_0\text{ is false})=1-\beta

Power generally increases when the sample size increases, the true effect is larger, measurement variability decreases, the increases, or a justified one-tailed test is used instead of a two-tailed test. With sample size fixed, lowering α\alpha usually lowers power because the rejection region becomes smaller. A power analysis before data collection can identify the sample size needed to detect a scientifically meaningful effect with a target power, often 80% or 90%.

Takeaway: Every decision involves possible errors, and power describes the ability to detect a specified real effect.

A complete testing workflow

A complete hypothesis test follows a consistent sequence:

  1. Define the population parameter and state H0H_0 and HaH_a.

  2. Choose α\alpha before analyzing the data.

  3. Select the and verify the relevant assumptions.

  4. Calculate the from the sample.

  5. Find the or determine whether the statistic lies in the rejection region.

  6. Reject or fail to reject H0H_0.

  7. Interpret the result in the context of the research question, including direction and practical size.

Example: battery lifetime

A company claims that its batteries last an average of 10 hours. An independent sample gives n=64n=64, xˉ=10.4\bar{x}=10.4 hours, and known population standard deviation σ=1.6\sigma=1.6 hours. To test whether the true mean lifetime is greater than 10 hours, use

H0:μ=10,Ha:μ>10H_0:\mu=10,\qquad H_a:\mu>10

At α=0.05\alpha=0.05, calculate

z=10.4−101.6/64=2.00z=\frac{10.4-10}{1.6/\sqrt{64}}=2.00

For a right-tailed standard normal test, z=2.00z=2.00 gives a of about 0.0230.023. Since 0.023<0.050.023<0.05, reject H0H_0. The data provide statistically significant evidence that the population mean battery life exceeds 10 hours. This does not establish that every battery lasts more than 10 hours or that the additional 0.40.4 hours is practically important.

Final checklist: State the parameter, specify the hypotheses and direction, set α\alpha, verify assumptions, calculate the statistic, determine the or rejection region, make the decision, and interpret both statistical and practical meaning.