6 Hypothesis Testing

Learn how to formulate and evaluate hypotheses about population parameters, interpret test results, and understand the limits and error risks of statistical decisions.

Formulating hypotheses

Hypothesis testing uses sample data to evaluate a claim about a population parameter, such as a mean, proportion, or difference between groups. It assesses whether the data are sufficiently inconsistent with a specified ; it does not prove a claim true or false.

The H0H_0 represents a benchmark or no-effect claim. The , written HaH_a or H1H_1, represents the effect or difference being investigated. Both hypotheses concern a population parameter, not a sample statistic, and the equality belongs in the .

For example, to ask whether a revised checkout reduces mean service time below 44 minutes, let μ\mu be the population mean service time after the revision:

  • H0:μ=4H_0: \mu = 4 minutes

  • Ha:μ<4H_a: \mu < 4 minutes

A tests a specified direction. A tests for any difference, such as Ha:μ≠4H_a: \mu \ne 4. Choose the direction before examining results, based on the question and the consequences of errors.

Conducting a test

A standard test follows an ordered process:

  1. Define the business question, the population parameter, and the hypotheses.

  2. Choose a α\alpha in advance and select an appropriate test for the parameter, data, and design.

  3. Check the test's assumptions, including whether observations are independent and whether the method's distributional conditions are reasonable.

  4. Use the sample to calculate a test statistic, which measures how far the result is from what the predicts.

  5. Find the and compare it with α\alpha. Reject H0H_0 if p≤αp \leq \alpha; otherwise, fail to reject H0H_0.

  6. State the conclusion in context and consider both the estimated effect and its practical importance.

Choosing the test, its direction, and the before evaluating results helps make the decision process explicit. A failure to reject is not proof that the is true; the data may be inconclusive.

Interpreting the

The is calculated under the assumption that the and the test assumptions are true. It is the probability of observing a test statistic at least as extreme as the one obtained, in the direction or directions specified by the alternative.

A is not the probability that the is true, and it is not the probability that chance alone caused the result. It measures how extreme the observed result is under the null model.

For example, suppose a company tests whether a new checkout process lowers average service time from 44 minutes. If the test gives p=0.032p = 0.032 and the prespecified is α=0.05\alpha = 0.05, then 0.032<0.050.032 < 0.05, so the is rejected. The sample provides evidence that mean service time is lower.

Significance and practical importance

The α\alpha is the test's chosen maximum probability of a under the . A result is called when its is at or below this threshold.

Statistical significance is a decision convention, not a measure of business value. A small effect can be , especially with a large sample, while a consequential effect may not reach significance in a small or noisy sample.

In the checkout example, rejecting the null does not establish how large the reduction is or whether it is worth the implementation cost. The estimated reduction and its uncertainty should also be examined.

Errors and

Because a test uses sample data, its decision can be wrong:

  • : Rejecting a true , also called a false positive. Under the test conditions, its long-run rate is controlled by α\alpha.

  • : Failing to reject a false , also called a false negative. Its probability is denoted by β\beta.

The probability of rejecting the when a specified alternative is true is the test's , 1−β1-\beta. generally increases with larger samples, less variability, and larger true effects.

Lowering α\alpha can reduce Type I errors but, all else equal, can also make real effects harder to detect. Test decisions therefore involve risks of both false positives and false negatives.