9. Confidence Intervals
A practical guide to constructing, interpreting, checking, and applying confidence intervals for population proportions and means.
Core idea and interval structure
A uses sample data to estimate an unknown population parameter while acknowledging sampling variability. Instead of giving only one number, it reports a range of plausible values.
The basic structure is
or equivalently, an interval from a lower bound to an upper bound. The is the sample statistic at the center of the interval, and the determines how far the endpoints extend from that estimate.
For a , the is the sample proportion . For a population mean , the is the sample mean .
The has the general form
This relationship explains the main factors affecting interval width:
Greater variability increases the and widens the interval.
A smaller sample size generally increases the and widens the interval.
A higher requires a larger critical value and widens the interval.
Increasing the sample size generally narrows the interval because standard errors decrease as .
Takeaway: A combines an estimate with a quantified description of sampling uncertainty.
and interpretation
The describes how well the interval-building method performs over many repeated random samples. A 95% confidence procedure produces intervals that contain the true parameter in about 95% of repeated samples of the same size and design.
A suitable interpretation of one computed interval is: We are 95% confident that the population parameter lies between the reported endpoints. The does not mean that there is a 95% probability that one already-computed interval contains the parameter. The parameter is treated as fixed, while the interval would change from sample to sample.
For two-sided intervals, commonly used critical values are approximately:
90% confidence:
95% confidence:
99% confidence:
A 99% interval based on the same data is wider than a 95% interval because it uses a larger critical value. This is a trade-off: more confidence gives greater coverage in repeated sampling but less precision in each individual interval.
Takeaway: Always distinguish the long-run meaning of the from an incorrect probability claim about one fixed interval.
Intervals for a
For a categorical outcome, the sample proportion is
where is the number of observed successes and is the sample size. For a sufficiently large random sample, a one-proportion interval is
The estimated is
and the is
Before using this normal-approximation formula, check:
Randomness: The data come from a random sample or randomized process.
Independence: Observations are independent. When sampling without replacement, a common guideline is that the sample is no more than 10% of the population.
Large counts: Both observed success and failure counts are at least 10:
For example, suppose of sampled adults report using a transportation app. Then
Using a 95% critical value of , the estimated is approximately , and the is approximately . The interval is
In context, we are 95% confident that between 53.2% and 62.8% of adults in the target population use the transportation app. The interval can also be reported as 58.0% with a of about 4.8 percentage points.
Takeaway: A proportion interval requires an appropriate sampling process, independence, and enough observed successes and failures for the approximation to be reliable.
Intervals for a population mean
For a quantitative variable, the sample mean is
When the population standard deviation is unknown, the one-mean interval uses the :
where is the sample standard deviation, is the sample size, and is based on degrees of freedom. The is , and the is
If the population standard deviation is known, a z-interval may instead be used:
For a one-mean t-interval, check randomness, independence, and the shape of the data. With a small sample, the quantitative variable should be approximately normal and free of extreme outliers or severe skewness. With a larger sample, the generally makes the procedure more reliable, although extreme outliers can still distort the result.
For example, a random sample of packages has hours and hours. The degrees of freedom are
Using , the is hours and the is approximately hours. Thus,
We are 95% confident that the mean shipping time for all packages handled by the company is between 68.89 hours and 75.91 hours.
Takeaway: Use a t-interval when estimating a mean with an unknown population standard deviation, and verify that the data support the procedure.
Precision, decisions, and common errors
The width of an interval determines how precisely the parameter has been estimated. For a fixed , increasing the sample size decreases the . Because contains in the denominator, quadrupling the sample size approximately halves the .
For a mean, greater sample variability increases and therefore increases the . For a proportion, the quantity is largest when is near . When no prior estimate is available for planning a study, using gives a conservative sample-size calculation.
An interval should be translated into the context of the question. A statement containing only numerical endpoints is incomplete. A useful conclusion identifies:
The population being studied.
The parameter being estimated.
The .
The interval in meaningful units or percentage terms.
Interval width also matters for decisions. If a company wants at least 60% customer approval and the interval is from 53.2% to 62.8%, the interval contains values both below and above 60%. The study therefore does not establish that approval exceeds 60%; more data or other evidence may be needed.
A cannot repair a biased sample, nonresponse, dependence, poor measurement, or other flaws in study design. It addresses sampling variability, not every possible source of error. Avoid these common mistakes:
Treating the as the probability for one fixed interval.
Forgetting to check randomization, independence, sample size, or data shape.
Using a z-interval for a mean when is unknown and is used.
Confusing the with the .
Reporting percentages without identifying whether they are decimals or percentages, or reporting means without units.
Takeaway: A statistically correct interval must also be communicated in context and judged for practical usefulness.