11 Simple Linear Regression
A progressive guide to modeling quantitative relationships with simple linear regression, interpreting fitted lines, checking assumptions, making predictions, and conducting inference for the population slope.
Model Purpose and Setup
A model studies the relationship between two quantitative variables:
The predictor, or explanatory variable, is denoted by .
The response variable is denoted by .
The population model is:
where is the population intercept, is the population , and is the random error for observation . The model concerns the mean response at each predictor value:
The phrase refers to using one predictor and a linear mean relationship. Regression summarizes an association and can support prediction, but it does not automatically establish causation.
Takeaway: Regression connects a quantitative predictor to the mean of a quantitative response; the data-collection design determines whether a causal conclusion is justified.
The Least-Squares Regression Line
For paired observations , the estimated regression equation is:
The least-squares method chooses and to minimize the sum of squared vertical errors:
The coefficient estimates are:
and:
The fitted line always passes through . The gives the estimated change in predicted for a one-unit increase in . The intercept is the predicted response when , but that interpretation is useful only when zero is meaningful and within the relevant data range.
For example, if:
then each additional hour studied is associated with an estimated increase of exam-score points. The predicted score at zero hours is , although this interpretation may be questionable if zero is outside the observed range or the model is not appropriate there.
Takeaway: Least squares produces the straight line with the smallest possible sum of squared residuals, and its coefficients must be interpreted in the original units.
Fitted Values and Residuals
For observation , the fitted value is:
and the is:
A positive means the observed response is above the fitted line.
A negative means the observed response is below the fitted line.
A of zero means the observation lies exactly on the line.
Suppose the fitted equation is . At :
If the observed score is , then:
Thus, the model underpredicted the score by points. plots help determine whether a linear model is reasonable: random scatter around zero is desirable, while systematic curves or changing spread indicate possible problems.
Takeaway: A measures observed minus predicted response, and its pattern across the data is more informative than any single .
Explained Variation and \(R^2\)
The measures the proportion of sample variation in the response explained by the fitted regression model. Define:
as total variation,
as variation explained by the regression, and:
as unexplained variation. Then:
If , then of the observed sample variation in is explained by its linear relationship with , while is not explained by this model. A high does not prove causation, guarantee that the model is appropriate, or ensure accurate predictions outside the observed range. The value of itself lies between and .
Takeaway: describes explained sample variation; it is not a complete measure of model quality or evidence of a causal relationship.
Inference for the Population
The sample estimates the population . A common uses:
against:
The test statistic is:
which follows a -distribution with:
degrees of freedom under the model assumptions and the null hypothesis. A small p-value provides evidence that the population has a nonzero linear association, but it does not demonstrate causation or practical importance. A directional alternative, such as , should be selected only when justified in advance.
A confidence interval for the population is:
where is based on degrees of freedom. For example, an interval of supports a positive population because every value in the interval is above zero.
Takeaway: Use the test or confidence interval to assess evidence about the population relationship, while keeping statistical significance separate from practical importance and causation.
Prediction and Its Uncertainty
For a predictor value , the fitted response is:
Two intervals answer different questions:
A confidence interval for the mean response estimates the average value of for all units with .
A estimates the response for one new unit with .
The is wider because it includes both uncertainty in estimating the mean and individual random variation. For a new individual response, its form is:
where:
and:
Predictions are usually most precise near and less precise as moves farther from the center of the observed predictor values. occurs when a prediction is made outside the observed range of -values; it is risky because the linear pattern may not continue.
Takeaway: Distinguish mean-response intervals from individual prediction intervals, and make predictions only within a scientifically reasonable range.
Assumptions and Diagnostics
Reliable regression analysis depends on the :
Linearity: The mean response changes linearly with the predictor.
Independence: Errors are independent of one another.
Normality: Errors are approximately normally distributed for each predictor value.
Equal variance: Errors have the same variance across predictor values.
Use the following diagnostics:
A scatterplot of versus should show an approximately straight-line pattern. Curvature suggests nonlinearity.
A residuals-versus-fitted-values plot should look like a random horizontal band centered near zero. A curve suggests nonlinearity, while a funnel or megaphone suggests unequal variance.
A residuals-versus-order or time plot can reveal dependence or autocorrelation.
A normal probability plot or histogram can assess approximate normality, which is especially relevant for small-sample inference and prediction intervals.
Outlier and leverage checks identify unusual observations. A large indicates an outlier in the response direction; a point far from the mean predictor value may have high leverage and strongly affect the line.
Minor nonnormality may have limited impact in large samples, but serious violations, dependence, strong outliers, nonlinearity, or unequal variance require caution or alternative methods.
Takeaway: Check the assumptions with the data rather than treating a fitted line as automatically trustworthy.
A Complete Regression Analysis
A complete analysis can follow this sequence:
Identify the predictor and response variables.
Examine a scatterplot for direction, form, strength, and unusual observations.
Fit the least-squares line .
Interpret the using the variables and their units.
Examine plots and assess linearity, independence, normality, and equal variance.
Report as a description of the proportion of sample variation explained.
When appropriate, test and report a p-value or confidence interval for .
Make predictions only within a scientifically reasonable range, distinguishing mean-response confidence intervals from individual prediction intervals.
Avoid causal claims unless the study design supports them.
Final takeaway: A sound analysis combines interpretation of the fitted line, -based diagnostics, measures of explained variation, inference for the , and cautious prediction.