7 Correlation and Regression
Learn how correlation and linear regression describe numerical relationships, assess model fit, and support cautious business predictions.
: direction and strength
A scatterplot is a useful first step when examining two numerical variables: place one variable on each axis and look for a pattern. and regression can describe relationships and support forecasts, but a relationship alone does not establish that one variable causes another.
, usually written as , summarizes the direction and strength of a linear relationship. Its possible values range from to :
If , the is positive: larger values of one variable tend to occur with larger values of the other.
If , the is negative: larger values of one variable tend to occur with smaller values of the other.
A value near indicates little linear association, though a curved relationship may still exist.
A value near or indicates a strong negative or positive linear association, respectively.
has no units, and unusual observations can strongly affect its value. It describes linear association, not cause and effect. For example, ice-cream sales and swimming-pool visits may rise together because both are influenced by warm weather.
Linear regression: describing a trend
models the average relationship between a response variable and an explanatory variable with a straight line:
Here, is the intercept and is the slope. The chooses the line that minimizes the sum of squared residuals. A is the observed value minus the value predicted by the line.
For example, suppose a retailer models weekly sales, measured in thousands of dollars, from advertising spending, also measured in thousands of dollars:
Within the range and context of the data, the slope of means that each additional in advertising is associated with an estimated increase in weekly sales. The intercept of is the model’s predicted sales when advertising is zero. That intercept is meaningful only if zero advertising is plausible and represented by the data. The estimated relationship alone does not prove that increasing advertising will cause sales to rise by that amount.
Assessing model fit
The , denoted by , describes the proportion of variation in the response variable accounted for by the fitted model in the data used to fit it. For example, means that the model accounts for of the observed variation in sales around their sample mean. It does not mean that predictions are accurate, and a high does not prove that the model is appropriate or causal.
Inspect the scatterplot and a plot, which displays residuals against predicted values or against , as well as summary statistics. A suitable straight-line model generally has residuals scattered around zero without a systematic curve or changing spread. Curvature may indicate that a straight line misses a pattern; a widening or narrowing band may indicate nonconstant error variance. Investigate influential observations, which can substantially change the fitted line.
For inference about coefficients and conventional uncertainty estimates, assumptions typically include a suitable linear form, independent errors, and roughly constant error variance. Approximately normal errors are also an assumption, especially for small-sample tests and intervals.
Prediction and responsible use
Substitute a relevant value into the fitted equation to obtain a point prediction. A for the mean response at that value describes uncertainty about the average outcome. A for an individual future outcome is wider because it also accounts for individual variation.
Use predictions cautiously:
Stay within the data’s range. Predicting beyond observed values is ; the relationship may change outside the range studied.
Check predictive performance. When possible, evaluate predictions on data not used to fit the model, such as a holdout sample, and compare the sizes of prediction errors with a practical benchmark.
Consider context and missing factors. Seasonality, prices, promotions, customer mix, or other variables may affect both the explanatory variable and the outcome. or regression alone does not establish causation.
Include uncertainty in decisions. A forecast is an estimate, not a guarantee. Compare plausible outcomes, costs, benefits, and risks rather than relying on a single predicted value.
A regression forecast can inform a business decision, such as whether a proposed advertising budget is worthwhile. The decision should also account for uncertainty, costs, and evidence about whether the change itself is likely to produce the expected result.