10 Correlation
A structured guide to using scatterplots, covariance, and Pearson correlation to describe linear associations, test population relationships, and distinguish association from causation.
Reading Relationships in Scatterplots
A relationship between two quantitative variables can be described by its direction, form, strength, and unusual features. Start with a before calculating a numerical summary.
Direction: A pattern that rises from left to right indicates a positive ; a pattern that falls indicates a negative .
Form: Decide whether the pattern is approximately linear, curved, or otherwise structured.
Strength: Assess how closely the points follow the overall pattern.
Unusual features: Look for outliers, clusters, gaps, or changes in spread.
The explanatory variable is usually placed on the horizontal axis, and the response variable is placed on the vertical axis. A shapeless cloud suggests little or no linear , but a numerical correlation alone may conceal curvature or an influential outlier.
Takeaway: Visual inspection comes first because the shape of the data determines whether a linear summary is appropriate.
Measuring Joint Variation with
describes whether two variables tend to move together. For paired observations, calculate each variable's deviation from its mean and multiply the deviations:
If both deviations are positive or both are negative, their product is positive.
If one deviation is positive and the other is negative, their product is negative.
A positive value suggests a positive , while a negative value suggests a negative .
The main limitation is that depends on measurement units. Converting a measurement from meters to centimeters changes its numerical . This makes difficult to compare across data sets.
Takeaway: identifies the direction of joint variation, but it is not a unit-free measure of strength.
Interpreting Pearson Correlation
The standardizes by the standard deviations of the two variables:
The coefficient always satisfies
Interpret the sign and absolute value separately:
indicates a positive linear .
indicates a negative linear .
Values of closer to indicate a stronger linear pattern.
Values of closer to indicate a weaker linear pattern.
For example, indicates a strong positive linear , while indicates a weak negative linear . The context and should support the description; there is no universal cutoff separating strong from weak. A value of means no linear , not necessarily no relationship at all.
For a population, the corresponding parameter is denoted by , whereas describes a sample.
Takeaway: The sign of gives direction, and gives the strength of the linear .
Testing a Population Correlation
A sample correlation can differ from zero because of random sampling variation. To test whether the population has a nonzero linear , use
and choose the alternative hypothesis based on the research question:
For a sample of size , the usual test statistic is
which is compared with a -distribution having degrees of freedom. A complete conclusion should state the significance level, the decision about , and the interpretation in context.
A sufficiently small provides evidence against . Failing to reject does not prove that ; it means that the sample does not provide sufficient evidence of a nonzero population linear under the test's assumptions.
Before carrying out the test, check that observations are correctly paired and independent, the relationship is reasonably linear, no extreme outlier dominates the result, the sampling or study design supports the intended inference, and the variation around the linear pattern is reasonably consistent.
Takeaway: Statistical testing evaluates evidence for a population linear ; it does not turn a sample correlation into proof of a relationship.
From to Causation
Correlation describes , not causation. If two variables are related, several explanations remain possible:
Direct causation: causes changes in .
Reverse causation: causes changes in .
Confounding: a third variable influences both and .
Chance: the observed results from sampling variation.
Selection or measurement effects: the way observations are selected or measured creates an apparent relationship.
For example, ice-cream sales and drowning incidents may increase during the same months. Warm weather can increase both swimming activity and ice-cream purchases, so the correlation does not show that ice cream causes drowning.
A can support more credible causal conclusions because random assignment helps balance confounding variables between treatment groups. Observational correlation can identify useful patterns and generate hypotheses, but additional design or evidence is needed to establish causation.
Takeaway: A statistically significant correlation can still be noncausal, and causal claims require appropriate study design or additional evidence.