10 Correlation

A structured guide to using scatterplots, covariance, and Pearson correlation to describe linear associations, test population relationships, and distinguish association from causation.

Reading Relationships in Scatterplots

A relationship between two quantitative variables can be described by its direction, form, strength, and unusual features. Start with a before calculating a numerical summary.

  • Direction: A pattern that rises from left to right indicates a positive ; a pattern that falls indicates a negative .

  • Form: Decide whether the pattern is approximately linear, curved, or otherwise structured.

  • Strength: Assess how closely the points follow the overall pattern.

  • Unusual features: Look for outliers, clusters, gaps, or changes in spread.

The explanatory variable is usually placed on the horizontal axis, and the response variable is placed on the vertical axis. A shapeless cloud suggests little or no linear , but a numerical correlation alone may conceal curvature or an influential outlier.

Takeaway: Visual inspection comes first because the shape of the data determines whether a linear summary is appropriate.

Measuring Joint Variation with

describes whether two variables tend to move together. For paired observations, calculate each variable's deviation from its mean and multiply the deviations:

sxy=∑i=1n(xi−xˉ)(yi−yˉ)n−1.s_{xy}=\frac{\sum_{i=1}^{n}(x_i-\bar{x})(y_i-\bar{y})}{n-1}.
  • If both deviations are positive or both are negative, their product is positive.

  • If one deviation is positive and the other is negative, their product is negative.

  • A positive value suggests a positive , while a negative value suggests a negative .

The main limitation is that depends on measurement units. Converting a measurement from meters to centimeters changes its numerical . This makes difficult to compare across data sets.

Takeaway: identifies the direction of joint variation, but it is not a unit-free measure of strength.

Interpreting Pearson Correlation

The standardizes by the standard deviations of the two variables:

r=sxysxsy=∑i=1n(xi−xˉ)(yi−yˉ)∑i=1n(xi−xˉ)2∑i=1n(yi−yˉ)2.r=\frac{s_{xy}}{s_xs_y} =\frac{\sum_{i=1}^{n}(x_i-\bar{x})(y_i-\bar{y})} {\sqrt{\sum_{i=1}^{n}(x_i-\bar{x})^2}\sqrt{\sum_{i=1}^{n}(y_i-\bar{y})^2}}.

The coefficient always satisfies

−1≤r≤1.-1\le r\le 1.

Interpret the sign and absolute value separately:

  • r>0r>0 indicates a positive linear .

  • r<0r<0 indicates a negative linear .

  • Values of ∣r∣|r| closer to 11 indicate a stronger linear pattern.

  • Values of ∣r∣|r| closer to 00 indicate a weaker linear pattern.

For example, r=0.85r=0.85 indicates a strong positive linear , while r=−0.20r=-0.20 indicates a weak negative linear . The context and should support the description; there is no universal cutoff separating strong from weak. A value of r=0r=0 means no linear , not necessarily no relationship at all.

For a population, the corresponding parameter is denoted by ρ\rho, whereas rr describes a sample.

Takeaway: The sign of rr gives direction, and ∣r∣|r| gives the strength of the linear .

Testing a Population Correlation

A sample correlation can differ from zero because of random sampling variation. To test whether the population has a nonzero linear , use

H0:ρ=0H_0:\rho=0

and choose the alternative hypothesis based on the research question:

Ha:ρ≠0,Ha:ρ>0,orHa:ρ<0.H_a:\rho\ne0,\qquad H_a:\rho>0,\qquad\text{or}\qquad H_a:\rho<0.

For a sample of size nn, the usual test statistic is

t=rn−21−r2,t=\frac{r\sqrt{n-2}}{\sqrt{1-r^2}},

which is compared with a tt-distribution having n−2n-2 degrees of freedom. A complete conclusion should state the significance level, the decision about H0H_0, and the interpretation in context.

A sufficiently small provides evidence against H0H_0. Failing to reject H0H_0 does not prove that ρ=0\rho=0; it means that the sample does not provide sufficient evidence of a nonzero population linear under the test's assumptions.

Before carrying out the test, check that observations are correctly paired and independent, the relationship is reasonably linear, no extreme outlier dominates the result, the sampling or study design supports the intended inference, and the variation around the linear pattern is reasonably consistent.

Takeaway: Statistical testing evaluates evidence for a population linear ; it does not turn a sample correlation into proof of a relationship.

From to Causation

Correlation describes , not causation. If two variables are related, several explanations remain possible:

  1. Direct causation: XX causes changes in YY.

  2. Reverse causation: YY causes changes in XX.

  3. Confounding: a third variable influences both XX and YY.

  4. Chance: the observed results from sampling variation.

  5. Selection or measurement effects: the way observations are selected or measured creates an apparent relationship.

For example, ice-cream sales and drowning incidents may increase during the same months. Warm weather can increase both swimming activity and ice-cream purchases, so the correlation does not show that ice cream causes drowning.

A can support more credible causal conclusions because random assignment helps balance confounding variables between treatment groups. Observational correlation can identify useful patterns and generate hypotheses, but additional design or evidence is needed to establish causation.

Takeaway: A statistically significant correlation can still be noncausal, and causal claims require appropriate study design or additional evidence.