4. Exploring Relationships in Data
A structured guide to matching variable types with appropriate displays, interpreting quantitative relationships, using correlation and regression responsibly, and distinguishing association from causation.
Match the Analysis to the Variable Types
The first step in analyzing a relationship is to classify the variables. A quantitative variable records numerical measurements for which differences have meaningful interpretations, such as age, income, temperature, or exam score. A categorical variable records labels or groups, such as treatment group, region, or transportation type.
Choose the display and summary that match the variable pairing:
For one quantitative variable and one categorical variable, compare group distributions with side-by-side boxplots, grouped dotplots, or grouped histograms. Means or medians can summarize the groups.
For two quantitative variables, use a , , or .
For two categorical variables, use a two-way table, conditional percentages, or a clustered or stacked bar chart.
Category codes are labels, not automatically meaningful numerical measurements. Therefore, do not calculate a between arbitrary category codes and a quantitative outcome.
Takeaway: Identify the variable types before selecting a graph, numerical summary, or model.
Compare Groups with Quantitative and Categorical Variables
When one variable is quantitative and the other is categorical, compare the distribution of the quantitative variable across the groups. For example, hours of sleep is quantitative, while class standing is categorical. The relevant question is how sleep differs among first-year students, sophomores, juniors, and seniors—not whether numerical codes assigned to class standing correlate with sleep.
Side-by-side boxplots allow comparison of:
center, such as the median;
spread, such as the interquartile range;
shape and possible skewness; and
unusual observations or outliers.
A responsible conclusion might be: “In this sample, first-year students tended to report more hours of sleep than seniors, although the distributions overlapped.” This describes an observed group difference without claiming that progressing through college causes students to sleep less.
A categorical variable can also identify groups within a . Different colors or symbols might distinguish online and in-person sections in a plot of study hours and exam scores. Examine both the combined pattern and the within-group patterns, because differences between groups can conceal, weaken, or reverse a relationship.
Read Relationships from Scatterplots
A displays paired quantitative observations as points . Each point represents one case, such as a student, household, day, or business.
Construct the graph carefully:
Put the on the horizontal axis.
Put the response variable on the vertical axis.
Label both axes and include units.
Choose a scale that makes the pattern visible without exaggerating it.
Describe the graph using four features:
Direction: A positive association means larger values of tend to occur with larger values of ; a negative association means larger values of tend to occur with smaller values of .
Form: Decide whether the pattern is approximately linear, curved, clustered, or otherwise structured.
Strength: Assess how closely the points follow the general pattern.
Unusual observations: Look for outliers or influential points.
For example, advertising expenditure and sales might show a moderately strong, positive, roughly linear association. That description does not imply that every business follows the same pattern or that advertising alone caused higher sales.
Takeaway: Inspect the graph before calculating a numerical summary; curvature, clusters, and influential points can be hidden by a single number.
Interpret and Explained Variation
measures the direction and strength of a linear association between two quantitative variables. The sample coefficient is written , and it satisfies
Interpret its value as follows:
: positive linear association;
: negative linear association;
: little or no linear association;
near : strong linear association; and
near : weak linear association.
For example, indicates a strong positive linear association, whereas indicates a weak negative linear association.
has important limits:
It describes linear association only. A strong curved relationship can have a near zero.
It is not appropriate when one variable is merely a set of category labels.
It can be strongly affected by outliers.
It does not identify which variable influences the other.
It does not establish causation.
The is . If , then
The appropriate interpretation is that about of the variation in the response is accounted for by its linear association with the in the sample. The remaining variation may reflect other variables, random variation, measurement error, or departures from the linear model.
Use Linear Regression to Predict
predicts a quantitative response from one quantitative . The fitted regression equation is
where is the predicted response, is the fitted intercept, is the fitted slope, and is the explanatory-variable value.
The slope describes the predicted average change in for a one-unit increase in . For example, in
the model predicts an average increase of exam-score points for each additional hour studied, within the range of study hours represented in the data.
The intercept is the predicted response when . Interpret it only when zero is meaningful and is within, or reasonably near, the observed range. If the observed study times range from to hours, interpreting the intercept as a realistic prediction for zero hours may be inappropriate.
The usual fitted line is the least-squares regression line. It minimizes the sum of squared vertical differences between observed and predicted responses. These differences are values:
A positive places the observation above the line, a negative places it below the line, and a zero means the prediction is exact. A plot should show points scattered randomly around zero. Curvature suggests that a straight-line model is inadequate; a funnel shape suggests that variability changes as the changes.
Recognize the Risks of
A regression equation is generally most trustworthy for explanatory-variable values within the range used to fit the model. occurs when a model is used outside that range.
For example, a model based on temperatures from to may reasonably predict ice-cream sales at . It should not automatically be used to predict sales at or , because the relationship may change outside the observed temperatures.
A strong or statistically significant regression slope does not make an extrapolated prediction safe. Check the range, the graph, the residuals, and the practical context before using a model for prediction.
Takeaway: A model can describe the observed range well and still give unreliable predictions beyond that range.
Separate Association from Causation
An association means that two variables vary together in a systematic way. Causation means that changing one variable produces a change in the other, with other relevant conditions held constant. Association alone does not establish causation.
An observed relationship may result from:
direct causation;
reverse causation, in which the response influences the ;
, in which a third variable affects both variables;
coincidence; or
sampling variation.
For example, people who carry larger umbrellas may be more likely to wear raincoats. Carrying an umbrella does not cause someone to wear a raincoat; rainy weather is a lurking variable that influences both behaviors.
Similarly, students who study more may earn higher scores, but the association could also reflect prior preparation, access to tutoring, motivation, course difficulty, or available time. An observational study can establish an association and provide evidence consistent with causation, but it generally cannot eliminate every alternative explanation. A well-designed randomized experiment provides stronger evidence for a causal claim because treatments are assigned randomly and systematic group differences are reduced.
Match the conclusion to the design:
In an observational study, say that the variables were associated or that one group had a higher average.
In a randomized experiment, a causal statement may be justified when the design and analysis support it.
For a regression model, say that the model predicts rather than claiming that the causes the response.
Takeaway: Neither nor regression alone proves causation; the study design and possible determine how strong a conclusion is justified.
Apply a Complete Relationship-Analysis Strategy
Use this sequence when analyzing relationships in a dataset:
Classify each variable as quantitative or categorical.
Clarify whether the goal is group comparison, description of association, or prediction.
Select a suitable display, such as grouped boxplots for one quantitative and one categorical variable or a for two quantitative variables.
Describe the visual pattern, including direction, form, strength, clusters, group differences, and unusual observations.
Calculate appropriate summaries, such as group means or medians, , or a regression equation.
Check conditions by looking for nonlinearity, unequal spread, outliers, dependence, and a restricted range.
Interpret the result in context, including units and the population or sample under discussion.
Consider the study design and possible before making a causal statement.
Avoid overreach, including unsafe and treating statistical detectability as practical importance.
For example, suppose a sample of homes produces a positive, roughly linear relationship between house size and annual energy use, with and the fitted model
A suitable interpretation is: In this sample, larger homes tended to use more energy, and the association was moderately strong and positive. The model predicts an additional units of annual energy use for each additional square feet of house size, on average. Because the data are observational, this association alone does not show that increasing house size causes increased energy use; insulation, climate, occupancy, and heating system may also matter.
This approach connects the graph, numerical summaries, units, model limitations, and study design in one defensible conclusion.