2. Exploring Categorical and Quantitative Data
A progressive guide to classifying data, choosing appropriate statistical displays, interpreting distributions and time patterns, and writing evidence-based conclusions.
Classifying Variables and Organizing Data
Exploratory data analysis starts by organizing observations so that meaningful patterns become visible. The first decision is to identify the type of variable.
Categorical variables place individuals or objects into groups, such as transportation method.
Quantitative variables record numerical amounts for which arithmetic comparisons are meaningful, such as waiting time or height.
Quantitative variables may be discrete, meaning they count separate values such as the number of pets, or continuous, meaning they can take values along a numerical scale such as height or time.
A useful display should make it possible to identify the values or categories, how often they occur, the center or most common values, the variability or spread, unusual observations, and trends when the data are recorded over time.
Organizing observations
A is often the first step. It lists possible values or categories and their counts. A expresses those counts as proportions or percentages. For any category,
For example, if of students usually take the bus, the relative frequency is . The frequencies should add to the total number of observations, and the relative frequencies should add to , or .
When a quantitative variable has many distinct values, group the observations into classes or intervals. The intervals should cover every observation, should not overlap, and should have a clear interpretation.
Takeaway: Classify the variable before choosing a display, and use tables to organize counts and proportions.
Comparing Categories with Bar Charts
A is designed for categorical data. Put the category names on one axis and the frequencies or relative frequencies on the other. Each category receives its own bar, and the bars are separated by spaces because the categories are distinct groups.
For transportation data with the categories walk, bicycle, bus, and car, compare bar heights to determine which method is most common, which is least common, and how large the differences are. A grouped can compare the same categories across another variable, such as grade level. A segmented or stacked shows how subgroups contribute to a total.
The separation between bars is important. A 's categories are labels, so their order may often be changed. A , in contrast, uses connected numerical intervals whose order is fixed by the number line. Using a for named categories or treating separated categorical bars as if they represented numerical intervals can lead to an incorrect interpretation.
Takeaway: Use a to compare categories, and interpret the heights as counts or percentages in context.
Displaying and Describing Quantitative Data
Quantitative data can be displayed in several ways, depending on the size of the data set and the detail needed.
Dotplots
A places one dot above every quantitative value on a number line. Repeated observations appear as stacks. For the waiting times
there are three observations at , two at , and one at each of the other listed values. The values range from to minutes. Most observations are near or minutes, while may be an unusually high observation that deserves attention in context.
Dotplots preserve individual observations, making them useful for small or moderately sized data sets. They can reveal clusters, gaps, repeated values, and possible outliers.
Describing a distribution with
Use to organize a description:
Shape: Is the distribution symmetric, skewed left, skewed right, uniform, or multimodal?
Outliers: Are any observations unusually far from the rest?
Center: What value is typical? The mean or median may be appropriate.
Spread: How variable are the observations? Possible measures include the range or .
Takeaway: A strong description goes beyond listing values; it explains shape, unusual observations, center, and spread.
Understanding Histograms
A groups quantitative observations into numerical intervals called bins. The horizontal axis shows the intervals, and the height of each adjacent bar shows the frequency or relative frequency in that interval.
To construct a :
Identify the quantitative variable and its units.
Choose equal-width intervals that cover the full range.
Count observations in each interval.
Place the intervals on the horizontal axis.
Place frequency or relative frequency on the vertical axis.
Draw adjacent bars with heights matching the counts or percentages.
For commute times grouped into , , , , , and minutes, the largest frequencies occur from to minutes. If the smaller frequencies extend farther toward larger times, the distribution has a longer right tail and is described as right-skewed.
Bin width affects the appearance of a . Very wide bins can hide clusters or gaps, while very narrow bins can make random fluctuations look important. Equal-width bins are generally preferred. If unequal widths are used, bar area rather than height must represent frequency.
A differs from a in four main ways:
A is for quantitative data; a is for categorical data.
A uses numerical intervals; a uses named categories.
bars usually touch; bar-chart bars are separated.
intervals have a fixed numerical order; category order may often be changed.
Takeaway: Use a to study the shape of a larger quantitative distribution, and interpret its tails, clusters, gaps, and concentration in context.
Summarizing Distributions with Boxplots
A , also called a box-and-whisker plot, summarizes a quantitative distribution with five values:
the minimum;
the first quartile, , or 25th percentile;
the median, , or 50th percentile;
the third quartile, , or 75th percentile; and
the maximum.
The box extends from to , so it contains the middle of observations. Its width is the :
For example, if the five-number summary is minimum , , median , , and maximum , then the middle half of the observations lies from to minutes, the median is minutes, the is minutes, and the overall range is minutes.
Some conventions flag values below
or above
as potential outliers. A longer upper whisker or a median positioned closer to the lower edge of the box may suggest right-skewness, but the context and the full data should guide the conclusion.
Boxplots are efficient for comparing several quantitative distributions. They show differences in median, spread, and possible outliers, but they do not preserve individual observations or show detailed clusters as clearly as dotplots and histograms.
Takeaway: Use a when five-number summaries and comparisons among quantitative groups are the main goals.
Finding Patterns Over Time
A , also called a time-series graph, displays a variable measured at successive times. Time belongs on the horizontal axis, and the measured quantity belongs on the vertical axis. Points are usually connected in chronological order.
For monthly library visits of , , , , , and from January through June, the visits generally increase during the first five months and then decrease slightly in June. This conclusion includes both the direction of change and the time context.
Look for these features:
Trend: a general upward or downward movement;
Seasonality or cycles: patterns that repeat at regular intervals;
Changing variability: fluctuations that become larger or smaller over time;
Unusual observations: points that differ sharply from nearby times; and
Time dependence: nearby observations may be related, so they should not automatically be treated as independent.
Chronological order is essential. Reordering the values from smallest to largest would destroy the pattern over time.
Takeaway: Use a whenever the order in which observations occur is part of the question.
Choosing Displays and Drawing Conclusions
The appropriate display depends on both the variable type and the purpose of the analysis.
Use a or to summarize how observations are distributed among categories.
Use a when a small quantitative data set's individual values, repeated observations, clusters, or gaps matter.
Use a to study the shape of a larger quantitative distribution.
Use a to compare medians, quartiles, spreads, and possible outliers across quantitative groups.
Use side-by-side boxplots or dotplots when comparing several quantitative distributions.
Use a when the variable is measured in chronological order and trends or cycles are important.
A graph can mislead when it does not match the data type. For example, a is inappropriate for named categories, while a can conceal the numerical order of a quantitative variable.
Writing a statistical conclusion
A practical conclusion should identify the variable, describe the important pattern, and provide numerical evidence when possible. Include:
the context and units;
the distribution's shape or time trend;
a reasonable measure of center;
a measure of spread; and
clusters, gaps, trends, cycles, or possible outliers.
Instead of writing, “The graph looks unusual,” write a conclusion such as: “The commute-time is right-skewed: most students travel for fewer than minutes, while a small number have commutes longer than minutes.”
Such a conclusion is descriptive. It summarizes the observations collected, but it does not by itself establish that one variable causes another or that the same pattern must occur in a larger population.
Final takeaway: Match the display to the data and question, then describe the resulting pattern with context, shape, center, spread, and evidence.