2. Exploring Categorical and Quantitative Data

A progressive guide to classifying data, choosing appropriate statistical displays, interpreting distributions and time patterns, and writing evidence-based conclusions.

Classifying Variables and Organizing Data

Exploratory data analysis starts by organizing observations so that meaningful patterns become visible. The first decision is to identify the type of variable.

  • Categorical variables place individuals or objects into groups, such as transportation method.

  • Quantitative variables record numerical amounts for which arithmetic comparisons are meaningful, such as waiting time or height.

  • Quantitative variables may be discrete, meaning they count separate values such as the number of pets, or continuous, meaning they can take values along a numerical scale such as height or time.

A useful display should make it possible to identify the values or categories, how often they occur, the center or most common values, the variability or spread, unusual observations, and trends when the data are recorded over time.

Organizing observations

A is often the first step. It lists possible values or categories and their counts. A expresses those counts as proportions or percentages. For any category,

relative frequency=frequencytotal number of observations.\text{relative frequency}=\frac{\text{frequency}}{\text{total number of observations}}.

For example, if 1818 of 4040 students usually take the bus, the relative frequency is 1840=0.45=45%\frac{18}{40}=0.45=45\%. The frequencies should add to the total number of observations, and the relative frequencies should add to 11, or 100%100\%.

When a quantitative variable has many distinct values, group the observations into classes or intervals. The intervals should cover every observation, should not overlap, and should have a clear interpretation.

Takeaway: Classify the variable before choosing a display, and use tables to organize counts and proportions.

Comparing Categories with Bar Charts

A is designed for categorical data. Put the category names on one axis and the frequencies or relative frequencies on the other. Each category receives its own bar, and the bars are separated by spaces because the categories are distinct groups.

For transportation data with the categories walk, bicycle, bus, and car, compare bar heights to determine which method is most common, which is least common, and how large the differences are. A grouped can compare the same categories across another variable, such as grade level. A segmented or stacked shows how subgroups contribute to a total.

The separation between bars is important. A 's categories are labels, so their order may often be changed. A , in contrast, uses connected numerical intervals whose order is fixed by the number line. Using a for named categories or treating separated categorical bars as if they represented numerical intervals can lead to an incorrect interpretation.

Takeaway: Use a to compare categories, and interpret the heights as counts or percentages in context.

Displaying and Describing Quantitative Data

Quantitative data can be displayed in several ways, depending on the size of the data set and the detail needed.

Dotplots

A places one dot above every quantitative value on a number line. Repeated observations appear as stacks. For the waiting times

3, 4, 4, 5, 6, 6, 6, 7, 8, 10, 11, 15,3,\ 4,\ 4,\ 5,\ 6,\ 6,\ 6,\ 7,\ 8,\ 10,\ 11,\ 15,

there are three observations at 66, two at 44, and one at each of the other listed values. The values range from 33 to 1515 minutes. Most observations are near 66 or 77 minutes, while 1515 may be an unusually high observation that deserves attention in context.

Dotplots preserve individual observations, making them useful for small or moderately sized data sets. They can reveal clusters, gaps, repeated values, and possible outliers.

Describing a distribution with

Use to organize a description:

  • Shape: Is the distribution symmetric, skewed left, skewed right, uniform, or multimodal?

  • Outliers: Are any observations unusually far from the rest?

  • Center: What value is typical? The mean or median may be appropriate.

  • Spread: How variable are the observations? Possible measures include the range or .

Takeaway: A strong description goes beyond listing values; it explains shape, unusual observations, center, and spread.

Understanding Histograms

A groups quantitative observations into numerical intervals called bins. The horizontal axis shows the intervals, and the height of each adjacent bar shows the frequency or relative frequency in that interval.

To construct a :

  1. Identify the quantitative variable and its units.

  2. Choose equal-width intervals that cover the full range.

  3. Count observations in each interval.

  4. Place the intervals on the horizontal axis.

  5. Place frequency or relative frequency on the vertical axis.

  6. Draw adjacent bars with heights matching the counts or percentages.

For commute times grouped into 0–90\text{--}9, 10–1910\text{--}19, 20–2920\text{--}29, 30–3930\text{--}39, 40–4940\text{--}49, and 50–5950\text{--}59 minutes, the largest frequencies occur from 1010 to 2929 minutes. If the smaller frequencies extend farther toward larger times, the distribution has a longer right tail and is described as right-skewed.

Bin width affects the appearance of a . Very wide bins can hide clusters or gaps, while very narrow bins can make random fluctuations look important. Equal-width bins are generally preferred. If unequal widths are used, bar area rather than height must represent frequency.

A differs from a in four main ways:

  • A is for quantitative data; a is for categorical data.

  • A uses numerical intervals; a uses named categories.

  • bars usually touch; bar-chart bars are separated.

  • intervals have a fixed numerical order; category order may often be changed.

Takeaway: Use a to study the shape of a larger quantitative distribution, and interpret its tails, clusters, gaps, and concentration in context.

Summarizing Distributions with Boxplots

A , also called a box-and-whisker plot, summarizes a quantitative distribution with five values:

  1. the minimum;

  2. the first quartile, Q1Q_1, or 25th percentile;

  3. the median, Q2Q_2, or 50th percentile;

  4. the third quartile, Q3Q_3, or 75th percentile; and

  5. the maximum.

The box extends from Q1Q_1 to Q3Q_3, so it contains the middle 50%50\% of observations. Its width is the :

IQR=Q3−Q1.\mathrm{IQR}=Q_3-Q_1.

For example, if the five-number summary is minimum =3=3, Q1=5Q_1=5, median =7=7, Q3=10Q_3=10, and maximum =15=15, then the middle half of the observations lies from 55 to 1010 minutes, the median is 77 minutes, the is 10−5=510-5=5 minutes, and the overall range is 15−3=1215-3=12 minutes.

Some conventions flag values below

Q1−1.5(IQR)Q_1-1.5(\mathrm{IQR})

or above

Q3+1.5(IQR)Q_3+1.5(\mathrm{IQR})

as potential outliers. A longer upper whisker or a median positioned closer to the lower edge of the box may suggest right-skewness, but the context and the full data should guide the conclusion.

Boxplots are efficient for comparing several quantitative distributions. They show differences in median, spread, and possible outliers, but they do not preserve individual observations or show detailed clusters as clearly as dotplots and histograms.

Takeaway: Use a when five-number summaries and comparisons among quantitative groups are the main goals.

Finding Patterns Over Time

A , also called a time-series graph, displays a variable measured at successive times. Time belongs on the horizontal axis, and the measured quantity belongs on the vertical axis. Points are usually connected in chronological order.

For monthly library visits of 820820, 790790, 860860, 940940, 1,0201{,}020, and 980980 from January through June, the visits generally increase during the first five months and then decrease slightly in June. This conclusion includes both the direction of change and the time context.

Look for these features:

  • Trend: a general upward or downward movement;

  • Seasonality or cycles: patterns that repeat at regular intervals;

  • Changing variability: fluctuations that become larger or smaller over time;

  • Unusual observations: points that differ sharply from nearby times; and

  • Time dependence: nearby observations may be related, so they should not automatically be treated as independent.

Chronological order is essential. Reordering the values from smallest to largest would destroy the pattern over time.

Takeaway: Use a whenever the order in which observations occur is part of the question.

Choosing Displays and Drawing Conclusions

The appropriate display depends on both the variable type and the purpose of the analysis.

  • Use a or to summarize how observations are distributed among categories.

  • Use a when a small quantitative data set's individual values, repeated observations, clusters, or gaps matter.

  • Use a to study the shape of a larger quantitative distribution.

  • Use a to compare medians, quartiles, spreads, and possible outliers across quantitative groups.

  • Use side-by-side boxplots or dotplots when comparing several quantitative distributions.

  • Use a when the variable is measured in chronological order and trends or cycles are important.

A graph can mislead when it does not match the data type. For example, a is inappropriate for named categories, while a can conceal the numerical order of a quantitative variable.

Writing a statistical conclusion

A practical conclusion should identify the variable, describe the important pattern, and provide numerical evidence when possible. Include:

  1. the context and units;

  2. the distribution's shape or time trend;

  3. a reasonable measure of center;

  4. a measure of spread; and

  5. clusters, gaps, trends, cycles, or possible outliers.

Instead of writing, “The graph looks unusual,” write a conclusion such as: “The commute-time is right-skewed: most students travel for fewer than 3030 minutes, while a small number have commutes longer than 4040 minutes.”

Such a conclusion is descriptive. It summarizes the observations collected, but it does not by itself establish that one variable causes another or that the same pattern must occur in a larger population.

Final takeaway: Match the display to the data and question, then describe the resulting pattern with context, shape, center, spread, and evidence.