3. Numerical Summaries and Comparing Distributions
A structured guide to summarizing quantitative distributions, interpreting center, spread, position, and shape, identifying outliers, and comparing groups responsibly.
A Framework for Describing Distributions
A quantitative distribution should be described through four connected features:
Center: a typical or representative value.
Spread: how much the observations vary.
Position: where observations lie relative to the rest of the data.
Shape: the overall pattern, including symmetry, skewness, clusters, gaps, and outliers.
A numerical summary should be interpreted alongside an appropriate graph, such as a dotplot, histogram, or boxplot. No single number can reveal every important feature of a distribution.
Takeaway: Begin with the full structure of the distribution rather than choosing a statistic in isolation.
Center: and
The is the arithmetic average. For observations ,
The uses every observation, so unusually large or small values can pull it toward them. For example, the of is , even though most observations are between and .
The is the middle observation after the data are ordered. For , the is . The is resistant to extreme observations because it depends primarily on the order of the data.
Use the when a distribution is reasonably symmetric and has no influential outliers. Use the when a distribution is skewed or contains outliers. In a roughly symmetric distribution, the two measures are often similar. In a , the is typically greater than the ; in a , the is typically less than the .
Takeaway: Choose the measure of center according to the distribution’s shape and unusual values.
Spread: Measuring Variability
Center alone does not describe how much observations vary. The is the maximum minus the minimum, so it is simple but can be dominated by one extreme value.
The , or IQR, describes the middle of observations:
Because it excludes the most extreme portions of the data, the IQR is resistant to outliers and is often paired with the .
The measures the typical distance of observations from the . For a sample,
Its units are the same as those of the original variable. A small indicates that observations tend to cluster near the , while a large indicates more variability. Since it is based on deviations from the , it is sensitive to outliers and is generally paired with the .
For an approximately bell-shaped distribution, the empirical rule gives the approximation that about of observations lie within one of the , about lie within two, and about lie within three. Do not apply this rule automatically to strongly skewed or irregular distributions.
Takeaway: Pair the with the and the with the IQR when those pairings fit the distribution.
Position: Percentiles, Quartiles, and
A describes relative position. The -th is a value at or below which approximately of observations fall. Thus, is the 25th , the is the 50th , and is the 75th .
The consists of the minimum, , , , and maximum. It provides a compact view of location and spread. In a boxplot, the box extends from to , so it represents the middle of the observations.
A z-score measures position in standard-deviation units. For a sample,
For example, if the exam score is , the is , and a score is , then
The score is two standard deviations above the . can help compare values measured on different scales when the reference distributions are appropriate.
Takeaway: Percentiles and quartiles describe relative location, while describe location relative to a and .
Shape and Outliers
Shape is usually assessed from a graph rather than from one numerical summary. A symmetric distribution has left and right sides with approximately similar shape and spread around the center. In a roughly symmetric, unimodal distribution, the and are often close.
A has a longer tail toward larger values, while a has a longer tail toward smaller values. The is generally pulled in the direction of the longer tail.
Also look for peaks or modes, clusters, gaps, and unusual features. A cluster may indicate a group of observations separated from another group; a gap may indicate an interval with few or no observations. These features can suggest that the data combine different populations or that the measurement process deserves investigation.
An outlier is an observation unusually distant from the general pattern. Possible causes include a data-entry error, equipment problem, unusual but genuine case, or observation from a different population. Investigate an outlier rather than removing it automatically.
The is a screening method. The lower fence is
and the upper fence is
Values beyond these fences are potential outliers, not necessarily incorrect observations. Outliers can increase the , , and , while usually having less effect on the and IQR.
Takeaway: Use graphs to identify shape and unusual features, then choose resistant or nonresistant summaries accordingly.
Comparing Distributions with
When comparing groups, use the same four-part structure for each distribution. means:
Shape: Decide whether each distribution is symmetric, skewed, unimodal, or multimodal, and note clusters or gaps.
Outliers: Identify unusually high or low observations.
Center: Compare typical values using the or that fits each distribution.
Spread: Compare variability using standard deviations or IQRs as appropriate.
For example, suppose Neighborhood A has a commute of minutes and an IQR of minutes, with slight right skew and one high value. Suppose Neighborhood B has a commute of minutes and a of minutes, with an approximately symmetric shape. Neighborhood B has the higher reported typical commute, but the comparison requires caution because the groups use different measures of center and spread. Neighborhood A appears more variable according to its IQR, while Neighborhood B’s is smaller; however, these are not the same type of spread measure.
If both distributions are reasonably symmetric and lack important outliers, compare means and standard deviations directly. If both are skewed or contain outliers, compare medians and IQRs instead. Finally, interpret numerical differences in context: a difference of minutes may have very different practical importance depending on the setting.
Takeaway: A responsible comparison discusses shape, outliers, center, and spread, uses compatible summaries, and explains practical meaning in context.
A Practical Description Template
Use this sequence to write a complete description of a quantitative distribution:
Identify the variable and its units.
Describe the shape using a graph.
Report unusual features or outliers.
Choose a resistant or nonresistant measure of center.
Report a compatible measure of spread.
Interpret the values in context.
When comparing groups, discuss shape, outliers, center, and spread for each group.
For example:
The distribution of household travel times is right-skewed, with one unusually long trip. The travel time is minutes, and the IQR is minutes, so half of the trips fall between approximately and minutes. The and IQR are appropriate because the distribution is skewed and contains an outlier.
This approach connects the numerical values to the graph, the data’s shape, and the practical setting.