02 Descriptive Statistics
A practical guide to organizing, visualizing, summarizing, and interpreting data with measures of center, variation, position, standardized scores, and outlier rules.
The purpose of descriptive statistics
Descriptive statistics turns a collection of observations into an interpretable description. A strong description usually combines a table or graph with numerical summaries.
Begin by identifying the variable, its units, and the population represented. Then inspect the observations for missing, duplicated, or impossible values before calculating summaries.
A useful description addresses five features:
Shape: whether the distribution is symmetric, skewed, uniform, or bimodal.
Center: a typical or middle value.
Spread: how widely the observations vary.
Clusters and gaps: concentrations or empty regions.
Possible outliers: observations unusually far from the rest.
Takeaway: Numerical summaries are most meaningful when interpreted alongside a visual display and the context of the data.
Organizing observations in tables
Tables organize observations so that frequencies and patterns can be compared.
A gives each value or interval its count. The relative frequency expresses the count as a proportion of all observations:
Cumulative frequency is the running total of observations at or below a specified value. When quantitative data contain many distinct values, group them into class intervals such as , , and . Intervals should not overlap, should cover the relevant observations, and generally should have equal widths.
A two-way table summarizes two categorical variables at once. Row totals, column totals, and clearly labeled percentages can reveal relationships. For example, “percent of first-year students” answers a different question from “percent of bicycle owners.”
Takeaway: Check what each count or percentage uses as its denominator before interpreting a table.
Choosing and reading graphs
Match the display to the type of variable and the purpose of the analysis.
Bar graph: compares categories; bars are separated because categories are distinct.
Pie chart: shows relative shares of a whole; it works best with a small number of categories.
Dot plot or stem-and-leaf plot: displays individual quantitative values and their clusters.
: displays quantitative observations grouped into numerical intervals; adjacent bars normally touch.
Box plot: summarizes the , quartiles, spread, and possible outliers.
Line or time-series graph: shows changes in values measured over time.
Scatterplot: shows the association between two quantitative variables.
When reading a graph, describe its shape, center, spread, clusters, gaps, and possible outliers. Two data sets can have the same mean and while having very different shapes, so a graph can reveal information that a numerical summary hides.
Takeaway: The spacing between bars matters: separated bars represent distinct categories, while touching bars represent adjacent numerical intervals.
Measuring center
Measures of center describe a typical or central value.
The is calculated by adding all observations and dividing by their number:
For a population, the corresponding formula is:
The mean uses every observation and is sensitive to extreme values. The is the middle ordered value; with an even number of observations, it is the mean of the two middle values. Because it is resistant to extreme values, the is often preferable for skewed variables such as household income or home prices.
The mode is the most frequently occurring value or category. A data set can have one mode, several modes, or no mode, and the mode is especially useful for categorical data.
For the ordered data set , the mean is approximately , the is , and the mode is . The relatively large value pulls the mean above the , suggesting some right-skewness.
Takeaway: Use the mean when extreme values do not distort the center; use the when skewness or extreme values make the mean less representative.
Measuring variation
Measures of variation describe how far observations lie from one another or from the center.
The range is the difference between the maximum and minimum:
It is easy to calculate but depends on only two observations and can be strongly affected by outliers.
The population variance and sample variance are:
The sample formula uses because it estimates population variation from a sample. The is the square root of the variance:
A small indicates concentration near the mean; a large indicates greater spread. The measures the middle half of the data:
The IQR is resistant to extreme values and is particularly useful for skewed distributions.
Takeaway: Choose variation measures that fit the distribution: for roughly symmetric data and IQR for skewed data or data with extreme values.
Position and quartiles
Percentiles describe relative position in an ordered data set. The 80th percentile is a value below which approximately of observations lie and at or above which approximately lie. Methods for calculating percentiles can differ slightly, especially for small samples, so state the method when precision matters.
Quartiles divide ordered data into four sections:
is approximately the 25th percentile.
is the , approximately the 50th percentile.
is approximately the 75th percentile.
For the ordered data , using the common method that excludes the when finding quartiles, , , and . Therefore:
A box plot displays the five-number summary: minimum, , , , and maximum. It makes it easier to compare center, spread, and possible outliers across groups.
Takeaway: Quartiles and percentiles locate observations, while the IQR summarizes the spread of the middle half.
Standardizing observations
A standardizes an observation by expressing its distance from the mean in standard-deviation units. For a population:
For a sample, a commonly used standardized score is:
Interpretation is direct:
means the observation equals the mean.
means it is two standard deviations above the mean.
means it is one and a half standard deviations below the mean.
For the example data, using a sample mean of approximately and sample of approximately , the value has:
Thus, is about sample standard deviations above the mean. Standardization also permits comparisons across variables measured in different units.
Takeaway: A describes relative standing, not the original measurement units.
Investigating unusual observations
An is unusually distant from the other observations. It may be caused by a recording or measurement error, membership in a different population or process, or genuine rare variation. Investigate it rather than deleting it automatically.
The defines potential- fences:
For the example data, , , and , so:
Every observation lies between and , so this rule identifies no potential outliers.
For approximately symmetric, bell-shaped data, an absolute around or greater is often used as a screening convention. This is not a universal definition, and z-scores can be misleading for small samples or strongly skewed distributions.
When an unusual value appears:
Check the original record and measurement units.
Determine whether it is a data-entry or measurement error.
Check whether it belongs to the intended population.
Examine how conclusions change with and without it.
Report the treatment transparently.
Takeaway: rules flag observations for investigation; context determines how they should be handled.
A practical analysis workflow
Use this workflow to describe a new quantitative data set:
Identify the variable and its units.
Inspect raw observations for missing, duplicated, or impossible values.
Create a table or graph suited to the variable and question.
Describe shape, center, spread, clusters, gaps, and possible outliers.
Calculate suitable numerical summaries.
Compare groups using the same summaries and scales.
Interpret the results in context, including the units and the population represented.
For roughly symmetric data, the mean and often work well together. For skewed data or data containing extreme values, the and IQR are usually more representative. Effective analysis combines visual evidence, numerical summaries, and contextual interpretation rather than relying on any single statistic.
Final takeaway: Descriptive statistics is most useful when the display, summary measures, and conclusions all match the structure of the data.