1 Descriptive Statistics and Data Visualization
Learn how to organize data, choose clear displays, summarize quantitative distributions, and communicate findings without overstating what the data show.
Organizing data
Descriptive statistics organize and summarize observed data. Tables, graphs, and numerical measures help reveal patterns and communicate what the data show. They describe the data at hand; by themselves, they do not establish cause and effect or prove that a pattern will hold in a larger population.
A dataset is usually arranged so that each row represents one observation, such as an order, customer, or store, and each column represents a variable recorded about it. Give each variable a clear name and, where applicable, a unit, such as delivery_days or revenue_USD.
Before summarizing, check for duplicate records, inconsistent categories or units, missing values, and implausible entries. Keep missing values distinct from zero: a blank delivery-time entry does not mean that delivery took zero days.
Variables may be categorical, such as sales channel or product type, or quantitative, such as order value or delivery time. Counts and proportions are useful for categories; the mean and require quantitative values.
Tables and graphs
A frequency table lists categories or value intervals and the number of observations in each. is the count divided by the total. For the example of twenty orders, the sales-channel counts and percentages are:
Online: twelve orders, .
Store: five orders, .
Phone: three orders, .
Total: twenty orders, .
Choose a graph to match both the type of data and the question:
A compares counts or percentages across categories. Its bars are separated because the categories are distinct.
A shows the distribution of a by grouping values into intervals, or bins. Its bars touch because the intervals form a continuous numerical scale. Bin width can affect the pattern, so choose it clearly and avoid implying more precision than the data support.
A line chart shows change over time. Put time in chronological order on the horizontal axis.
A scatterplot shows the relationship between two quantitative variables, such as advertising spending and sales. A visible association does not, by itself, show that one variable caused the other.
A compactly summarizes a quantitative distribution using its , quartiles, and spread. Side-by-side boxplots help compare groups.
Label axes, units, groups, and time periods; provide a useful title; and identify the data source when relevant. Use consistent scales when comparing groups. Truncated axes, uneven intervals, crowded labels, or selective time windows can exaggerate or obscure differences. A graph should clarify the evidence, not distort it.
Describing quantitative distributions
A useful description of a quantitative distribution covers its shape, center, spread, and unusual observations. A or dotplot can show whether the distribution is roughly symmetric or skewed, whether it has one or several peaks, how widely values vary, and whether any values stand apart. These features help determine which numerical summaries are most informative.
Center
The mean is calculated by adding the values and dividing by the number of observations. It uses every value, so unusually high or low observations can pull it toward the tail.
The is the middle value after sorting the data. When there is an even number of observations, it is the average of the two middle values. It is less affected by extreme values than the mean.
The mode is the most frequent value or category. A dataset can have more than one mode, or none that is useful.
For a roughly symmetric distribution without influential outliers, the mean is often a helpful measure of center. For a skewed distribution or one with extreme values, the is often more representative. Reporting both can reveal that the distribution is pulled toward one side.
Spread
The range is the maximum minus the minimum. It is simple to calculate but depends only on the two endpoints.
Quartiles divide ordered data into four parts. The is the third quartile minus the first; it describes the spread of the middle half of the data and is less sensitive to extremes than the range.
The measures the typical distance of values from the mean. It is expressed in the variable’s original units and, like the mean, is sensitive to extreme values. For a sample with observations and sample mean , the sample is calculated by taking the square root after dividing the sum of squared deviations from the sample mean by :
Interpreting delivery times
Suppose five orders took two, three, three, four, and eight days to deliver. Their mean is four days, is three days, and range is six days. The mean is larger than the because the eight-day delivery pulls the mean upward.
A clear report might state both the mean and and include a graph or measure of spread, rather than presenting the mean alone. Before treating the eight-day value as an error or an outlier, check the original record and the business context; an unusual value may be real and important.
Communicating results
Make the scope of a descriptive summary clear: state how many observations were included, what period and population they represent, and how the variables were defined. Report units and use sensible rounding. When comparing groups, show group sizes as well as their summaries; the same percentage based on ten orders has a different context from one based on ten thousand orders.
For example, a manager reviewing delivery performance might use a to see the distribution, report the and IQR to describe a typical order and the spread of the middle half, and compare those summaries across shipping methods. That description can guide further investigation or planning, but it does not by itself establish that the shipping method caused the difference or predict how all future orders will perform.
Organize records consistently, distinguish categorical from quantitative variables, and select tables and graphs suited to the data. Describe quantitative distributions by shape, center, spread, and unusual values, choosing summaries that suit the distribution. Clear labels, context, and honest scales help readers use the results without overstating what the data show.