7 Basic Statistical Applications

A structured introduction to probability, descriptive statistics, random variables, sampling, simulation, and responsible interpretation of statistical results.

Connecting Probability, Samples, and Populations

Probability provides a language for uncertainty, while statistics uses data to summarize observations and draw conclusions. Their connection comes from sampling: a random sample produces data that vary from sample to sample, and probability helps describe that variation.

A population is the complete group of interest, while a sample is a subset selected from that population. A describes a population; a describes a sample. For example, the true average commute time for all employees might be μ=31\mu=31 minutes, whereas the average commute time calculated from 100 sampled employees is a used to estimate μ\mu.

A probability model begins with a sample space, the set of all possible outcomes. An event is a collection of outcomes. Probabilities satisfy

0≤P(A)≤1,P(S)=1.0\le P(A)\le 1,\qquad P(S)=1.

For equally likely outcomes,

P(A)=number of outcomes in Anumber of outcomes in S.P(A)=\frac{\text{number of outcomes in }A}{\text{number of outcomes in }S}.

Observed relative frequency provides an empirical estimate of probability:

P^(A)=number of observations in Atotal number of observations.\widehat{P}(A)=\frac{\text{number of observations in }A}{\text{total number of observations}}.

As repeated observations increase, relative frequencies often become more stable and may approach the underlying probability. This tendency does not guarantee that a small sample will closely match the theoretical probability.

Takeaway: Probability models possible outcomes, while statistics uses sample data to learn about populations and quantify uncertainty.

Counting Outcomes and Displaying Data

Counting methods determine how many outcomes are possible. The multiplication principle says that if one task can be completed in mm ways and a second task in nn ways, then the combined task can be completed in mnmn ways. Thus, three yes-or-no questions have 23=82^3=8 possible response patterns.

After data are collected, displays reveal distributional structure:

  • Bar charts summarize counts or percentages for categorical variables.

  • Histograms group quantitative values into intervals.

  • Dot plots display small quantitative data sets.

  • Box plots show center, spread, and possible outliers.

  • Scatterplots show the relationship between two quantitative variables.

A useful graph helps identify the distribution's shape, center, spread, and unusual observations. It can reveal patterns that a single numerical summary hides.

Takeaway: Count possible outcomes before calculating probabilities, and inspect a suitable graph before interpreting numerical summaries.

Measuring Center and Spread

The is defined by

xˉ=x1+x2+⋯+xnn.\bar{x}=\frac{x_1+x_2+\cdots+x_n}{n}.

It uses every observation but is sensitive to extreme values. The is the middle ordered observation, or the average of the two middle observations when the sample size is even. It is generally more resistant to outliers.

For the data set 4,5,5,6,204,5,5,6,20, the mean is 88, while the is 55. The value 2020 pulls the mean upward, so the better represents a typical observation in this small, right-skewed sample.

The range is the maximum minus the minimum. Sample variance and sample standard deviation are calculated by

s2=∑i=1n(xi−xˉ)2n−1,s^2=\frac{\sum_{i=1}^{n}(x_i-\bar{x})^2}{n-1},

and

s=s2.s=\sqrt{s^2}.

Standard deviation measures the typical distance of observations from the in the original units. A small standard deviation indicates that observations are concentrated near the mean; a large standard deviation indicates greater dispersion.

The mean and standard deviation are most informative for distributions that are reasonably symmetric and do not contain severe outliers. For skewed data, the and are often more appropriate. The is

IQR⁡=Q3−Q1,\operatorname{IQR}=Q_3-Q_1,

and measures the spread of the middle 50 percent of observations.

Takeaway: Match the summary to the distribution: use mean and standard deviation for reasonably symmetric data, and and IQR when skewness or outliers are important.

Conditional Probability, Independence, and Association

A conditional probability describes the chance of an event given that another event has occurred:

P(A∣B)=P(A∩B)P(B),P(B)>0.P(A\mid B)=\frac{P(A\cap B)}{P(B)},\qquad P(B)>0.

For example, the probability of renewal among application users is

P(renew∣app user)=P(renew and app user)P(app user).P(\text{renew}\mid\text{app user})=\frac{P(\text{renew and app user})}{P(\text{app user})}.

This value may differ from the overall renewal probability. Comparing conditional percentages is a basic way to investigate associations between variables.

Events are when learning that one occurred does not change the probability of the other. This can be expressed as

P(A∣B)=P(A),P(A\mid B)=P(A),

or equivalently,

P(A∩B)=P(A)P(B).P(A\cap B)=P(A)P(B).

An observed association does not by itself establish causation. A randomized experiment can support causal conclusions because tends to balance other factors across treatment groups. An observational association may instead be explained by confounding variables.

Takeaway: Conditional probabilities describe groups under specified conditions, while independence describes whether learning one event changes the probability of another.

Random Variables and Probability Distributions

A random variable assigns a numerical value to each outcome of a random process. A discrete random variable has countable possible values, such as the number of defective items. A continuous random variable can take values throughout an interval, such as temperature or waiting time.

For a discrete random variable XX, the probability mass function is

p(x)=P(X=x),p(x)=P(X=x),

with

p(x)≥0and∑xp(x)=1.p(x)\ge 0\quad\text{and}\quad\sum_x p(x)=1.

A continuous random variable is described by a probability density function. Probabilities are areas under the density curve. For a continuous variable, the probability of one exact value is typically zero, while an interval can have positive probability.

Important probability distributions include:

  • The Bernoulli distribution models one trial with success probability pp.

  • The binomial distribution counts successes in nn Bernoulli trials with constant success probability pp.

  • The normal distribution is a symmetric, bell-shaped model described by mean μ\mu and standard deviation σ\sigma.

  • Sampling distributions describe the behavior of statistics such as Xˉ\bar{X} or p^\hat{p} across repeated samples.

A probability distribution is a model, not a guarantee that every sample will look exactly like the model. Its assumptions should be checked before analysis.

Takeaway: Random variables convert random outcomes into numerical quantities, and probability distributions describe how those quantities behave.

Expected Value and Variability

The expected value of a discrete random variable is its long-run average:

E(X)=∑xxP(X=x).E(X)=\sum_x xP(X=x).

Suppose a game pays $10\$10 with probability 0.200.20 and pays nothing otherwise. Its expected payoff is

E(X)=10(0.20)+0(0.80)=2.E(X)=10(0.20)+0(0.80)=2.

The expected payoff is $2\$2 per play. It does not mean that every play pays $2\$2; over many plays, the average payoff tends to approach $2\$2 before any entry fee is considered.

For a constant cc and random variables XX and YY, expectation has these properties:

E(c)=c,E(c)=c,
E(cX)=cE(X),E(cX)=cE(X),
E(X+Y)=E(X)+E(Y).E(X+Y)=E(X)+E(Y).

The variance and standard deviation of a random variable are

Var⁡(X)=E[(X−E(X))2],\operatorname{Var}(X)=E\left[(X-E(X))^2\right],

and

σX=Var⁡(X).\sigma_X=\sqrt{\operatorname{Var}(X)}.

Expectation describes location, while variance and standard deviation describe uncertainty or spread.

Takeaway: Expected value describes long-run location; variance and standard deviation describe the variability around that location.

Sampling Design and Sampling Distributions

A probability sample gives population members known chances of selection. In a simple random sample, every member has an equal chance of being selected. Probability-based sampling helps reduce selection bias, whereas a convenience sample may overrepresent people who are easy to reach.

Random sampling and serve different purposes:

  • Random sampling determines which units enter a study and supports generalizing from a sample to a population.

  • determines which treatment or condition units receive and supports causal comparisons in experiments.

A is the distribution of a across repeated samples of the same size from the same population. For example, repeatedly selecting samples of 50 employees and calculating each produces a of Xˉ\bar{X}.

Sampling distributions explain why statistics vary. Even when a population mean is fixed, different samples generally produce different sample means. For many statistics, larger samples produce less variability. Under suitable conditions, the states that the of a becomes approximately normal as sample size becomes sufficiently large, even when the population distribution is not normal.

Takeaway: Sampling design determines what can be generalized, assignment design determines whether causal comparisons are justified, and sampling distributions describe the resulting statistical variability.

and Randomization

A imitates a random process using random numbers or repeated sampling. It is useful when an exact calculation is difficult, assumptions are uncertain, or sampling variability needs to be visualized.

A general procedure is:

  1. Define the random process and its probability model.

  2. Generate random outcomes according to that model.

  3. Calculate the or outcome of interest.

  4. Repeat the process many times.

  5. Summarize the simulated results with a graph, proportion, mean, or percentile.

Suppose a store claims that 60%60\% of customers use its loyalty card. To investigate variation in sample proportions, simulate many samples of 100 customers with success probability p=0.60p=0.60. The resulting distribution of simulated sample proportions shows values that could reasonably occur through random sampling alone.

can also support a randomization test. For two groups, repeatedly reassign group labels to represent what differences might occur if there were no real treatment effect. The observed difference can then be compared with the simulated distribution.

Takeaway: approximates probability and sampling behavior by repeating a clearly defined random process.

Interpreting Statistical Results Responsibly

A sound statistical interpretation identifies the population, variable, sample, method, and uncertainty. A numerical value without context is incomplete. For example, instead of saying that the average is 72, report that in a sample of 250 surveyed students, the mean reported study time was 7.2 hours per week.

Use the following checklist:

  • Identify what was measured and include its units.

  • Distinguish the sample from the target population.

  • Describe how observations were collected, including possible voluntary-response, convenience-sampling, or nonresponse bias.

  • State what the summary measures: mean, , proportion, standard deviation, or another .

  • Include an appropriate measure of spread or uncertainty.

  • Check for outliers and skewness by comparing the mean and and inspecting a graph.

  • Do not treat association as proof of causation.

  • Consider both statistical evidence and .

For example, waiting times of 3,4,4,5,6,7,8,153,4,4,5,6,7,8,15 minutes have mean

xˉ=3+4+4+5+6+7+8+158=6.5 minutes,\bar{x}=\frac{3+4+4+5+6+7+8+15}{8}=6.5\text{ minutes},

and

median⁡=5+62=5.5 minutes.\operatorname{median}=\frac{5+6}{2}=5.5\text{ minutes}.

The 15-minute wait is a high outlier that creates right skew, so the may better describe a typical customer's wait. The mean remains important when estimating the total average waiting time across many customers. A graph such as a box plot or histogram would help reveal the shape.

If customers are selected randomly, the sample can provide information about the broader customer population. If observations are collected only during a quiet weekday morning, the sample may not represent customers at other times. A of repeated random samples can show how much the or varies.

A result can be statistically unusual without being practically important. A very large sample might detect a difference of only 0.10.1 percentage points. Interpretation should address both the strength of evidence and the size and real-world meaning of the effect.

Takeaway: Strong statistical communication combines context, study design, appropriate summaries, uncertainty, and a careful distinction between association, causation, and practical importance.