7. Random Variables and Probability Distributions
A structured guide to modeling uncertainty with random variables, probability distributions, expected value, variability, and common discrete and continuous models.
Identify and Classify Random Variables
Probability provides a mathematical language for uncertainty. The first step is to identify the random process and define the numerical quantity that will be recorded.
A assigns a number to each relevant outcome of a random phenomenon. For example, might be the number of defective products in a sample of , while might be a bus's waiting time. The outcome itself and the numerical record are not necessarily the same: a coin toss produces heads or tails, but a variable could record the number of heads in five tosses.
Classifying the possible values
A has countable possible values. Counts of customers, correct answers, accidents, or successes are typical examples. A discrete variable is modeled with a probability mass function, which assigns to each possible value. Valid probabilities satisfy:
They must also satisfy:
A can take any value in an interval. Time, height, mass, temperature, and distance are common examples. A continuous variable is modeled with a probability density function. Probabilities are areas over intervals:
The total area is , and an exact point has probability . This does not make the value impossible; it means that the continuous model assigns probability to intervals rather than individual points.
Takeaway
Classify the variable before calculating: counts are usually discrete, while measurements are usually continuous. The classification determines whether to add probabilities at individual values or find areas over intervals.
Read Probability Distributions
A connects possible outcomes to their likelihoods. For a discrete variable, a distribution can be displayed in a probability table. For example, if can be , , , or , the probabilities might be , , , and , respectively. The probabilities must be nonnegative and add to . In this example, .
For a continuous variable, the is represented by a density curve. The probability that the variable falls in an interval is the area under the curve across that interval. The model's form should reflect the process: a count of successes may fit a binomial model, a count of events in an interval may fit a Poisson model, and a roughly symmetric measurement may fit a normal model.
A useful model answers two questions:
What values can the take?
How much probability is assigned to each value or interval?
The model should be selected from the data-generating process rather than from a formula that happens to be convenient.
Takeaway
A is both a description of possible values and a rule for assigning probability. Always check that the model matches how the observations arise.
Calculate and Interpret
The describes the center of a as a long-run average. For a :
This is a weighted average: an outcome with a larger probability contributes more to the average. Suppose a game pays with probability and with probability . If is the payout, then:
The expected payout is per game in the long run. It does not mean that every individual play pays exactly . If the entry cost is , the expected net gain is:
is linear. For constants and :
For two random variables, independence is not required for addition of expected values:
Takeaway
Use to describe long-run center, not to predict the exact result of one repetition.
Measure Variability
The gives the center, but two distributions can have the same center and very different spread. and measure this variability.
For a :
The is:
An equivalent formula for is:
is expressed in squared units, while has the same units as the original variable. Consequently, is often easier to interpret. A small means that outcomes tend to cluster near the mean; a large one means that they are more dispersed.
For example, two delivery services can each have an average delivery time of minutes. If one has a of minutes and the other has a of minutes, the first service is more predictable even though their means are equal.
Adding a constant shifts the center but does not change spread:
Multiplication changes according to:
Takeaway
Report the mean to describe typical location and the to describe typical spread. Interpret spread in the original units whenever possible.
Model Bernoulli Trials and Binomial Counts
A has exactly two possible outcomes, such as defective or not defective, response or no response, or heads or tails. If success has probability , failure has probability . For a Bernoulli coded as for success and for failure:
A extends this idea to the number of successes in several trials. The conditions are:
There is a fixed number of trials.
Each trial has two relevant outcomes.
The success probability is the same on every trial.
The trials are independent.
When these conditions hold:
The probability of exactly successes is:
The mean, , and are:
For a quality-control sample of circuit boards with defect probability , the number of defective boards is modeled as . Its expected number of defects is \(25(0.04)=1), and its is approximately .
Translate wording into events
Exactly :
At most :
Fewer than :
At least :
More than :
The binomial model is inappropriate when the number of trials is not fixed, there are more than two outcomes, the success probability changes substantially, or dependence changes later trial probabilities.
Takeaway
Before using a binomial formula, verify the fixed-trial, two-outcome, constant-probability, and independence conditions. Pay special attention to whether boundary values are included.
Compare Other Discrete Models
Other discrete models are defined by different stopping rules or event-generating processes.
A geometric counts the number of independent Bernoulli trials needed to obtain the first success. If the success probability is , then:
This model is appropriate when the process continues until the first success rather than stopping after a fixed number of trials.
A counts events in a fixed interval of time or space when events occur independently at a stable average rate . Its probability rule is:
For this model:
The number of help-desk calls in an hour or flaws in a fixed length of material may be modeled this way when the rate and independence assumptions are reasonable.
Takeaway
Choose among discrete models by examining the process: fixed trials and successes suggest binomial, waiting for a first success suggests geometric, and event counts over an interval suggest Poisson.
Use the and Z-Scores
The is continuous, symmetric, and bell-shaped. It is determined by the mean , which sets the center, and the , which sets the spread. One common notation is:
The normal density is:
Probabilities are areas under the curve. The mean, median, and mode are equal; the total area is ; and the tails extend indefinitely while approaching the horizontal axis.
For a reasonably , approximately of observations lie within of the mean, within , and within . This empirical rule is not appropriate for strongly skewed or multimodal data.
Standardize with z-scores
A converts an observation to a common scale:
For commute times with mean minutes and minutes, a -minute commute has:
It is therefore standard deviations above the mean. Under a reasonable normal model, the empirical rule suggests that approximately of observations are above this value.
The standard is:
To convert a known back to the original scale, use:
For scores with mean and , a of gives:
Takeaway
Use areas to find normal probabilities, z-scores to compare relative positions, and the inverse formula to return to the original units.
Choose, Check, and Apply a Probability Model
A probability model should be selected and checked before calculations begin. Use the following sequence:
Define the clearly: state what is being counted or measured.
Decide whether it is discrete or continuous.
Identify the possible values and their range.
Examine the process that generated the data, including fixed trials, success probabilities, event rates, measurement variation, and dependence.
Select a model whose assumptions fit that process.
Check the assumptions using context and, when data are available, graphs or diagnostic checks.
Interpret the resulting probability in the original situation and units.
For example, adult height may be reasonably modeled with a when the population is approximately symmetric and the observations are appropriately independent. The number of defective items in a batch is not naturally modeled as normal because it is a discrete count with a lower bound of zero.
Probability models also support statistical inference. Probability asks what could happen under an assumed model; inference uses observed sample data to learn about an unknown population or process. Expected values describe long-run behavior, standard deviations quantify typical variation, and probability models help calculate tail probabilities. Sampling distributions then provide a foundation for confidence intervals and significance tests.
A precise calculation does not guarantee a sound conclusion. If the data-collection design or model assumptions are inappropriate, the result may be misleading even when the arithmetic is correct.
Final checklist
Define the variable.
Classify it as discrete or continuous.
Match the model to the generating process.
Verify important assumptions.
Translate wording carefully, especially inclusive phrases such as “at least.”
Interpret the answer in context.
Takeaway
Model selection is part of the solution, not a step to skip before applying a formula.