4 Random Variables and Probability Distributions
A structured guide to modeling uncertainty with random variables, probability distributions, distribution functions, expected value, variability, and common discrete and continuous models.
Random Variables and Their Types
A random variable converts outcomes of a random experiment into numerical values. It is usually written with an uppercase letter such as , while a lowercase letter such as denotes a particular value.
For three fair coin tosses, let be the number of heads observed. The possible values are , , , and . The individual sequences of coin tosses are outcomes in the sample space; the random variable summarizes each sequence by its number of heads.
A random variable is discrete when its possible values are finite or countably infinite. Counts of defects, customers, or heads are discrete. It is continuous when it can take any value in an interval of real numbers, as with height, temperature, waiting time, or exact weight. The distinction concerns the values the variable can take, not merely the precision used to record a measurement.
Takeaway: First define what is being measured and determine whether its possible values are discrete or continuous.
Probability Distributions
A probability distribution specifies how probability is assigned to the possible values of a random variable. It may be represented by a table, formula, graph, or .
Every valid distribution obeys two basic requirements: probabilities cannot be negative, and the total probability must equal one. For discrete variables, probability is assigned to individual values. For continuous variables, probability is assigned to intervals through areas under a density curve.
This distinction determines which mathematical tool to use. A count such as the number of defective items calls for a , whereas a measurement such as waiting time is modeled with a .
Takeaway: The type of random variable determines how probabilities are represented and calculated.
Discrete Probabilities
For a discrete random variable , the is
It must satisfy
and
If an event contains several possible values, add their masses:
For three fair coin tosses, if counts heads, the probabilities for are , , , and , respectively. Therefore,
Takeaway: For discrete variables, calculate event probabilities by adding the relevant point probabilities.
Continuous Probabilities
For a continuous random variable, the probability of any single exact value is zero:
Instead, probabilities are assigned to intervals using a . If is the density, then
and
The interval probability is the area under the density:
For example, if is uniform from to , then on that interval. Consequently,
Because individual points have probability zero in a continuous model, including or excluding an endpoint does not change an interval probability.
Takeaway: For continuous variables, probability is area over an interval, not the height of the density at one point.
Cumulative Probability
The works for both discrete and continuous variables. It is defined by
A CDF is nondecreasing, remains between and , approaches as approaches negative infinity, and approaches as approaches positive infinity.
For , subtracting CDF values gives an interval probability:
For a discrete variable, the CDF accumulates point masses:
For a continuous variable with density , it accumulates area:
When the CDF is differentiable, its derivative is the PDF:
Takeaway: A CDF provides a unified way to describe cumulative probability and calculate intervals for either type of random variable.
Common Discrete Models
Discrete models are selected according to the process being counted.
A Bernoulli distribution describes one trial with two outcomes. If the success probability is , then and .
A counts successes in a fixed number of independent Bernoulli trials with common success probability . The number of correct guesses on ten independent true-false questions is an example with and .
A geometric distribution counts the trials needed to obtain the first success. Its probability rule is
A Poisson distribution counts events in a fixed interval when they occur independently at a constant average rate . Its probability rule is
The Poisson distribution can also approximate a when the number of trials is large and the success probability is small, provided that approximation is reasonable.
Takeaway: Match the model to the trial structure, stopping rule, or event-rate assumptions.
Common Continuous Models
Several continuous distributions describe measurements and waiting times.
A uniform distribution on assigns equal probability to intervals of equal length. Its density on the interval is
A is symmetric and bell-shaped, with parameters and . Standardization uses
An exponential distribution often models waiting time until an event in a Poisson-process setting. With rate , its density and CDF for are
and
A model is appropriate only when its assumptions fit the context. For example, a normal model should not be adopted simply because a variable is numerical; the shape and process should support it.
Takeaway: Use continuous distributions to model measurements or waiting times, while checking the assumptions that justify each choice.
Center and Spread
The summarizes the probability-weighted center of a distribution. For a discrete variable,
and for a continuous variable,
The need not be an attainable outcome. A fair six-sided die has
even though one roll cannot equal .
measures spread around the mean:
The standard deviation is
For a linear transformation , the center and spread transform as
Takeaway: describes location, while and standard deviation describe variability.
Applying Probability Models
A practical modeling workflow connects the mathematical distribution to the real setting:
Define the random variable and state its units.
List its possible values and classify it as discrete or continuous.
Choose a distribution whose assumptions fit the process.
Specify the distribution parameters.
Calculate probabilities with a PMF, PDF, or CDF.
Interpret the result in the original context.
Check whether independence, constant rates, fixed trial counts, or approximate normality are reasonable.
Distributions can model counts of successes, arrivals, or defects; estimate the chance of exceeding a threshold; calculate expected costs or waiting times; and describe sampling behavior. A formula by itself does not validate a model. The assumptions behind the formula must be appropriate for the situation.
Final takeaway: Sound probability modeling requires both correct calculations and a defensible connection between the distribution's assumptions and the process being studied.