7. Random Variables and Probability Distributions

A structured guide to modeling uncertainty with random variables, probability distributions, expected value, variability, and common discrete and continuous models.

Identify and Classify Random Variables

Probability provides a mathematical language for uncertainty. The first step is to identify the random process and define the numerical quantity that will be recorded.

A assigns a number to each relevant outcome of a random phenomenon. For example, XX might be the number of defective products in a sample of 2020, while YY might be a bus's waiting time. The outcome itself and the numerical record are not necessarily the same: a coin toss produces heads or tails, but a variable could record the number of heads in five tosses.

Classifying the possible values

A has countable possible values. Counts of customers, correct answers, accidents, or successes are typical examples. A discrete variable is modeled with a probability mass function, which assigns P(X=x)P(X=x) to each possible value. Valid probabilities satisfy:

0≤P(X=x)≤1.0\leq P(X=x)\leq 1.

They must also satisfy:

∑xP(X=x)=1.\sum_x P(X=x)=1.

A can take any value in an interval. Time, height, mass, temperature, and distance are common examples. A continuous variable is modeled with a probability density function. Probabilities are areas over intervals:

P(a≤X≤b)=area under the density curve from a to b.P(a\leq X\leq b)=\text{area under the density curve from }a\text{ to }b.

The total area is 11, and an exact point has probability P(X=c)=0P(X=c)=0. This does not make the value impossible; it means that the continuous model assigns probability to intervals rather than individual points.

Takeaway

Classify the variable before calculating: counts are usually discrete, while measurements are usually continuous. The classification determines whether to add probabilities at individual values or find areas over intervals.

Read Probability Distributions

A connects possible outcomes to their likelihoods. For a discrete variable, a distribution can be displayed in a probability table. For example, if XX can be 00, 11, 22, or 33, the probabilities might be 0.100.10, 0.300.30, 0.400.40, and 0.200.20, respectively. The probabilities must be nonnegative and add to 11. In this example, P(X=2)=0.40P(X=2)=0.40.

For a continuous variable, the is represented by a density curve. The probability that the variable falls in an interval is the area under the curve across that interval. The model's form should reflect the process: a count of successes may fit a binomial model, a count of events in an interval may fit a Poisson model, and a roughly symmetric measurement may fit a normal model.

A useful model answers two questions:

  1. What values can the take?

  2. How much probability is assigned to each value or interval?

The model should be selected from the data-generating process rather than from a formula that happens to be convenient.

Takeaway

A is both a description of possible values and a rule for assigning probability. Always check that the model matches how the observations arise.

Calculate and Interpret

The describes the center of a as a long-run average. For a :

E(X)=μX=∑xxP(X=x).E(X)=\mu_X=\sum_x xP(X=x).

This is a weighted average: an outcome with a larger probability contributes more to the average. Suppose a game pays $10\$10 with probability 0.200.20 and $0\$0 with probability 0.800.80. If XX is the payout, then:

E(X)=(10)(0.20)+(0)(0.80)=2.E(X)=(10)(0.20)+(0)(0.80)=2.

The expected payout is $2\$2 per game in the long run. It does not mean that every individual play pays exactly $2\$2. If the entry cost is $3\$3, the expected net gain is:

E(net gain)=2−3=−1.E(\text{net gain})=2-3=-1.

is linear. For constants aa and bb:

E(a+bX)=a+bE(X).E(a+bX)=a+bE(X).

For two random variables, independence is not required for addition of expected values:

E(X+Y)=E(X)+E(Y).E(X+Y)=E(X)+E(Y).

Takeaway

Use to describe long-run center, not to predict the exact result of one repetition.

Measure Variability

The gives the center, but two distributions can have the same center and very different spread. and measure this variability.

For a :

Var⁡(X)=σX2=∑x(x−μX)2P(X=x).\operatorname{Var}(X)=\sigma_X^2=\sum_x(x-\mu_X)^2P(X=x).

The is:

σX=Var⁡(X).\sigma_X=\sqrt{\operatorname{Var}(X)}.

An equivalent formula for is:

Var⁡(X)=E(X2)−[E(X)]2.\operatorname{Var}(X)=E(X^2)-[E(X)]^2.

is expressed in squared units, while has the same units as the original variable. Consequently, is often easier to interpret. A small means that outcomes tend to cluster near the mean; a large one means that they are more dispersed.

For example, two delivery services can each have an average delivery time of 3030 minutes. If one has a of 33 minutes and the other has a of 1212 minutes, the first service is more predictable even though their means are equal.

Adding a constant shifts the center but does not change spread:

Var⁡(X+c)=Var⁡(X).\operatorname{Var}(X+c)=\operatorname{Var}(X).

Multiplication changes according to:

SD⁡(aX)=∣a∣SD⁡(X).\operatorname{SD}(aX)=|a|\operatorname{SD}(X).

Takeaway

Report the mean to describe typical location and the to describe typical spread. Interpret spread in the original units whenever possible.

Model Bernoulli Trials and Binomial Counts

A has exactly two possible outcomes, such as defective or not defective, response or no response, or heads or tails. If success has probability pp, failure has probability 1−p1-p. For a Bernoulli coded as 11 for success and 00 for failure:

E(X)=p.E(X)=p.
Var⁡(X)=p(1−p).\operatorname{Var}(X)=p(1-p).

A extends this idea to the number of successes in several trials. The conditions are:

  1. There is a fixed number nn of trials.

  2. Each trial has two relevant outcomes.

  3. The success probability pp is the same on every trial.

  4. The trials are independent.

When these conditions hold:

X∼Binomial⁡(n,p).X\sim\operatorname{Binomial}(n,p).

The probability of exactly xx successes is:

P(X=x)=(nx)px(1−p)n−x,x=0,1,2,…,n.P(X=x)=\binom{n}{x}p^x(1-p)^{n-x},\qquad x=0,1,2,\ldots,n.

The mean, , and are:

μX=np,σX2=np(1−p),σX=np(1−p).\mu_X=np,\qquad \sigma_X^2=np(1-p),\qquad \sigma_X=\sqrt{np(1-p)}.

For a quality-control sample of 2525 circuit boards with defect probability 0.040.04, the number XX of defective boards is modeled as Binomial⁡(25,0.04)\operatorname{Binomial}(25,0.04). Its expected number of defects is \(25(0.04)=1), and its is approximately 0.980.98.

Translate wording into events

  • Exactly 33: P(X=3)P(X=3)

  • At most 33: P(X≤3)P(X\leq3)

  • Fewer than 33: P(X<3)=P(X≤2)P(X<3)=P(X\leq2)

  • At least 33: P(X≥3)P(X\geq3)

  • More than 33: P(X>3)=P(X≥4)P(X>3)=P(X\geq4)

The binomial model is inappropriate when the number of trials is not fixed, there are more than two outcomes, the success probability changes substantially, or dependence changes later trial probabilities.

Takeaway

Before using a binomial formula, verify the fixed-trial, two-outcome, constant-probability, and independence conditions. Pay special attention to whether boundary values are included.

Compare Other Discrete Models

Other discrete models are defined by different stopping rules or event-generating processes.

A geometric counts the number of independent Bernoulli trials needed to obtain the first success. If the success probability is pp, then:

P(X=x)=(1−p)x−1p,x=1,2,3,…P(X=x)=(1-p)^{x-1}p,\qquad x=1,2,3,\ldots

This model is appropriate when the process continues until the first success rather than stopping after a fixed number of trials.

A counts events in a fixed interval of time or space when events occur independently at a stable average rate λ\lambda. Its probability rule is:

P(X=x)=e−λλxx!,x=0,1,2,3,…P(X=x)=\frac{e^{-\lambda}\lambda^x}{x!},\qquad x=0,1,2,3,\ldots

For this model:

E(X)=λ,Var⁡(X)=λ.E(X)=\lambda,\qquad \operatorname{Var}(X)=\lambda.

The number of help-desk calls in an hour or flaws in a fixed length of material may be modeled this way when the rate and independence assumptions are reasonable.

Takeaway

Choose among discrete models by examining the process: fixed trials and successes suggest binomial, waiting for a first success suggests geometric, and event counts over an interval suggest Poisson.

Use the and Z-Scores

The is continuous, symmetric, and bell-shaped. It is determined by the mean μ\mu, which sets the center, and the σ\sigma, which sets the spread. One common notation is:

X∼N(μ,σ2).X\sim N(\mu,\sigma^2).

The normal density is:

f(x)=1σ2πe−12(x−μσ)2,σ>0.f(x)=\frac{1}{\sigma\sqrt{2\pi}}e^{-\frac{1}{2}\left(\frac{x-\mu}{\sigma}\right)^2},\qquad \sigma>0.

Probabilities are areas under the curve. The mean, median, and mode are equal; the total area is 11; and the tails extend indefinitely while approaching the horizontal axis.

For a reasonably , approximately 68%68\% of observations lie within 11 of the mean, 95%95\% within 22, and 99.7%99.7\% within 33. This empirical rule is not appropriate for strongly skewed or multimodal data.

Standardize with z-scores

A converts an observation to a common scale:

z=x−μσ.z=\frac{x-\mu}{\sigma}.

For commute times with mean 3030 minutes and 88 minutes, a 4646-minute commute has:

z=46−308=2.z=\frac{46-30}{8}=2.

It is therefore 22 standard deviations above the mean. Under a reasonable normal model, the empirical rule suggests that approximately 2.5%2.5\% of observations are above this value.

The standard is:

Z∼N(0,1).Z\sim N(0,1).

To convert a known back to the original scale, use:

x=μ+zσ.x=\mu+z\sigma.

For scores with mean 500500 and 100100, a of 1.281.28 gives:

x=500+(1.28)(100)=628.x=500+(1.28)(100)=628.

Takeaway

Use areas to find normal probabilities, z-scores to compare relative positions, and the inverse formula x=μ+zσx=\mu+z\sigma to return to the original units.

Choose, Check, and Apply a Probability Model

A probability model should be selected and checked before calculations begin. Use the following sequence:

  1. Define the clearly: state what is being counted or measured.

  2. Decide whether it is discrete or continuous.

  3. Identify the possible values and their range.

  4. Examine the process that generated the data, including fixed trials, success probabilities, event rates, measurement variation, and dependence.

  5. Select a model whose assumptions fit that process.

  6. Check the assumptions using context and, when data are available, graphs or diagnostic checks.

  7. Interpret the resulting probability in the original situation and units.

For example, adult height may be reasonably modeled with a when the population is approximately symmetric and the observations are appropriately independent. The number of defective items in a batch is not naturally modeled as normal because it is a discrete count with a lower bound of zero.

Probability models also support statistical inference. Probability asks what could happen under an assumed model; inference uses observed sample data to learn about an unknown population or process. Expected values describe long-run behavior, standard deviations quantify typical variation, and probability models help calculate tail probabilities. Sampling distributions then provide a foundation for confidence intervals and significance tests.

A precise calculation does not guarantee a sound conclusion. If the data-collection design or model assumptions are inappropriate, the result may be misleading even when the arithmetic is correct.

Final checklist

  • Define the variable.

  • Classify it as discrete or continuous.

  • Match the model to the generating process.

  • Verify important assumptions.

  • Translate wording carefully, especially inclusive phrases such as “at least.”

  • Interpret the answer in context.

Takeaway

Model selection is part of the solution, not a step to skip before applying a formula.