05 Common Statistical Distributions

A structured guide to recognizing, interpreting, calculating with, and approximating common discrete and continuous probability distributions.

Foundations: Discrete and Continuous Models

Probability distributions are mathematical models for random variables. Begin by identifying whether the variable is discrete or continuous.

  • A discrete random variable takes countable values, such as the number of defective items. Its probabilities are assigned to individual outcomes with a .

  • A continuous random variable can take any value in an interval, such as a waiting time. Its probabilities are areas over intervals under a .

  • For a continuous variable, the probability of one exact value is zero; meaningful probabilities concern ranges of values.

A distribution is commonly described by its:

  • Mean or expected value, which indicates the center;

  • Variance and standard deviation, which describe spread;

  • Shape, including symmetry, skewness, and tail behavior;

  • Support, the set of values the variable can take.

Takeaway: First determine the type of random variable and the event being measured. Then check whether the distribution's assumptions fit the process.

Counting Successes:

Use the binomial model when you count successes across a predetermined number of trials. If XX counts successes in nn trials with constant success probability pp, write

X∼Binomial⁡(n,p).X\sim\operatorname{Binomial}(n,p).

The four required conditions are:

  1. The number of trials nn is fixed.

  2. Each trial has two possible outcomes, designated success and failure.

  3. The trials are independent.

  4. The success probability pp is the same on every trial.

The probability of exactly xx successes is

P(X=x)=(nx)px(1−p)n−x,x=0,1,…,n.P(X=x)=\binom{n}{x}p^x(1-p)^{n-x}, \qquad x=0,1,\ldots,n.

The mean, variance, and standard deviation are

μ=np,σ2=np(1−p),σ=np(1−p).\mu=np, \qquad \sigma^2=np(1-p), \qquad \sigma=\sqrt{np(1-p)}.

For example, if each of 50 components independently has defect probability 0.020.02, the number of defects follows X∼Binomial⁡(50,0.02)X\sim\operatorname{Binomial}(50,0.02). The probability of exactly two defects is

P(X=2)=(502)(0.02)2(0.98)48,P(X=2)=\binom{50}{2}(0.02)^2(0.98)^{48},

and the expected number of defects is 50(0.02)=150(0.02)=1.

Do not use this model when the number of trials is not fixed, trials are dependent, or the success probability changes. Sampling without replacement from a small population may instead require a hypergeometric model.

Takeaway: “How many successes in a fixed number of independent, identical trials?” points to the .

Waiting for the First Success

Use the geometric model when the experiment continues until the first success. If each trial has success probability pp, then

X∼Geometric⁡(p),X=1,2,3,…X\sim\operatorname{Geometric}(p), \qquad X=1,2,3,\ldots

The probability that the first success occurs on trial xx is

P(X=x)=(1−p)x−1p.P(X=x)=(1-p)^{x-1}p.

The first x−1x-1 trials must fail, followed by a success on trial xx. Useful cumulative probabilities are

P(X≤x)=1−(1−p)x,P(X>x)=(1−p)x.P(X\le x)=1-(1-p)^x, \qquad P(X>x)=(1-p)^x.

Its mean and variance are

E(X)=1p,Var⁡(X)=1−pp2.E(X)=\frac{1}{p}, \qquad \operatorname{Var}(X)=\frac{1-p}{p^2}.

For example, if a customer makes a purchase independently with probability 0.200.20 on each visit, the probability of a first purchase on the fourth visit is

P(X=4)=(0.80)3(0.20)=0.1024,P(X=4)=(0.80)^3(0.20)=0.1024,

and the expected number of visits is 10.20=5\frac{1}{0.20}=5.

The has the :

P(X>s+t∣X>s)=P(X>t).P(X>s+t\mid X>s)=P(X>t).

Thus, after any number of failures, the chance of needing more than an additional number of trials is unchanged.

Takeaway: “How many trials until the first success?” points to the .

Equal Likelihood Across an Interval

A continuous variable follows a uniform model on [a,b][a,b] when all subintervals of equal length are equally likely. Its probability density is

f(x)={1b−a,a≤x≤b,0,otherwise.f(x)= \begin{cases} \dfrac{1}{b-a}, & a\le x\le b,\\ 0, & \text{otherwise}. \end{cases}

For a≤x≤ba\le x\le b, the cumulative probability is

F(x)=P(X≤x)=x−ab−a.F(x)=P(X\le x)=\frac{x-a}{b-a}.

For an interval [c,d]⊆[a,b][c,d]\subseteq[a,b],

P(c≤X≤d)=d−cb−a.P(c\le X\le d)=\frac{d-c}{b-a}.

The mean and variance are

E(X)=a+b2,Var⁡(X)=(b−a)212.E(X)=\frac{a+b}{2}, \qquad \operatorname{Var}(X)=\frac{(b-a)^2}{12}.

For a waiting time uniformly distributed from zero to 10 minutes, the probability of waiting at most three minutes is

P(X≤3)=3−010−0=0.30.P(X\le3)=\frac{3-0}{10-0}=0.30.

Takeaway: Uniform models are appropriate when probability is proportional to interval length throughout a finite range.

Symmetric Continuous Measurements

The normal model is continuous, symmetric, and bell-shaped. It is determined by its mean μ\mu and standard deviation σ>0\sigma>0:

X∼N(μ,σ2).X\sim N(\mu,\sigma^2).

Its density is

f(x)=1σ2πexp⁡[−12(x−μσ)2].f(x)=\frac{1}{\sigma\sqrt{2\pi}}\exp\left[-\frac{1}{2}\left(\frac{x-\mu}{\sigma}\right)^2\right].

For this distribution, the mean, median, and mode are equal. Approximately 68.27 percent of observations lie within one standard deviation of the mean, 95.45 percent lie within two standard deviations, and 99.73 percent lie within three standard deviations. These are useful benchmarks, while exact probabilities require the normal CDF or statistical software.

For example, if adult height is modeled by X∼N(68,32)X\sim N(68,3^2), the probability of a height no greater than 71 inches is found by standardizing:

P(X≤71)=P(Z≤71−683)=P(Z≤1).P(X\le71)=P\left(Z\le\frac{71-68}{3}\right)=P(Z\le1).

Takeaway: The is a natural model for symmetric, bell-shaped continuous measurements, but its fit must be assessed rather than assumed.

Standardization and Normal Probabilities

The has mean zero and standard deviation one:

Z∼N(0,1).Z\sim N(0,1).

Its density and cumulative distribution function are

ϕ(z)=12πe−z2/2,Φ(z)=P(Z≤z).\phi(z)=\frac{1}{\sqrt{2\pi}}e^{-z^2/2}, \qquad \Phi(z)=P(Z\le z).

To convert X∼N(μ,σ2)X\sim N(\mu,\sigma^2) into a standard normal variable, calculate the :

z=x−μσ.z=\frac{x-\mu}{\sigma}.

Then

P(X≤x)=Φ(x−μσ).P(X\le x)=\Phi\left(\frac{x-\mu}{\sigma}\right).

For an interval,

P(a≤X≤b)=Φ(b−μσ)−Φ(a−μσ).P(a\le X\le b)=\Phi\left(\frac{b-\mu}{\sigma}\right)-\Phi\left(\frac{a-\mu}{\sigma}\right).

Because the standard normal curve is symmetric about zero,

Φ(−z)=1−Φ(z).\Phi(-z)=1-\Phi(z).

For instance,

P(−1≤Z≤1)=Φ(1)−Φ(−1)≈0.8413−0.1587=0.6826.P(-1\le Z\le1)=\Phi(1)-\Phi(-1)\approx0.8413-0.1587=0.6826.

A reliable workflow is to identify the mean and standard deviation, standardize each boundary, obtain CDF values from a table or calculator, and subtract for an interval. Use complements or symmetry for upper-tail probabilities.

Takeaway: Standardization puts different normal variables on the same reference scale, making probability calculations manageable.

Approximating Binomial Probabilities

A can sometimes be approximated by a when both expected success and expected failure counts are sufficiently large. A common guideline is

np≥10andn(1−p)≥10.np\ge10 \qquad\text{and}\qquad n(1-p)\ge10.

For X∼Binomial⁡(n,p)X\sim\operatorname{Binomial}(n,p), use a normal variable YY with

Y∼N(np,np(1−p)),Y\sim N\left(np,np(1-p)\right),

so its standard deviation is np(1−p)\sqrt{np(1-p)}. The approximation is less dependable when pp is very close to zero or one or when nn is small.

Because the binomial variable is discrete and the normal variable is continuous, apply a :

  • P(X≤k)≈P(Y≤k+0.5)P(X\le k)\approx P(Y\le k+0.5);

  • P(X≥k)≈P(Y≥k−0.5)P(X\ge k)\approx P(Y\ge k-0.5);

  • P(X=k)≈P(k−0.5≤Y≤k+0.5)P(X=k)\approx P(k-0.5\le Y\le k+0.5).

The interval from k−0.5k-0.5 to k+0.5k+0.5 represents the discrete value kk. After making the correction, standardize the normal boundary or boundaries and use Φ\Phi.

Takeaway: Check the approximation conditions, match the binomial event to the correct half-unit boundary, and then use normal-probability methods.

Selecting the Right Distribution

Choose a distribution by matching the question to the random mechanism:

  • Count successes in a fixed number of independent trials: .

  • Count trials until the first success: .

  • Model a value that is equally likely anywhere in a finite interval: .

  • Model a continuous, symmetric, bell-shaped measurement: .

  • Work with a standardized normal value: .

Before calculating, ask:

  1. Is the variable discrete or continuous?

  2. What is being counted or measured?

  3. Is the number of trials fixed, or does the process stop at the first success?

  4. Are trials independent and is the success probability constant?

  5. What are the support, center, spread, and shape?

  6. Does the model's mechanism match the data-generating process?

Probability distributions are models, not guarantees. Their assumptions should be checked before use. They provide foundations for confidence intervals, hypothesis tests, simulation, and broader statistical modeling.

Final summary: Discrete models assign probability to countable outcomes, while continuous models assign probability to intervals. Binomial and geometric distributions describe success-based trials; uniform and normal distributions describe continuous measurements. Standardization and connect these models to practical probability calculations.