6 Common Probability Models

A structured guide to identifying, applying, and comparing common discrete and continuous probability models, including their assumptions, formulas, and key measures of center and spread.

1. Discrete and Continuous Foundations

A probability model combines a random variable with a distribution that assigns probabilities to its possible outcomes. Begin by asking whether the variable is a count or a measurement.

A discrete random variable has a finite or countably infinite set of possible values, such as the number of defective items. Its probabilities are described by a probability mass function:

pX(x)=P(X=x),∑xpX(x)=1.p_X(x)=P(X=x), \qquad \sum_x p_X(x)=1.

A continuous random variable can take any value in an interval, such as time, length, or temperature. Its probabilities are areas under a probability density function rather than probabilities assigned to individual points:

P(a≤X≤b)=∫abfX(x) dx.P(a\le X\le b)=\int_a^b f_X(x)\,dx.

For a continuous variable, P(X=a)=0P(X=a)=0 for every individual value aa. The works for both types of variables and is defined by

FX(x)=P(X≤x).F_X(x)=P(X\le x).

Takeaway: Counts usually require discrete models, while measured quantities usually require continuous models; the mechanism generating the outcomes determines the specific distribution.

2. Fixed Numbers of Successes

A Bernoulli trial has exactly two outcomes, usually called success and failure. If the probability of success is pp, the probability of failure is 1−p1-p. A Bernoulli random variable can be written as

X={1,success,0,failure.X= \begin{cases} 1,&\text{success},\\ 0,&\text{failure}. \end{cases}

Use the when all of the following conditions are reasonable:

  • There is a fixed number nn of trials.

  • Each trial has two outcomes.

  • The trials are independent, or close enough to independent for the context.

  • The success probability pp is constant.

The number of successes XX then follows

X∼Binomial⁡(n,p),X\sim\operatorname{Binomial}(n,p),

with

P(X=k)=(nk)pk(1−p)n−k.P(X=k)=\binom{n}{k}p^k(1-p)^{n-k}.

For example, if ten components are inspected independently and each has probability 0.020.02 of being defective, then X∼Binomial⁡(10,0.02)X\sim\operatorname{Binomial}(10,0.02). The probability of exactly one defective component is

P(X=1)=(101)(0.02)(0.98)9≈0.166.P(X=1)=\binom{10}{1}(0.02)(0.98)^9\approx 0.166.

The binomial mean and standard deviation are

E(X)=np,SD⁡(X)=np(1−p).E(X)=np, \qquad \operatorname{SD}(X)=\sqrt{np(1-p)}.

Takeaway: Choose the for a fixed number of comparable, independent success-or-failure trials.

3. Waiting for Successes

When the number of trials is not fixed in advance, focus on how many trials are needed to reach a target number of successes.

The counts trials until the first success. If the success probability is pp, then

P(X=k)=(1−p)k−1p.P(X=k)=(1-p)^{k-1}p.

The first k−1k-1 trials must fail, followed by success on trial kk. For example, if a call has probability 0.250.25 of resolving an issue, the probability that the first resolution occurs on the fourth call is

P(X=4)=(0.75)3(0.25)=0.1055.P(X=4)=(0.75)^3(0.25)=0.1055.

The is memoryless:

P(X>s+t∣X>s)=P(X>t).P(X>s+t\mid X>s)=P(X>t).

The negative extends this idea by counting trials until the rrth success. If XX is the trial number of the rrth success, then

P(X=k)=(k−1r−1)pr(1−p)k−r,k=r,r+1,…P(X=k)=\binom{k-1}{r-1}p^r(1-p)^{k-r}, \qquad k=r,r+1,\ldots

For example, the negative binomial model can describe how many sales calls are needed to obtain five purchases.

Takeaway: Use the for the first success and the negative for a specified later success.

4. Sampling Without Replacement

The applies when a sample is drawn without replacement from a finite population. Let NN be the population size, KK the number of successes in the population, nn the sample size, and XX the number of successes selected. Then

P(X=k)=(Kk)(N−Kn−k)(Nn).P(X=k)=\frac{\binom{K}{k}\binom{N-K}{n-k}}{\binom{N}{n}}.

The key distinction from the binomial model is dependence: after one item is selected, the composition of the remaining population changes. For example, if a shipment contains 20 items, 5 of which are defective, and 4 items are selected without replacement, then

P(X=1)=(51)(153)(204).P(X=1)=\frac{\binom{5}{1}\binom{15}{3}}{\binom{20}{4}}.

When the population is very large relative to the sample, often with a sample no more than about 5% of the population, the dependence may be small enough that a binomial approximation is reasonable.

Takeaway: Sampling without replacement points to the ; sampling with effectively independent trials may point to the .

5. Event Counts and Rare Events

Use the for event counts in a fixed interval of time, length, area, or volume when events occur independently at a stable average rate. If the average number of events per interval is λ\lambda, then

X∼Poisson⁡(λ)X\sim\operatorname{Poisson}(\lambda)

and

P(X=k)=e−λλkk!.P(X=k)=e^{-\lambda}\frac{\lambda^k}{k!}.

The mean and are both λ\lambda, so the standard deviation is λ\sqrt{\lambda}. If a help desk receives an average of 3 urgent requests per hour, the probability of exactly 5 requests in one hour is

P(X=5)=e−3355!≈0.101.P(X=5)=e^{-3}\frac{3^5}{5!}\approx 0.101.

The interval matters. A rate of 3 requests per hour gives an expected count of 66 in two hours.

The can also approximate a when nn is large, pp is small, and the expected count npnp is moderate. Set

λ=np.\lambda=np.

Takeaway: Use Poisson for stable-rate event counts, and consider it as a rare-event approximation to the binomial when its assumptions are plausible.

6. Continuous Waiting-Time Models

Continuous waiting times and measurements require models based on density and area.

A on [a,b][a,b] treats all equal-length subintervals as equally likely. Its density is

fX(x)={1b−a,a≤x≤b,0,otherwise.f_X(x)= \begin{cases} \dfrac{1}{b-a},&a\le x\le b,\\ 0,&\text{otherwise}. \end{cases}

For a≤c≤d≤ba\le c\le d\le b,

P(c≤X≤d)=d−cb−a.P(c\le X\le d)=\frac{d-c}{b-a}.

For example, if a waiting time is equally likely to fall anywhere in a 12-minute interval, then X∼Uniform⁡(0,12)X\sim\operatorname{Uniform}(0,12), and

P(X≤3)=312=0.25.P(X\le 3)=\frac{3}{12}=0.25.

The models the continuous waiting time until the next event in a Poisson process. With rate λ\lambda,

P(X>x)=e−λxP(X>x)=e^{-\lambda x}

and

P(X≤x)=1−e−λx.P(X\le x)=1-e^{-\lambda x}.

It is memoryless, just like the :

P(X>s+t∣X>s)=P(X>t).P(X>s+t\mid X>s)=P(X>t).

If customers arrive at an average rate of 4 per hour, the probability of waiting more than 20 minutes, or 13\frac{1}{3} hour, is

P(X>13)=e−4/3≈0.264.P\left(X>\frac{1}{3}\right)=e^{-4/3}\approx 0.264.

Takeaway: Use uniform for genuinely equal likelihood across an interval and exponential for waiting times between events occurring at a stable rate.

7. Normal Models and Sample Means

The is symmetric and bell-shaped. It is determined by its mean μ\mu and standard deviation σ\sigma, and is written as

X∼N(μ,σ2).X\sim N(\mu,\sigma^2).

To calculate probabilities, standardize a value using the z-score:

Z=X−μσ.Z=\frac{X-\mu}{\sigma}.

Then Z∼N(0,1)Z\sim N(0,1). For example, if X∼N(68,32)X\sim N(68,3^2), a value of 74 inches has

z=74−683=2.z=\frac{74-68}{3}=2.

It is therefore two standard deviations above the mean. The empirical rule gives approximate coverage of 68%, 95%, and 99.7% within 1, 2, and 3 standard deviations of the mean, respectively.

The explains why normal models are useful for sample means. If independent observations have mean μ\mu and standard deviation σ\sigma, then for a sufficiently large sample size nn,

Xˉ≈N(μ,σ2n),\bar X\approx N\left(\mu,\frac{\sigma^2}{n}\right),

so

SD⁡(Xˉ)=σn.\operatorname{SD}(\bar X)=\frac{\sigma}{\sqrt n}.

For a sum Sn=X1+⋯+XnS_n=X_1+\cdots+X_n,

E(Sn)=nμ,SD⁡(Sn)=σn.E(S_n)=n\mu, \qquad \operatorname{SD}(S_n)=\sigma\sqrt n.

A normal model is less appropriate for strongly skewed data, data with an important hard lower bound, or clearly multimodal data.

Takeaway: Use normal models for suitable symmetric measurements and for sufficiently large-sample means when the central limit theorem conditions are defensible.

8. Comparing Models and Checking Assumptions

describes the long-run location of a random variable. For a discrete variable,

E(X)=∑xxP(X=x),E(X)=\sum_x xP(X=x),

and for a continuous variable,

E(X)=∫−∞∞xfX(x) dx.E(X)=\int_{-\infty}^{\infty}x f_X(x)\,dx.

describes spread around the mean:

Var⁡(X)=E[(X−μ)2]=E(X2)−[E(X)]2.\operatorname{Var}(X)=E[(X-\mu)^2]=E(X^2)-[E(X)]^2.

The standard deviation is the square root of the . Linear transformations obey

E(aX+b)=aE(X)+b,E(aX+b)=aE(X)+b,

and

Var⁡(aX+b)=a2Var⁡(X).\operatorname{Var}(aX+b)=a^2\operatorname{Var}(X).

If XX and YY are independent, then means and variances add as follows:

E(X+Y)=E(X)+E(Y),E(X+Y)=E(X)+E(Y),
Var⁡(X+Y)=Var⁡(X)+Var⁡(Y).\operatorname{Var}(X+Y)=\operatorname{Var}(X)+\operatorname{Var}(Y).

Independence is essential for the simple addition rule. Without independence, covariance terms may be required.

To select a model, identify the variable first, then check whether it is discrete or continuous. Next ask whether there is a fixed number of trials, a waiting target, sampling without replacement, or a stable event rate. Finally, compare the model assumptions with the actual context.

Takeaway: A formula is useful only when the model's mechanism and assumptions match the process being described.