5 Expectation and Variability

A progressive guide to expectation, variance, standard deviation, covariance, and sums of random variables, with formulas, examples, applications, and common errors to avoid.

The Center:

summarizes the center of a probability distribution. It is a probability-weighted average and represents the long-run average over repeated observations.

For a discrete random variable XX with possible values xx and probability mass function pX(x)p_X(x),

E[X]=∑xxpX(x).E[X]=\sum_x x p_X(x).

For a continuous random variable with density fX(x)f_X(x),

E[X]=∫−∞∞xfX(x) dx,E[X]=\int_{-\infty}^{\infty}x f_X(x)\,dx,

provided the is well-defined. The expected value does not have to be a value that the variable can actually take. For a fair six-sided die,

E[X]=1(16)+2(16)+⋯+6(16)=3.5.E[X]=1\left(\frac{1}{6}\right)+2\left(\frac{1}{6}\right)+\cdots+6\left(\frac{1}{6}\right)=3.5.

No individual roll equals 3.53.5, but the average approaches this value over many rolls.

For a function g(X)g(X), average the function values with the same probabilities:

E[g(X)]=∑xg(x)pX(x)E[g(X)]=\sum_x g(x)p_X(x)

in the discrete case, with the corresponding integral for a continuous variable.

Takeaway: describes the center or long-run average, not necessarily a directly observable outcome.

Adding Expected Values

makes totals easier to analyze. For constants aa and bb,

E[aX+b]=aE[X]+b.E[aX+b]=aE[X]+b.

More generally,

E[∑i=1naiXi+b]=∑i=1naiE[Xi]+b.E\left[\sum_{i=1}^n a_iX_i+b\right]=\sum_{i=1}^n a_iE[X_i]+b.

This rule holds whether or not the random variables are independent. For example, let XiX_i indicate whether trial ii is a success:

Xi={1,if trial i is a success,0,otherwise.X_i=\begin{cases} 1,&\text{if trial }i\text{ is a success},\\ 0,&\text{otherwise}. \end{cases}

If the success probability on trial ii is pip_i, then E[Xi]=piE[X_i]=p_i. For the total number of successes S=X1+⋯+XnS=X_1+\cdots+X_n,

E[S]=p1+⋯+pn.E[S]=p_1+\cdots+p_n.

The trials may be dependent; the formula still applies.

Takeaway: Add expected contributions directly, and do not impose an unnecessary assumption.

Spread Around the Mean:

measures spread around the mean. If μ=E[X]\mu=E[X], its definition is

Var⁡(X)=E[(X−μ)2].\operatorname{Var}(X)=E[(X-\mu)^2].

A useful computational form is

Var⁡(X)=E[X2]−(E[X])2.\operatorname{Var}(X)=E[X^2]-\bigl(E[X]\bigr)^2.

For a fair die, E[X]=3.5E[X]=3.5 and

E[X2]=12+22+32+42+52+626=916.E[X^2]=\frac{1^2+2^2+3^2+4^2+5^2+6^2}{6}=\frac{91}{6}.

Therefore,

Var⁡(X)=916−(3.5)2=3512≈2.917.\operatorname{Var}(X)=\frac{91}{6}-(3.5)^2=\frac{35}{12}\approx 2.917.

is never negative. It is zero exactly when the random variable is constant with probability one. Transformations behave as follows:

Var⁡(c)=0,Var⁡(aX+b)=a2Var⁡(X).\operatorname{Var}(c)=0, \qquad \operatorname{Var}(aX+b)=a^2\operatorname{Var}(X).

Adding a constant changes the center but not the spread. Multiplying by aa scales the by a2a^2.

Takeaway: quantifies squared spread and must be interpreted in squared units.

Interpreting

converts back to the original measurement scale:

SD⁡(X)=σX=Var⁡(X).\operatorname{SD}(X)=\sigma_X=\sqrt{\operatorname{Var}(X)}.

For the fair die,

σX=3512≈1.71.\sigma_X=\sqrt{\frac{35}{12}}\approx 1.71.

Unlike , has the same units as the random variable. It is not necessarily the exact average distance from the mean, but it is a useful measure of typical spread. Interpret its size relative to the scale and units of the variable.

For independent measurements with standard deviations 33 and 44, the variances are 99 and 1616. The of their sum is found by adding variances first:

Var⁡(X+Y)=9+16=25,SD⁡(X+Y)=25=5.\operatorname{Var}(X+Y)=9+16=25, \qquad \operatorname{SD}(X+Y)=\sqrt{25}=5.

The standard deviations do not generally add.

Takeaway: Use for algebraic calculations, then take its square root to obtain an interpretable spread in the original units.

How Variables Move Together

describes how two variables move together relative to their means:

Cov⁡(X,Y)=E[(X−E[X])(Y−E[Y])].\operatorname{Cov}(X,Y)=E[(X-E[X])(Y-E[Y])].

An equivalent form is

Cov⁡(X,Y)=E[XY]−E[X]E[Y].\operatorname{Cov}(X,Y)=E[XY]-E[X]E[Y].

Positive means that large values of XX tend to occur with large values of YY, while small values tend to occur together. Negative means that large values of one variable tend to occur with small values of the other. near zero indicates little linear co-movement, but it does not rule out nonlinear dependence.

is the special case

Cov⁡(X,X)=Var⁡(X).\operatorname{Cov}(X,X)=\operatorname{Var}(X).

Scaling follows the rule

Cov⁡(aX+b,cY+d)=acCov⁡(X,Y).\operatorname{Cov}(aX+b,cY+d)=ac\operatorname{Cov}(X,Y).

Adding constants has no effect, whereas multiplying the variables changes according to the product of the scaling factors.

implies zero when the relevant expectations exist. However, zero generally does not imply .

Takeaway: records joint linear movement, including whether co-movement increases or decreases the spread of a total.

Combining Uncertain Quantities

The depends on both individual variability and the relationships among components. For two variables,

Var⁡(X+Y)=Var⁡(X)+Var⁡(Y)+2Cov⁡(X,Y).\operatorname{Var}(X+Y)=\operatorname{Var}(X)+\operatorname{Var}(Y)+2\operatorname{Cov}(X,Y).

Positive increases the of the sum. Negative reduces it. For a difference,

Var⁡(X−Y)=Var⁡(X)+Var⁡(Y)−2Cov⁡(X,Y).\operatorname{Var}(X-Y)=\operatorname{Var}(X)+\operatorname{Var}(Y)-2\operatorname{Cov}(X,Y).

For constants aa and bb,

Var⁡(aX+bY)=a2Var⁡(X)+b2Var⁡(Y)+2abCov⁡(X,Y).\operatorname{Var}(aX+bY)=a^2\operatorname{Var}(X)+b^2\operatorname{Var}(Y)+2ab\operatorname{Cov}(X,Y).

If XX and YY are independent, is zero, so their variances add. For example, if Var⁡(X)=9\operatorname{Var}(X)=9, Var⁡(Y)=16\operatorname{Var}(Y)=16, and Cov⁡(X,Y)=2\operatorname{Cov}(X,Y)=2, then

Var⁡(X+Y)=9+16+2(2)=29.\operatorname{Var}(X+Y)=9+16+2(2)=29.

Takeaway: Never add standard deviations directly; account for and add variances when permits it.

Many-Variable Sums

For a total S=X1+X2+⋯+XnS=X_1+X_2+\cdots+X_n, linearity gives

E[S]=∑i=1nE[Xi].E[S]=\sum_{i=1}^nE[X_i].

The includes every pairwise :

Var⁡(S)=∑i=1nVar⁡(Xi)+2∑1≤i<j≤nCov⁡(Xi,Xj).\operatorname{Var}(S)=\sum_{i=1}^n\operatorname{Var}(X_i)+2\sum_{1\leq i<j\leq n}\operatorname{Cov}(X_i,X_j).

If the variables are mutually independent, all terms vanish:

Var⁡(S)=∑i=1nVar⁡(Xi).\operatorname{Var}(S)=\sum_{i=1}^n\operatorname{Var}(X_i).

If they are also identically distributed with mean μ\mu and σ2\sigma^2, then

E[S]=nμ,Var⁡(S)=nσ2,SD⁡(S)=σn.E[S]=n\mu, \qquad \operatorname{Var}(S)=n\sigma^2, \qquad \operatorname{SD}(S)=\sigma\sqrt{n}.

Thus, the expected total grows proportionally to nn, while the of an independent total grows proportionally to n\sqrt{n}.

Takeaway: Repeated independent contributions accumulate their means linearly but their standard deviations more slowly, at a square-root rate.

From Distributions to Data

Observed data use sample statistics to estimate or summarize the corresponding distributional quantities. For data x1,…,xnx_1,\ldots,x_n, the sample mean is

xˉ=1n∑i=1nxi.\bar{x}=\frac{1}{n}\sum_{i=1}^n x_i.

The commonly used and sample are

s2=1n−1∑i=1n(xi−xˉ)2,s=s2.s^2=\frac{1}{n-1}\sum_{i=1}^n(x_i-\bar{x})^2, \qquad s=\sqrt{s^2}.

The denominator n−1n-1 is used when the estimates the of a larger population. When describing an entire finite population rather than estimating a larger one, a denominator of nn is often used.

For paired observations (xi,yi)(x_i,y_i), the sample is

sXY=1n−1∑i=1n(xi−xˉ)(yi−yˉ).s_{XY}=\frac{1}{n-1}\sum_{i=1}^n(x_i-\bar{x})(y_i-\bar{y}).

These statistics help summarize center and spread, quantify uncertainty, analyze relationships, and evaluate risk. For example, a portfolio's expected return is a weighted sum of expected returns, while its depends on both individual asset variances and covariances. Assets with low or negative can reduce overall variability.

Keep population quantities distinct from sample quantities: μ\mu, σ2\sigma^2, and σ\sigma describe a distribution or population, whereas xˉ\bar{x}, s2s^2, and ss are calculated from observed data.

Takeaway: Choose the statistic and denominator according to whether the goal is to describe the observed population or estimate a larger one.

Common Errors and Final Checklist

Several mistakes recur when working with and variability.

  1. Adding standard deviations instead of variances: For independent variables, variances add; standard deviations generally do not.

  2. Assuming is needed for expected values to add: holds without .

  3. Assuming zero means : implies zero , but zero does not generally imply .

  4. Ignoring units: has squared units, while has the original units.

  5. Confusing population parameters with sample statistics: μ\mu, σ2\sigma^2, and σ\sigma describe a distribution or population; xˉ\bar{x}, s2s^2, and ss are calculated from observed data.

Final takeaway: First identify whether the task concerns a center, individual spread, joint movement, or the variability of a total. Then select the corresponding , , , or formula and check the assumptions, units, and denominator.