5 Expectation and Variability
A progressive guide to expectation, variance, standard deviation, covariance, and sums of random variables, with formulas, examples, applications, and common errors to avoid.
The Center:
summarizes the center of a probability distribution. It is a probability-weighted average and represents the long-run average over repeated observations.
For a discrete random variable with possible values and probability mass function ,
For a continuous random variable with density ,
provided the is well-defined. The expected value does not have to be a value that the variable can actually take. For a fair six-sided die,
No individual roll equals , but the average approaches this value over many rolls.
For a function , average the function values with the same probabilities:
in the discrete case, with the corresponding integral for a continuous variable.
Takeaway: describes the center or long-run average, not necessarily a directly observable outcome.
Adding Expected Values
makes totals easier to analyze. For constants and ,
More generally,
This rule holds whether or not the random variables are independent. For example, let indicate whether trial is a success:
If the success probability on trial is , then . For the total number of successes ,
The trials may be dependent; the formula still applies.
Takeaway: Add expected contributions directly, and do not impose an unnecessary assumption.
Spread Around the Mean:
measures spread around the mean. If , its definition is
A useful computational form is
For a fair die, and
Therefore,
is never negative. It is zero exactly when the random variable is constant with probability one. Transformations behave as follows:
Adding a constant changes the center but not the spread. Multiplying by scales the by .
Takeaway: quantifies squared spread and must be interpreted in squared units.
Interpreting
converts back to the original measurement scale:
For the fair die,
Unlike , has the same units as the random variable. It is not necessarily the exact average distance from the mean, but it is a useful measure of typical spread. Interpret its size relative to the scale and units of the variable.
For independent measurements with standard deviations and , the variances are and . The of their sum is found by adding variances first:
The standard deviations do not generally add.
Takeaway: Use for algebraic calculations, then take its square root to obtain an interpretable spread in the original units.
How Variables Move Together
describes how two variables move together relative to their means:
An equivalent form is
Positive means that large values of tend to occur with large values of , while small values tend to occur together. Negative means that large values of one variable tend to occur with small values of the other. near zero indicates little linear co-movement, but it does not rule out nonlinear dependence.
is the special case
Scaling follows the rule
Adding constants has no effect, whereas multiplying the variables changes according to the product of the scaling factors.
implies zero when the relevant expectations exist. However, zero generally does not imply .
Takeaway: records joint linear movement, including whether co-movement increases or decreases the spread of a total.
Combining Uncertain Quantities
The depends on both individual variability and the relationships among components. For two variables,
Positive increases the of the sum. Negative reduces it. For a difference,
For constants and ,
If and are independent, is zero, so their variances add. For example, if , , and , then
Takeaway: Never add standard deviations directly; account for and add variances when permits it.
Many-Variable Sums
For a total , linearity gives
The includes every pairwise :
If the variables are mutually independent, all terms vanish:
If they are also identically distributed with mean and , then
Thus, the expected total grows proportionally to , while the of an independent total grows proportionally to .
Takeaway: Repeated independent contributions accumulate their means linearly but their standard deviations more slowly, at a square-root rate.
From Distributions to Data
Observed data use sample statistics to estimate or summarize the corresponding distributional quantities. For data , the sample mean is
The commonly used and sample are
The denominator is used when the estimates the of a larger population. When describing an entire finite population rather than estimating a larger one, a denominator of is often used.
For paired observations , the sample is
These statistics help summarize center and spread, quantify uncertainty, analyze relationships, and evaluate risk. For example, a portfolio's expected return is a weighted sum of expected returns, while its depends on both individual asset variances and covariances. Assets with low or negative can reduce overall variability.
Keep population quantities distinct from sample quantities: , , and describe a distribution or population, whereas , , and are calculated from observed data.
Takeaway: Choose the statistic and denominator according to whether the goal is to describe the observed population or estimate a larger one.
Common Errors and Final Checklist
Several mistakes recur when working with and variability.
Adding standard deviations instead of variances: For independent variables, variances add; standard deviations generally do not.
Assuming is needed for expected values to add: holds without .
Assuming zero means : implies zero , but zero does not generally imply .
Ignoring units: has squared units, while has the original units.
Confusing population parameters with sample statistics: , , and describe a distribution or population; , , and are calculated from observed data.
Final takeaway: First identify whether the task concerns a center, individual spread, joint movement, or the variability of a total. Then select the corresponding , , , or formula and check the assumptions, units, and denominator.