6. Probability and Randomness: Models, Rules, and Independence

A progressive guide to modeling uncertainty, applying core probability rules, and interpreting dependence and independence in real-world processes.

6.1 Modeling Outcomes and Uncertainty

Probability provides a mathematical language for uncertainty. It describes what tends to happen across many repetitions or within a population rather than guaranteeing the result of one trial. For example, a forecast of a 70%70\% chance of rain means that comparable situations would be expected to produce rain about 70%70\% of the time; it does not mean that rain will last 70%70\% of the day.

A useful model begins by identifying three components:

  1. The possible outcomes.

  2. The events, which are collections of outcomes.

  3. The probability assigned to each outcome or .

Probabilities must lie between 00 and 11, inclusive, and the probability of the entire must equal 11. When outcomes are equally likely, the probability of an is

P(A)=number of outcomes in Anumber of outcomes in S.P(A)=\frac{\text{number of outcomes in }A}{\text{number of outcomes in }S}.

For a fair six-sided die, the is S={1,2,3,4,5,6}S=\{1,2,3,4,5,6\}. The of rolling an even number is A={2,4,6}A=\{2,4,6\}, so P(A)=3/6=1/2P(A)=3/6=1/2. The equally likely formula does not automatically apply to real-world outcomes such as equipment failures or customer arrivals, which may have unequal probabilities.

Probabilities can come from a theoretical assumption, long-run relative frequencies, or a statistical model. In every case, the model is an approximation whose quality depends on whether its assumptions represent the process well.

Takeaway: A must describe the relevant outcomes completely and assign probabilities that are consistent with the process.

6.2 Complements, Intersections, and Unions

The rule handles outcomes that do not belong to an :

P(Ac)=1−P(A).P(A^c)=1-P(A).

If the probability that a package arrives late is 0.080.08, then the probability that it is not late is 1−0.08=0.921-0.08=0.92. The same idea gives the useful rule

P(at least one)=1−P(none).P(\text{at least one})=1-P(\text{none}).

For “and” questions, use the A∩BA\cap B. For “or” questions, use the A∪BA\cup B. The general addition rule is

P(A∪B)=P(A)+P(B)−P(A∩B).P(A\cup B)=P(A)+P(B)-P(A\cap B).

The is subtracted because it is counted in both P(A)P(A) and P(B)P(B). For example, if 40%40\% of households subscribe to Service A, 30%30\% subscribe to Service B, and 10%10\% subscribe to both, then

P(A∪B)=0.40+0.30−0.10=0.60.P(A\cup B)=0.40+0.30-0.10=0.60.

Thus, 60%60\% subscribe to at least one service. If two events cannot happen together, their has probability 00, and their probabilities can be added directly.

Takeaway: Translate the wording first: “and” means an , “or” means a , and overlapping events require a double-counting correction.

6.3 Conditional and Sequential Probability

A restricts attention to cases in which a specified condition is known to hold. If P(B)>0P(B)>0, then

P(A∣B)=P(A∩B)P(B).P(A\mid B)=\frac{P(A\cap B)}{P(B)}.

For a standard deck, there are 1313 hearts and exactly one ace of hearts. Therefore,

P(ace∣heart)=113.P(\text{ace}\mid\text{heart})=\frac{1}{13}.

The denominator is the number of hearts because the condition changes the reference group from all 5252 cards to the 1313 hearts.

Two-way tables make this denominator choice visible. Suppose 9696 people use a fitness app, and 7272 of those people exercise regularly. Then

P(regular exercise∣app use)=7296=0.75.P(\text{regular exercise}\mid\text{app use})=\frac{72}{96}=0.75.

The denominator is 9696, not the total of 200200, because the question is about app users.

Rearranging the formula produces the :

P(A∩B)=P(B)P(A∣B)=P(A)P(B∣A).P(A\cap B)=P(B)P(A\mid B)=P(A)P(B\mid A).

For a box containing 55 red and 33 blue markers, the probability of selecting two red markers without replacement is

P(two red)=58⋅47=514≈0.357.P(\text{two red})=\frac{5}{8}\cdot\frac{4}{7}=\frac{5}{14}\approx 0.357.

The second factor is 4/74/7, not 5/85/8, because the first red marker was not replaced.

Takeaway: In a , identify the condition and use the size or probability of that restricted group as the denominator.

6.4 Independence and Dependence

Independence means that learning one occurred does not change the probability of the other. The main tests are

P(A∣B)=P(A)P(A\mid B)=P(A)

and

P(A∩B)=P(A)P(B).P(A\cap B)=P(A)P(B).

For two independent fair coin tosses,

P(two heads)=12⋅12=14.P(\text{two heads})=\frac{1}{2}\cdot\frac{1}{2}=\frac{1}{4}.

Sampling without replacement usually creates dependence. If the first card drawn from a standard deck is an ace, then only 33 aces remain among 5151 cards, so

P(second ace∣first ace)=351,P(\text{second ace}\mid\text{first ace})=\frac{3}{51},

which differs from the unconditional probability 4/524/52. Replacing the first card and reshuffling restores independence in the usual model.

Mutual exclusivity and independence are not the same. cannot occur together. If both have positive probability, knowing that one occurred makes the other impossible, so the events are dependent rather than independent.

Apparent association also does not establish a direct causal relationship. Umbrella purchases and traffic accidents may both increase during rainy weather, while weather acts as a common influence on both.

Takeaway: Check how the process works before multiplying probabilities. Replacement, shared conditions, and common causes can create dependence.

6.5 Trees for Sequential Processes

Sequential processes can be organized with a . Each branch represents a possible next outcome, and the probability of a complete path is found by multiplying the probabilities along that path.

Suppose a machine produces acceptable items with probability 0.900.90, and two selections are independent. The probability that both items are acceptable is

0.90×0.90=0.81.0.90\times 0.90=0.81.

The probability that the first is acceptable and the second is defective is

0.90×0.10=0.09.0.90\times 0.10=0.09.

For a dependent process, later branch probabilities change after earlier outcomes. This occurs in sampling without replacement, repeated equipment failures, and other processes in which earlier results alter the situation. After constructing a tree, verify that the probabilities of all complete paths add to 11.

Takeaway: Trees make sequential structure explicit: multiply along paths, and account for changing branch probabilities when events are dependent.

6.6 Building and Evaluating Real-World Models

A real-world should be built and interpreted in context. Use the following progression:

  1. Define the random process and what counts as one trial.

  2. Identify a that is complete for the question.

  3. Define events that represent the practical questions.

  4. Assign probabilities using theory, historical data, or a statistical model.

  5. Check assumptions about equal likelihood, independence, and stability over time.

  6. Calculate the result and explain what it means, including important limitations.

For example, a coffee shop may model the number of customers arriving between 8:00 and 8:15 a.m. with outcomes S={0,1,2,3,…}S=\{0,1,2,3,\ldots\}. If historical data suggest that the probability of at least 1010 arrivals is 0.200.20, then

P(fewer than 10 arrivals)=1−0.20=0.80.P(\text{fewer than }10\text{ arrivals})=1-0.20=0.80.

This result can support staffing decisions, but it is dependable only when the historical data represent similar conditions. A holiday, severe weather, or nearby road construction could change the arrival pattern.

Similarly, if a factory estimates that 2%2\% of products are defective, an independence assumption may support calculations for a sample of 1010 products. If a malfunctioning machine causes defects to cluster, that assumption may underestimate the chance of observing several defects together.

Common errors include reversing P(A∣B)P(A\mid B) and P(B∣A)P(B\mid A), adding overlapping probabilities without subtracting the , assuming independence without justification, and treating a probability as a guarantee. A probability of 0.900.90 does not ensure an occurs in one trial, and a probability of 0.010.01 does not make an impossible.

Final takeaway: Probability calculations are only as trustworthy as the assumptions about the process. Interpret every result in the setting that produced it.