Join · scan

OTTO-PR1

ottolearn.org/join

Module 2 · Probability Foundations · PR-1

Basic Probability
"Tonight I'm due a good night… right?"

The position piece runs, and one reply sticks: "Six short nights in a row. Statistically, tonight has to be a good one — I'm due." A Toronto casino once banned a player for betting $200,000 on exactly that logic. Before Maya can answer either of them, she needs the mathematics of chance — starting from zero.

Live session · join with code OTTO-PR1 · ottolearn.org/join

Same rules as always: vote before I explain, every poll is anonymous, and changing your mind after discussion is the goal.

Today · the reply Maya owes her reader

One reader's streak, five questions about chance.

Five questions, one mailbag reply. Every one starts with your vote.

Recall · last class · 45 seconds

Warm-up · 60 seconds · anonymous

Hold your vote. Today is about where probability numbers legitimately come from — and the counting shortcut has fine print.

Core concept 1 · C1

The sample space: what could happen at all?

Before any number: list the outcomes. An experiment is any process whose outcome is uncertain — a die roll, a coin flip, spin #10 of a roulette wheel.

The sample space

The set of every outcome the experiment could produce. Roll a die: \(S = \{1, 2, 3, 4, 5, 6\}\). Nothing outside the list can happen.

\(S\)

An event

Any collection of outcomes you care about — a subset of \(S\). "Roll an even number" is the event \(A = \{2, 4, 6\}\).

\(A \subseteq S\)

The event occurs

The die shows 4 → the outcome landed inside \(A\), so \(A\) occurred. One roll can make several events occur at once: 4 is even and greater than 3.

Every tool today is built on this picture — outcomes in a set, events as subsets. Get the sample space right and the rest is counting.

Core concept 2 · C2

Classical probability: count the shares

The counting formula

When every outcome in \(S\) is equally likely, probability is a counting ratio: outcomes in \(A\) over outcomes in \(S\). Even number on a fair die: \(P(A) = 3/6 = 1/2\).

\(P(A) = \dfrac{|A|}{|S|}\)

The fine print

One condition: equally likely. Fair dice, fair coins, shuffled decks qualify. A bent coin still has \(S = \{H, T\}\) — but \(P(H) \neq 1/2\). The formula fails the moment the outcomes aren't uniform.

The warm-up, settled

"Rain or no rain — 50%" applies the counting formula to outcomes that are not equally likely. Two outcomes never means \(1/2\) — that's the fine print doing its job.

When the outcomes aren't equally likely, counting is worthless — you need data. That's the next tool, and it's the one the reader's streak needs.

Practice · classical probability · fresh numbers on demand

Commit first · the reader's logic, tested

Core concept 3 · C3

Empirical probability: let the data speak

Estimate from trials

No equal-likelihood to lean on? Run the experiment \(n\) times, count how often \(A\) happened (\(f\)), and estimate the probability by the relative frequency — the same ratio you've computed since the first data class.

\(P(A) \approx \dfrac{f}{n}\)

The Law of Large Numbers

As \(n\) grows, \(f/n\) settles toward the true probability. It converges by averaging over ever more trials — not by tails "catching up." No debt is ever repaid.

Maya's own survey

Of her 300 sleep responses, 81 reported under 6 hours. Last month that ratio described her data; today it predicts: pick a respondent at random and \(P(\text{under 6 h}) = 0.27\).

\(\tfrac{81}{300} = 0.27\)

Watch both claims at once: the proportion converges — and the heads-tails gap does not shrink. Convergence comes from averaging, not from evening out.

See it · C3

The coin has no memory

Predict first: after a long run of tails, does the running proportion snap back? Flip in bursts, watch the streak strip — and notice the predictor never budges from \(0.50\).

Two stacked line charts over the same series of coin flips. The top chart tracks the running proportion of heads, which converges toward the dashed line at 0.50. The bottom chart tracks the absolute gap between the number of heads and tails on a fixed axis, with a dashed √n reference curve; the gap wanders and tends to grow rather than shrink — illustrating that proportions converge while counts do not catch up.
Recent flips:

Click a button to start flipping.

Core concept 4 · C4 · enrichment

Subjective probability: a degree of belief

When there's no experiment

Some events can't be repeated: this start-up succeeding, this election, this exam. A number can still express an informed degree of belief — "I'd put it at 30%."

Belief isn't exempt

A subjective probability must still obey every rule of probability: between 0 and 1, complements summing to 1. A gut feeling that breaks the axioms isn't a probability — it's a mistake with confidence.

Three sources, one test

Counting (classical), data (empirical), judgment (subjective) — every probability you'll ever meet comes from one of the three. Ask: could I repeat this? Are the outcomes equally likely?

For the curious: classifying the type is context, not an exam skill — but it sharpens how you read every "70% chance" in the news.

Try it · C4

Classify the claim

Decide the type before flipping each card. Where does Maya's 0.27 sit? And the forum's "50% rain"?

Pick the type for each scenario below.

Read the scenario, then pick the type of probability it is:

Core concept 5 · C5

The axioms: rules every probability obeys

Stay on the scale

Every probability lives between 0 (impossible) and 1 (certain). A claimed probability of 1.05 — or −0.2 — is broken on arrival, whatever its source.

\(0 \leq P(A) \leq 1\)

Something must happen

The full sample space has probability exactly 1. This is what makes complements work: an event and its opposite split the whole 1 between them.

\(P(S) = 1\)

Nothing from nothing

The empty event — no outcome at all — has probability 0. And when two events can't overlap, their "or" is a plain sum. (The fine print on that arrives in a few slides.)

\(P(\emptyset) = 0\)

The axioms are the fraud detector: any claimed set of probabilities that breaks them is wrong before you check anything else.

See it · C5

Sums can leave the scale — probabilities can't

Predict first: push \(P(A)\) and \(P(B)\) until their raw sum passes 1 — which axiom, if any, breaks? Then drag the overlap and watch the union needle pull back inside the valid zone.

0 Impossible 0.5 Fair coin 1 Certain
0.60
0.50
0.30

Quick check · spot the broken claim

Core concept 6 · C6

The complement: the subtraction shortcut

The rule

Everything not in \(A\) is the complement \(A'\). Together they fill \(S\) and never overlap — so they split 1 between them: whatever \(A\) doesn't take, \(A'\) gets.

\(P(A') = 1 - P(A)\)

The strategy

When an event is hard to count directly, count what it isn't. "At least one…" questions almost always fall to this: compute the probability of none, subtract from 1.

The move is a decision, not just a formula: spot the complement, then subtract instead of recounting. Watch me make that call out loud:

I do · watch me choose the shortcut

  1. 1 I do
  2. 2 We do
  3. 3 You do

Full support — watch me decide to flip to the complement, then subtract.

Count what it isn't, out loud

Step 1 · Spot the complement

The forecast gives \(P(\text{rain}) = 0.35\); Maya needs \(P(\text{no rain})\) for a survey-day plan. "Rain" and "no rain" fill every outcome and can't overlap — textbook complements.

Step 2 · Subtract, don't recount

Rather than tally every dry scenario, I take the whole \(1\) and remove the slice I don't want: \(1 - 0.35\).

Step 3 · Compute and sanity-check

\(= 0.65\). Check: it sits in \([0,1]\) ✓, and \(0.35 + 0.65 = 1\) ✓. Bigger than the rain chance — which fits a mostly-dry forecast.

The habit: name the complement, subtract from 1, check the sum. Now Maya's own survey number — you take the last step.

We do · you finish it

  1. 1 I do
  2. 2 We do
  3. 3 You do

Support fades — the setup is given; you supply the subtraction and the warning.

The other side of 0.27

From the survey: \(P(\text{a random respondent sleeps under 6 h}) = 0.27\).

  1. "Under 6 h" and "6 h or more" cover every respondent and can't both happen — they are complements, so their probabilities must split 1.
  2. So \(P(\text{6 h or more}) = ?\) — and what should Maya warn the reader this number is not?

    \(1 - 0.27 = 0.73\). And it is not a promise about tonight after a streak — it's a long-run share of respondents. The streak doesn't move it.

Practice · the complement rule · fresh numbers on demand

  1. 1 I do
  2. 2 We do
  3. 3 You do

On your own now — fresh numbers, every step yours. Re-roll as many as the room needs.

Core concept 7 · C7

"Or" and "and": two ways to combine events

Union — "or"

\(A \cup B\): at least one of the two happens — \(A\), \(B\), or both. In a two-way table the union covers three cells: each "only" cell plus the overlap.

\(A \cup B\)

Intersection — "and"

\(A \cap B\): both happen at once. In the table it is a single cell — the one where the row and the column cross.

\(A \cap B\)

You've been reading these since the first data class — every cell of a two-way table is an intersection, every row total a marginal. Today the counts become probabilities.

See it · C7

Every cell is an event

Predict first: click the top-left cell — which Venn region lights up? Then flip to probabilities and find the cell whose value is \(P(A \cap B)\).

Survey of 120 students — click any cell

Passed B Failed B′ Total
Groups A 48 12 60
Alone A′ 42 18 60
Total 90 30 120

Venn diagram (n = 120)

An area-proportional Euler diagram with two overlapping circles. Circle A (studied in groups, 60 students) is smaller than circle B (passed, 90 students), and their overlap area is proportional to the 48 students who did both. Click any count to highlight its region; the matching table cell highlights too.

Click any highlighted cell or total to see the corresponding Venn region.

Which number? · mixes DS-1 + PR-1

Core concept 8 · C8

The General Addition Rule: subtract the overlap once

Why plain adding fails

Add \(P(A) + P(B)\) and every outcome in the overlap is counted twice — once inside each event. The raw sum can even sail past 1, as the thermometer just showed.

The repair

Subtract the overlap once and the double-count is exactly undone. This works for any two events — overlapping or not.

\(P(A \cup B) = P(A) + P(B) - P(A \cap B)\)

Watch the decision, not just the formula. The expert move is a check that happens before any arithmetic:

See it · C8

Watch the double-count happen

Predict first: with the overlap at zero, how do \(P(A) + P(B)\) and \(P(A \cup B)\) compare? Then grow the overlap and watch the two pull apart by exactly \(P(A \cap B)\).

A 10 by 10 grid of 100 dots representing the sample space S. Dots are grouped into columns by region — A-only (orange), the A∩B overlap (green), B-only (blue), and neither (faint). Two translucent rounded rectangles mark event A and event B; they overlap over the green intersection columns, showing that those dots belong to both events. Sliders control region sizes.

Each dot is 1 of 100 students. A = plays a sport, B = is in a club.

30
20
10

I do · watch the whole decision

  1. 1 I do
  2. 2 We do
  3. 3 You do

Full support — watch me make every decision, especially the overlap check.

Hearts or face cards, out loud

Step 1 · Name the events

Draw one card from a fair 52-card deck. \(A\) = heart, \(B\) = face card (J, Q, K). Fair deck → equally likely → classical: \(P(A) = \tfrac{13}{52}\), \(P(B) = \tfrac{12}{52}\).

Step 2 · Check the overlap first

Before I add anything: can one card be both? Yes — J♥, Q♥, K♥. So \(P(A \cap B) = \tfrac{3}{52} \neq 0\), and plain adding would count those three cards twice.

Step 3 · Apply the rule

\(P(A \cup B) = \tfrac{13}{52} + \tfrac{12}{52} - \tfrac{3}{52} = \tfrac{22}{52}\). Subtracting once removes the second copy of each overlap card — they stay counted exactly once.

Step 4 · Sanity-check

\(\tfrac{22}{52} \approx 0.42\): on the 0–1 scale, larger than either event alone, smaller than the raw sum \(\tfrac{25}{52}\). Exactly what a repaired double-count should look like.

The habit: overlap first, arithmetic second. Now Maya's own table — you take the last step.

We do · you finish it

  1. 1 I do
  2. 2 We do
  3. 3 You do

Support fades — the setup is done for you; you commit to the final step.

"Part-time or short on sleep?"

Same crosstab as the poll: \(P(\text{part-time}) = 0.40\), \(P(\text{under 6 h}) = 0.27\), \(P(\text{both}) = 0.15\).

  1. Overlap check first: \(P(\text{both}) = 0.15 \neq 0\) — the events overlap, so the full rule applies.
  2. Finish it: \(P(\text{part-time OR under 6 h}) = ?\)

    \(0.40 + 0.27 - 0.15 = 0.52\). Just over half — and without the subtraction Maya would have printed 0.67, counting all 45 double-dippers twice.

Practice · the General Addition Rule · fresh numbers on demand

  1. 1 I do
  2. 2 We do
  3. 3 You do

On your own now — fresh numbers, every step yours. Re-roll as many as the room needs.

Core concept 9 · C9

Mutually exclusive: when the overlap is empty

The definition

Two events are mutually exclusive when they cannot both happen on the same trial — the overlap is empty. One die roll can't show both a 1 and a 6.

\(P(A \cap B) = 0\)

The shortcut it buys

Nothing to double-count → the addition rule sheds its correction: \(P(A \cup B) = P(A) + P(B)\). A special case — the general rule still holds; the subtraction is just \(-\,0\).

Verify, never assume

"They're about different things" is not a proof. "Even" and "greater than 4" sound unrelated — yet 6 is both. The only test is the sample space: is \(P(A \cap B) = 0\)?

A student ran the shortcut without the check. One line of the work below is wrong — find it before the votes reveal it.

Find the flaw · one line is wrong

Commit first · the semester's stickiest trap

Core concept 10 · C10

Mutually exclusive is not independent

ME: about overlap

A statement about the sample space: the two events share no outcomes. In the Venn picture, the circles don't touch.

\(P(A \cap B) = 0\)

Independent: about information

A statement about knowledge: learning that \(B\) occurred leaves the probability of \(A\) untouched. The coin's five tails told us nothing about flip #6 — that's independence at work. (Made precise next class.)

\(P(A \mid B) = P(A)\)

ME forces dependence

If both events have positive probability and can't co-occur, one happening slams the other to 0. Mutually exclusive events are never independent — they are maximally dependent.

The reader's fallacy and this trap are cousins: both invent a relationship ("due") or deny one ("separate") instead of checking what the sample space actually says.

See it · C10

One click, two different worlds

Predict first: in the mutually-exclusive panel, what happens to \(P(B)\) the instant \(A\) occurs? Then run the independent panel and compare.

Mutually Exclusive

Die roll — A = {1, 2}, B = {4, 5, 6}

A unit square split into columns A and A-prime. The blue B-band has zero height in column A and is taller in column A-prime, and there is no green A-and-B cell — the events cannot co-occur. When A occurs, column A-prime dims and column A is outlined, showing P(B given A) equals 0: the events are mutually exclusive and therefore dependent.

Independent

Two coins — A = Flip 1 = H, B = Flip 2 = H

A unit square split into four cells. A vertical line at P(A) separates column A from column A-prime; a horizontal dashed blue line at P(B) marks the B band. Because the blue line runs straight across at the same height in both columns, B fills the same fraction of column A as of the whole square. When A occurs, the A-prime column dims and the A column is outlined, showing P(B given A) equals P(B): A and B are independent.

Practice · mutually exclusive or not? · fresh scenarios on demand

Production task · write the reply

Exit ticket · 60 seconds · anonymous

The reply Maya files

What Maya can now write

The wheel has no memory: a streak changes nothing about the next trial. Probability is a property of the process, not a debt the past owes you.

A valid probability is a share of the sample space: between 0 and 1, complements splitting 1, overlaps subtracted exactly once.

"Can't both happen" is the opposite of "unrelated" — mutually exclusive events are maximally dependent.

Next class, the reader's follow-up: "But my nights aren't coin flips — doesn't one bad night cause the next?" Fair question. That's conditional probability, and it gets its own machinery.

Vote counts on today's slides are simulated for rehearsal; live sessions show the room's real votes.