Join · scan
OTTO-PR6
ottolearn.org/join
Module 3 · Probability Distributions · PR-6
Last class ended on a cliffhanger: raise \(n\) and every binomial — tilted or not — smooths into the same bell. Today that bell gets its name, the normal distribution — and Maya puts it on trial. Her survey histogram has looked bell-shaped since the winter. If the shape is real, one small table can answer questions her 300 raw responses can't.
Live session · join with code OTTO-PR6 · ottolearn.org/join
Same rules as always: vote before I explain, every poll is anonymous, and changing your mind after discussion is the goal.
Today · one curve, one table, five moves
Properties + the Empirical Rule (68–95–99.7)
The z-table's three moves: read, complement, subtract
Standardization: \(z = (x - \mu)/\sigma\), the bridge
Inverse normal: table body → \(z^*\) → unstandardize
Recognizing non-normal shapes — skew, counts, bounds
Learn the shape, read the table, build the bridge that connects them to any normal distribution, run the bridge backwards — and know when to refuse the bell entirely. Every step starts with your vote.
Recall · last class · 45 seconds
Warm-up · 60 seconds · anonymous
That one reading — probability = area = proportion of the population — is the entire lesson. Everything today is a smarter way to compute areas. And keep the 27% in your pocket: it comes back with a starring role.
Core concept 1 · C1
The bell isn't one curve — it's a family, written \(X \sim N(\mu, \sigma^2)\). The mean \(\mu\) slides the center; the standard deviation \(\sigma\) sets the width. Everything else about the shape is fixed, which is why one rule of thumb works for all of them.
Symmetric around \(\mu\), where mean = median = mode all coincide; tails that approach the axis but never touch it; total area exactly 1. Know \(\mu\) and \(\sigma\) and you know everything about the curve.
\(X \sim N(\mu,\ \sigma^2)\)About 68% of values fall within \(1\sigma\) of the mean, 95% within \(2\sigma\), 99.7% within \(3\sigma\). For the sleep model: 68% of students in \([5.5,\ 8.1]\) h, 95% in \([4.2,\ 9.4]\) h.
\(\mu \pm 1\sigma,\ 2\sigma,\ 3\sigma\)The notation \(N(\mu, \sigma^2)\) carries the variance. Maya's sleep model is \(N(6.8,\ 1.69)\) because \(1.3^2 = 1.69\) — writing \(N(6.8,\ 1.3)\) silently claims \(\sigma \approx 1.14\). Always check which one the second number is.
\(N(6.8,\ 1.69)\)Remember the two mailbag readers from the position class? The 5.5-h worrier sits at exactly \(\mu - 1\sigma\), the 9.4-h mega-sleeper at \(\mu + 2\sigma\) — the Empirical Rule already places them both. See the ruler at work:
See it · C1
Predict first: drag \(\sigma\) wider — does the 68% band keep its area, its width, or both? Then slide \(\mu\): what moves and what doesn't?
Quick check · the mega-sleeper
Commit first · today's summit
Core concept 2 · C2
The z-table speaks exactly one sentence: "the area to the left of \(a\) is …". Every normal probability question — every one, all semester — becomes one of three moves built on that sentence.
\(P(Z < a)\) is the table's native language — the entry IS the answer. \(P(Z < 1.23)\): row 1.2, column 0.03, done: \(0.8907\).
\(P(Z < a)\)The area above \(a\) is whatever the left tail didn't take: \(P(Z > a)\) \(= 1 - P(Z < a)\). The complement rule from the first probability class, back on duty.
\(1 - P(Z < a)\)The strip between \(a\) and \(b\) is the big left area minus the small one: \(P(Z < b) - P(Z < a)\). Never add — adding double-counts everything left of \(a\) (and can even top 1).
\(P(Z < b) - P(Z < a)\)Read, complement, subtract — the whole toolkit. The discipline: name which area the question asks for before touching the table, and never report a raw entry unconverted. Watch the areas do the arithmetic:
See it · C2
Predict first: shade \(P(Z > 1)\) — how much of the curve survives? Then shade between \(-1\) and \(1\): watch the subtraction happen in the picture, big left area minus small.
Practice · reading the z-table · fresh numbers on demand
Core concept 3 · C3
The z-table covers exactly one distribution — \(N(0, 1)\). Maya's sleep model is \(N(6.8,\ 1.69)\). The standardization formula is the bridge between them, and it's a face you already know: the z-score from the position class, promoted to a new job.
\(z = (x - \mu)/\sigma\) — how many standard deviations \(x\) sits above (+) or below (−) the mean. In the position class z-scores compared values; today they unlock a probability table.
\(z = \dfrac{x - \mu}{\sigma}\)Standardizing doesn't move the curve — it relabels the axis. The area below \(x = 6\) h on Maya's curve and the area below its z-score on the standard curve are the same region: \(P(X < 6) = P(Z < z)\).
\(P(X < x) = P(Z < z)\)Looking up a raw \(x\) — say 6 — asks the table "what's below 6 standard deviations above the mean?" and gets \(\approx 1\), nonsense for this question. Cross the bridge first, every single time.
\(P(Z < 6) \neq P(X < 6)\)Standardize, then it's yesterday's news: read, complement, or subtract. See the bridge in action, then watch me run a whole crossing out loud:
See it · C3
Predict first: the top curve is \(X \sim N(\mu, \sigma^2)\), the bottom is \(Z \sim N(0,1)\). Slide \(x\) — the correspondence line stays vertical. What does that say about the two shaded areas?
I do · watch the whole crossing
Watch me decide every step out loud — which area the phrase asks for, then standardize, read, and convert only if needed.
The question is \(P(X < 6)\) — a less-than area. I decide the ending before any arithmetic: once I'm on the Z scale, the table entry will BE the answer. No complement, no subtraction.
\(z = \dfrac{6 - 6.8}{1.3}\) \(= \dfrac{-0.8}{1.3} \approx -0.62\). Negative — 6 h sits below the mean, so I expect an area under 0.5. Writing the expectation down is the cheapest error alarm there is.
Row \(-0.6\), column 0.02: \(P(Z < -0.62) = 0.2676\). Under 0.5, as predicted ✓. So the model says: about 26.8% of students sleep under 6 hours.
The survey counted it directly: 81 of 300 under 6 h \(= 0.27\). Model: \(0.268\). Count: \(0.270\). The bell just passed its audition — the curve reproduces what the raw data say.
The habit: name the area the phrase asks for, standardize, read, convert only if needed — then sanity-check against the picture, and, when you can, against the data.
We do · you finish it
Same pipeline, one blank left for the room: you spot the conversion and supply the final probability.
Same model, \(X \sim N(6.8,\ 1.69)\). The 9.4-h reader wants her exact rarity: \(P(X > 9.4)\). The Empirical Rule guessed about 2.5% earlier — now the table sharpens it.
Finish it: \(P(X > 9.4) = ?\)
Practice · standardize and find the area · fresh numbers on demand
Now you run the whole bridge solo — fresh numbers every re-roll, the table entries you need on screen, support at its minimum.
Core concept 4 · C4
So far: value in, probability out. Maya's editor asks the reverse: "Where's the cutoff for the least-rested 10% of students?" — probability in, value out. Same table, same bridge, driven in reverse.
Forward reading: margins → entry. Inverse reading: search the table body for the target area, then read \(z^*\) off the margins. For the bottom 10%: the body entry nearest \(0.1000\) sits at \(z^* = -1.28\).
\(P(Z < z^*) = p\)The bridge run backwards: \(x = \mu + z^* \sigma\). The 10% cutoff: \(x = 6.8 + (-1.28)(1.3)\) \(\approx 5.14\) h. Anyone under about 5 h 8 min is in the least-rested tenth — a printable sentence.
\(x = \mu + z^* \sigma\)The classic inverse error: hunting for \(0.10\) in the margins — where the z-values live — and reading off a number that means nothing. Areas live in the body. Body → \(z^*\) → unstandardize, always in that order.
Forward: \(x \to z \to\) area. Backward: area \(\to z^* \to x\). Two directions, one bridge. Watch the reverse reading happen:
See it · C4
Predict first: set the target area to \(0.90\) — will \(z^*\) land left or right of zero? Watch the reading direction: area in the body first, then \(z^*\), then back into real units.
Practice · percentile to value · fresh numbers on demand
Core concept 5 · C5
The table only tells the truth about data that actually follow the bell. Maya's sleep data earned the model — bell-shaped histogram, and the 27% check passed. Not every dataset on her desk gets the same license.
Roughly symmetric, one peak, continuous measurements, and the Empirical Rule roughly holds against the data. Heights, blood pressure, machine fills — and Maya's sleep hours, now audited and approved.
Strong skew (incomes, wait times), discrete counts (complaints per day — binomial or Poisson territory), bounded proportions, two peaks. Each failure points to a different proper tool; none of them is the z-table.
Her winter rent piece: mean \$1{,}650, median \$920 — a mean nearly double the median screams right skew. A normal model there would park half the curve above \$1{,}650 while most real listings cluster near \$920: confidently, precisely wrong.
Symmetric bell first, table second — in that order, or the precision is fake. The shapes below are the field guide; then a which-tool poll that reaches back to last class:
See it · C5
Predict first: which of these shapes would the Empirical Rule lie about the most — the skewed one, the two-peaked one, or the discrete bars? Check each against the bell.
Which tool applies? · mixes last class + today
Practice · does the bell apply? · fresh numbers on demand
Production task · answer the editor
Exit ticket · 60 seconds · anonymous
What Maya files
Last class's cliffhanger resolved: the shape every binomial smooths into is the normal distribution — and Maya's survey wears it honestly. The histogram looks like the bell, and the model reproduced a number it was never told: \(P(X < 6) \approx 0.268\) vs the counted \(81/300 = 0.27\).
The toolkit: probabilities are areas; the z-table gives left tails only; \(z = (x - \mu)/\sigma\) bridges any normal to it (read, complement, or subtract); and the bridge runs backwards — body \(\to z^* \to x = \mu + z^* \sigma\) put the least-rested 10% below about 5.1 h.
The refusal skill survives the new tool: rents (median \$920, mean \$1{,}650) get no bell, and counts stay with last class's binomial. A model is earned — by shape, and by checks like the 27% audit — never assumed.
Module 3 closes at module 4's door. Maya's 6.8 h is one sample of 300 drawn from 48,000 students — so what would a different 300 have said? Next class the bell follows her there: the distribution of sample means — and the question she's carried since the very first survey finally becomes answerable.
Vote counts on today's slides are simulated for rehearsal; live sessions show the room's real votes.