Join · scan
OTTO-DS5
ottolearn.org/join
Module 1 · Descriptive Statistics · DS-5
Maya's spread follow-up runs — and the inbox fills up. One email sticks: "I sleep 5.5 hours. Should I be worried?" The average is 6.8. But "below average" isn't an answer — today Maya learns to say exactly where one reader stands among 48,000 students.
Live session · join with code OTTO-DS5 · ottolearn.org/join
Same rules as always: vote before I explain, every poll is anonymous, and changing your mind after discussion is the goal.
Today · the mailbag: "should I be worried?"
Percentiles, quartiles, deciles
The z-score
Reading z-scores: sign, size, units
Skewness — the tail tells
The 68–95–99.7 rule
Each question is one step toward Maya's reply — and a poll you'll answer first.
Recall · last class · 45 seconds
Warm-up · 60 seconds · anonymous
Hold your vote. By the end of class you'll answer the reader three ways: a rank, a distance — and a percentage.
Core concept 1 · C1
Tool #1 turns "below average" into a rank. Line up all 300 sleepers from shortest to longest — where in the queue does the reader stand?
Sort the data. \(P_k\) is the value with about \(k\%\) of the data below it. "You're at the 16th percentile" = about 16% of students sleep less than you do.
\(P_k\)Last class's quartiles are percentiles wearing familiar names: \(Q_1 = P_{25}\), the median \(= P_{50}\), \(Q_3 = P_{75}\). Deciles cut by tens: \(D_1 = P_{10}\) up to \(D_9 = P_{90}\).
\(Q_1 = P_{25},\ Q_3 = P_{75}\)"83rd percentile" does not mean a score of 83. The \(k\) counts who's below you; the value sitting at that position is a different number entirely.
By hand: \(P_k\) is the value at rank \(L = \lceil \tfrac{k}{100} \times n \rceil\) of the sorted list. Watch the rule run in both directions:
Watch it · C1
Predict first: for these 12 values, \(P_{75}\) needs rank \(L = \lceil 0.75 \times 12 \rceil\) — which position is that? Then reverse the map: pick a value and read off its rank.
Practice · C1 · find the percentile
Core concept 2 · C2
Subtract the mean, divide by the standard deviation. The result counts standard deviations: \(z = -2\) means two spreads below the mean — whatever the original units were.
\(z = \dfrac{x - \bar{x}}{s}\)Data for a whole population → \(\mu\) and \(\sigma\). A sample — like Maya's 300 — → \(\bar{x}\) and \(s\). Same recipe; match the symbols to the data you actually have.
\(z = \dfrac{x - \mu}{\sigma}\)Hours minus hours, divided by hours: nothing survives but a pure number. That's the superpower — sleep, rents and exam scores all land on one common scale.
Time to answer the reader properly. Watch me do the first one — out loud, every decision:
I do · watch me place the reader
Full support — watch me choose the symbols, subtract, divide, and say it in words.
The 300 responses are a sample of 48,000 students — so \(\bar{x} = 6.8\), \(s = 1.3\), and the recipe is \(z = (x - \bar{x}) \div s\).
\(5.5 - 6.8 = -1.3\) h. The sign is already talking: negative → below the mean.
\(-1.3 \div 1.3 = -1.00\). The hours cancel — what remains is a count of standard deviations.
The reader sits exactly one typical distance below the mean. Below average, yes — but by an ordinary amount, not an alarming one.
\(z = -1.00\). Then a second email lands: "I sleep 9.4 hours. Basically superhuman, right?" Your turn —
We do · you finish it
Support fades — the setup is done; you take the last division and read the result.
Same survey: \(\bar{x} = 6.8\) h, \(s = 1.3\) h. The new email claims \(x = 9.4\) h.
Divide by \(s = 1.3\): what is \(z\) — and how far out is that, really?
Practice · C2 · compute the z-score
On your own now — fresh numbers, every step yours. Re-roll as many as the room needs.
See it · C2
Predict first: drop \(x\) exactly on the mean — what does \(z\) read? Then slide \(x\) up by one full \(\sigma\) and watch the bracket count it.
Core concept 3 · C3
Positive → above the mean; negative → below. \(z = -1.8\) isn't an error or a bad score — just 1.8 spreads below average. Direction, nothing more.
\(z = 0\) is the mean itself. Within \(\pm 1\): ordinary. Around \(\pm 2\): far out. Beyond \(\pm 3\): rare in most data — worth a second look, the same job last class's fences did.
Unitless means portable: a rent and a sleep duration can meet on the same scale. Farther from its own mean, in its own spreads = more extreme — regardless of units.
Two famous values from Maya's own reporting are about to meet on that scale. Commit to a read before anyone computes:
Which is more extreme? · mixes DS-3 + DS-4
See it · C3
Predict first: can a smaller raw gap win the bigger \(|z|\)? Rebuild the penthouse-vs-sleeper duel — set each panel's mean, spread and value — and let the shared ruler settle it.
Practice · C3 · interpret z-scores
C4 · Predict first · ConcepTest · vote → discuss → re-vote
Core concept 4 · C4 · why that happened
The tail is the long, thin side — and it names the skew. Tail right → right-skewed. Tail left → left-skewed. The peak usually sits on the opposite side.
Tail values are extreme, and the sensitive mean gets dragged toward them; the resistant median barely moves. Right tail: mean > median. Left tail: mean < median.
Salaries: mean \(\$112{,}000\), median \(\$84{,}000\). No histogram needed — the mean sits far above the median, so the tail is on the right, and the median belongs in the report.
Three reference shapes, side by side. Find the TAIL and the BULK in each:
See it · C4
Predict first: in the right-skewed panel, does the mean marker sit left or right of the median? Check yourself on all three shapes.
Explore · C4
Predict first: pushing the skew slider right — which marker chases the new tail, mean or median? Drag, watch them separate, then re-converge at symmetric.
Practice · C4 · name the shape
Core concept 5 · C5
Bell-shaped data keeps a schedule: about 68% of it within 1 SD of the mean, 95% within 2, 99.7% within 3. Memorize the trio.
\(\mu \pm 1\sigma,\ \mu \pm 2\sigma,\ \mu \pm 3\sigma\)The percentages are properties of the bell shape — that's why the shape check (C4) comes first. Skewed, bimodal or outlier-heavy data: the rule's numbers go badly wrong.
Say "approximately 68%", never "exactly 68%". These are benchmarks for real, roughly-normal data — not laws.
Maya pulls up her survey histogram from the charting class: one hump, roughly symmetric. The rule is available — watch it lay its bands:
See it · C5
Predict first: to cover 95% of the data, how many SDs wide must the net be? Move \(\mu\) and \(\sigma\) — the boundary values change, the three percentages refuse to.
We do · answer the reader
Survey: \(\bar{x} = 6.8\) h, \(s = 1.3\) h, shape approximately bell. The reader's \(z = -1\) puts her exactly on the lower edge of the 68% band (\(5.5\) to \(8.1\) h).
Approximately what percentage of students sleep less than the reader — and what should Maya write back?
Quick check · C5 · 45 seconds
Practice · C5 · 68–95–99.7
Production · partners · one sentence
Exit ticket · muddiest point · anonymous
Wrap · Module 1 complete
Locate — a percentile gives the rank (\(P_k\); quartiles and deciles are special cases); the z-score gives the distance: \(z = (x - \bar{x}) \div s\), counted in standard deviations.
Read the shape — the tail names the skew, and the mean chases it: mean > median → right tail; mean < median → left tail. Skewed data → resistant pair (median + IQR).
Benchmark — bell-shaped only: approximately 68 / 95 / 99.7% within 1 / 2 / 3 SDs of the mean. Check the shape before quoting a percentage.
That completes Module 1: Maya can sample, chart, and summarize — honestly. Next class (PR-1): a reader shrugs off the all-nighter cluster as "just chance". Is it? That question needs probability.
The survey figures (\(\bar{x} = 6.8\) h, \(s = 1.3\) h) and reader emails are illustrative, constructed for this class — not survey results.