Join · scan
OTTO-DS3
ottolearn.org/join
Module 1 · Descriptive Statistics · DS-3
Maya's sleep story ran. Her editor's next assignment: a rival paper claims the average rent near campus is $1,650 — but every student Maya asks pays about $900. Someone's number is misleading. Today we find out whose, and why.
Live session · join with code OTTO-DS3 · ottolearn.org/join
Same rules as always: vote before I explain, every poll is anonymous, and changing your mind after discussion is the goal.
Today · Maya's rent fact-check
The sample mean
Outlier sensitivity
The median
The mode
Choosing the measure
Every step is a decision Maya has to make before her rent story runs — and a poll you'll answer first.
Recall · last class · 45 seconds
Warm-up · 60 seconds · anonymous
Keep your answer in mind. By the end of class you'll be able to name exactly what's happening — and fix the headline.
Core concept 1 · C1
First tool: the one everybody calls "the average". Add everything up, divide by how many. But the number you get is best understood as a balance point:
Add all \(n\) values, divide by \(n\). \(\bar{x}\) ("x-bar") is the sample mean — the population version, \(\mu\), arrives in Module 3.
\(\bar{x} = \dfrac{\sum x_i}{n}\)Put the values as weights on a number line: \(\bar{x}\) is exactly where the beam balances. Deviations cancel — what's above offsets what's below.
\(\sum (x_i - \bar{x}) = 0\)The mean of \(\{2, 3, 5\}\) is \(10 \div 3 \approx 3.33\) — a value in no dataset. "2.4 children per family" is a summary, not a family.
One number carries the whole dataset — which is exactly why we need to know when it can be trusted.
See it · C1
Predict first: where must the fulcrum go to level the beam? Then drag one weight far to the right — where does the balance point go?
Drag the fulcrum (▲) — or focus it and use the arrow keys — until the beam is level. It balances at exactly one spot: the mean. Drag the blue weights to change the data.
Practice · C1 · compute the mean
I do · watch me fact-check
Maya pulls 6 ordinary listings from the street the story names ($/month): 820, 860, 900, 920, 960, 990.
\(\sum x_i = 820 + 860 + 900 + 920 + 960 + 990 = 5450\)
\(\bar{x} = \dfrac{5450}{6} \approx 908.33\) — note: not itself a data value, and that's fine.
Typical rent on that street ≈ $908. Nowhere near the headline's $1,650.
So where does $1,650 come from? There is a 7th listing on that street — a penthouse at $6,100. Predict what it does:
C2 · Predict first · ConcepTest · vote → discuss → re-vote
Core concept 2 · C2 · why that happened
Every dollar of every value pulls on \(\bar{x}\). The penthouse alone contributes \(6100 \div 7 \approx \$871\) of the $1,650 average — more than the other six listings combined.
Only the order matters. Swap the penthouse's $6,100 for $60,000 and the median doesn't move a single dollar.
The median is resistant to outliers; the mean is sensitive. That one word is why Statistics Canada reports median household income.
One penthouse: mean +$742, median +$10. That asymmetry is today's whole story.
See it · C2
Predict first: how far right must you drag the outlier before the median jumps the way the mean just did?
Drag the orange dot left or right — or focus it and use the arrow keys. The mean (▲) chases the outlier; the median (◆) barely moves.
Core concept 3 · C3
The median is the middle of the sorted list. The middle of the raw list is a wrong answer delivered with confidence — the #1 median error.
The median sits at position \(\frac{n+1}{2}\) of the sorted list. Seven listings → position 4.
\(\dfrac{n+1}{2}\)Average the values at positions \(\frac{n}{2}\) and \(\frac{n}{2}+1\). The result may not be a data value — that's fine.
\(\dfrac{n}{2},\ \dfrac{n}{2}+1\)Position, not magnitude — that's the entire source of the median's resistance.
I do · watch the algorithm run
Full support — watch me sort the listings and count to the middle position.
The full median algorithm, one step at a time — watch the #1 error (skipping the sort) become visible in step 1 vs step 2. Then toggle to even \(n\) to see the algorithm change.
We do · you finish it
Support fades — the values are sorted; you locate the middle for an even count.
Maya finds an 8th listing at $1,040. Sorted: 820, 860, 900, 920, 960, 990, 1040, 6100 — now \(n = 8\).
What is the median rent now?
Practice · C3 · find the median
On your own now — fresh lists to order. Re-roll as many as the room needs.
Core concept 4 · C4
The mode is whatever occurs most often — and it's the only measure that works for categories. You can't average "Instagram" and "TikTok", but you can count which appears most.
One winner → unimodal. Two values tie → bimodal (both are modes). Every value appears equally often → no mode at all.
Two modes often mean two groups hiding in one dataset — men's and women's shoe sizes recorded together. Ask: is this really one population?
Mean and median need numbers. The mode just needs a tally.
See it · C4
Predict first: which bar wins? Then switch to the bimodal example — what should we say when two bars tie?
Quick check · C4 · 45 seconds
Practice · C4 · name the mode
Core concept 5 · C5
The DS-1 question. Categories → the mode is the only option — done. Numbers → keep going. Rent is quantitative → keep going.
The DS-2 question — look at the histogram. Symmetric, no outliers → the mean is fine (it uses every value). Skewed or outliers → one more step.
The mean chases the tail; the median holds the centre. Maya's rents are right-skewed → she files the median: $920, not $1,650.
Type → shape → decision. Every "which average?" call runs this same chain. Your turn:
We do · you make the call
Ten finish times — nine recreational runners between 3.5 h and 5.0 h, and one elite runner at 2.1 h.
Which measure does she report — and which direction is the mean being pulled?
See it · C5
Predict first: in a right-skewed distribution, which is bigger — mean or median? Switch tabs and check your rule against all three shapes.
Which tool applies? · mixes DS-1 + DS-2
Practice · C5 · which measure?
Production · partners · one sentence
Exit ticket · muddiest point · anonymous
Wrap
Compute — mean \(\bar{x} = \sum x_i / n\), median (sort → middle), mode (tally): three different questions about "typical".
Watch the tail — the mean chases outliers and skew; the median holds its position.
Choose by shape — type → shape → outliers. The data picks the measure; the measure keeps the headline honest.
Next class (DS-4): two datasets can share the same centre and be wildly different — how spread out is the data?
Rents and scenarios are illustrative, constructed for this class — not a cited market survey.