Join · scan
OTTO-DS1
ottolearn.org/join
Module 1 · Descriptive Statistics · DS-1
Follow one journalist’s investigation — and meet the language behind every statistical claim.
Live session · join with code OTTO-DS1 · ottolearn.org/join
How today works
You’ll often vote before I explain. Predicting incorrectly is doing it right — committing first is what makes the idea stick.
Every poll is anonymous. No one sees who answered what — not even me.
Changing your mind after talking it over is the goal, not a failure.
Today · Maya’s investigation
Population & sample
Parameter & statistic
Variable types
Sampling methods
Bias
Each step is a real decision Maya has to make — and a poll you’ll answer first.
Warm-up · 60 seconds · anonymous
Keep that instinct — by the end you’ll know exactly what went wrong, and how Maya fixes it.
C1 · Predict first · ConcepTest · vote → discuss → re-vote
Core concept 1 · C1 · now let’s name it
The complete set Maya wants conclusions about — all 48,000.
size = NThe subset she actually studies — her 300.
size = nIf you picked “the 300” — that’s the sample. The population is who Maya wants to generalize to: “Who do I want my conclusions to apply to?”
See it · a sample drawn from a population
Predict first: will two of Maya’s samples give the same average? Then I’ll draw a few.
Practice · population vs sample · new version anytime
Core concept 2 · C2
Describes the population. Usually unknown — Maya can’t reach all 48,000.
μ σComputed from the sample. Known, but varies sample to sample.
x̄ sParameter → Population · Statistic → Sample. Greek for the population (μ, σ); Roman for the sample (x̄, s).
Model · watch me label all four · think aloud
Maya wants the true mean nightly sleep of all 48,000 CEGEP students. She surveys 300 and gets x̄ = 6.8 h.
Who she wants to conclude about — all 48,000. (Not the 300.)
Who she actually measured — the 300 surveyed.
The number she wants but can’t see: true mean sleep of all 48,000 → μ. Unknown.
The number she computed from the 300 → x̄ = 6.8 h. Known — her best estimate of μ.
Two groups, two numbers: x̄ is what we know; μ is what we’re after.
Together · we label one · your turn on the two numbers
New study: a campus gym wants the average weekly visits of all 5,000 members. A random sample of 120 averages 2.4 visits/week.
Your turn: name the parameter and the statistic — with notation.
See it · the estimation bridge
Predict: as Maya’s sample grows, does x̄ drift toward the true μ — or wander away?
Activity · classify it · single anonymous vote
Practice · parameter vs statistic · new version anytime
Core concept 3 · C3
Labels, not numbers you do arithmetic with.
Numbers where differences and ratios make sense.
The deciding question is never “is it written with digits?” — it’s “does arithmetic make sense?”
Model · watch me classify · think aloud
Maya’s form also records each respondent’s phone number. What type is it?
It’s written with digits — but I don’t stop there. Digits alone never decide it. On to the real test.
Averaging two phone numbers is meaningless; …0199 isn’t “more” than …0188. Arithmetic fails → qualitative.
No number ranks above another → no order → nominal.
Phone number → qualitative, nominal. The digits were a decoy; the arithmetic test settled it.
Together · we finish this one · your turn on the last step
Next field: each respondent’s number of caffeinated drinks yesterday (0, 1, 2, …).
3 · Your turn: counted or measured? Name the sub-type — then I’ll reveal.
See it · trace a variable to its type
Pick a ⚠ trap example — predict its type before you trace the path.
Activity · classify it · single anonymous vote
Practice · variable types · new version anytime
Core concept 4 · C4 · sampling methods (1/2)
Every sample of size n equally likely. Gold standard — but needs a full list.
Split into homogeneous strata; sample from each. Guarantees representation.
Random start, then every k-th. Easy — biased if the list has a periodic pattern.
All three give every student a known chance of selection — the basis for valid inference.
Core concept 4 · C4 · sampling methods (2/2)
Split into heterogeneous clusters; sample a few entirely. Cheap when spread out.
Combine methods in stages. Practical for big national surveys; errors can compound.
Whoever is easiest to reach. Fast, cheap — and almost always biased.
Cluster ≠ stratified. Strata are similar inside (sample all); clusters are diverse inside (sample some).
See it · who actually gets selected?
Before I switch methods — predict which dots stratified will pick.
Model · watch me choose · think aloud
Maya wants the typical sleep of all CEGEP students — so every student needs a real chance of being picked. That alone kills convenience: the cafeteria crowd isn’t all students.
She has a full enrolment list. With a complete list, simple random is on the table — the gold standard. Systematic is easier, but I pause: is the list ordered by program or cohort? If it repeats, every k-th could lock onto a pattern.
Could a small program vanish in a random draw? If that risk mattered I’d switch to stratified to guarantee each appears. Here it doesn’t — so I commit to SRS, and I can say why I rejected the rest.
Let’s run those three steps together on a new constraint →
Together · we pick one · your turn on the final call
New constraint: Maya has one list of all 9,000 students at her campus, ordered by student ID (effectively random order). She wants a quick, evenly spread sample of 300 — and has no software to draw random numbers.
3 · Your turn: which method fits — and why is the “periodic pattern” risk low here?
Activity · which tool applies? · single anonymous vote
Practice · sampling methods · new version anytime
C5 · Predict first · ConcepTest · vote → discuss → re-vote
Core concept 5 · C5 · now let’s name them
Some groups have little or no chance of selection. an online-only survey skips the offline
Only those who feel strongly reply. Maya’s gaming-Discord poll
Selected people don’t answer — and differ from those who do.
Easiest-to-reach people are systematically different. the 11 p.m. library crowd
Bias is directional, not random — it does not shrink with sample size. A biased survey of 10,000 can beat nothing.
See it · bias vs. precision
Predict: does a bigger biased sample land closer to the truth, or just tighter around the wrong answer?
1936 — the original cautionary tale. The Literary Digest mailed 10 million ballots and got 2.4 million back, then predicted the wrong U.S. president by a landslide. Its list came from car and phone owners — wealthier than most in the Depression. A giant sample, sunk by who it left out.
Practice · spot the bias · new version anytime
Build it · with your neighbour · constructed response
Beyond the sample · C6
Wording that pushes an answer.
Two questions in one.
People answer to look good.
“Do you sleep regularly?”
Earlier questions sway later ones.
See it · bad-question autopsy
Spot the loaded word in each question first — then compare the original against the fix.
Activity · spot the flaw · tap one line
Exit ticket · shapes next class
Maya files her story
Population vs. sample — Maya studies 300 to learn about 48,000.
Statistic estimates parameter — her x̄ → the unknown μ, never assume they’re equal.
Bias ≠ small sample — direction, not size, is the danger.
Maya’s headline: a stratified sample of 300 → x̄ ≈ 6.8 h (vs. 8 h recommended). Her friend’s Discord poll said 5.1 — voluntary response, don’t trust it. Next: DS-2 — Graphs & numerical summaries.
Sleep figures are inspired by real college-sleep research — illustrative, not a specific study.