Join · scan
OTTO-INF4
ottolearn.org/join
Module 4 · Statistical Inference · INF-4
The athlete piece ran. The next reader letter changes the question: "Enough averages — what percentage of students pull all-nighters?" Maya's survey had the yes/no item all along: item 7, "Did you pull at least one all-nighter last month?" — 75 of her 300 said yes. Today: turning a count into an honest range — and pricing the sharper poll the editor will inevitably demand.
Live session · join with code OTTO-INF4 · ottolearn.org/join
Same rules as always: vote before I explain, every poll is anonymous, and changing your mind after discussion is the goal.
Today · one count, five questions
The sample proportion \(\hat{p} = x/n\) — a statistic aiming at the fixed \(p\)
The SE of \(\hat{p}\): the center sets the spread — \(z\) keeps its label
Margin of error \(E = z^* \cdot \text{SE}\) — and the full recipe \(\hat{p} \pm E\)
Ten of each: \(n\hat{p} \geq 10\) and \(n(1-\hat{p}) \geq 10\), plus the quiet two
Sample size: \(n = (z^*/E)^2 p^*(1-p^*)\), rounded up, safest at \(p^* = 0.5\)
Meet the sample proportion, settle which yardstick it deserves (the day's big vote), build and license the interval, keep the warranty honest — and price the sharper poll the editor is already drafting. Every step starts with your vote.
Recall · last class · 45 seconds
Warm-up · commit first · 60 seconds · anonymous
Hold your answer — today's slides settle it piece by piece, and we return to this tally before the practice.
Core concept 1 · C1
One of them is on Maya's screen; the other is the reason she's writing. Confusing them is the module's original sin — same distinction as \(\bar{x}\) and \(\mu\), new cast.
\(\hat{p}\) — "p-hat," the sample proportion: yeses over sample size. Known, exact for these 300 — and it would land somewhere else for a different 300. A statistic: it varies sample to sample.
\(\hat{p} = \dfrac{x}{n}\) \(= \dfrac{75}{300}\) \(= 0.25\)\(p\) — the true fraction of all 48,000 students who pulled an all-nighter. Fixed, unknown, doesn't wobble. A parameter. Every claim Maya prints is a claim about \(p\), built from \(\hat{p}\).
\(p = \ ?\)25% is not \(p\) — it's the estimate. And the reverse trap matters more today: every formula this class runs on \(\hat{p}\), never on \(p\) — because we don't have \(p\). If we did, there'd be nothing to estimate.
\(\hat{p} \leftrightarrow \bar{x},\ \ p \leftrightarrow \mu\)Same cast as always, third production: \(\bar{x}\) estimated \(\mu\), now \(\hat{p}\) estimates \(p\). Watch the estimate scatter around the truth it's aiming at:
See it · C1
Each "Draw a sample" drops a tick at that sample's \(\hat{p}\). Predict: where will the ticks pile up — and does the dashed \(p\) line ever move? Then: what happens to the pile's spread when \(n\) goes from 25 to 100?
Commit first · today's summit
Core concept 2 · C2
This is why proportions get their own lesson: for yes/no data, one number does two jobs — and that's why the yardstick vote has a different answer than last class.
Each answer is a yes/no trial with success chance \(p\); the yes-count is binomial, and its variance — \(np(1-p)\), from the distributions module — is set by \(p\) itself. The center decides the spread. There is no second dial for \(t\) to pay for.
\(\sigma_{\hat{p}} = \sqrt{\dfrac{p(1-p)}{n}}\)The theoretical spread uses \(p\), which nobody has — so \(\hat{p}\) stands in. Maya's: \(\sqrt{\dfrac{0.25 \times 0.75}{300}}\) \(= \sqrt{0.000625}\) \(= 0.025\). About 2.5 points of wobble on a one-in-four estimate.
\(\text{SE} = \sqrt{\dfrac{\hat{p}(1-\hat{p})}{n}}\)\(\hat{p}\) standing in for \(p\) is an estimate inside an estimate — a real (small) liberty. The fix is not heavier tails: the substitution is safe exactly when the sample holds plenty of both answers, and that gets its own checkpoint right after the build.
The anchor, resolved: no separate \(\sigma\), no degrees of freedom spent, \(z\) keeps its label. Next ingredient: how far the interval must reach.
Core concept 3 · C3
Third appearance of the third ingredient: critical value × standard error — how far from \(\hat{p}\) the interval must stretch to earn its warranty.
Maya's: \(E = 1.96 \times 0.025\) \(= 0.049\) — ±4.9 points. The "margin of error ±3 points" in every news-poll footnote is exactly this \(E\), for that poll's \(n\).
\(E = z^* \cdot \sqrt{\dfrac{\hat{p}(1-\hat{p})}{n}}\)The warranty dial is untouched: 90% → 1.645, 95% → 1.96, 99% → 2.576. Same normal table as the mean interval — because the yardstick stayed \(z\).
\(1.645,\ 1.96,\ 2.576\)\(E\) is the half-width. "±4.9 points" spans 9.8 points end to end — nearly a ten-point window. Reporting \(E\) as the full width claims double the precision the data bought.
Center, yardstick, reach — all three ingredients on the desk. Time to assemble:
Core concept 4 · C4
Point estimate, ±, critical value × SE. You've now built this skeleton with a known-\(\sigma\) mean and a small-sample mean — today a count. That's not repetition; that's the design.
Everything structural: best single answer in the middle, reach out a margin \(E\), print the range, say the claim in words. Three classes, one skeleton.
\(\text{estimate} \pm E\)The center is now \(\hat{p}\); the SE is now \(\sqrt{\hat{p}(1-\hat{p})/n}\). And the critical value does not move: \(z^*\) stays, because nothing separate was estimated — that was the big vote.
\(\bar{x} \to \hat{p},\ \ z^* \to z^*\)Read it left to right as a sentence: best guess, then honest wobble, scaled by the warranty you promised. Conditions first, though — the license is two slides away.
\(\hat{p} \pm z^* \cdot \sqrt{\dfrac{\hat{p}(1-\hat{p})}{n}}\)Before building Maya's for real — feel how the width answers to \(n\), \(\hat{p}\), and the warranty:
See it · C4
Predict first: to cut the width in HALF, how many times more \(n\)? And where does \(\hat{p}\) make the interval widest — 0.2, 0.5, or 0.8? Check both with the sliders.
I do · watch a full build
Watch me decide every step out loud — check the license (counts, not shape), get center and spread from one number, compute the margin, assemble the claim and say it in words.
Random sample? The 300 came from the survey's random draw — ✓. Enough of both answers? 75 yeses and 225 nos, both far past 10 — ✓ (the full rule is two slides away). Small slice? 300 is well under 10% of 48,000 — ✓. And notice what I did not check: no \(\sigma\), no \(df\), no t-table.
Center: \(\hat{p} = 75/300\) \(= 0.25\). Spread, from the same number: \(\text{SE} = \sqrt{0.25 \times 0.75 / 300}\) \(= \sqrt{0.000625}\) \(= 0.025\). Sanity check: 2.5 points of wobble from 300 yes/no answers — believable.
The paper's standard warranty is 95%, so \(z^* = 1.96\). Reach: \(E = 1.96 \times 0.025\) \(= 0.049\) — ±4.9 points. (If the editor demanded 99%, only this step changes: \(2.576 \times 0.025\) \(= 0.0644\).)
\(0.25 \pm 0.049\) \(\Rightarrow (0.201,\ 0.299)\). In words: "We are 95% confident that between 20.1% and 29.9% of students pulled at least one all-nighter last month." One in four, give or take five points.
The reader asked for one number; the honest answer is a ten-point window — and the window itself is news: even its floor says one student in five. What the window can't do is come cheap. The editor's reply is already in: "Pin it to ±2." Pricing that demand is the last stop today.
We do · you finish it
New survey, new count: you supply the margin and the interval — the numbers change, the habit doesn't.
Same week, next question: "Fine — all-nighters happen. Do students at least catch up? Who naps?" The residence association's random survey of 400 students has the count: 80 report a daily nap. Build the 95% interval.
Finish it: \(E = ?\) — and the interval is…? (Then say the sentence.)
Practice · build the interval, solo · fresh counts on demand
Now you build whole proportion intervals solo — fresh counts every re-roll, support at its minimum.
Core concept 5 · C5
The z-label is earned by a bell — and the bell only shows up when the sample holds enough of both answers. For proportions, the license is a pair of counts.
\(n\hat{p}\) and \(n(1-\hat{p})\) are literally the yes-count and the no-count — \(x\) and \(n - x\). Both must reach 10. Maya's: 75 and 225 — licensed with room to spare.
\(n\hat{p} \geq 10,\ \ n(1-\hat{p}) \geq 10\)Push \(\hat{p}\) near an edge — say 19 yeses out of 20 — and \(\hat{p}\)'s distribution squashes against the wall: strongly skewed, no bell, and the "95%" label goes quietly false. Ten of each keeps room on both sides for the count's distribution to turn symmetric.
Random sample — independence comes from the draw, the rule since the very first survey — and \(n\) at most 10% of the population: 300 of 48,000 clears it easily. Break the draw and no formula rescues you.
\(n \leq 10\%\ \text{of pop.}\)Two counts, two quiet checks — cheaper than last class's dot-plot stare, and just as load-bearing. Watch the bell fail when the counts don't clear:
See it · C5
Predict first: \(n = 30\), \(p = 0.1\) — only 3 expected yeses. What shape will the bars take, and will the overlaid bell fit? Then find settings where the fit turns good — and read the two counts at that moment.
Practice · check the license, solo · fresh counts on demand
Which tool applies? · mixes earlier classes + today
Core concept 6 · C6
Third interval, same contract. The words survive unchanged — and so do the misreads, which is why this slide exists every single time.
"We are 95% confident that between 20.1% and 29.9% of students pulled an all-nighter last month." — the accepted shorthand, crediting the method that built the interval.
A statement about the procedure: run 1,000 surveys like hers and about 950 of the resulting intervals capture the true \(p\). Hers is one of the batch — built by a method with a 95% track record.
Not a probability on \(p\) — a fixed constant that either sits in \((0.201,\ 0.299)\) or doesn't. Not a claim that 20–30% of students are "in the interval." Not a promise about the next survey.
The warranty is on the factory, not the unit — third class running. Watch the batch:
See it · C6
Twenty intervals, one dashed \(p\). Predict: at 95%, how many of the twenty will miss — and does the dashed line ever move? Then drop the level to 90% and predict again.
Quick check · print it or fix it?
Core concept 7 · C7
The editor reads the ten-point window and sends the inevitable reply: "Pin it to ±2. What does that cost me?" Decide the precision first — then solve for \(n\).
The margin formula, inverted. \(p^*\) is your best prior guess at \(p\) — last month's poll, a pilot, or (if you truly have nothing) the built-in insurance value. Decide \(E\) and the warranty; the formula quotes the sample.
\(n = \left(\dfrac{z^*}{E}\right)^{2} p^*(1-p^*)\)±2 points at 95%, no prior: \(n = 9604 \times 0.25\) \(= 2401\) people. Using last month's \(p^* = 0.25\): \(9604 \times 0.1875\) \(= 1800.75\) \(\to 1801\). The 600-person gap is the premium for admitting you don't know \(p\). (±3 points? 1068 — why national polls hover near a thousand.)
Always round up: 1800.75 means 1801 — rounding down quietly breaks the promised \(E\). And \(p^* = 0.5\) is not a guess that \(p\) is 50% — it's the worst case: \(p(1-p)\) peaks at 0.5, so budgeting there can never undershoot.
\(p^*(1-p^*) \leq 0.25\)Precision is bought in people, and the price grows with the square of the demand — halving \(E\) still quadruples \(n\). One picture explains the 0.5 rule:
See it · C7
The curve is \(p^*(1-p^*)\). Predict its peak. Then I'll set \(p^* = 0.3\) — a shaded risk zone appears: what does it mean, and which single choice of \(p^*\) makes it vanish for good?
Practice · price the poll, solo · fresh targets on demand
Production task · the reply she publishes
Exit ticket · 60 seconds · anonymous
What Maya files
The reader's question got an honest answer: between 20% and 30% of students pulled an all-nighter last month — one in four, give or take five points. The skeleton from the last two classes carried a brand-new kind of data without bending.
The fork grew a second prong: measurements → mean tools (\(z\) if \(\sigma\) is known, \(t\) if \(s\) stands in); yes/no counts → the proportion interval, where \(z\) keeps its label because \(\hat{p}\) prices its own spread — licensed by ten of each.
And the formula runs backwards: decide the precision, quote the sample — 2,401 for ±2 points with the 0.5 insurance. Precision is a budget item now, not a wish.
Next on the desk: the campus wellness office reads Maya's numbers and pushes back — "students get eight hours; your survey must be off." A claim, her data, and a referee to decide between them… that's a hypothesis test, and it starts next class.
Vote counts on today's slides are simulated for rehearsal; live sessions show the room's real votes.