Join · scan

OTTO-INF4

ottolearn.org/join

Module 4 · Statistical Inference · INF-4

Confidence Intervals for a Proportion
"One in four?"

The athlete piece ran. The next reader letter changes the question: "Enough averages — what percentage of students pull all-nighters?" Maya's survey had the yes/no item all along: item 7, "Did you pull at least one all-nighter last month?"75 of her 300 said yes. Today: turning a count into an honest range — and pricing the sharper poll the editor will inevitably demand.

Live session · join with code OTTO-INF4 · ottolearn.org/join

Same rules as always: vote before I explain, every poll is anonymous, and changing your mind after discussion is the goal.

Today · one count, five questions

From 75-of-300 to a printable range.

Meet the sample proportion, settle which yardstick it deserves (the day's big vote), build and license the interval, keep the warranty honest — and price the sharper poll the editor is already drafting. Every step starts with your vote.

Recall · last class · 45 seconds

Warm-up · commit first · 60 seconds · anonymous

Hold your answer — today's slides settle it piece by piece, and we return to this tally before the practice.

Core concept 1 · C1

Two numbers wearing one percent sign

One of them is on Maya's screen; the other is the reason she's writing. Confusing them is the module's original sin — same distinction as \(\bar{x}\) and \(\mu\), new cast.

What she computed

\(\hat{p}\) — "p-hat," the sample proportion: yeses over sample size. Known, exact for these 300 — and it would land somewhere else for a different 300. A statistic: it varies sample to sample.

\(\hat{p} = \dfrac{x}{n}\) \(= \dfrac{75}{300}\) \(= 0.25\)

What she's after

\(p\) — the true fraction of all 48,000 students who pulled an all-nighter. Fixed, unknown, doesn't wobble. A parameter. Every claim Maya prints is a claim about \(p\), built from \(\hat{p}\).

\(p = \ ?\)

The mix-up to refuse

25% is not \(p\) — it's the estimate. And the reverse trap matters more today: every formula this class runs on \(\hat{p}\), never on \(p\) — because we don't have \(p\). If we did, there'd be nothing to estimate.

\(\hat{p} \leftrightarrow \bar{x},\ \ p \leftrightarrow \mu\)

Same cast as always, third production: \(\bar{x}\) estimated \(\mu\), now \(\hat{p}\) estimates \(p\). Watch the estimate scatter around the truth it's aiming at:

See it · C1

The estimate wobbles; the truth doesn't

Each "Draw a sample" drops a tick at that sample's \(\hat{p}\). Predict: where will the ticks pile up — and does the dashed \(p\) line ever move? Then: what happens to the pile's spread when \(n\) goes from 25 to 100?

Figure: The dashed line is the true proportion p — one fixed value that never moves (and is usually unknown, so it starts hidden). Each dot is a different sample's p̂ = x/n. Draw repeatedly: the dots scatter and cluster around the line. changes every sample; p does not.

Commit first · today's summit

Core concept 2 · C2

The spread comes free with the center

This is why proportions get their own lesson: for yes/no data, one number does two jobs — and that's why the yardstick vote has a different answer than last class.

One number, two jobs

Each answer is a yes/no trial with success chance \(p\); the yes-count is binomial, and its variance — \(np(1-p)\), from the distributions module — is set by \(p\) itself. The center decides the spread. There is no second dial for \(t\) to pay for.

\(\sigma_{\hat{p}} = \sqrt{\dfrac{p(1-p)}{n}}\)

The working formula

The theoretical spread uses \(p\), which nobody has — so \(\hat{p}\) stands in. Maya's: \(\sqrt{\dfrac{0.25 \times 0.75}{300}}\) \(= \sqrt{0.000625}\) \(= 0.025\). About 2.5 points of wobble on a one-in-four estimate.

\(\text{SE} = \sqrt{\dfrac{\hat{p}(1-\hat{p})}{n}}\)

The honest caveat

\(\hat{p}\) standing in for \(p\) is an estimate inside an estimate — a real (small) liberty. The fix is not heavier tails: the substitution is safe exactly when the sample holds plenty of both answers, and that gets its own checkpoint right after the build.

The anchor, resolved: no separate \(\sigma\), no degrees of freedom spent, \(z\) keeps its label. Next ingredient: how far the interval must reach.

Core concept 3 · C3

The reach: margin of error

Third appearance of the third ingredient: critical value × standard error — how far from \(\hat{p}\) the interval must stretch to earn its warranty.

The reach

Maya's: \(E = 1.96 \times 0.025\) \(= 0.049\) — ±4.9 points. The "margin of error ±3 points" in every news-poll footnote is exactly this \(E\), for that poll's \(n\).

\(E = z^* \cdot \sqrt{\dfrac{\hat{p}(1-\hat{p})}{n}}\)

The price list survives

The warranty dial is untouched: 90% → 1.645, 95% → 1.96, 99% → 2.576. Same normal table as the mean interval — because the yardstick stayed \(z\).

\(1.645,\ 1.96,\ 2.576\)

Half, not whole

\(E\) is the half-width. "±4.9 points" spans 9.8 points end to end — nearly a ten-point window. Reporting \(E\) as the full width claims double the precision the data bought.

Center, yardstick, reach — all three ingredients on the desk. Time to assemble:

Core concept 4 · C4

The recipe: same skeleton, third filling

Point estimate, ±, critical value × SE. You've now built this skeleton with a known-\(\sigma\) mean and a small-sample mean — today a count. That's not repetition; that's the design.

What stays

Everything structural: best single answer in the middle, reach out a margin \(E\), print the range, say the claim in words. Three classes, one skeleton.

\(\text{estimate} \pm E\)

What moves — and what doesn't

The center is now \(\hat{p}\); the SE is now \(\sqrt{\hat{p}(1-\hat{p})/n}\). And the critical value does not move: \(z^*\) stays, because nothing separate was estimated — that was the big vote.

\(\bar{x} \to \hat{p},\ \ z^* \to z^*\)

The formula

Read it left to right as a sentence: best guess, then honest wobble, scaled by the warranty you promised. Conditions first, though — the license is two slides away.

\(\hat{p} \pm z^* \cdot \sqrt{\dfrac{\hat{p}(1-\hat{p})}{n}}\)

Before building Maya's for real — feel how the width answers to \(n\), \(\hat{p}\), and the warranty:

See it · C4

What the width answers to

Predict first: to cut the width in HALF, how many times more \(n\)? And where does \(\hat{p}\) make the interval widest — 0.2, 0.5, or 0.8? Check both with the sliders.

Figure: Drag the sliders to see how n and affect the interval width. Switch confidence levels to see the width–confidence tradeoff. The row at the bottom always shows what happens when you quadruple the sample size: E is halved.

I do · watch a full build

  1. 1 I do
  2. 2 We do
  3. 3 You do

Watch me decide every step out loud — check the license (counts, not shape), get center and spread from one number, compute the margin, assemble the claim and say it in words.

The all-nighter interval — built aloud

Step 1 · Check the license

Random sample? The 300 came from the survey's random draw — ✓. Enough of both answers? 75 yeses and 225 nos, both far past 10 — ✓ (the full rule is two slides away). Small slice? 300 is well under 10% of 48,000 — ✓. And notice what I did not check: no \(\sigma\), no \(df\), no t-table.

Step 2 · Center and spread — one number

Center: \(\hat{p} = 75/300\) \(= 0.25\). Spread, from the same number: \(\text{SE} = \sqrt{0.25 \times 0.75 / 300}\) \(= \sqrt{0.000625}\) \(= 0.025\). Sanity check: 2.5 points of wobble from 300 yes/no answers — believable.

Step 3 · Compute the margin

The paper's standard warranty is 95%, so \(z^* = 1.96\). Reach: \(E = 1.96 \times 0.025\) \(= 0.049\) — ±4.9 points. (If the editor demanded 99%, only this step changes: \(2.576 \times 0.025\) \(= 0.0644\).)

Step 4 · Assemble — and say it

\(0.25 \pm 0.049\) \(\Rightarrow (0.201,\ 0.299)\). In words: "We are 95% confident that between 20.1% and 29.9% of students pulled at least one all-nighter last month." One in four, give or take five points.

The reader asked for one number; the honest answer is a ten-point window — and the window itself is news: even its floor says one student in five. What the window can't do is come cheap. The editor's reply is already in: "Pin it to ±2." Pricing that demand is the last stop today.

We do · you finish it

  1. 1 I do
  2. 2 We do
  3. 3 You do

New survey, new count: you supply the margin and the interval — the numbers change, the habit doesn't.

Second letter: who naps?

Same week, next question: "Fine — all-nighters happen. Do students at least catch up? Who naps?" The residence association's random survey of 400 students has the count: 80 report a daily nap. Build the 95% interval.

  1. License: random draw ✓; 80 yeses and 320 nos, both ≥ 10 ✓; 400 is under 10% of the residence population ✓.
  2. Center: \(\hat{p} = 80/400\) \(= 0.20\).
  3. Spread, from the same number: \(\text{SE} = \sqrt{0.20 \times 0.80 / 400}\) \(= \sqrt{0.0004}\) \(= 0.02\).
  4. Finish it: \(E = ?\) — and the interval is…? (Then say the sentence.)

    \(E = 1.96 \times 0.02\) \(= 0.0392\) \(\Rightarrow 0.20 \pm 0.0392\) \(= (0.1608,\ 0.2392)\). "We are 95% confident that between about 16% and 24% of residence students nap daily." New count, same habit — the habit is the deliverable.

Practice · build the interval, solo · fresh counts on demand

  1. 1 I do
  2. 2 We do
  3. 3 You do

Now you build whole proportion intervals solo — fresh counts every re-roll, support at its minimum.

Core concept 5 · C5

The license: ten of each

The z-label is earned by a bell — and the bell only shows up when the sample holds enough of both answers. For proportions, the license is a pair of counts.

Ten of each

\(n\hat{p}\) and \(n(1-\hat{p})\) are literally the yes-count and the no-count — \(x\) and \(n - x\). Both must reach 10. Maya's: 75 and 225 — licensed with room to spare.

\(n\hat{p} \geq 10,\ \ n(1-\hat{p}) \geq 10\)

Why ten

Push \(\hat{p}\) near an edge — say 19 yeses out of 20 — and \(\hat{p}\)'s distribution squashes against the wall: strongly skewed, no bell, and the "95%" label goes quietly false. Ten of each keeps room on both sides for the count's distribution to turn symmetric.

The quiet two

Random sample — independence comes from the draw, the rule since the very first survey — and \(n\) at most 10% of the population: 300 of 48,000 clears it easily. Break the draw and no formula rescues you.

\(n \leq 10\%\ \text{of pop.}\)

Two counts, two quiet checks — cheaper than last class's dot-plot stare, and just as load-bearing. Watch the bell fail when the counts don't clear:

See it · C5

Where the bell earns its keep

Predict first: \(n = 30\), \(p = 0.1\) — only 3 expected yeses. What shape will the bars take, and will the overlaid bell fit? Then find settings where the fit turns good — and read the two counts at that moment.

Figure: Exact binomial probabilities for p̂ (bars) vs. the normal approximation (curve). When n·p or n·(1−p) falls below 10, the bars are visibly asymmetric and the curve is a poor fit — the conditions badges flip red. Increase n or move p toward 0.5 to watch them flip green.

Practice · check the license, solo · fresh counts on demand

Which tool applies? · mixes earlier classes + today

Core concept 6 · C6

The sentence: what 95% still means

Third interval, same contract. The words survive unchanged — and so do the misreads, which is why this slide exists every single time.

What Maya may print

"We are 95% confident that between 20.1% and 29.9% of students pulled an all-nighter last month." — the accepted shorthand, crediting the method that built the interval.

What the 95% is

A statement about the procedure: run 1,000 surveys like hers and about 950 of the resulting intervals capture the true \(p\). Hers is one of the batch — built by a method with a 95% track record.

What it still is not

Not a probability on \(p\) — a fixed constant that either sits in \((0.201,\ 0.299)\) or doesn't. Not a claim that 20–30% of students are "in the interval." Not a promise about the next survey.

The warranty is on the factory, not the unit — third class running. Watch the batch:

See it · C6

Twenty surveys, one fixed truth

Twenty intervals, one dashed \(p\). Predict: at 95%, how many of the twenty will miss — and does the dashed line ever move? Then drop the level to 90% and predict again.

Figure: Each bar is a confidence interval built from a different random sample of size n. Teal bars capture the fixed true p (dashed line); red bars miss it. The line never moves — only the intervals vary. Over many repetitions, approximately C% of all such intervals will contain p.

Quick check · print it or fix it?

Core concept 7 · C7

Run the formula backwards: pricing a poll

The editor reads the ten-point window and sends the inevitable reply: "Pin it to ±2. What does that cost me?" Decide the precision first — then solve for \(n\).

Solve for n

The margin formula, inverted. \(p^*\) is your best prior guess at \(p\) — last month's poll, a pilot, or (if you truly have nothing) the built-in insurance value. Decide \(E\) and the warranty; the formula quotes the sample.

\(n = \left(\dfrac{z^*}{E}\right)^{2} p^*(1-p^*)\)

The quote

±2 points at 95%, no prior: \(n = 9604 \times 0.25\) \(= 2401\) people. Using last month's \(p^* = 0.25\): \(9604 \times 0.1875\) \(= 1800.75\) \(\to 1801\). The 600-person gap is the premium for admitting you don't know \(p\). (±3 points? 1068 — why national polls hover near a thousand.)

Two traps

Always round up: 1800.75 means 1801 — rounding down quietly breaks the promised \(E\). And \(p^* = 0.5\) is not a guess that \(p\) is 50% — it's the worst case: \(p(1-p)\) peaks at 0.5, so budgeting there can never undershoot.

\(p^*(1-p^*) \leq 0.25\)

Precision is bought in people, and the price grows with the square of the demand — halving \(E\) still quadruples \(n\). One picture explains the 0.5 rule:

See it · C7

Why 0.5 is the insurance

The curve is \(p^*(1-p^*)\). Predict its peak. Then I'll set \(p^* = 0.3\) — a shaded risk zone appears: what does it mean, and which single choice of \(p^*\) makes it vanish for good?

Figure: The curve shows p*(1−p*) — the factor that drives sample size — as a function of the prior estimate p*. The teal dot at p* = 0.5 marks the peak (0.25): choosing p* = 0.5 guarantees the largest required n, covering all possible true p. The red dot shows your current p*. The shaded zone shows every true p value that would require more samples than your current choice provides — drag the slider toward 0.5 to watch the zone disappear.

Practice · price the poll, solo · fresh targets on demand

Production task · the reply she publishes

Exit ticket · 60 seconds · anonymous

What Maya files

One in four — and the price of sharper

The reader's question got an honest answer: between 20% and 30% of students pulled an all-nighter last month — one in four, give or take five points. The skeleton from the last two classes carried a brand-new kind of data without bending.

The fork grew a second prong: measurements → mean tools (\(z\) if \(\sigma\) is known, \(t\) if \(s\) stands in); yes/no counts → the proportion interval, where \(z\) keeps its label because \(\hat{p}\) prices its own spread — licensed by ten of each.

And the formula runs backwards: decide the precision, quote the sample — 2,401 for ±2 points with the 0.5 insurance. Precision is a budget item now, not a wish.

Next on the desk: the campus wellness office reads Maya's numbers and pushes back — "students get eight hours; your survey must be off." A claim, her data, and a referee to decide between them… that's a hypothesis test, and it starts next class.

Vote counts on today's slides are simulated for rehearsal; live sessions show the room's real votes.