Join · scan

OTTO-INF2

ottolearn.org/join

Module 4 · Statistical Inference · INF-2

Confidence Intervals for a Mean
"What can we actually print?"

Maya's reply to the skeptic ended with "a different 300 lands within about ±0.15 h of the truth." Her editor circles the word about: "I can't print 'about'. Give me an exact range — and tell me how sure you are of it." Today Maya builds the sentence she's been chasing since her very first survey: an interval around 6.8 with a stated warranty, good for all 48,000 students.

Live session · join with code OTTO-INF2 · ottolearn.org/join

Same rules as always: vote before I explain, every poll is anonymous, and changing your mind after discussion is the goal.

Today · one demand from the editor, five answers

From "about 6.8" to a sentence the paper can print.

Turn a point estimate into an interval, learn the recipe and where its critical value comes from, pin down what "95% confident" really promises, see what makes intervals wide — then price the editor's precision demand in students. Every step starts with your vote.

Recall · last class · 45 seconds

Warm-up · commit first · 60 seconds · anonymous

Hold your answer — it is today's lesson. By the end of class you'll have the exact range, the exact warranty, and the price list for making the range tighter.

Core concept 1 · C1

A point estimate answers "what?" — an interval answers "how sure?"

The editor's complaint names a real gap. \(\bar{x} = 6.8\) is Maya's single best guess at \(\mu\) — her point estimate — but alone it says nothing about how far off it might be. The fix: attach a margin of error and print a range.

The point estimate

The sample's best single guess at \(\mu\): \(\bar{x} = 6.8\) h. Unbiased (it aims at the truth — last class) but almost surely not exactly right. A different 300 gives a different 6.8-ish number.

\(\bar{x} = 6.8\) h

The margin of error

The new object: \(E\), the half-width of the claim. Built from last class's yardstick — \(E = z^* \times \text{SE}\) — it says how far out from \(\bar{x}\) the claim reaches. Small \(E\) = a precise claim.

\(E = z^* \cdot \text{SE}\)

The interval estimate

The printable object: \(\bar{x} \pm E\), a range that owns up to its own uncertainty. Instead of "the average is 6.8", the paper prints "the average is between here and here" — with a warranty attached.

\(\bar{x} \pm E\)

The warm-up, resolved: the strongest honest sentence is a range with a warranty. Two questions remain — how wide is the range (the recipe), and what exactly does the warranty promise (the day's summit).

Core concept 2 · C2

The recipe: three ingredients, one line

Every confidence interval for a mean is the same assembly: start at the point estimate, step out \(z^*\) standard errors in both directions. Two ingredients are already on Maya's desk; only one is new.

Ingredient 1 · the center

The sample mean \(\bar{x}\) — the interval is built around it. Maya's: 6.8 h from her 300 responses. This is the only ingredient the data hands you directly.

\(\bar{x} = 6.8\) h

Ingredient 2 · the yardstick

The standard error \(\text{SE} = \sigma/\sqrt{n}\) — last class's object, unchanged. It sets the scale of the reach: how far averages typically wander from \(\mu\). Maya's: \(\approx 0.075\) h.

\(\text{SE} = \dfrac{\sigma}{\sqrt{n}}\)

Ingredient 3 · the warranty dial

The critical value \(z^*\) — how many SEs to reach out. You choose it by choosing a confidence level: more confidence, bigger \(z^*\), wider interval. Where its numbers come from is next.

\(z^*\)

The whole recipe: \(\bar{x} \pm z^* \cdot \dfrac{\sigma}{\sqrt{n}}\) — center, dial, yardstick. License conditions first, though: a random sample, and \(n \geq 30\) (or a normal population) so the CLT vouches for the bell. One ingredient still needs unpacking: \(z^*\).

Core concept 3 · C3

Where \(z^*\) comes from: the middle-C% of the bell

The warranty dial isn't magic — every \(z^*\) is an ordinary inverse table lookup, the skill from the normal-distribution class. The only subtlety is which area you look up.

The definition

\(z^*\) is the z-score that fences in the middle C% of the standard normal. For 95%: leave \(2.5\%\) in each tail, so look up the value whose left-tail area is \(0.975\) — the table answers \(1.96\).

\(\Phi(z^*) = 1 - \alpha/2\)

The three worth memorizing

90% \(\rightarrow 1.645\) · 95% \(\rightarrow 1.96\) · 99% \(\rightarrow 2.576\). These three run every inference problem in this module — and any other level is just another inverse lookup away.

\(1.645,\ 1.96,\ 2.576\)

The lookup trap

For a 95% interval, do not look up the z with left-tail 0.95 — that's 1.645, the 90% dial. The confidence level is the area in the middle; the table wants the area to the left. Convert first: \(0.95 + 0.025 = 0.975\).

\(z^* \neq 0.95\)

One picture holds all of it: the level is the tinted middle, the tails split what's left, and \(z^*\) is the fence post. See the anatomy:

See it · C3

The anatomy of \(z^*\)

Predict first: as I slide the confidence level from 90% up to 99%, which way do the fence posts move? And where would they sit for 100% confidence?

Figure 4: The anatomy of z*. The blue region contains the middle C% of the standard normal distribution; the orange tails together hold the remaining (1−C)%. The dashed lines mark ±z* — a distance along the z-axis, not an area. Toggle confidence levels to see how z* is derived from the inverse cumulative normal.

Quick check · the lookup trap

See it · C2

Watch the interval assemble

Predict first: I'll set \(\bar{x} = 6.8\), \(\sigma = 1.3\), \(n = 300\), level 95%. Roughly how wide will the interval be — hours, tenths of an hour, or hundredths? Then: which control will stretch it fastest?

Figure 6: Build a confidence interval. The green tick marks the point estimate x̄; the blue bracket reaches x̄ ± E in both directions, where the margin of error is E = z* · σ/√n. Raise the confidence level and the bracket widens; raise the sample size n and it narrows.

I do · watch a full build

  1. 1 I do
  2. 2 We do
  3. 3 You do

Watch me decide every step out loud — check the license, choose the warranty, compute the margin, assemble the claim and say it in words.

Maya's interval: from 6.8 to a printable range

Step 1 · Check the license

Random sample? Yes — the 300 were drawn by a plan, not by whoever answered. Size? \(n = 300 \geq 30\), so the CLT vouches for the bell. Spread? The survey's \(s = 1.3\) h stands in for \(\sigma\) — fine at this \(n\), and I'm noting the IOU out loud: paying it properly is next class.

Step 2 · Choose the warranty

The paper's standard is 95%. Middle 95% of the bell means 2.5% left out in each tail — so \(z^*\) is the value with left-tail area 0.975: \(z^* = 1.96\). I chose this number; the data didn't hand it to me.

Step 3 · Compute the margin

Yardstick first: \(\text{SE} = 1.3/\sqrt{300}\) \(\approx 0.075\) h. Then the reach: \(E = 1.96 \times 0.075\) \(\approx 0.147\) h — about 9 minutes. Before I assemble, a sanity check: two yardsticks wide, so a tenth-of-an-hour-ish margin. It is.

Step 4 · Assemble — and say it

\(6.8 \pm 0.147\) \(\Rightarrow (6.65,\ 6.95)\) h. In words: "We are 95% confident the campus average sits between 6.65 and 6.95 hours." Notice what the range does: even its top is more than an hour below the 8-hour recommendation.

The habit: check the license, choose the warranty (\(z^*\)), compute \(E = z^* \times \text{SE}\), assemble \(\bar{x} \pm E\) — then say the claim in words.

We do · you finish it

  1. 1 I do
  2. 2 We do
  3. 3 You do

Same survey, bigger warranty: you turn the dial to 99% and supply the new interval — and watch what the extra confidence costs.

Legal wants more certainty: "make it 99%"

Before printing, the paper's lawyer asks what the range would be at 99% confidence. Same survey, same numbers: \(\bar{x} = 6.8\), \(\text{SE} \approx 0.075\) h. Re-run the build.

  1. License: unchanged — same sample, same \(n = 300\). Only the warranty dial moves.
  2. Choose the warranty: 99% leaves 0.5% in each tail, so \(z^* = 2.576\).
  3. Finish it: \(E = ?\) — and the 99% interval is… wider or narrower than (6.65, 6.95)?

    \(E = 2.576 \times 0.075\) \(\approx 0.193\) h \(\Rightarrow 6.8 \pm 0.193\) \(= (6.61,\ 6.99)\) h. Wider. More warranty costs more width — the claim gets safer and vaguer at the same time. (Why, and what it costs to undo, is exactly where today goes next.)

Practice · build the interval, solo · fresh numbers on demand

  1. 1 I do
  2. 2 We do
  3. 3 You do

Now you build whole intervals solo — fresh numbers every re-roll, the three critical values on screen, support at its minimum.

Commit first · today's summit

Core concept 4 · C4

The warranty is on the method, not the interval

The trap has one root: treating \(\mu\) as if it moves. It doesn't. The true campus average is some fixed number — unknown, but not random. What's random is the interval, because it's built from a random sample.

\(\mu\) doesn't move

The true mean is a constant. Once (6.65, 6.95) is computed, \(\mu\) is either inside it or not — probability 1 or 0, we just can't see which. Probability language about this interval has nothing left to attach to.

\(\mu\) is a constant

What the 95% is NOT

Not the chance that \(\mu\) is in this interval. Not the fraction of students inside the range. Not the chance the data is right. Every one of these puts the randomness on the wrong object — on \(\mu\) or the students, instead of on the sampling.

What Maya may print

"We are 95% confident the campus average is between 6.65 and 6.95 h." — the accepted shorthand, read as: we used a method that captures the truth in 95% of surveys like this one. Same words as her draft, one crucial repair: the 95% credits the method.

One picture settles it: run the survey many times, build many intervals, and watch — the intervals jump around, \(\mu\) stands still, and about 95 in 100 of them catch it. Watch it live:

See it · C4 · the anchor, resolved

One hundred surveys, one hundred intervals, one fixed \(\mu\)

Predict first: at 95%, roughly how many of 100 intervals will miss \(\mu\) entirely? Then: when I drop the level to 90%, do the intervals get wider or narrower — and do more or fewer miss?

Figure 1: CI Coverage Explorer — each horizontal bar is one confidence interval built from a different random sample. Solid green bars capture the true population mean μ (orange dashed line); hatched red bars marked ✗ miss it. Watch the running tally approach the stated confidence level as you draw more samples.

Practice · what may the report claim? · fresh numbers on demand

Core concept 5 · C5

Three dials set the width — and you only control one

The 99% re-build already showed it: width is not fate, it's arithmetic. \(\text{width} = 2E = 2\,z^* \sigma/\sqrt{n}\) — three quantities, three dials, but they are not equally available.

The warranty dial

Raise the level and \(z^*\) grows: Maya's margin runs 0.123 h at 90%, 0.147 at 95%, 0.193 at 99%. More confidence is more width — a safer claim is a vaguer claim, mechanically.

\(z^* \uparrow\ \Rightarrow\ E \uparrow\)

The sample-size dial

The one dial you truly own — and it fights back with a square root: to halve the width you must quadruple \(n\). Precision is bought in bulk, at √-scale prices (last class's diminishing returns, now with a price tag).

\(4n\ \Rightarrow\ E/2\)

Not a dial: \(\sigma\)

The population's own spread. More variable populations give wider intervals — but \(\sigma\) is a fact about the students, not a setting on the survey. You can't ask a campus to sleep more consistently.

\(\sigma\)

So the honest routes to a tighter printed range are exactly two: accept a weaker warranty, or collect more data. Watch the three dials move one interval:

See it · C5

The width, dial by dial

Predict first: which dial changes the width fastest — the level or \(n\)? And can any setting of \(n\) undo a jump from 95% to 99%?

Figure 5: All three intervals share the same centre x̄ on one common scale, so a bar that is twice as wide is drawn twice as long. Drag the n slider to watch every bar shrink together — doubling n reduces each width by 29%; quadrupling n halves all widths (use the 4×n button to see it). Drag σ to see how population variability scales all three proportionally. The relative widths (90% < 95% < 99%) never change because they depend only on z*.

Quick check · the price of half

Core concept 6 · C6

Planning ahead: how many students does a spec cost?

So far the data came first and the width fell out. Real surveys run backwards: the editor names a margin of error first, and Maya must find the smallest \(n\) that delivers it — before collecting anything.

Solve the margin for \(n\)

Set \(E = z^* \sigma/\sqrt{n}\) and rearrange: \(n = (z^* \sigma / E)^2\), rounded up. Every symbol on the right is known before the survey: the level picks \(z^*\), history supplies a \(\sigma\), the demand sets \(E\).

\(n = \left\lceil \left(\dfrac{z^* \sigma}{E}\right)^{\!2} \right\rceil\)

Always round UP

If the formula says 61.47, the answer is 62 — never 61. Rounding down leaves \(n\) short and the margin over spec: the ceiling isn't a convention, it's the difference between meeting the demand and missing it.

\(\lceil 61.47 \rceil = 62\)

The spec is a pre-data choice

\(E\) and the confidence level are decisions, made before any student is surveyed — the formula just prices them. That's the planning move: negotiate the spec while it's still cheap to change.

The editor's demand is about to get a price tag. First, see the whole trade-off as one curve:

See it · C6

The price curve: \(n\) against the margin you demand

Predict first: as the demanded margin \(E\) shrinks toward zero, does the required \(n\) grow steadily — or explode? Find where "a little tighter" starts costing hundreds of students.

Figure 7: The cost of precision. Drag the green dot, click anywhere on the chart, or focus it and use the arrow keys to set your desired margin of error E. The gold dot shows what happens if you halve E — the required sample size approximately quadruples (half the error means roughly 4× the sample), because n scales as 1/E². Notice how the curve accelerates steeply for small E: asking for twice the precision is far more expensive than it looks.

I do · watch a full pricing

  1. 1 I do
  2. 2 We do
  3. 3 You do

Watch me decide every step out loud — name the demand (\(E\), level), rearrange the margin formula for \(n\), compute, and round UP to the price.

The editor: "make the margin ±3 minutes"

Step 1 · Name the demand

The editor wants the printed range tight: a margin of \(\pm 3\) minutes \(= 0.05\) h, at the paper's standard 95%. So the spec is \(E = 0.05\), \(z^* = 1.96\) — and history supplies the spread: \(\sigma \approx 1.3\) h from the survey. All three knowns, no data needed.

Step 2 · Rearrange for \(n\)

The margin formula \(E = z^*\sigma/\sqrt{n}\), solved for the one unknown: \(n = \left(\dfrac{z^*\sigma}{E}\right)^{\!2}\). I say the units out loud as a check: SEs per demanded margin, squared — a count of students. It's the width machine, run in reverse.

Step 3 · Compute

Inside first: \(\dfrac{1.96 \times 1.3}{0.05}\) \(= \dfrac{2.548}{0.05}\) \(= 50.96\). Then square: \(50.96^2\) \(\approx 2596.9\). Expectation check before I round: 3 minutes is a third of her current 9-minute margin, so the bill should be about \(3^2 = 9\) times her 300. It is.

Step 4 · Round UP — and say the price

\(n = \lceil 2596.9 \rceil = \mathbf{2{,}597}\) students. Not 2,596 — that one falls just short of spec. In words: "±3 minutes costs nearly nine surveys' worth of students." Now the editor can decide if the tightness is worth the field work.

The habit: name the demand (\(E\), level, \(\sigma\)), rearrange \(n = (z^*\sigma/E)^2\), compute, round UP — then say the price in words.

We do · you finish it

  1. 1 I do
  2. 2 We do
  3. 3 You do

Same formula, softer demand: you re-run it at ±6 minutes and supply the new \(n\) — and watch what relaxing the spec refunds.

The editor blinks: "fine — ±6 minutes"

Twenty-six hundred students is too many. The editor relaxes the demand to \(\pm 6\) minutes \(= 0.1\) h, same 95%. Re-run the pricing: \(z^* = 1.96\), \(\sigma = 1.3\), \(E = 0.1\).

  1. Name the demand: \(E = 0.1\) h now — the only symbol that moved. Rearrange: \(n = (z^*\sigma/E)^2\), same as before.
  2. Inside the square: \(\dfrac{1.96 \times 1.3}{0.1}\) \(= \dfrac{2.548}{0.1}\) \(= 25.48\).
  3. Finish it: \(n = ?\) — and check it against the quadruple rule: the demand doubled, so the bill should…

    \(25.48^2 \approx 649.2\) \(\Rightarrow n = \lceil 649.2 \rceil = \mathbf{650}\) students. The demand doubled (0.05 → 0.1), the bill fell to a quarter (2,597 → 650) — the quadruple rule read backwards. Still more than double her current 300: even the softened spec isn't free.

Practice · price the spec, solo · fresh numbers on demand

  1. 1 I do
  2. 2 We do
  3. 3 You do

Now you price whole specs solo — fresh numbers every re-roll, the formula on screen, support at its minimum.

Which tool applies? · mixes earlier classes + today

Production task · the sentence she files

Exit ticket · 60 seconds · anonymous

What Maya files

The claim she's been chasing since the first survey — printed

The sentence the editor wanted exists, and it's exact: "We are 95% confident the campus average is between 6.65 and 6.95 hours." Built as \(\bar{x} \pm z^* \cdot \sigma/\sqrt{n}\) — center from the data, yardstick from last class, warranty dial chosen and owned.

The 95% is a warranty on the method: across many surveys like hers, 95 in 100 such intervals capture \(\mu\). This one either did or didn't — \(\mu\) is a constant, and no probability attaches to it. That one repair is what makes the sentence printable.

And width has a price list: 99% would stretch the range to (6.61, 6.99); halving the width costs \(4\times\) the students; the editor's ±3-minute wish priced out at 2,597 students and was renegotiated to ±6 minutes for 650. Precision is a budget line now.

One IOU is still open: Maya used her survey's \(s = 1.3\) as if it were the true \(\sigma\) — a stand-in the large sample forgives. Next class pays that debt properly: what happens when \(\sigma\) is honestly unknown and \(n\) isn't large — the t-distribution, and intervals that own up to one more layer of uncertainty.

Vote counts on today's slides are simulated for rehearsal; live sessions show the room's real votes.