Join · scan

OTTO-DS5

ottolearn.org/join

Module 1 · Descriptive Statistics · DS-5

Position & Shape
Is 5.5 hours unusual?

Maya's spread follow-up runs — and the inbox fills up. One email sticks: "I sleep 5.5 hours. Should I be worried?" The average is 6.8. But "below average" isn't an answer — today Maya learns to say exactly where one reader stands among 48,000 students.

Live session · join with code OTTO-DS5 · ottolearn.org/join

Same rules as always: vote before I explain, every poll is anonymous, and changing your mind after discussion is the goal.

Today · the mailbag: "should I be worried?"

One email, five questions about position.

Each question is one step toward Maya's reply — and a poll you'll answer first.

Recall · last class · 45 seconds

Warm-up · 60 seconds · anonymous

Hold your vote. By the end of class you'll answer the reader three ways: a rank, a distance — and a percentage.

Core concept 1 · C1

Percentiles: where in the line-up?

Tool #1 turns "below average" into a rank. Line up all 300 sleepers from shortest to longest — where in the queue does the reader stand?

The percentile

Sort the data. \(P_k\) is the value with about \(k\%\) of the data below it. "You're at the 16th percentile" = about 16% of students sleep less than you do.

\(P_k\)

You already know three

Last class's quartiles are percentiles wearing familiar names: \(Q_1 = P_{25}\), the median \(= P_{50}\), \(Q_3 = P_{75}\). Deciles cut by tens: \(D_1 = P_{10}\) up to \(D_9 = P_{90}\).

\(Q_1 = P_{25},\ Q_3 = P_{75}\)

Rank \(\neq\) score

"83rd percentile" does not mean a score of 83. The \(k\) counts who's below you; the value sitting at that position is a different number entirely.

By hand: \(P_k\) is the value at rank \(L = \lceil \tfrac{k}{100} \times n \rceil\) of the sorted list. Watch the rule run in both directions:

Watch it · C1

Rank → value, value → rank

Predict first: for these 12 values, \(P_{75}\) needs rank \(L = \lceil 0.75 \times 12 \rceil\) — which position is that? Then reverse the map: pick a value and read off its rank.

Practice · C1 · find the percentile

Core concept 2 · C2

The z-score: distance, measured in spreads

Distance, in spreads

Subtract the mean, divide by the standard deviation. The result counts standard deviations: \(z = -2\) means two spreads below the mean — whatever the original units were.

\(z = \dfrac{x - \bar{x}}{s}\)

Population or sample

Data for a whole population → \(\mu\) and \(\sigma\). A sample — like Maya's 300 — → \(\bar{x}\) and \(s\). Same recipe; match the symbols to the data you actually have.

\(z = \dfrac{x - \mu}{\sigma}\)

The units cancel

Hours minus hours, divided by hours: nothing survives but a pure number. That's the superpower — sleep, rents and exam scores all land on one common scale.

Time to answer the reader properly. Watch me do the first one — out loud, every decision:

I do · watch me place the reader

  1. 1 I do
  2. 2 We do
  3. 3 You do

Full support — watch me choose the symbols, subtract, divide, and say it in words.

The reader's 5.5 hours, out loud

Step 1 · Pick the symbols

The 300 responses are a sample of 48,000 students — so \(\bar{x} = 6.8\), \(s = 1.3\), and the recipe is \(z = (x - \bar{x}) \div s\).

Step 2 · Subtract the mean

\(5.5 - 6.8 = -1.3\) h. The sign is already talking: negative → below the mean.

Step 3 · Divide by the spread

\(-1.3 \div 1.3 = -1.00\). The hours cancel — what remains is a count of standard deviations.

Step 4 · Say it in words

The reader sits exactly one typical distance below the mean. Below average, yes — but by an ordinary amount, not an alarming one.

\(z = -1.00\). Then a second email lands: "I sleep 9.4 hours. Basically superhuman, right?" Your turn —

We do · you finish it

  1. 1 I do
  2. 2 We do
  3. 3 You do

Support fades — the setup is done; you take the last division and read the result.

The 9.4-hour reader

Same survey: \(\bar{x} = 6.8\) h, \(s = 1.3\) h. The new email claims \(x = 9.4\) h.

  1. Subtract the mean: \(9.4 - 6.8 = +2.6\) h — positive, so above the mean.
  2. Divide by \(s = 1.3\): what is \(z\) — and how far out is that, really?

    \(z = 2.6 \div 1.3 = +2.00\) — exactly two standard deviations above the mean. Genuinely far out; whether it earns "superhuman" is a percentage question. That's C5.

Practice · C2 · compute the z-score

  1. 1 I do
  2. 2 We do
  3. 3 You do

On your own now — fresh numbers, every step yours. Re-roll as many as the room needs.

See it · C2

One value on the z ruler

Predict first: drop \(x\) exactly on the mean — what does \(z\) read? Then slide \(x\) up by one full \(\sigma\) and watch the bracket count it.

Core concept 3 · C3

Reading a z-score: sign, size, and no units

Read the sign

Positive → above the mean; negative → below. \(z = -1.8\) isn't an error or a bad score — just 1.8 spreads below average. Direction, nothing more.

Read the size

\(z = 0\) is the mean itself. Within \(\pm 1\): ordinary. Around \(\pm 2\): far out. Beyond \(\pm 3\): rare in most data — worth a second look, the same job last class's fences did.

Compare anything

Unitless means portable: a rent and a sleep duration can meet on the same scale. Farther from its own mean, in its own spreads = more extreme — regardless of units.

Two famous values from Maya's own reporting are about to meet on that scale. Commit to a read before anyone computes:

Which is more extreme? · mixes DS-3 + DS-4

See it · C3

Two distributions, one ruler

Predict first: can a smaller raw gap win the bigger \(|z|\)? Rebuild the penthouse-vs-sleeper duel — set each panel's mean, spread and value — and let the shared ruler settle it.

Practice · C3 · interpret z-scores

C4 · Predict first · ConcepTest · vote → discuss → re-vote

Core concept 4 · C4 · why that happened

Skewness: the tail names the shape

Name the thin side

The tail is the long, thin side — and it names the skew. Tail right → right-skewed. Tail left → left-skewed. The peak usually sits on the opposite side.

The mean chases the tail

Tail values are extreme, and the sensitive mean gets dragged toward them; the resistant median barely moves. Right tail: mean > median. Left tail: mean < median.

Diagnose without a picture

Salaries: mean \(\$112{,}000\), median \(\$84{,}000\). No histogram needed — the mean sits far above the median, so the tail is on the right, and the median belongs in the report.

Three reference shapes, side by side. Find the TAIL and the BULK in each:

See it · C4

Left-skewed, symmetric, right-skewed

Predict first: in the right-skewed panel, does the mean marker sit left or right of the median? Check yourself on all three shapes.

Explore · C4

Morph the shape, watch the measures

Predict first: pushing the skew slider right — which marker chases the new tail, mean or median? Drag, watch them separate, then re-converge at symmetric.

Practice · C4 · name the shape

Core concept 5 · C5

The Empirical Rule: what bell-shaped buys you

Three benchmarks

Bell-shaped data keeps a schedule: about 68% of it within 1 SD of the mean, 95% within 2, 99.7% within 3. Memorize the trio.

\(\mu \pm 1\sigma,\ \mu \pm 2\sigma,\ \mu \pm 3\sigma\)

Bell-shaped ONLY

The percentages are properties of the bell shape — that's why the shape check (C4) comes first. Skewed, bimodal or outlier-heavy data: the rule's numbers go badly wrong.

Approximately, always

Say "approximately 68%", never "exactly 68%". These are benchmarks for real, roughly-normal data — not laws.

Maya pulls up her survey histogram from the charting class: one hump, roughly symmetric. The rule is available — watch it lay its bands:

See it · C5

The 68–95–99.7 bands, live

Predict first: to cover 95% of the data, how many SDs wide must the net be? Move \(\mu\) and \(\sigma\) — the boundary values change, the three percentages refuse to.

We do · answer the reader

So IS 5.5 hours unusual?

Survey: \(\bar{x} = 6.8\) h, \(s = 1.3\) h, shape approximately bell. The reader's \(z = -1\) puts her exactly on the lower edge of the 68% band (\(5.5\) to \(8.1\) h).

  1. Within \(\pm 1\) SD lies approximately 68% of the survey — so approximately 32% falls outside the band.
  2. The bell is symmetric: that 32% splits evenly between the two tails.
  3. Approximately what percentage of students sleep less than the reader — and what should Maya write back?

    About 16%. The reply: "You're at roughly the 16th percentile — about 84% of students sleep more than you. Lower than typical, but one student in six is right there with you."

Quick check · C5 · 45 seconds

Practice · C5 · 68–95–99.7

Production · partners · one sentence

Exit ticket · muddiest point · anonymous

Wrap · Module 1 complete

You can now place any value in its distribution

Locate — a percentile gives the rank (\(P_k\); quartiles and deciles are special cases); the z-score gives the distance: \(z = (x - \bar{x}) \div s\), counted in standard deviations.

Read the shape — the tail names the skew, and the mean chases it: mean > median → right tail; mean < median → left tail. Skewed data → resistant pair (median + IQR).

Benchmark — bell-shaped only: approximately 68 / 95 / 99.7% within 1 / 2 / 3 SDs of the mean. Check the shape before quoting a percentage.

That completes Module 1: Maya can sample, chart, and summarize — honestly. Next class (PR-1): a reader shrugs off the all-nighter cluster as "just chance". Is it? That question needs probability.

The survey figures (\(\bar{x} = 6.8\) h, \(s = 1.3\) h) and reader emails are illustrative, constructed for this class — not survey results.