Join · scan

OTTO-DS3

ottolearn.org/join

Module 1 · Descriptive Statistics · DS-3

Central Tendency
Which number is "typical"?

Maya's sleep story ran. Her editor's next assignment: a rival paper claims the average rent near campus is $1,650 — but every student Maya asks pays about $900. Someone's number is misleading. Today we find out whose, and why.

Live session · join with code OTTO-DS3 · ottolearn.org/join

Same rules as always: vote before I explain, every poll is anonymous, and changing your mind after discussion is the goal.

Today · Maya's rent fact-check

Five calls before the story goes to print — yours first.

Every step is a decision Maya has to make before her rent story runs — and a poll you'll answer first.

Recall · last class · 45 seconds

Warm-up · 60 seconds · anonymous

Keep your answer in mind. By the end of class you'll be able to name exactly what's happening — and fix the headline.

Core concept 1 · C1

The sample mean: the balance point of the data

First tool: the one everybody calls "the average". Add everything up, divide by how many. But the number you get is best understood as a balance point:

The formula

Add all \(n\) values, divide by \(n\). \(\bar{x}\) ("x-bar") is the sample mean — the population version, \(\mu\), arrives in Module 3.

\(\bar{x} = \dfrac{\sum x_i}{n}\)

A balance point

Put the values as weights on a number line: \(\bar{x}\) is exactly where the beam balances. Deviations cancel — what's above offsets what's below.

\(\sum (x_i - \bar{x}) = 0\)

Not always a data value

The mean of \(\{2, 3, 5\}\) is \(10 \div 3 \approx 3.33\) — a value in no dataset. "2.4 children per family" is a summary, not a family.

One number carries the whole dataset — which is exactly why we need to know when it can be trusted.

See it · C1

Balance the beam

Predict first: where must the fulcrum go to level the beam? Then drag one weight far to the right — where does the balance point go?

Practice · C1 · compute the mean

I do · watch me fact-check

Fact-check the headline: compute \(\bar{x}\) ourselves

Step 1 · List the data

Maya pulls 6 ordinary listings from the street the story names ($/month): 820, 860, 900, 920, 960, 990.

Step 2 · Sum with \(\Sigma\)

\(\sum x_i = 820 + 860 + 900 + 920 + 960 + 990 = 5450\)

Step 3 · Divide by \(n\)

\(\bar{x} = \dfrac{5450}{6} \approx 908.33\) — note: not itself a data value, and that's fine.

Step 4 · Interpret

Typical rent on that street ≈ $908. Nowhere near the headline's $1,650.

So where does $1,650 come from? There is a 7th listing on that street — a penthouse at $6,100. Predict what it does:

C2 · Predict first · ConcepTest · vote → discuss → re-vote

Core concept 2 · C2 · why that happened

The mean is sensitive; the median is resistant

The mean absorbs magnitudes

Every dollar of every value pulls on \(\bar{x}\). The penthouse alone contributes \(6100 \div 7 \approx \$871\) of the $1,650 average — more than the other six listings combined.

The median counts positions

Only the order matters. Swap the penthouse's $6,100 for $60,000 and the median doesn't move a single dollar.

The vocabulary

The median is resistant to outliers; the mean is sensitive. That one word is why Statistics Canada reports median household income.

One penthouse: mean +$742, median +$10. That asymmetry is today's whole story.

See it · C2

Drag the outlier

Predict first: how far right must you drag the outlier before the median jumps the way the mean just did?

Core concept 3 · C3

The median: sort, locate, read

Sort first — always

The median is the middle of the sorted list. The middle of the raw list is a wrong answer delivered with confidence — the #1 median error.

Odd \(n\): one middle

The median sits at position \(\frac{n+1}{2}\) of the sorted list. Seven listings → position 4.

\(\dfrac{n+1}{2}\)

Even \(n\): average two

Average the values at positions \(\frac{n}{2}\) and \(\frac{n}{2}+1\). The result may not be a data value — that's fine.

\(\dfrac{n}{2},\ \dfrac{n}{2}+1\)

Position, not magnitude — that's the entire source of the median's resistance.

I do · watch the algorithm run

  1. 1 I do
  2. 2 We do
  3. 3 You do

Full support — watch me sort the listings and count to the middle position.

Raw data → sorted → located → read

The full median algorithm, one step at a time — watch the #1 error (skipping the sort) become visible in step 1 vs step 2. Then toggle to even \(n\) to see the algorithm change.

We do · you finish it

  1. 1 I do
  2. 2 We do
  3. 3 You do

Support fades — the values are sorted; you locate the middle for an even count.

An 8th listing appears

Maya finds an 8th listing at $1,040. Sorted: 820, 860, 900, 920, 960, 990, 1040, 6100 — now \(n = 8\).

  1. \(n = 8\) is even → average the values at positions \(\frac{8}{2} = 4\) and \(\frac{8}{2} + 1 = 5\).
  2. Position 4 holds \(920\); position 5 holds \(960\).
  3. What is the median rent now?

    \(\text{Median} = \frac{920 + 960}{2} = \$940\). Still an ordinary apartment — the penthouse still can't touch it.

Practice · C3 · find the median

  1. 1 I do
  2. 2 We do
  3. 3 You do

On your own now — fresh lists to order. Re-roll as many as the room needs.

Core concept 4 · C4

The mode: what shows up most

Most frequent value

The mode is whatever occurs most often — and it's the only measure that works for categories. You can't average "Instagram" and "TikTok", but you can count which appears most.

One, two, or none

One winner → unimodal. Two values tie → bimodal (both are modes). Every value appears equally often → no mode at all.

Bimodal is a clue

Two modes often mean two groups hiding in one dataset — men's and women's shoe sizes recorded together. Ask: is this really one population?

Mean and median need numbers. The mode just needs a tally.

See it · C4

Build the tally, watch the mode emerge

Predict first: which bar wins? Then switch to the bimodal example — what should we say when two bars tie?

Quick check · C4 · 45 seconds

Practice · C4 · name the mode

Core concept 5 · C5

Which number goes in the headline? The chain, out loud

Ask: what TYPE?

The DS-1 question. Categories → the mode is the only option — done. Numbers → keep going. Rent is quantitative → keep going.

Ask: what SHAPE?

The DS-2 question — look at the histogram. Symmetric, no outliers → the mean is fine (it uses every value). Skewed or outliers → one more step.

Skewed → median

The mean chases the tail; the median holds the centre. Maya's rents are right-skewed → she files the median: $920, not $1,650.

Type → shape → decision. Every "which average?" call runs this same chain. Your turn:

We do · you make the call

Next assignment: the marathon story

Ten finish times — nine recreational runners between 3.5 h and 5.0 h, and one elite runner at 2.1 h.

  1. Type: finish time is quantitative → mean or median both possible.
  2. Shape: one extreme LOW value → the long tail points left (left-skewed).
  3. Which measure does she report — and which direction is the mean being pulled?

    The median. The mean is dragged left, toward the elite 2.1 h — it understates what a typical recreational runner takes.

See it · C5

Three shapes, one rule

Predict first: in a right-skewed distribution, which is bigger — mean or median? Switch tabs and check your rule against all three shapes.

Which tool applies? · mixes DS-1 + DS-2

Practice · C5 · which measure?

Production · partners · one sentence

Exit ticket · muddiest point · anonymous

Wrap

You can now put one honest number in a headline

Compute — mean \(\bar{x} = \sum x_i / n\), median (sort → middle), mode (tally): three different questions about "typical".

Watch the tail — the mean chases outliers and skew; the median holds its position.

Choose by shape — type → shape → outliers. The data picks the measure; the measure keeps the headline honest.

Next class (DS-4): two datasets can share the same centre and be wildly different — how spread out is the data?

Rents and scenarios are illustrative, constructed for this class — not a cited market survey.