Probability and statistics · GCSE Maths

Averages and data

Choose and calculate mean, median, mode and range, including from frequency tables, and estimate the mean of grouped data.

UNDERSTANDRETRIEVEREMEMBER
THE MEMORY HOOK
Mean is the share-out, median is the middle, mode is the most, range is the spread. Grouped data can only estimate the mean.

The important bits

What you need to know

  1. 1

    Mean = total of values ÷ number of values. In a frequency table, mean = Σ(fx) ÷ Σf, where x is the value or the midpoint.

  2. 2

    Median is the middle value once the data are ordered. For n values, it sits at position (n + 1) / 2.

  3. 3

    Mode is the most frequent value. A data set may have more than one mode, or none that is useful.

  4. 4

    Range = largest − smallest. It is a crude measure of spread, easily distorted by a single outlier.

  5. 5

    For grouped data, you cannot find the exact mean. Use midpoints to estimate. The estimate depends on the assumption that data sit at the midpoint.

  6. 6

    A frequency polygon plots frequencies against midpoints. A cumulative frequency graph plots running total against the upper class boundary.

  7. 7

    The interquartile range is Q3 − Q1. It measures the spread of the middle half and is more resistant to outliers than the range.

  8. 8

    Compare two data sets with a pair: an average and a spread. “Higher mean, smaller IQR” is a complete comparison; a lone mean is not.

Go deeper

Frequency tables are totals in disguise

If 2, 3 and 5 appear with frequencies 4, 6 and 10, you do not have three numbers; you have twenty. The total is 2×4 + 3×6 + 5×10 = 8 + 18 + 50 = 76. The mean is 76 ÷ 20 = 3.8. Students who divide 76 by 3 have averaged the labels, not the data. The median is the 10.5th value in an ordered list of 20, so it sits between the 10th and 11th. Both of those lie in the value 5 if the running frequencies are 4, then 10, then 20. Show a cumulative frequency column: 4, 10, 20. That column is how you find median and quartiles from a table without rewriting twenty numbers. The mode is 5, the value with frequency 10.

Go deeper

Grouped means are estimates, and should be named as such

A class 10 ≤ t < 20 has midpoint 15. That 15 stands in for every value in the class. If the frequencies are large, the estimate is often close; if a class is wide, it may not be. Write “estimated mean” on the answer line when the data are grouped. For 10 ≤ t < 20, 20 ≤ t < 30, 30 ≤ t < 40 with frequencies 5, 8, 7, the midpoints are 15, 25, 35, Σfx = 75 + 200 + 245 = 520, Σf = 20, estimated mean = 26. Modal class is 20 ≤ t < 30, not “25”. The median class is the one containing the 10.5th value, here also the middle class. A histogram uses frequency density, frequency ÷ class width, when widths are unequal. Bars of equal width can use frequency; unequal widths cannot.

WORKED EXAMPLE

See the idea in action

Five test scores: 4, 7, 7, 8, 9. Mean = 35 ÷ 5 = 7. Median = 7 (third value). Mode = 7. Range = 9 − 4 = 5. If a sixth score of 10 is added, the mean becomes 45 ÷ 6 = 7.5, the median becomes 7.5, and the mode is still 7.

Exam technique

Turn knowledge into marks

When comparing two distributions, write one sentence about average and one about spread, using the actual calculated figures. “Class A has a higher mean (62 vs 55) but a larger IQR (18 vs 10)” scores; “Class A did better” does not.

Common mistakes

Do not give these marks away

  1. 01

    Dividing Σfx by the number of rows in the table instead of by the total frequency.

  2. 02

    Calling the midpoint of the modal class the mode, or giving a single number as the exact mean of grouped data.

  3. 03

    Forgetting to order the list before claiming a median.

QUICK RETRIEVAL

The mean of 4, 7, 7, 8 and 9 is

A7

B7.5

C8

D5

Show the answer

7. The total is 35 and there are 5 values, so the mean is 7. 7 is also the median and the mode in this particular list, which is coincidence, not a rule.

Quick questions

If this is the bit you searched

Which average should I use in a comment question?

The mean uses all the data but is pulled by outliers. The median is safer when a single extreme value would distort the mean. The mode is useful for categories, such as shoe size.

Why is a grouped mean only an estimate?

Because you no longer know each original value. Replacing a whole class by its midpoint is a modelling choice, so the answer must be described as an estimate.