Probability and statistics · GCSE Maths

Scatter graphs and correlation

GCSE Maths scatter graphs: describe positive, negative or no correlation, draw a line of best fit by eye, and interpolate inside the data range rather than extrapolating beyond it.

UNDERSTANDRETRIEVEREMEMBER
THE MEMORY HOOK
Correlation is a trend, not a proof of cause. Draw a line of best fit through the cloud, interpolate between the smallest and largest x, and refuse to extrapolate unless the question forces it with a warning.

The important bits

What you need to know

  1. 1

    A scatter graph plots paired data (x, y). Each cross is one pair: hours revised against test score, temperature against ice-cream sales.

  2. 2

    Positive correlation: as x increases, y tends to increase. Negative: as x increases, y tends to decrease. No correlation: a shapeless cloud, no useful line.

  3. 3

    Strength is how tightly the points hug a line: strong, moderate, weak. A strong negative correlation is still a clear trend, just downhill.

  4. 4

    A line of best fit goes through the middle of the cloud, roughly as many points above as below, and should follow the trend, not join the first point to the last, and not be forced through the origin unless the context demands it.

  5. 5

    Interpolation is reading a y from the line for an x inside the range of the data. Extrapolation is reading outside that range and is unreliable because the trend may not continue.

  6. 6

    Correlation is not causation. Ice-cream sales and drowning both rise in summer; heat is a lurking variable. Do not write “x causes y” from a scatter graph alone.

  7. 7

    An outlier sits far from the cloud. It may be an error or a genuine unusual point. Draw the line using the main trend; mention the outlier if you are asked to describe the data.

  8. 8

    The line of best fit can estimate a gradient (rate) in context: extra mark per extra hour of revision, if the scale is read correctly. Units belong in the sentence.

Quotations worth analysing

Short evidence. Real method.

Line of best fit
GCSE scatter-graph construction

Drawn by eye through the main cloud, not from origin to the furthest point, and not joining every cross. Use a ruler. The estimate you read is only as good as this line.

Interpolation is more reliable than extrapolation
Standard GCSE comment on estimates from a scatter graph

Inside the data range, the trend was actually observed. Beyond it, you are guessing that the same gradient continues, which the diagram cannot show.

Correlation does not imply causation
Interpretation mark on scatter-graph comments

A strong positive correlation is a description of the plot. “Therefore x causes y” is a different claim and usually scores zero unless an experiment is described.

Go deeper

Describe the trend in two words, then a sentence of context

Look at the cloud. Uphill left to right: positive. Downhill: negative. Random: none. Then add strength: the tighter the cigar shape, the stronger. “Strong positive correlation between temperature and sales” is a full description. Do not write “positive correlation” if half the points wander with no slope. An outlier at (2, 90) when everyone else sits near a line from (2, 20) to (10, 70) should be mentioned: it does not destroy a clear positive trend, but it would drag a mean, and it might drag a carelessly drawn line. Draw the line ignoring a single wild point unless the question says to use all points. The origin: time against distance from rest might go through (0, 0); height against age for teenagers probably should not. Let the data, not a habit, decide.

Go deeper

Interpolate inside the range; say why extrapolation is weak

If hours of revision in the table run from 1 to 8, reading the score at 5 hours from the line is interpolation. Reading the score at 20 hours is extrapolation: nobody in the sample revised for 20 hours, and the trend might flatten (there are only so many marks available). The exam often asks “why this estimate might be unreliable” and wants “the value is outside the range of the data” or “correlation is not causation” depending on the question. When you read from the line, show the dashed lines on the graph: up from x, across from the line, to y. The method mark can sit on those construction lines even if the scale is misread by one square. Give y to a sensible accuracy given the scale, not to five decimal places.

Go deeper

Gradient in context, and the causation trap

If the line rises 15 marks for every 2 extra hours, the gradient is 7.5 marks per hour. That is a rate of association in this sample, not a promise that one more hour of revision will manufacture 7.5 marks for every student. Lurking variables (prior ability, sleep, the difficulty of the paper) are still there. Papers love two statements: (1) describe the correlation, (2) explain that you cannot conclude causation. Another favourite: two scatter graphs, one with a tighter cloud. The tighter one has stronger correlation; the slopes might still differ. Strength is scatter about the line; gradient is the steepness of the line. Do not call a steep line “strong” if the points are wildly spread around it. Strong means close to the line; steep means a large rate.

WORKED EXAMPLE

See the idea in action

A scatter graph of hours revised (from 1 to 6 hours) against score is a fairly tight uphill cloud. A line of best fit passes through (2, 40) and (6, 72). Step 1: Correlation: strong positive correlation between hours revised and score. Step 2: Gradient = (72 − 40) / (6 − 2) = 32/4 = 8, so about 8 extra marks per extra hour along this line. Step 3: Estimate at 5 hours (interpolation, since 5 is between 1 and 6): from 2 hours, 3 extra hours × 8 = 24, score ≈ 40 + 24 = 64. Step 4: Estimate at 12 hours would be extrapolation beyond 6 hours and is unreliable: the trend may not continue, and a score cannot exceed the paper’s maximum. Step 5: You cannot conclude that revising causes the higher scores from the graph alone; stronger students may revise more.

Exam technique

Turn knowledge into marks

Draw the line of best fit with a ruler through the main cloud. When you estimate, state whether the x-value is inside the data range. For “unreliable”, “outside the range of the data” is the sentence they want for extrapolation.

Common mistakes

Do not give these marks away

  1. 01

    Forcing the line of best fit through the origin, or joining the leftmost point to the rightmost regardless of the cloud.

  2. 02

    Extrapolating far beyond the data and treating the number as equally reliable as an interpolated value.

  3. 03

    Writing that correlation proves causation, or calling a steep but widely scattered cloud “strong correlation”.

QUICK RETRIEVAL

Hours of revision in a sample run from 1 to 8. Using the line of best fit to estimate the score after 5 hours is

Ainterpolation

Bextrapolation

Ccausation

Dno correlation

Show the answer

interpolation. 5 lies inside the data range 1 to 8, so the reading is interpolation. Extrapolation would be a value outside 1 to 8, such as 15 hours. Causation is a separate interpretation claim. No correlation would mean a line of best fit was not useful.

Quick questions

If this is the bit you searched

What is the difference between interpolation and extrapolation?

Interpolation estimates within the range of the plotted x-values. Extrapolation estimates outside that range and is less reliable because the trend may not continue.

How do you draw a line of best fit?

By eye, with a ruler, through the middle of the cloud, following the trend, with a roughly even split of points above and below. Do not join the first point to the last.

Does a strong correlation mean one variable causes the other?

No. Correlation describes the scatter graph. Causation needs a reason or an experiment. A lurking variable may be driving both x and y.

How do you describe correlation in an exam?

State the type (positive, negative, or none) and the strength (strong, moderate, weak), then put it in context: “strong negative correlation between speed and time taken”.