MathsStatistics › Correlation and regression

Correlation and regression

A scatter graph asks whether two variables move together. The correlation coefficient scores the answer and the regression line turns it into predictions. The skill worth having is knowing what each number licenses you to say, and what it never can.

Builds on Representing and interpreting data and Straight lines.

IN THIS TOPIC

  • Describe correlation from a scatter diagram and interpret a product moment correlation coefficient r between −1 and 1.
  • Use a regression line y = a + bx to make predictions, interpreting a and b in context.
  • Say when a prediction is trustworthy, and name the lurking variable when correlation is being mistaken for cause.
  • Take logs to straighten y = axn and y = kbx, then read the constants off the regression line.

COMMON MISCONCEPTION

A correlation of r = 0.9 proves that changing one variable causes the other to change.

Reading a scatter graph

Plot one variable against another and the cloud of points tells a story. Rising together is positive correlation. One falling as the other rises is negative. A shapeless blob is no correlation at all. The product moment correlation coefficient r puts a number on it, always between −1 and 1. Sign gives direction, size gives strength, and the two extremes mean the points lie exactly on a straight line.

Your calculator produces r from the data and the exam asks you to interpret it. “r = 0.96 shows strong positive linear correlation, so taller plants do tend to carry more fruit.” Strength, direction, and the actual context. Every time.

Look at the picture before you trust the number. A scatter diagram sometimes shows two separate clusters, boys and girls, or summer and winter, and a single r calculated across both describes neither of them. Saying so is a mark.

The regression line

A scatter graph of eight points with the least squares regression line y equals 2 plus 0.5 x passing through the mean pointmean pointy = 2 + 0.5x
FIG. 1The least squares line minimises the total squared vertical miss, and it always passes through the mean point of the data.

The least squares regression line y = a + bx is the straight line making the sum of the squared vertical distances from the points as small as possible. It always passes through the mean point of the data. The two letters earn marks when they are read in context. Here b is the change in y for each extra unit of x, and a is the predicted y when x is zero, which may or may not describe a real situation.

WORKED EXAMPLE

Interpreting a and b

For ice cream sales s (pounds) against temperature t (°C), a calculator gives s = 42 + 15.3t. Interpret the numbers.

The gradient: each extra degree is associated with about £15.30 more in daily sales.

The intercept: £42 of sales predicted at 0 °C. Plausible as a baseline, though 0 °C may sit outside the temperatures collected.

The wording earns the marks. Write “associated with”, never “causes”, and always in the units of the problem.

What a prediction is worth

A regression line drawn solid across the range of the data and dashed beyond it, where predictions become extrapolationdata lives hereguesswork
FIG. 2Inside the data's range the line is evidence; beyond it, the dashes are a warning that nothing supports the model out there.

Predicting inside the range of observed x values is interpolation and is usually safe. Predicting outside it is extrapolation. The line will still produce a number, but no data supports the model out there, so the correct comment is that the prediction is unreliable.

Say what interpolation justifies, and no more. It removes the risk of relying on a pattern nobody observed, because there is data across that stretch of the line. It does not certify the answer. A weak r, a visible curve, wide scatter or shaky measurement all spoil a prediction at an x sitting comfortably in the middle of the data.

Direction matters too. The line of y on x is built to predict y from x, with x the explanatory variable under some control and y the response. Running it backwards to fish out x from y is invalid. And however large r is, correlation does not establish cause. Ice cream sales and drowning rates climb together because summer drives both of them.

Straightening a curve with logs

Not every relationship is linear, and two curved ones straighten under logarithms. For y = axn, taking logs of both sides gives log y = log a + n log x, so plotting log y against log x should give a straight line of gradient n and intercept log a. For y = kbx, log y = log k + x log b, so log y against x is the straight one.

Fit the regression line to the transformed data, then convert back. A line log y = 0.42 + 1.6 log x gives n = 1.6 and a = 100.42 ≈ 2.63, so y ≈ 2.63x1.6. The commonest error is converting the gradient when only the intercept needed converting, so write down which axis was logged before you touch the powers of ten.

ASSESSMENT FOCUS

  • Interpret r with strength, direction and context in one sentence. Repeating the number back adds nothing.
  • Read b as “for each extra unit of x, y changes by b” in the question's own units, and read a as the prediction at x = 0, then say whether that is meaningful here.
  • Predict only for x inside the data range, and only with the line of y on x. Anything else earns the word “unreliable” and a reason.
  • “Correlation does not imply causation” needs a third factor named in context before it scores.
  • For log models, state which variable you logged. Marks are lost in the conversion back, not in the regression.

CHECK YOURSELF

A regression line for crop yield y (tonnes) against rainfall x (cm), from data with 20 ≤ x ≤ 60, is y = 1.2 + 0.08x. A farmer asks for the predicted yield when x = 100. What do you say?

Show a hint

Where does 100 sit relative to the data?

Show the answer

The line gives 1.2 + 0.08 × 100 = 9.2 tonnes, but x = 100 is far outside the observed range 20 to 60.

That is extrapolation, so the prediction is unreliable and should not be trusted. The linear pattern need not continue out there.

r scores the direction and strength of a linear link, never cause.

Regress y on x, predict only inside the data, and read a and b in the question's own units.

WORKBOOK

Printable practice for this topic: original exam-style questions with room to work, and a fully worked answer book. Free to use; please do not redistribute or sell.

7 questions on this topicAnswer them one at a time and mark yourself against the worked answer.Practise this topic

Or read them with their worked answers on the correlation and regression questions page.

CHECK YOUR PROGRESS

Rate how confident you feel with each objective for this lesson. Ratings are saved in this browser, on this device, unless you sign in.

  • Describe correlation from a scatter diagram and interpret a product moment correlation coefficient r between −1 and 1.
  • Use a regression line y = a + bx to make predictions, interpreting a and b in context.
  • Say when a prediction is trustworthy, and name the lurking variable when correlation is being mistaken for cause.
  • Take logs to straighten y = axn and y = kbx, then read the constants off the regression line.

Open the full revision checklist to see every objective in the course in one place.