MathsFurther Statistics 2 › Testing a correlation coefficient

Testing a correlation coefficient

A sample correlation is almost never exactly zero, even when the population one is. Tables of critical values say how far from zero counts as evidence, and the threshold depends on how much data you have.

Builds on Correlation coefficients and Hypothesis testing: correlation.

IN THIS TOPIC

  • State hypotheses about a population correlation coefficient correctly.
  • Read a critical value from the tables and reach a conclusion in context.
  • Explain the condition the product moment test needs and why the rank test avoids it.

COMMON MISCONCEPTION

A sample correlation of 0.5 is moderately strong, so it is significant evidence of a real relationship.

Hypotheses about the population

The sample coefficient r estimates a population coefficient, written ρ for the product moment version and ρs for the rank version. The null hypothesis is always that the population value is zero. The alternative is one-tailed or two-tailed according to what the question asks. Write the hypotheses in terms of ρ and never in terms of r, because the sample value is what you measured, not what you are testing.

Whether a given r counts as evidence depends entirely on the sample size. With n = 10 a coefficient of 0.5 falls short of the 5% critical value. With n = 30 the same 0.5 is comfortably significant. On its own, r is a number stripped of the one piece of information that gives it meaning.

Testing a correlation coefficient with n = 10 at 5% in one tail: the critical value is 0.5494−1010.5494r = 0.680.68 beats 0.5494, so reject the null hypothesis
FIG. 1The critical region for n = 10 at 5% in one tail, with an observed 0.68 falling inside it.

WORKED EXAMPLE

A one-tailed test

A sample of 10 pairs gives r = 0.68. Test at the 5% level whether there is positive correlation in the population.

H0: ρ = 0; H1: ρ > 0. One-tailed, 5%, n = 10.

From the tables the critical value is 0.5494.

Since 0.68 > 0.5494 the result lies in the critical region, so reject H0. There is evidence at the 5% level of positive correlation between the two variables.

Sample size changes everything

Critical values fall steadily as n grows, because a large sample makes a chance correlation less likely. At 5% in one tail the threshold drops from about 0.73 at n = 6 to about 0.31 at n = 30. So a strong-looking coefficient from five points proves very little, and a modest one from thirty proves a good deal.

The product moment test carries a condition. Its critical values are worked out on the assumption that the pairs come from a bivariate normal population, which in practice means both variables are roughly normal and the scatter is an elliptical cloud. Formal checking is not required, but the condition should be stated. Spearman's test makes no assumption about the shape of the distributions, since it uses only the ranks, and that is a second reason to reach for it with skewed data or with a small sample containing an outlier.

Critical values for the product moment correlation coefficient at 5% in one tail, falling as the sample grows0.7293n = 60.5494n = 100.3783n = 200.3061n = 30a bigger sample convicts on weaker evidence
FIG. 2Critical values at 5% in one tail for four sample sizes, falling from 0.73 to 0.31 as n grows.

GUIDED PRACTICE

A rank test

Two judges rank eight competitors and Spearman's coefficient is 0.905. Test at the 5% level whether there is positive agreement, given a critical value of 0.6429.

Show the working

H0: ρs = 0; H1: ρs > 0. One-tailed, 5%, n = 8.

The observed 0.905 exceeds the critical 0.6429, so the result is in the critical region.

Reject H0: there is evidence at the 5% level that the two judges agree in their rankings.

Nothing had to be assumed about the shape of the underlying distributions, since only the orders were used.

ASSESSMENT FOCUS

  • Write the hypotheses using ρ or ρs, not r or rs. The sample value belongs in the comparison, not the hypothesis.
  • State the tail, the level and the sample size before quoting a critical value.
  • Conclude twice, once about the null hypothesis and once about the variables in context.
  • For the product moment test, mention the bivariate normal condition; for the rank test, say that no distributional assumption is needed.
  • Significant does not mean large. A significant r of 0.31 from 30 pairs still leaves most of the variation unexplained.

CHECK YOURSELF

A sample of 20 pairs gives r = 0.35. The 5% one-tailed critical value is 0.3783. What do you conclude?

Show a hint

Compare, then say what it means for the variables.

Show the answer

0.35 < 0.3783, so the result is not in the critical region. Do not reject H₀: there is insufficient evidence at the 5% level of positive correlation in the population.

Hypotheses are about the population coefficient ρ or ρs, tested against a critical value that depends on n, the tail and the level.

Critical values fall as n grows; the product moment test assumes a bivariate normal population, while the rank test assumes nothing about shape.

WORKBOOK

Printable practice for this topic: original exam-style questions with room to work, and a fully worked answer book. Free to use; please do not redistribute or sell.

7 questions on this topicAnswer them one at a time and mark yourself against the worked answer.Practise this topic

Or read them with their worked answers on the testing a correlation coefficient questions page.

CHECK YOUR PROGRESS

Rate how confident you feel with each objective for this lesson. Ratings are saved in this browser, on this device, unless you sign in.

  • State hypotheses about a population correlation coefficient correctly.
  • Read a critical value from the tables and reach a conclusion in context.
  • Explain the condition the product moment test needs and why the rank test avoids it.

Open the full revision checklist to see every objective in the course in one place.