MathsStatistics › Hypothesis testing: correlation and the normal

Hypothesis testing: correlation and the normal

Two more claims meet their data. That a correlation seen in a sample is a fluke, and that a normal population still has the mean it used to. Both tests run on last lesson's logic, and only the distribution under the null changes.

Builds on Hypothesis testing with the binomial and Correlation and regression.

IN THIS TOPIC

  • Test H₀: ρ = 0 against a critical value read from the table, using the right column for the number of tails.
  • Use the fact that the mean of n observations of N(μ, σ²) is N(μ, σ²/n).
  • Carry out a hypothesis test for the mean of a normal distribution with known σ, concluding in context.

COMMON MISCONCEPTION

The sample mean of n readings has the same standard deviation as a single reading.

Is a correlation real?

A sample scatter can look correlated when the population it came from is not, so the question is whether this sample's r is too large to be explained by chance. The hypotheses are about ρ, the population correlation coefficient. Test H₀: ρ = 0 against H₁: ρ > 0, ρ < 0 or ρ ≠ 0. Writing the hypotheses in r instead of ρ is an instant lost mark.

An observed correlation of 0.62 beyond a critical value of 0.55 on the r number line from minus one to one-1-0.500.51critical value 0.55r = 0.62beyond the critical value: reject H₀
FIG. 1The paper supplies the critical value; the test is one comparison on the r number line.

The booklet's table of critical values is indexed by sample size and by one-tail probability, and choosing the column is half the skill. A one-tailed test at 5% uses the 0.05 column. A two-tailed test at 5% uses the 0.025 column, because the level has been split.

Then it is one comparison. For a one-tailed test against H₁: ρ > 0, reject when r exceeds the positive critical value. Against H₁: ρ < 0, reject when r falls below its negative. Only the two-tailed test compares the size of r and ignores its sign. With n = 20 and r = 0.48 against H₁: ρ ≠ 0 at 5%, the table gives 0.4438, and 0.48 > 0.4438, so reject H₀ and conclude there is significant evidence of correlation between the two named variables.

The sample mean's distribution

mean of n observations: N(μ,σ2/n)\text{mean of } n \text{ observations: } \, N(μ, \, σ^{2}/n)NOT IN THE BOOKLET — LEARN IT
One observation against the mean of twenty-five: the sample mean's distribution is five times narrowerone value: σ = 4mean of 25: σ/√n = 0.8
FIG. 2Averaging 25 readings divides the spread by five: sample means are far better behaved than single values.

Averages wobble less than individuals. The mean of n independent observations of N(μ, σ²) is itself normal, with the same centre and the spread divided by √n. That standard deviation σ/√n is what every normal-mean test standardises against. Divide by n instead of √n and every z-value in the question collapses.

Learn that distribution. All the booklet gives you, under Sampling distributions, is the standardised version, the sample mean minus μ over σ/√n, which is N(0, 1). The N(μ, σ²/n) form itself is printed only in the Further Mathematics section, which you may not use, so write it out from memory and then standardise.

Testing a normal mean

WORKED EXAMPLE

Has the machine drifted?

Bags from a machine have masses N(μ, 4²) grams. The machine is set to μ = 30. A sample of 25 bags has mean 31.6 g. Test at the 5% level whether the mean has increased.

H₀: μ = 30, H₁: μ > 30. Under H₀ the sample mean is N(30, 4²/25), standard deviation 4/5 = 0.8.

z = (31.6 − 30)/0.8 = 2.0, and P(Z ≥ 2.0) = 0.0228.

0.0228 < 0.05, so reject H₀. There is significant evidence at the 5% level that the mean mass has increased.

Sense check: 31.6 sits two standard errors above 30, and two-sigma events are rare enough to be worth noticing.

The critical-value route works identically. At the 5% level one-tailed, reject when z > 1.6449; two-tailed, when |z| > 1.96. Two-tailed versions split the level exactly as the binomial test did, and the conclusion vocabulary is word for word the same. The only genuinely new content in this lesson is the √n.

ASSESSMENT FOCUS

  • Correlation tests: hypotheses in ρ, comparison against the table's critical value, conclusion naming both variables.
  • Pick the table column by the number of tails. Two-tailed at 5% means the 0.025 column, and half the candidates who lose this mark never notice.
  • Normal-mean tests: write the distribution of the sample mean with σ/√n before standardising. That line carries the method marks.
  • One tail or two comes from the wording. “Changed” is two-tailed. “Increased” or “decreased” is one.
  • Finish in context every time: reject or not, at what level, about which quantity.

CHECK YOURSELF

Readings are N(μ, 9) and H₀ sets μ = 50. A sample of 36 has mean 49.2. Find the z-value for testing H₁: μ < 50.

Show a hint

The sample mean's standard deviation is σ/√n.

Show the answer

σ/√n = 3/6 = 0.5.

z = (49.2 − 50)/0.5 = −1.6.

Correlation: hypotheses in ρ, and compare the sample r against the table value for the right number of tails.

A mean of n readings lives on N(μ, σ²/n), so standardise with σ/√n and test as usual.

WORKBOOK

Printable practice for this topic: original exam-style questions with room to work, and a fully worked answer book. Free to use; please do not redistribute or sell.

8 questions on this topicAnswer them one at a time and mark yourself against the worked answer.Practise this topic

Or read them with their worked answers on the hypothesis testing: correlation and the normal questions page.

CHECK YOUR PROGRESS

Rate how confident you feel with each objective for this lesson. Ratings are saved in this browser, on this device, unless you sign in.

  • Test H₀: ρ = 0 against a critical value read from the table, using the right column for the number of tails.
  • Use the fact that the mean of n observations of N(μ, σ²) is N(μ, σ²/n).
  • Carry out a hypothesis test for the mean of a normal distribution with known σ, concluding in context.

Open the full revision checklist to see every objective in the course in one place.