MathsFurther Statistics 2 › Comparing two normal means

Comparing two normal means

Two samples, two means, and a question about whether the populations behind them differ. The difference of the sample means is itself normal, and everything follows from its standard error.

Builds on Combinations of normal random variables and Estimators, standard error and confidence intervals.

IN THIS TOPIC

  • Find the standard error of a difference of sample means.
  • Carry out a two-sample z test and state the conclusion in context.
  • Build a confidence interval for the difference and link it to the test.

COMMON MISCONCEPTION

To compare two sample means, subtract their standard errors to get the standard error of the difference.

The difference has its own distribution

Each sample mean is normal about its population mean, with variance σ²/n. The two samples are independent, so the difference of the means is normal as well, and the variances add:

(difference of means)-(μx-μy)σx2nx+σy2ny is N(0,1)\frac{(\text{difference of means}) - (μ_{x} - μ_{y})}{\sqrt{\frac{σ_{x}^{2}}{n_{x}} + \frac{σ_{y}^{2}}{n_{y}}}} \text{ is N}(0, 1)IN THE FORMULAE BOOKLET

It is printed in the booklet under Sampling distributions, among the tests for the mean when σ is known, so read it off and substitute. Hypotheses are about the population means, so write H₀: μx = μy. Under that null the second bracket in the numerator vanishes and you are left with a plain z score. Standard errors combine by adding their squares, never by subtracting. Subtraction would make the spread of a difference smaller than the spread of a single mean, which is the wrong way round.

The standard error of a difference of means built from both samples: 2.5 apart is 2.16 standard errorssample 152.3, 25/40sample 249.8, 36/50the variances of the two means addz = 2.5 / 1.160 = 2.16
FIG. 1The two samples each contributing a variance, added under the square root to give the standard error of the difference.

WORKED EXAMPLE

A two-sample test

Sample 1: n = 40, mean 52.3, from a population with variance 25. Sample 2: n = 50, mean 49.8, variance 36. Test at 5% whether the population means differ.

H0: μx = μy; H1: μx ≠ μy. Two-tailed at 5%, so the critical values are ±1.96.

Standard error = √(25/40 + 36/50) = √1.345 = 1.160.

z = 2.5/1.160 = 2.16. Since 2.16 > 1.96, reject H0. There is evidence at the 5% level of a difference between the population means.

When the variances are unknown

Large samples rescue the method twice over. The Central Limit Theorem makes each sample mean approximately normal whatever the population shape, and the sample variances s² are close enough to the population values to stand in for them. The statistic is unchanged with s² replacing σ², and the conclusion becomes approximate. Say so, because a mark usually hangs on it.

A confidence interval for the difference follows from the same standard error, and it answers the same question as the test. An interval excluding zero corresponds exactly to a two-tailed test at the matching level rejecting equality. Quoting the interval as well as the verdict is worth doing, since it says how large the difference might be where the test says only that one exists.

A 95% interval for the difference of two means, from 0.23 to 4.77: it misses zero, so the means differ0240.234.77zero is outsidean interval missing zero and a significant test say the same thing
FIG. 2The 95% interval for the difference, running from 0.23 to 4.77 and missing zero, which is the same verdict as the test.

GUIDED PRACTICE

An interval for the difference

For the samples above, find a 95% confidence interval for the difference in population means, and comment.

Show the working

The point estimate is 52.3 − 49.8 = 2.5, and the standard error is 1.160 as before.

The interval is 2.5 ± 1.96 × 1.160 = 2.5 ± 2.27, that is (0.23, 4.77).

Zero lies outside, so the two-tailed test at 5% rejects equality, agreeing with the test already carried out.

The interval is wide. The size of the difference stays poorly pinned down even though its existence is established.

ASSESSMENT FOCUS

  • Set out the standard error as a separate calculation, showing both variance-over-n terms.
  • Say whether the test is one-tailed or two-tailed before quoting a critical value.
  • State the independence of the two samples. The variance rule depends on it.
  • With unknown variances, say that large samples make the result approximate.

CHECK YOURSELF

Two independent samples of size 25 have means 80 and 74, from populations with variances 50 and 30. Find the standard error of the difference.

Show a hint

Divide each variance by its own n, add, then take the root.

Show the answer

50/25 + 30/25 = 2 + 1.2 = 3.2, so the standard error is √3.2 = 1.789 to three decimal places.

The difference of two independent sample means is normal, with the two variance-over-n terms added under the square root.

With large samples the sample variances may replace the population ones, and an interval missing zero matches a significant two-tailed test.

WORKBOOK

Printable practice for this topic: original exam-style questions with room to work, and a fully worked answer book. Free to use; please do not redistribute or sell.

6 questions on this topicAnswer them one at a time and mark yourself against the worked answer.Practise this topic

Or read them with their worked answers on the comparing two normal means questions page.

CHECK YOUR PROGRESS

Rate how confident you feel with each objective for this lesson. Ratings are saved in this browser, on this device, unless you sign in.

  • Find the standard error of a difference of sample means.
  • Carry out a two-sample z test and state the conclusion in context.
  • Build a confidence interval for the difference and link it to the test.

Open the full revision checklist to see every objective in the course in one place.