Practise › Questions › Hypothesis testing: correlation and the normal
Hypothesis testing: correlation and the normal questions
Two more claims meet their data. That a correlation seen in a sample is a fluke, and that a normal population still has the mean it used to. Both tests run on last lesson's logic, and only the distribution under the null changes.
8 original questions · 31 marks · the hypothesis testing: correlation and the normal notes · Statistics
Every question here is written for this library rather than taken from a past paper. Write your answer out before opening the worked one: the answers award marks point by point, and the marks are easier to see when you have something of your own to compare against.
Observations are drawn independently from N(μ, σ2). State the distribution of the mean of n such observations.
Worked answer
The sample mean is normal with the same centre and a tighter spread, so X̄ ~ N(μ, σ2/n). B1 for the normal distribution with mean μ, B1 for the variance. Averaging cancels wobble, and the variance shrinks by a factor of n, so the standard deviation of the mean is σ/√n rather than σ/n.The masses, in grams, of items from a production line are normally distributed with mean μ and standard deviation 3. A test of H0: μ = 50 against H1: μ < 50 uses a random sample of 36 items, whose mean mass is 48.8 g. Find the value of the test statistic.
Worked answer
The standard error of the sample mean is 3/√36 = 0.5, so z = (48.8 − 50)/0.5 = −2.4. M1 for the standard error, M1 for the z formula, A1 for −2.4. The denominator is the standard error, not the population standard deviation. Dividing by 3 instead of 0.5 gives −0.4 and turns a decisive result into a limp one, and that is the single most common error on this topic.The masses, in grams, of bags of flour are normally distributed with mean μ and standard deviation 4. A supplier claims that μ = 100. A random sample of 25 bags has mean mass 98.6 g. Test, at the 5% significance level, whether the mean mass is less than the supplier claims. State your hypotheses clearly.
Worked answer
H0: μ = 100 and H1: μ < 100. Under H0, X̄ ~ N(100, 42/25), so the standard error is 0.8 and z = (98.6 − 100)/0.8 = −1.75. Then P(Z ≤ −1.75) = 0.0401. Since 0.0401 < 0.05, reject H0. There is significant evidence at the 5% level that the mean mass of the bags is below 100 g. The five marks are B1 hypotheses, M1 distribution of X̄, M1 probability, A1 0.0401, A1 comparison and conclusion in context. All three parts of the conclusion are needed: the comparison, the decision, and a sentence about flour rather than about z.Bags filled by a machine have masses, in grams, distributed as N(μ, 42). A test of H0: μ = 30 against H1: μ > 30 at the 5% significance level uses the mean of a random sample of 16 bags. Find the critical region for the test, and state the conclusion when the sample mean is 31.2 g.
Worked answer
Under H0, X̄ ~ N(30, 16/16), so the standard error is 1. A one-tailed 5% test rejects when z > 1.6449, that is when X̄ > 30 + 1.6449 × 1 = 31.6449 g. The critical region is X̄ > 31.6449. The observed mean 31.2 is well below that boundary, so it is not in the critical region. Do not reject H0. There is insufficient evidence at the 5% level that the mean mass has risen. M1 for the standard error, B1 for z = 1.6449, M1 for the boundary, A1 for 31.6449, A1 for the conclusion. As a cross-check, z = 1.2 gives a tail probability of 0.115, comfortably above 0.05, and the two methods agree.
Carry the boundary at full accuracy. Rounding it to 31.6 hands the critical region every mean between 31.6 and 31.6449, and a sample mean of 31.62 would then be declared significant when its tail probability is 0.0526, above 0.05. That is where the critical-region and p-value methods appear to disagree, and the rounding is always the culprit. Round the boundary only for reporting, and only when the observed mean is nowhere near it.A random sample of 20 pairs of observations gives a product moment correlation coefficient of r = 0.62. The critical value from the table for a sample of size 20 at the 0.01 level is 0.5155. Test, at the 1% significance level, whether there is positive correlation between the two variables.
Worked answer
H0: ρ = 0 and H1: ρ > 0, where ρ is the population correlation coefficient. The test is one-tailed, so the 0.01 column of the table is the right one and the critical value is 0.5155. Since 0.62 > 0.5155, reject H0. There is significant evidence at the 1% level of positive correlation between the two variables in the population. B1 for the hypotheses in terms of ρ, B1 for the critical value, M1 for the comparison, A1 for the conclusion in context. Hypotheses written about r rather than ρ lose the first mark every time, since r is only the sample's echo of the population quantity being tested.Explain why increasing the sample size makes a fixed difference between the sample mean and the hypothesised mean more likely to be judged significant.
Worked answer
The standard error σ/√n shrinks as n grows, so the same difference is measured against a tighter yardstick and the z value rises. B1 for the standard error shrinking, B1 for the effect on z or on the tail probability. A drift of 1 gram on 4 bags is ordinary noise; on 400 bags the mean has no business straying that far, and the test says so.A student measures two variables on a random sample of 30 subjects and obtains a product moment correlation coefficient of r = 0.42. The table of critical values for a sample of size 30 gives 0.3061 at the 0.05 level, 0.3610 at the 0.025 level and 0.4226 at the 0.01 level. Test, at the 5% significance level, whether there is any correlation between the two variables, and state whether the conclusion would change at the 2% level.
Worked answer
H0: ρ = 0 and H1: ρ ≠ 0. The alternative hypothesis names no direction, so the test is two-tailed and the 5% is split between the two ends. The table is indexed by one-tail probability, so the row must be read at 0.025, giving critical values ±0.3610. Since 0.42 > 0.3610, reject H0. There is significant evidence at the 5% level of correlation between the two variables. At the 2% level the relevant column is 0.01, where the critical value is 0.4226, and 0.42 < 0.4226, so H0 would not be rejected and the conclusion changes. B1 hypotheses, B1 identifying the test as two-tailed, B1 reading the 0.025 column, A1 first conclusion, A1 second conclusion. Reaching for the 0.05 column on a two-tailed 5% test is the mistake this question is built around, and it doubles the true significance level of the test.The lengths, in centimetres, of components are distributed as N(μ, 22). A two-tailed test of H0: μ = 50 at the 5% significance level uses the mean of a random sample of 16 components. Find the critical region for the test, and state, with a reason, whether a sample mean of 51.1 cm leads to rejection of H0.
Worked answer
Under H0, X̄ ~ N(50, 4/16), so the standard error is 0.5. A two-tailed 5% test puts 2.5% in each tail, so the critical z values are ±1.96 and the boundaries are 50 ± 1.96 × 0.5, that is 49.02 and 50.98. The critical region is X̄ < 49.02 or X̄ > 50.98. Since 51.1 > 50.98 it lies inside the critical region, so reject H0. There is significant evidence at the 5% level that the mean length is not 50 cm. M1 standard error, B1 z = ±1.96, M1 forming a boundary, A1 both boundaries, A1 conclusion. Using 1.6449 here would give a critical region starting at 50.82, and 51.1 would still be rejected, so the right answer arrives for the wrong reason and the accuracy marks are lost anyway.
The same practice on paper: the printable workbook for this topic, questions and a worked answer book.
Practise hypothesis testing: correlation and the normal one question at a time
The player marks nothing for you. It shows one question, waits, then shows the worked answer so you can mark yourself, and brings a question back sooner when it went badly.