Practise › Questions › Representing and interpreting data
Representing and interpreting data questions
A good chart is an argument you can see. Histograms make area mean frequency, box plots put five numbers on display, and a fence built from the quartiles decides, without sentiment, which values are outliers.
7 original questions · 21 marks · the representing and interpreting data notes · Statistics
Every question here is written for this library rather than taken from a past paper. Write your answer out before opening the worked one: the answers award marks point by point, and the marks are easier to see when you have something of your own to compare against.
In a histogram, the class 20 ≤ x < 25 has frequency 30. Find its frequency density.
Worked answer
Density = frequency/width = 30/5 = 6. M1 for frequency divided by width, A1 for 6. In a histogram the area of a bar carries the frequency, so its height has to be frequency per unit of width. That is why bars over classes of different widths can only be compared by area. A wide class spreads its frequency thinly and draws a low bar.Another class, 10 ≤ x < 30, is drawn with frequency density 1.5. Find its frequency.
Worked answer
Frequency = density × width = 1.5 × 20 = 30. M1 for density times width, A1 for 30. A low wide bar can hold exactly as much frequency as a tall narrow one. The height on its own tells you only how concentrated the data are.In a histogram, the class 0 ≤ x < 20 has frequency 40. The bar for 20 ≤ x < 25 is drawn 1.6 times as tall. Find the frequency of the second class.
Worked answer
The first bar's density is 40/20 = 2, so the second bar's density is 2 × 1.6 = 3.2 and its frequency is 3.2 × 5 = 16. M1 for the first bar's density, M1 for scaling it by 1.6, A1 for 16. Heights compare densities and never frequencies. The second bar is the taller of the two yet holds well under half the frequency, because it is only a quarter as wide. Answering 64 by scaling the frequency directly is the trap.A data set has Q1 = 20 and Q3 = 36. Using the 1.5 × IQR rule, determine whether the values 62 and 58 are outliers.
Worked answer
IQR = 36 − 20 = 16, so 1.5 × IQR = 24 and the fences sit at 20 − 24 = −4 and 36 + 24 = 60. The value 62 lies beyond the upper fence, so it is an outlier. The value 58 does not, so it is large but within bounds. M1 for 1.5 × IQR, A1 for both fences, B1 for the two verdicts. State both fences and then compare, because the fence is what justifies the verdict. The rule exists so that the same decision is reached every time, rather than by eye.On a box plot, the median line sits much closer to Q1 than to Q3. Describe the skew, and what it says about the data.
Worked answer
Positive skew. The lower half of the middle 50% is squashed and the upper half stretched, so values above the median straggle further than values below it. B1 for positive skew, B1 for what it says about the data. Most observations sit low with a long tail of high values, which is the classic shape of an income distribution.Two box plots show daily journey times for two routes. Route A: median 30 minutes, IQR 6 minutes. Route B: median 28 minutes, IQR 18 minutes. Compare the routes in context.
Worked answer
Route B is typically two minutes quicker, since its median is 28 minutes against route A's 30. Route B is far less predictable, though. Its middle half spans 18 minutes against route A's 6. A commuter who has to arrive on time would sensibly take route A and trade the two minutes for the consistency. B1 for comparing the medians, B1 for comparing the spread and B1 for the recommendation, all in minutes and all naming the routes.A histogram of 70 observations has four classes: 0 ≤ x < 10 with frequency density 1.2, 10 ≤ x < 25 with 2.4, 25 ≤ x < 40 with 1.0, and 40 ≤ x < 60 with 0.35. Estimate the number of observations below 20, and estimate the median.
Worked answer
Turn every density back into a frequency first, using frequency = density × width: 1.2 × 10 = 12, 2.4 × 15 = 36, 1.0 × 15 = 15 and 0.35 × 20 = 7. These total 70, which confirms the readings. For the count below 20, take the whole first class and the part of the second class up to 20, which is 10 of its 15 units of width: 12 + (10/15) × 36 = 12 + 24 = 36 observations. For the median, the position is 70/2 = 35. The first class carries 12, so the median lies 23 into the second class: 10 + (23/36) × 15 = 19.6. M1 for converting the densities, A1 for the four frequencies, M1 for the proportion of the second class below 20, A1 for 36 observations, M1 for interpolating for the median, A1 for 19.6. Both parts assume the data are spread evenly within each class, which is the standard assumption behind interpolation and worth stating. As a check, 36 of the 70 observations fall below 20 and that is just over half, so a median a little under 20 is exactly right.
The same practice on paper: the printable workbook for this topic, questions and a worked answer book.
Practise representing and interpreting data one question at a time
The player marks nothing for you. It shows one question, waits, then shows the worked answer so you can mark yourself, and brings a question back sooner when it went badly.