Maths › Statistics › Sampling and the large data set
Sampling and the large data set
You can rarely measure everyone, so you measure some and reason about the rest. This lesson compares the sampling methods in terms of bias, representativeness, practicality and the information each one needs before you can use it, and sets out what the exam's own large data set actually holds.
IN THIS TOPIC
- Define population, census, sample and sampling frame, and say when a census is impractical.
- Describe simple random, systematic, stratified, quota and opportunity sampling, and give an advantage and a limitation of each in a context.
- Criticise someone else's sampling method by naming the group it under-represents.
- Know the shape of the large data set: eight stations, two years, and the columns that are not numbers.
COMMON MISCONCEPTION
A larger sample automatically makes the results representative.
Why sample at all
A population is every member of the group you care about, and a census measures all of them. Sometimes that is easy. One teacher can poll one class before the bell goes. Scale the population up and a census turns slow and expensive, and for some measurements it defeats itself, because the measuring destroys the thing measured. Nobody crash-tests every car.
So you measure a sample, a subset measured fully and used to draw conclusions about the rest. Everything else in this unit is that argument running from subset back to population, and it is only as strong as the method that chose the subset. You choose from a sampling frame, a list naming every member. When the frame leaves people out, the bias has already happened before a single number is collected.
One warning before the methods. Two properly taken samples from one population will differ, and neither of them is wrong. Different samples support different conclusions. An answer that treats a sample mean as if it were the population mean has misunderstood what a sample is.
The three random methods
A simple random sample of size n gives every possible group of n an equal chance. Number the frame from 1 to N, then let a random number generator or a lottery draw pick. Fairness follows by construction, and every other method is judged against this one. The drawback is practical. You need a complete frame, and the names that come out may be scattered over half a county.
Systematic sampling takes every kth member of the frame after a random start. For 50 from 1000 the gap is k = 20, and the start is a random number from 1 to 20. Fast on a long list, and still random, unless the list carries a repeating pattern that lines up with k. Sample a four-week rota every 28th name and you collect the same day of the week for ever.
Stratified sampling splits the population into groups that matter, called strata, and takes a simple random sample from inside each one. The size of each piece is
so a school of 240 sampled 30 strong takes one student in eight from every stratum. Round sensibly and check the pieces still total the sample size. The sample then reproduces the population's structure automatically. The requirement is that you know the strata and their sizes before you start.
WORKED EXAMPLE
A stratified sample, both halves
A sixth form has 600 Year 12 and 400 Year 13 students. Explain how to take a stratified sample of 40 by year group.
Fractions first. Year 12 supplies 40 × 600/1000 = 24 students and Year 13 supplies 40 × 400/1000 = 16.
Then number each year group's names and draw a simple random sample of the right size from each, using random numbers.
Both halves earn marks. Candidates who do the arithmetic and stop there lose the second one every time.
When the randomness goes
Quota sampling keeps the proportions and throws away the randomness. An interviewer is told to find so many from each group and fills the quotas with whoever comes to hand. No frame needed, quick, cheap, and it is what most street surveys really are. The weakness is in the interviewer's choices, which can introduce a bias that no later arithmetic will remove.
Opportunity sampling, also called convenience sampling, takes whoever is available. The first twenty through the door. It is the quickest of the five and the weakest, because being available is rarely unrelated to what you are measuring. Ask about exercise habits at a gym entrance and your sample answers a different question from the one you asked.
The large data set
Edexcel's large data set is real Met Office weather data, and Paper 3 assumes you have handled it yourself. Reading about it is not the same thing. It holds daily records for May to October in two years, 1987 and 2015.
The variables are daily mean temperature, total rainfall, total sunshine, mean wind speed and direction, maximum gust, mean cloud cover, mean visibility, mean pressure and relative humidity. The 1987 records carry fewer of these than the 2015 ones, which is itself the answer to more than one past-paper question.
Five stations sit in the UK, at Camborne, Heathrow, Hurn, Leeming and Leuchars. Three sit overseas, at Beijing, Jacksonville and Perth. Perth is in the southern hemisphere, so its May-to-October window is winter while Heathrow's is summer. Examiners enjoy that one enormously.
Not every column is a number. Wind speed is in knots, with Edexcel's Beaufort conversion turning bands of it into words like calm and moderate, and wind direction arrives as a compass point. Rainfall below 0.05 mm is written tr, for trace, and gaps are marked n/a. Decide what to do with those before you calculate anything.
ASSESSMENT FOCUS
- Name the method the description matches, then give one advantage and one limitation about this sample and this population. Two memorised sentences with no context score close to nothing.
- Stratified arithmetic is worth two marks. Stratum over population, times sample size, and then random selection inside each stratum.
- Learn the large data set's frame: eight stations, two years, May to October, Perth's seasons the wrong way round.
- “Random” is technical here. Quota and opportunity samples are not random, however carefully somebody set the proportions.
- Criticising a method means naming who gets left out. “The sample is too small” is the answer that examiners see most and reward least.
CHECK YOURSELF
A factory's 1200 workers are 900 full-time and 300 part-time. Describe how to take a sample of 60, stratified by contract type.
Show a hint
Fractions of the population first, then random selection inside each stratum.
Show the answer
Full-time supplies 60 × 900/1200 = 45 workers; part-time supplies 60 × 300/1200 = 15.
Number each group's members and draw a simple random sample of the required size from each, for example with a random number generator.
The two counts total 60, which is the arithmetic check worth doing before you move on.
A sample represents its population only if chance, not convenience, chose it.
Stratify when you know the structure. The same fraction from every stratum keeps the sample's proportions matched to the population's.
WORKBOOK
Printable practice for this topic: original exam-style questions with room to work, and a fully worked answer book. Free to use; please do not redistribute or sell.
Or read them with their worked answers on the sampling and the large data set questions page.
CHECK YOUR PROGRESS
Rate how confident you feel with each objective for this lesson. Ratings are saved in this browser, on this device, unless you sign in.
- Define population, census, sample and sampling frame, and say when a census is impractical.
- Describe simple random, systematic, stratified, quota and opportunity sampling, and give an advantage and a limitation of each in a context.
- Criticise someone else's sampling method by naming the group it under-represents.
- Know the shape of the large data set: eight stations, two years, and the columns that are not numbers.
Open the full revision checklist to see every objective in the course in one place.