Understanding population groups from a random sample

An SAT Math micro-topic under Evaluating statistical claims (Problem solving and data analysis). Free to read — no account needed.

Surveying an entire population is usually impossible, so researchers study a smaller random sample and use it to estimate what the whole group is like.
But there is a catch: a sample only tells you about the exact group it was drawn from.
To see this, suppose a survey questions 5,0005{,}000 randomly chosen male cricket fans in India about which format of cricket they prefer.
Because the sample was limited to males, to cricket fans, and to India, its results describe only that specific group.
Picture the groups as circles inside circles: the male cricket fans in India sit inside the larger circle of all cricket fans in India, which sits inside the circle of all people in India.
The survey only reaches the innermost circle, so its conclusions cannot spread outward to the bigger circles.
sample_scope.png
Walk through what this survey can and cannot claim.
It cannot describe all people in India, because women and children were never asked.
It cannot even describe all cricket fans in India, because female fans were left out.
It can speak only for male cricket fans in India — the exact group that was sampled.
So read the description of the sample carefully and note every restriction — who, where, and what group.
For example, a survey of men in the city of Mumbai is restricted by place, so it cannot represent men across all of India.
And a survey of registered voters is restricted by group, so it cannot represent every resident of the city.
Each restriction shrinks the population your conclusion is allowed to cover.
When a sample is chosen randomly from an entire group, its results can be generalized to that whole group.
For instance, 400400 voters picked at random from across a city can fairly represent all of the city's registered voters.
But a sample from just one narrow slice cannot: if a study tests only the students at a single high school, its findings apply to that one school, not to all students everywhere.
On the test, the wrong answers almost always over-generalize.
They take a result about the narrow sampled group and stretch it to a much bigger population.
For example, a study of patients who all had one specific disease cannot be used to make claims about healthy people.
The correct answer always stays inside the exact group that was sampled.

Worked examples

A survey randomly questions 5,0005{,}000 male cricket fans in India. To which group can its results be reliably applied?
(A) All people in India.
(B) Male cricket fans in India.
The sample was restricted to males, to cricket fans, and to India.
So the results apply only to (B), male cricket fans in India — not to all Indians, who were not all represented.
A student surveys a random sample of students at her own high school and finds a link between eating breakfast and higher test scores. Can she conclude this holds for all students everywhere?
(A) No — it applies only to students at her school.
(B) Yes — it applies to all students.
The sample came from a single school, so it represents only that school.
Choice (A) is correct; the result cannot be stretched to all students.
A city randomly surveys 400400 registered voters. Which conclusion is valid?
(A) It represents all people living in the city.
(B) It represents all registered voters in the city.
The sample was drawn from registered voters, not from every resident.
So (B) is valid, while (A) over-generalizes to people who were never part of the sampled group.

Is this one of the topics costing you points?

Take the free 10-question diagnostic for a predicted SAT score and a breakdown of which domains are costing you points.

Predict your SAT score →

More in Evaluating statistical claims