Maths Statistics

Sample size calculator

Confidence level
%
Margin of error
%
Expected proportion
%
50% is the safest assumption
Population size
Leave 0 for a very large population
Sample size needed 385
z² × p(1−p) ÷ margin²
Before population adjustment 385
Critical z value 1.96
Margin of error 5 %
Confidence 95 %
385 gives ±5% at 95% confidence

A 95% confidence level with a ±5% margin and no assumption about the answer needs 385 responses, whatever the population. 384 is the figure most often quoted, from rounding the 384.15 rather than rounding it up, and it is why so many surveys land at "about 400 respondents". Assuming 50% is deliberately the worst case: any other expected proportion needs fewer.

Normal-approximation sample size for a single proportion. It is unreliable when the expected proportion is close to 0 or 100 per cent, where the interval it is sized for can run past the ends of the scale.

Sample size for a proportion is z² × p(1−p) ÷ margin², rounded up. At 95% confidence with a ±5% margin and no prior expectation, that is 385 responses, and the population size barely changes it unless the population is small.

How to work out sample size

1 Choose a confidence level. Ninety-five per cent is the usual starting point.
2 Choose the margin of error you can defend, in percentage points on the answer.
3 Leave the expected proportion at 50% unless earlier data gives you a better guess.
4 Enter a population size only if the group really is finite and small; leave it at 0 otherwise.
5 Compare the two size rows to see how much the population correction saved you.

The formula sizes one proportion at one point in time, using the normal approximation to the binomial. Read it as the number of responses that makes a 95 per cent interval about the answer no wider than the margin you asked for, on the assumption that the responses are a random draw from the group you care about. Each of those clauses is doing work.

The expected proportion is not a cosmetic setting

Fifty per cent maximises p(1−p), so leaving it there gives the largest sample any answer could require and the number is safe whatever comes back. Move it and the requirement drops fast: 10 per cent needs 139 responses for the same ±5 points, and 1 per cent needs 16.

Sixteen is where the approximation quietly stops being arithmetic anyone should act on. An interval of 1 per cent plus or minus 5 points runs from −4 to 6, which is not a range a proportion can occupy, and with sixteen people at a one-in-a-hundred rate the expected number of positive answers is 0.16. The usual guard is to require both n × p and n × (1 − p) to be at least about 5 before trusting a normal approximation for a proportion; at the extremes here they are nowhere near it. If you are measuring something rare, size the study on the count of rare events you need to observe, not on a percentage margin.

The population correction, in real numbers

The second size row applies the finite population correction, n ÷ (1 + (n − 1) ÷ N), which matters only when your sample would be a meaningful slice of the whole group. At ±5 per cent and 95 per cent confidence the unadjusted 385 becomes 80 in a population of 100, 218 in 500, 278 in 1,000, 370 in 10,000 and 383 in 100,000. Past a few tens of thousands the correction has nothing left to do.

This is the source of the standard surprise that a national poll and a town poll need roughly the same number of people. Precision comes from how many you asked, not from what fraction of the group they were. It also carries a warning in the other direction: when the correction bites, you are sampling a large share of a small group, and the people who decline are then a large share too.

Sampling error is the smaller problem

The number here prices exactly one thing, the randomness of who happened to land in the sample. It says nothing about non-response bias, a leading question, an unrepresentative sample frame, or a panel that has learned what researchers want to hear. Ten thousand self-selected website visitors are worse evidence than four hundred properly randomised responses, and no arithmetic on this page will reveal that. Treat it as the floor below which the sampling error alone makes the result useless, rather than as the size at which the result becomes trustworthy.

What people use it for

  • Planning a survey to a stated margin of error
  • Sizing one cell of an A/B test
  • Checking whether a published poll had the sample to say what it said
  • Deciding when to stop collecting responses
  • Costing fieldwork against the precision it buys

Questions

385 for ±5% at 95% confidence, 1,068 for ±3%, and 9,604 for ±1%. Each is the worst case, at an expected proportion of 50%.

NIST/SEMATECH e-Handbook of Statistical Methods
Was this tool any good?
Internal signal only · I use it to find the tools worth rebuilding