Confidence interval calculator
Margin of error scales with one over the square root of the sample size, so cutting it in half takes four times the data. That is why polls hover around a thousand respondents: it buys roughly ±3 points, and getting to ±1.5 would need four thousand at four times the cost. The relationship is also why a sample of 1,000 is nearly as good for a country of 60 million as for a town of 60,000: population size barely enters into it.
Uses the normal critical value at every sample size. Below about thirty observations, or with a standard deviation estimated from the same sample, a t-distribution gives the wider interval the data actually supports.
A confidence interval is the mean plus and minus a critical value times the standard error. A mean of 100 with a standard deviation of 15 and n = 100 gives a 95% interval of 97.06 to 102.94, a margin of error of ±2.94. This panel uses the normal critical value, 1.959964 at 95%.
How to calculate a confidence interval
Everything the panel does is visible in its own output rows. It looks up the normal critical value for the confidence level, divides the standard deviation by the square root of the sample size, and lays the product either side of the mean. At the defaults that is 1.959964 as the critical value, 1.5 as the standard error, and 2.939946 as the margin. Two things then need saying: what the resulting number claims, and where the choice of critical value stops being adequate.
The 95 per cent belongs to the procedure
Hoekstra and colleagues put six statements about a confidence interval to 120 researchers and 442 psychology students. All six were false. Both groups endorsed more than three of them on average, and the researchers barely outperformed students who had had no training in statistical inference at all. Whatever else this misreading is, it is not a beginner error.
So, plainly: for the default interval of 97.06 to 102.94, there is no 95 per cent probability that the true mean lies inside it. The sample has been drawn and the true mean either sits in that range or it does not; no probability remains to describe which. The 95 attaches to the recipe. Repeat the same sampling and the same arithmetic many times over and about 95 intervals in every hundred will contain the true value. The one on screen is a single draw from that collection and carries no record of which kind of draw it was.
One consequence is practical: an interval cannot be read as a probability distribution over the answer, with the middle more likely than the edges. Two values inside a 95 per cent interval have no ranking against each other that the interval itself supplies, and a value just outside it is not thereby ruled out.
Where the normal critical value stops being enough
The population standard deviation is almost never known, so in practice people feed a sample standard deviation into a formula written for a population one. Student’s t exists for that substitution. It widens the interval to pay for the uncertainty in an estimated spread, and how far it widens depends on the degrees of freedom.
At n = 100 the difference is negligible. The t critical value on 99 degrees of freedom is 1.9842 against the 1.9600 used here, so the interval on screen is about 1 per cent too narrow. Shrink the sample and the gap opens fast: at n = 30 the interval is 4 per cent narrow, at n = 10 it is 13 per cent narrow, and at n = 5 the panel returns ±13.15 where a t-interval on the same numbers gives ±18.62. That last one is not a rounding difference. It is close to a third of the interval.
Around thirty observations the correction fades below the noise in everything else, and the familiar folk rule about n = 30 is a rough statement of exactly that. The rule is about the correction becoming small, never about the t-distribution becoming wrong.
Looking first changes what the interval covers
Coverage of 95 per cent assumes the sampling plan was fixed before any data arrived. Collect some responses, compute the interval, and keep collecting until it stops straddling a value you would rather exclude, and the procedure being repeated is no longer the one the coverage was calculated for. The intervals produced by a stop-when-it-looks-right rule contain the true value far less than 95 per cent of the time, and with enough patience the rule terminates almost surely whatever the truth is.
Multiplicity leaks the same way. Twenty independent 95 per cent intervals have roughly a 64 per cent chance that at least one of them misses, since 0.95 to the twentieth power is 0.358. Reporting only the interval that came out interesting is a selection, and the confidence level printed beside it describes a procedure nobody actually ran.
What people use it for
- Reporting a survey result with a margin of error
- Judging how precise an estimate really is
- Deciding whether a sample was large enough to say anything
- Checking a published interval against the sample size behind it
- Deciding whether two published estimates really differ
Questions
That 95% of intervals built this way from repeated samples would contain the true value. It describes the procedure, not the particular interval in front of you.
No. The true value is fixed and the sample is drawn, so it is either inside or outside. The probability lives in the sampling, which has already happened.
The normal one, at every sample size: 1.6449 at 90%, 1.9600 at 95%, 2.5758 at 99%. The value is shown as its own output row.
Whenever the standard deviation came from the same sample, which is nearly always, and the sample is small. At n = 100 the difference is about 1%; at n = 10 it is 13%; at n = 5 it is 29%.
Collect more data. The margin falls with the square root of n, so halving it takes four times the sample.
Barely, unless the sample is a large fraction of it. Sampling error depends on how many you asked, not on how many you could have asked.
The standard deviation divided by the square root of the sample size. It describes the spread of the sample mean across repeated samples, not the spread of the data.
You can, and the result will not be a 95% interval. Coverage assumes the stopping rule was fixed before the data was seen.
Not reliably. Overlapping 95% intervals can still correspond to a significant difference between the two means; the comparison needs an interval built on the difference itself.