Maths Statistics

Standard deviation calculator

Your numbers
Sample standard deviation 6.324555
7 values, mean 21
The whole summary
Population standard deviation 5.8554
Mean 21
Median 22
Mode 25
Count 7
Sum 147
Minimum 12
Maximum 30
Range 18
Sample divides by n−1

Sample standard deviation divides the summed squared deviations by n−1 rather than n. The sample mean is itself fitted to the same data, so the deviations measured from it are systematically a little too small; dividing by one fewer compensates. Use the sample figure when your numbers are a sample of something larger, which is almost always the case. Use the population figure only when the list genuinely is the entire group.

Standard deviation measures spread around the mean. The sample figure divides the summed squared deviations by n−1; the population figure divides by n. The seven default values give 6.324555 and 5.855400, a gap of 8%, and the two converge as the list grows.

How to calculate standard deviation

1 Paste or type your numbers, separated by commas, spaces or new lines.
2 Read the sample standard deviation. It is the one nearly every use calls for.
3 Compare mean against median. A wide gap between them says the spread is not symmetric.
4 Take the population figure only when the list is the complete group rather than a sample of it.

The two deviation rows differ by a factor of √(n ÷ (n−1)) and nothing else. On the seven default values that is √(7/6), so 5.8554 becomes 6.3246, a gap of 8 per cent. At twenty values the gap is 2.6 per cent and at five hundred it is a tenth of one per cent, so the choice between them stops mattering long before most people stop worrying about it.

What n−1 actually repairs

It repairs the variance, and only the variance. Squared deviations are measured from the sample mean, and the sample mean is the value that makes their sum as small as it can possibly be for this data; measuring from the true mean would give a larger total. Dividing by n−1 instead of n scales that shortfall away exactly, and the resulting variance is unbiased.

A standard deviation is not, because taking a square root does not preserve the correction. On average the sample standard deviation lands below the true one, and the size of the shortfall is a known constant of n: about 6 per cent low at n = 5, 2.7 per cent at n = 10, 1 per cent at n = 25, and a twentieth of one per cent by n = 500. Quality engineering has tabulated that constant for a century under the name c₄, because a control chart built on five-piece subgroups would otherwise sit systematically too tight. For everyday summary work the bias is smaller than the sampling noise around it and can be left alone; the point is that n−1 is a correction with a specific target, not a general safety margin.

When the number stops describing the data

A standard deviation summarises a symmetric spread in one number, so it is only as good as that symmetry. Comparing mean against median is the cheapest test available. Close together, and the deviation figure describes the data honestly. Mean well above median, and a small number of large values are dragging it: household income is the standard example, where the mean, the standard deviation and every rule of thumb built on them describe a distribution nobody in the data actually lives in. Quartiles or an interquartile range say more in that situation than any single spread figure can.

Two smaller reading notes on the panel. The mode row shows a dash when every value appears exactly once, because a data set in which nothing repeats has no mode rather than having all of them. And the range row, being built from the two most extreme values, is the least stable summary here: it can only grow as you add data, so it says as much about how many numbers you collected as about how spread out they were.

What people use it for

  • Summarising a set of measurements
  • Checking whether a process is running consistently
  • Working out how unusual one value is
  • Reading a published mean and deviation critically

Questions

The denominator. Sample divides by n−1 to compensate for measuring deviations from a mean fitted to the same data; population divides by n and assumes the list is the whole group.

NIST/SEMATECH e-Handbook of Statistical MethodsNIST/SEMATECH, What are Variables Control Charts? (the c₄ factor)
Was this tool any good?
Internal signal only · I use it to find the tools worth rebuilding