Standard deviation calculator
Sample standard deviation divides the summed squared deviations by n−1 rather than n. The sample mean is itself fitted to the same data, so the deviations measured from it are systematically a little too small; dividing by one fewer compensates. Use the sample figure when your numbers are a sample of something larger, which is almost always the case. Use the population figure only when the list genuinely is the entire group.
Standard deviation measures spread around the mean. The sample figure divides the summed squared deviations by n−1; the population figure divides by n. The seven default values give 6.324555 and 5.855400, a gap of 8%, and the two converge as the list grows.
How to calculate standard deviation
The two deviation rows differ by a factor of √(n ÷ (n−1)) and nothing else. On the seven default values that is √(7/6), so 5.8554 becomes 6.3246, a gap of 8 per cent. At twenty values the gap is 2.6 per cent and at five hundred it is a tenth of one per cent, so the choice between them stops mattering long before most people stop worrying about it.
What n−1 actually repairs
It repairs the variance, and only the variance. Squared deviations are measured from the sample mean, and the sample mean is the value that makes their sum as small as it can possibly be for this data; measuring from the true mean would give a larger total. Dividing by n−1 instead of n scales that shortfall away exactly, and the resulting variance is unbiased.
A standard deviation is not, because taking a square root does not preserve the correction. On average the sample standard deviation lands below the true one, and the size of the shortfall is a known constant of n: about 6 per cent low at n = 5, 2.7 per cent at n = 10, 1 per cent at n = 25, and a twentieth of one per cent by n = 500. Quality engineering has tabulated that constant for a century under the name c₄, because a control chart built on five-piece subgroups would otherwise sit systematically too tight. For everyday summary work the bias is smaller than the sampling noise around it and can be left alone; the point is that n−1 is a correction with a specific target, not a general safety margin.
When the number stops describing the data
A standard deviation summarises a symmetric spread in one number, so it is only as good as that symmetry. Comparing mean against median is the cheapest test available. Close together, and the deviation figure describes the data honestly. Mean well above median, and a small number of large values are dragging it: household income is the standard example, where the mean, the standard deviation and every rule of thumb built on them describe a distribution nobody in the data actually lives in. Quartiles or an interquartile range say more in that situation than any single spread figure can.
Two smaller reading notes on the panel. The mode row shows a dash when every value appears exactly once, because a data set in which nothing repeats has no mode rather than having all of them. And the range row, being built from the two most extreme values, is the least stable summary here: it can only grow as you add data, so it says as much about how many numbers you collected as about how spread out they were.
What people use it for
- Summarising a set of measurements
- Checking whether a process is running consistently
- Working out how unusual one value is
- Reading a published mean and deviation critically
Questions
The denominator. Sample divides by n−1 to compensate for measuring deviations from a mean fitted to the same data; population divides by n and assumes the list is the whole group.
Sample, almost always. Having every member of a population is rare enough that the population figure is usually the wrong default.
By a factor of √(n ÷ (n−1)): 12% at n = 5, 2.6% at n = 20, 0.25% at n = 200 and nothing you would report beyond that.
The typical distance of a value from the mean. In roughly normal data about 68% of values sit within one deviation and 95% within two; in data that is not normal, neither figure holds.
The standard deviation squared. Variance is what the mathematics adds up cleanly; the deviation is what gets reported, because it is back in the original units.
Skew. A few extreme values move the mean and barely touch the median, so the size of the gap is a quick read on how lopsided the data is.
The variance is. The standard deviation is not, and it runs slightly low: around 6% at n = 5 and under 1% past n = 25.