P-value calculator
A p-value is the probability of seeing data at least this extreme if the null hypothesis were true. It is not the probability that the null hypothesis is true, it is not the probability your result is a fluke, and 0.049 is not meaningfully different from 0.051. The threshold of 0.05 is a convention Fisher suggested casually in 1925, and the replication crisis of the last decade is in large part a story about treating it as a bright line.
Assumes a normal (z) distribution. For small samples you need a t-distribution, which has heavier tails and gives larger p-values for the same statistic.
A p-value is the probability of a result at least this extreme under the null hypothesis. A z-score of 1.96 gives a two-tailed p of 0.05, which is why 1.96 is the critical value everyone quotes for a 95% confidence level.
How to find a p-value
Choosing one-tailed after seeing which direction the data went is a well-documented way to halve a p-value without doing any additional work, and it is not legitimate. The tail choice belongs to the hypothesis, decided before collection: a one-tailed test says you would treat an effect in the opposite direction as no more interesting than no effect at all, which is rarely true. When in doubt, two-tailed is the conservative and defensible choice.
The famous 1.96 is not exact
Enter 1.96 and the panel returns 0.04999565, not 0.05. The critical value that gives exactly 0.05 two-tailed is 1.959964, and 1.96 is the rounded version everyone quotes. The gap is far too small to change a decision and it is a useful reminder of where the round numbers in a statistics course come from: they were chosen for a printed table, not derived from anything.
Everything here comes from the standard normal distribution, so the p-value is only correct when the test statistic really is a z. That holds for a large sample with a known population standard deviation. Substitute a standard deviation estimated from a small sample and the correct reference is Student’s t, whose heavier tails give a larger p-value for the same statistic; the difference runs the same way as it does for a confidence interval, negligible past thirty observations and substantial below ten.
What people use it for
- Interpreting a test statistic
- Checking a reported significance claim
- Converting between one and two-tailed p-values
- Statistics coursework
Questions
Conventionally below 0.05, but that threshold is a convention rather than a law. Report the actual value.
One-tailed tests a direction, two-tailed tests for any difference. Two-tailed p is double the one-tailed p.
No. It gives the probability of data this extreme if the null were true. A different and often confused thing.
1.96 two-tailed, or 1.645 one-tailed. More precisely 1.959964 and 1.644854; the familiar figures are rounded.
With small samples or an unknown population standard deviation. The t-distribution has heavier tails and is more conservative.
No. Nothing about the evidence changes across the threshold, and treating the boundary as a decision point is most of what the replication debate has been about.
Not directly. A study with p = 0.01 and one with p = 0.04 may be reporting effects of very different sizes, and the difference between significant and non-significant is not itself significant.