What is a p-value? A p-value is the probability of seeing results at least as extreme as the ones observed, assuming the null hypothesis is true. Put simply, it measures how surprising your data are if nothing is really going on: not the probability that the null hypothesis is true, nor the probability your finding is a fluke. Below: what a p-value means, what it cannot tell you, and how to read a p-value in a paper.
What is a p-value? The one-sentence definition
The American Statistical Association’s 2016 statement on p-values (Wasserstein and Lazar) puts it this way:
“Informally, a p-value is the probability under a specified statistical model that a statistical summary of the data (for example, the sample mean difference between two compared groups) would be equal to or more extreme than its observed value.”
The logic runs one way: assume the null hypothesis first, then ask how unusual the data look under it.
A worked mini-example
Take a trial comparing a new painkiller with the standard drug, testing the null hypothesis that the new drug is no better. The analysis returns p = 0.045.
Here is what that p-value means. If the new drug truly worked no better than the standard, results at least this favorable would still appear in about 4.5 out of every 100 identical trials, purely through random variation.
Here is what it does not mean. It does not mean there is a 4.5 percent chance the null hypothesis is true. It does not say how much better the new drug is, or whether the difference matters to patients. Those questions are answered by the effect size and the confidence interval.
What a p-value is not: four classic mistakes
- It is not the probability that the null hypothesis is true. This is the most common error. The p-value is calculated on the premise that the null hypothesis is true, so it cannot tell you about the truth of that premise (du Prel et al., 2009).
- It is not the probability your result happened “by chance.” To compute the chance that a significant result is a false positive, you would need information the p-value alone does not contain.
- It does not measure the size of the effect. A huge study can produce a tiny p-value for a difference with no clinical value. Two studies can share the same p-value with wildly different effect sizes. Always look at the point estimate and the confidence interval.
- It does not make the decision for you. The ASA warns against basing conclusions only on whether a p-value passes a threshold. Context, study design, bias, and prior evidence all matter.
The ASA’s six principles
In 2016 the ASA issued its first statement on statistical practice, built on six principles:
- P-values can indicate how incompatible the data are with a specified statistical model.
- P-values do not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone.
- Scientific conclusions should not be based only on whether a p-value passes a specific threshold.
- Proper inference requires full reporting and transparency.
- A p-value, or statistical significance, does not measure the size of an effect or the importance of a result.
- By itself, a p-value does not provide a good measure of evidence regarding a model or hypothesis.
Print them out. Tape them to your monitor. They will save you from most of the errors that land in published papers.
Why is 0.05 the magic number?
A result with p = 0.049 is not meaningfully different from one with p = 0.051. Many journals now ask authors to report exact p-values instead of hiding behind the cutoff, and to interpret each p-value alongside the study design and effect size.
How to read a p-value in a paper: five steps
- Find the null hypothesis. What claim is the paper testing? No difference between groups? No association?
- Note the exact p-value. Prefer exact values over “p < 0.05,” which throws away information.
- Check the effect size and confidence interval. Ask how large the difference is and how precisely it is estimated.
- Check the sample size. Very large samples make even trivial differences “significant.”
- Ask whether the effect is clinically meaningful. Statistical and clinical significance are different things. A 1 mmHg drop in blood pressure can be statistically significant and medically irrelevant.
Frequently asked questions
Is a smaller p-value always better?
Not necessarily. A tiny p-value with a tiny effect size means the result is unlikely under the null, but the finding itself may not matter. Pair every p-value with an effect size.
Can a p-value be zero?
No. Software printing 0.000 just means the p-value is smaller than the display precision. Report it as p < 0.001.
What if p is greater than 0.05?
The data are not incompatible enough with the null to reject it at that threshold. A p-value above the threshold does not prove the null is true.
Sources
- Wasserstein RL, Lazar NA. The ASA’s statement on p-values. The American Statistician. 2016. Full text via PMC.
- du Prel JB et al. Confidence interval or p-value? Dtsch Arztebl Int. 2009. Full text via PMC.
- Andrade C. The P value and statistical significance: misunderstandings, explanations, challenges, and alternatives. Indian J Psychol Med. 2019. Full text via PMC.
- Tedeschi RG. P-values: a chronic conundrum. BMC Med Res Methodol. 2020. Full text.
This article explains statistics for educational purposes. It is not medical advice. Clinical decisions should follow established guidelines and professional judgment.
For more plain-English statistics guides, visit the Doctor With Data homepage.
