16 Jun 2026 · 6 min read

The Problem With P-Values in Biomedical Research

The p-value is the most widely misunderstood statistic in science. Understanding what it actually means changes how you read the literature, design your studies, and report your findings.

The Problem With P-Values in Biomedical Research

In 2016, the American Statistical Association took the unusual step of publishing a formal statement on the use and misuse of p-values. It was the first such statement in the organisation's 177-year history. The fact that it was deemed necessary, after decades of statistical education and methodology literature, tells you something about how persistent the problem is.

The p-value is the most widely reported statistic in biomedical research and one of the most consistently misunderstood. Not by students who have not yet learned statistics, but by experienced researchers, journal editors, grant reviewers, and clinicians who use it daily and have been using it wrong for years.

This is not a pedantic concern about technical definitions. The misuse of p-values has directly contributed to the replication crisis, to false confidence in findings that did not hold up, and to clinical decisions made on weaker evidence than the statistical presentation implied.

What a p-value actually means

A p-value is the probability of observing a test statistic at least as extreme as the one observed, assuming the null hypothesis is true. That is a technically precise statement that is not particularly intuitive, which is probably the root cause of the problem.

What a p-value does not mean, despite how it is almost universally interpreted:

It is not the probability that the null hypothesis is true. It is not the probability that your result occurred by chance. It is not a measure of the strength or importance of your finding. A p-value of 0.001 does not mean your effect is three times more likely to be real than a p-value of 0.003. And the threshold of 0.05, below which findings are conventionally called "statistically significant," is an arbitrary convention with no special mathematical status.

The 0.05 threshold was proposed by Ronald Fisher in the 1920s as a rough guide to when an effect might be worth investigating further. It was not intended as a binary criterion for distinguishing real findings from non-findings. The way it has been applied in biomedical research for the past half-century is a substantial distortion of Fisher's original intention, and the consequences are measurable in the reproducibility literature.

How p-value misuse entered the literature

The journal publication system created an incentive structure that turned p < 0.05 into a gate. Papers that crossed the threshold were more likely to be published. Papers that did not were more likely to end up in file drawers. Researchers, responding rationally to these incentives, designed studies to maximise the probability of crossing the threshold rather than to generate the most informative estimate of the effect.

Practices that predictably inflate false positive rates became common: testing multiple outcomes and reporting only the ones that reached significance (outcome switching), collecting data until p < 0.05 was reached (optional stopping), reclassifying analytic decisions as pre-specified when they were made post hoc (HARKing, or Hypothesising After Results are Known). None of these practices are technically fraudulent in all cases. All of them produce a literature where the reported p-values dramatically understate the actual rate of false positives.

A 2011 simulation study estimated that in fields with typical effect sizes, measurement variability, and publication bias, the false discovery rate among published significant findings could exceed 50 percent. That means in some areas of the literature, more than half of the significant findings you are reading are false positives. The p-value printed in the results section provides no information about which half.

Effect sizes, confidence intervals, and what they add

The alternatives are well-established and have been advocated in methodology literature for decades. They are still not routinely applied.

Effect sizes, the magnitude of the observed difference rather than just its statistical significance, are the fundamental quantity of interest in most research. A statistically significant effect that is clinically trivial is a finding of limited value. A clinically meaningful effect that does not reach statistical significance in an underpowered study is a finding that deserves follow-up, not dismissal. The effect size and the p-value tell different stories, and the effect size is usually the more important one.

Confidence intervals provide information about both the magnitude of the effect and the precision of its estimation. A 95% confidence interval that spans from a clinically trivial effect to a large effect communicates genuine uncertainty in a way that a single p-value cannot. The Gardner and Altman paper in the BMJ from 1986 made this case comprehensively nearly four decades ago. The argument has not improved with age because it did not need to.

Bayesian approaches offer a different philosophical framework that directly estimates the probability that a hypothesis is true given the data, which is what researchers usually want and what p-values are persistently misread as providing. Bayesian methods are computationally more demanding and require explicit prior specification, but they are increasingly accessible and are appearing more frequently in clinical research, particularly in adaptive trial designs where frequentist methods become awkward. The EMA's ICH E9 guideline on statistical principles for clinical trials explicitly permits Bayesian designs when properly pre-specified, and the FDA's guidance on the use of Bayesian statistics in medical device trials has been a reference document in the field since 2010.

Practical implications for reading the literature

Understanding what p-values actually mean changes how you read a paper. The threshold of p < 0.05 should not trigger automatic credibility. The appropriate questions are: what is the effect size, and is it clinically meaningful? How wide is the confidence interval, and what does the upper and lower bound imply? How many outcomes were tested, and was the primary outcome pre-specified? Was the sample size calculated prospectively, and was it achieved?

A p-value of 0.049 from a large adequately powered pre-registered trial with a pre-specified primary outcome is meaningfully different from a p-value of 0.049 from a small exploratory study that tested seven outcomes and reported the most significant one. Both appear identical when you read the abstract. The difference between them is the difference between evidence and noise.

When we help research teams synthesise literature at NousLab, this kind of methodological stratification, distinguishing findings by the quality of the statistical evidence behind them, is part of the work. A headline finding that rests on a p-value from an underpowered exploratory study should be weighted differently from a replicated effect confirmed across multiple adequately powered trials. Getting that distinction right requires reading beyond the abstract and understanding what the statistical numbers actually mean. See how our platform supports systematic evidence appraisal.

Jesus Arias
Jesus Arias
Founder & CEO at NousLab
Connect on LinkedIn

Join the Research Revolution

See how NousLab can accelerate your organization's research

AI