Why Most Biomedical Studies Fail to Replicate (And What to Do About It)
The replication crisis is not really a crisis of fraud. It is a crisis of incomplete evidence at the design stage, and that is a problem we can actually fix.
In 2011, researchers at Bayer HealthCare published a paper that caused considerable discomfort across the industry. They had attempted to reproduce findings from 67 published oncology and cardiovascular studies as part of their internal target validation process. They could fully confirm the results in fewer than 25 percent of cases. A similar exercise by Amgen a year later, across 53 landmark cancer studies, reached comparable numbers.
These were not obscure papers. Many were high-profile, heavily cited studies that had shaped research programmes and, in some cases, informed clinical practice. And the majority of them did not hold up when someone tried to repeat them under controlled conditions.
The replication problem has been discussed extensively since then, usually framed as a crisis of scientific culture, publication pressure, small sample sizes, and p-hacking. All of that is real. But there is a structural cause that gets less attention: most studies are designed on an incomplete picture of the existing evidence. The failures show up at the replication stage, but they are baked in at the design stage.
What replication failure actually looks like in practice
Not all failed replications are the same. It is worth being precise about what is failing and why, because the causes suggest different solutions.
The most common form is what methodologists call conceptual non-replication: a new study tests what appears to be the same hypothesis, in what appears to be a comparable population, and gets a substantially different result. The findings are not fabricated. The methods are defensible. The results simply diverge.
This happens for a cluster of reasons. Sample sizes in original studies are frequently too small to detect effects reliably, meaning published positive findings are often the result of chance variation. Effect sizes tend to be inflated in initial studies because only surprising results get published, what statisticians call the winner's curse. And the populations studied are often more restricted than the papers imply, meaning findings that held in a narrowly defined sample get generalised far beyond where the evidence supports.
A 2019 analysis in PLOS Biology estimated that the median statistical power of published biomedical studies, the probability that a study with a true effect would detect it, was around 23 percent. That means roughly three out of four published positive findings in under-powered literature may be false positives. You would not design an engineering system with those odds. But that is the evidence base clinical researchers are building on.
The design problem nobody talks about
Here is what we kept hearing when we spoke to researchers during the early stages of building NousLab: the studies they were designing were built on literature reviews that were thorough by any reasonable standard, but still incomplete in ways the researchers could not easily identify.
A researcher developing a hypothesis about a drug mechanism might search PubMed systematically, read the relevant reviews, follow the key citations, and end up with a coherent evidence base. What they would often miss were the studies that used different terminology, the findings from adjacent specialties that had relevant but non-obvious implications, and the grey literature, conference abstracts, dissertations, internal reports, that never made it into the indexed databases but contained negative results that would have changed the picture.
The gap between "the literature I found" and "the literature that exists" is where a lot of replication failures originate. A researcher who designs a study without knowing about three existing negative results is not being careless. They are working within the limits of the tools available to them. But those limits have consequences.
The publication bias amplifier
Publication bias compounds the problem in a specific way. Positive findings get published at higher rates than null results. This means the indexed literature is not a representative sample of what has been tested, it is a curated collection tilted toward results that confirmed hypotheses.
When a researcher draws on that literature to design a new study, they are drawing on a biased sample without knowing it. The effect estimates they use to calculate sample sizes are inflated. The mechanistic rationale they build on has been filtered through a publishing process that systematically underrepresents contrary evidence. The new study is designed to confirm something the literature only appears to support.
The EQUATOR Network has been working on this problem from the reporting side for years, and their guidelines for pre-registration and reporting of null results are an important partial solution. But reporting guidelines do not change what is already in the literature. They only affect what gets added going forward.
What a more complete evidence picture changes
We have seen this directly with researchers using NousLab during hypothesis development. The most consistent effect is not that they abandon their hypotheses, it is that they refine them in ways that make them more defensible.
A hypothesis that looked clean and original turns out to have a partially overlapping study in a different population. Rather than proceeding in ignorance and discovering this at peer review, the researcher can engage with that prior work explicitly, either by narrowing their question to something genuinely distinct, or by designing a replication with appropriate differences in population or methodology. Neither outcome is a failure. Both are better than a study that duplicates existing evidence without knowing it.
In one case, a team developing a clinical trial protocol for a cardiovascular intervention surfaced, through AI-assisted literature mapping, a Japanese study using a comparable compound that had stopped early due to a safety signal. The study had been published in a Japanese-language journal and had not been indexed under the English terminology the team was using. Finding it before trial registration rather than during it saved months of work and, more importantly, protected participants.
The systemic fix requires systemic change
There are things that individual researchers can do: pre-register studies on ClinicalTrials.gov or the ISRCTN registry, report null results, use more conservative sample size calculations, engage with grey literature more systematically. All of these help. None of them address the root problem, which is that the evidence landscape is too large and too fragmented for any individual to navigate completely using the tools most researchers currently have.
The infrastructure changes that would make a real difference include broader preprint adoption with formal linking to subsequent publications, mandatory deposition of negative results, and standardised data sharing that allows individual patient data meta-analyses on much larger scales than are currently feasible. The NIH's data sharing policy, which came into force in 2023 and requires data management and sharing plans for all funded research, is an important step in this direction. Progress is happening on all of these fronts, but slowly.
In the meantime, better tools for evidence synthesis and literature mapping at the design stage are the practical lever that is available right now. Not as a complete solution, but as a meaningful reduction in the information gap that sits between "the evidence I found" and "the evidence I needed."
The replication crisis will not be solved by any single intervention. But it will be reduced, study by study, by researchers who go into the design process knowing more about what already exists than their predecessors did. That is a problem the technology is genuinely starting to address. You can see how NousLab approaches evidence mapping here.