Contents
Insight
Biotech organizations rarely fail because they lack data. They often fail because they mistake one category of data for another, then make capital-allocation, partnering, or development decisions as if the evidentiary gap did not exist.
Around nine in ten drug programs entering clinical development fail to reach approval; the largest published analysis puts the Phase 1 to approval rate at 13.8%.1 The figure is familiar enough to have lost some of its force, but it carries a specific implication for anyone allocating capital. Plausible biology and clinical success are separated by several evidentiary filters, and a program can clear one while failing the next. Evidence tends to look stronger at the moment capital is committed than the following experiment proves it to be.
The familiar progression from in silico to in vitro to in vivo and ultimately clinical evidence is useful but incomplete. It leaves out a layer that sits alongside all four: propositions assembled from literature, analogies, selected figures, external precedent, and untested assumptions, then presented as though the conclusion itself had been demonstrated. We call this in deck, with "in PowerPoint" as the shorthand. A coherent deck is not a dataset.
The error is not the use of lower-level evidence. Development depends on it at each stage, and there is no alternative. The error is allowing one level of evidence to support a conclusion that only another can sustain, without acknowledging the leap.
Evidence is not a linear sequence in which each stage confirms the last. It is better understood as a set of increasingly demanding filters, each testing whether a biological proposition survives contact with a more complex system. In silico evidence asks whether a hypothesis is computationally or mechanistically plausible. In vitro evidence asks whether an intervention can produce the intended effect in a controlled experimental system. In vivo evidence asks whether that effect persists in an organism, amid exposure constraints, tissue context, compensatory biology, and immune function. Clinical evidence asks whether the intervention creates patient-relevant benefit at an acceptable risk in the intended population.
In-deck reasoning asks something else. It sits alongside the four rather than below them: the connective tissue between what was observed and what is being claimed, and where the leap usually happens.
The gap between these levels is measurable. Nelson and colleagues found that the proportion of drug mechanisms with direct human genetic support rose from 2.0% at the preclinical stage to 8.2% among mechanisms behind approved drugs.2 Minikel and colleagues, analyzing 29,476 target-indication pairs, estimated that genetically supported mechanisms carried 2.6 times the probability of success of unsupported ones.3 The two findings measure different things: Nelson describes how the prevalence of genetic support changes as programs advance, while Minikel estimates the effect of that support on the odds of success. Together they map the same terrain from opposite directions. The distance between "the biology is plausible" and "the biology holds up in patients" is where a great deal of portfolio value is destroyed, and it is where inference is most easily mistaken for validation.
The question to put to an evidence package is therefore not how much of it exists. It is what decision the evidence legitimately supports, and what the base rate is for propositions of this type surviving the next test.
| Evidence state | What it can establish | What it commonly cannot establish | Appropriate decision use |
|---|---|---|---|
| In silico | Target tractability, structural hypotheses, predicted binding, patient-stratification concepts, and PBPK or economic scenarios | Biological causality, clinically relevant potency, selectivity in native systems, safety, efficacy | Hypothesis prioritization and experiment design |
| In vitro | Direct interaction, cellular activity, pharmacology, preliminary selectivity, mechanism in defined systems | Whole-body PK/PD, biodistribution, tissue penetration, immunogenicity, organism-level toxicity, durability | Lead selection, mechanism testing, translational experiment planning |
| In vivo | Exposure-response relationships, tissue and immune context, organism-level pharmacology, preliminary tolerability, durability in a model | Human efficacy, human dose, clinical benefit, commercial product feasibility | Candidate selection, IND-enabling planning, initial clinical strategy |
| Clinical | Efficacy and safety in the population, setting, and conditions studied | Broader-population efficacy, comparative superiority, adoption, reimbursement, commercial attractiveness | Clinical, regulatory, portfolio, and transaction decisions |
Figure 1. The four empirical levels and the inference layer that spans them. Each figure sits on what it measures. Replication rates describe the reliability of published preclinical findings, not the step from in vitro to in vivo. The 13.8% is attrition within clinical development. Genetic support is tracked across the pipeline among well-studied indications; the 2.6x is an association, not a conversion rate. Sources: references 1 to 5.
The rows carry a caveat. They do not imply that in vivo evidence carries uniform weight, or that in vitro data is inherently weak. A human primary-cell system built around a disease-relevant phenotype may be more informative for a narrow mechanistic question than a poorly chosen mouse model. A transgenic or humanized in vivo system can generate false translational confidence when the critical human biology is absent. What matters is decision fitness rather than position in a hierarchy.
Computational evidence has become central to modern discovery. Structure prediction, molecular docking, machine-learning models, target prioritization, network biology, virtual screening, and PBPK modeling improve the efficiency of early research and portfolio choice. Their value lies in narrowing uncertainty at relatively low cost: identifying chemical matter worth testing, flagging potential off-target liabilities, modeling dose-exposure scenarios, or exposing the economic consequences of different target product profiles. None of that establishes that the underlying biological model is correct.
A docking score is not binding evidence. A predicted neoantigen is not an immunogenic neoantigen. A machine-learning-derived responder signature is not a clinically deployable biomarker. A PBPK projection is not a human PK dataset. In each case the computational output is a hypothesis about the world rather than an observation of it, and sophistication in the method does not change the evidentiary level of the claim it supports.
For BD and investment teams, the practical question is whether the proposition that matters has been independently challenged, and whether the sponsor can name the specific experiment that would falsify it. A team that cannot describe how its central claim could be wrong has usually not tested it.
In vitro work typically supplies the first direct evidence that an intervention can affect a defined biological system: target engagement, pathway modulation, cellular potency, cytotoxicity, vector transduction, receptor occupancy, or a preliminary selectivity window. It is indispensable, and it is also frequently presented as more decisive than it is.
The central limitation is that in vitro systems reduce biological complexity by design, which is often exactly what makes them useful for isolating a variable. The same reduction removes mechanisms that later destroy value in humans: inadequate delivery, altered tissue architecture, microenvironmental suppression, species-specific biology, active transport, target-mediated clearance, immune recognition, and treatment-emergent adaptation. Newer human-derived systems narrow this gap without closing it. A better model raises the ceiling on what can be claimed; it does not remove the obligation to say what has been established.
The appropriate use of in vitro evidence is to state the claim at the level the assay supports. "Compound X inhibits enzyme Y under defined assay conditions" is an evidence-based claim. "Compound X will provide a best-in-class therapy in patients with disease Z" is not, regardless of how many orthogonal cell-based assays support the former. The second proposition may ultimately prove correct, but the bridge between the two statements still has to be crossed, and someone has to say what crossing it will require.
In vivo evidence raises the bar because it tests an intervention amid pharmacokinetics, multicellular biology, tissue distribution, organism-level adaptation, and at least some safety-relevant context. It can convert an attractive molecular hypothesis into a development candidate, or expose that the hypothesis was too fragile to survive outside a controlled assay.
The translational record deserves attention. In the Amgen exercise described by Begley and Ellis, researchers attempted to reproduce findings from 53 landmark preclinical cancer studies; findings considered sufficiently reproducible to support the original conclusions were confirmed in six cases.4 Prinz and colleagues at Bayer reported a comparable pattern from an independent in-house effort, finding published data consistent with their own results in roughly a quarter of the projects examined.5 Neither exercise shows that preclinical research is useless or that animal models are unreliable in principle. Both show that "we have positive in vivo data" is a considerably weaker statement than it is usually treated as being in an investment or BD context.
Model choice determines what has been learned, and in vivo is not a single evidentiary category. An efficacy result in a subcutaneous xenograft may establish that a therapy can affect growth of a particular implanted tumor line under particular experimental conditions. It does not establish activity in heterogeneous patient tumors, treatment-refractory disease, metastatic compartments, an intact human immune system, or patients receiving the relevant standard of care.
The diligence question is therefore not "Do you have animal data?" It is which clinical failure mode the model interrogates, and how well that model has predicted for the indication in question.
Regulatory practice increasingly recognizes that a validated in vitro or computational method may answer a defined question more appropriately than a conventional animal study. The FDA's April 2025 roadmap set out a staged shift toward scientifically validated New Approach Methodologies in preclinical safety testing, beginning with monoclonal antibodies.6 In March 2026 the agency issued draft guidance describing a validation framework for NAMs in drug development, identifying context of use, human biological relevance, technical characterization, and fit-for-purpose as key features. The guidance also makes clear that full validation is not a prerequisite in every case, which is itself an evidentiary judgment rather than a relaxation of one.7
The operative word is validated, and the scope is narrow. The draft addresses validation principles rather than specific methods, and does not cover the use of NAMs in discovery. This is an evidentiary bridge established for a defined context of use rather than a general downgrade of organism-level evidence. The principle it embodies reinforces the framework: fitness for the decision, not position in a hierarchy, determines evidentiary weight.
Clinical evidence brings the proposition into patients, which makes it the strongest direct evidence available for questions of clinical efficacy and safety. Not every conclusion built from a clinical result is itself a clinical claim.
A response rate in a defined Phase II population does not establish comparative superiority. Benefit in a biomarker-defined subgroup does not establish equivalent benefit in a broader population. A statistically significant endpoint does not establish physician adoption. Clinical differentiation does not establish reimbursement. Reimbursement does not establish that the assumed price, treatment duration, or market share will survive contact with the market.
Evidentiary discipline therefore does not lapse when a program enters the clinic. The point at which direct evidence ends and inference begins moves downstream, and a deck can contain excellent clinical data while still making an in-deck commercial, regulatory, or competitive claim.
In deck describes assertions that derive from synthesis rather than from direct testing of the proposition by the program: published biology, competitor outcomes, market research, expert opinion, mechanistic analogy, regulatory precedent, and curated public data. None of these inputs is inherently weak. The limitation is that the combined proposition has not been tested in the relevant product, population, setting, or development context.
Consider five familiar examples.
| The evidence | The in-deck proposition |
|---|---|
| Another modality has shown clinical activity against the pathway | "The target is clinically validated." |
| Normal-tissue expression appears low in public transcriptomic datasets | "The asset should avoid on-target toxicity." |
| A construct feature improves persistence in preclinical models | "The platform should have superior clinical persistence." |
| Other products in the broad disease area have used surrogate endpoints | "The program has a credible accelerated-approval route." |
| Analogous modalities have been manufactured at commercial scale | "Manufacturing can scale." |
Each statement may be reasonable. None should be treated as established without identifying the missing bridge. This is not an argument against inference; there would be very little biotech investment, licensing, or development without it.
The problem is unpriced inference: assumptions presented with the confidence of observations.
That pattern is particularly common in three situations.
A platform may be validated while the next application built on it remains hypothetical. A delivery technology may work. A manufacturing process may be reproducible. An engineering principle may have been demonstrated. A screening system may generate credible candidates. None of that validates the disease-specific proposition sitting on top of it. Platform validation and product validation are separate claims requiring separate evidence.
Target validation travels further than the evidence more often than it should. A target validated with an antibody is sometimes treated as though it had been validated for a CAR-T, a radiopharmaceutical, or another modality, when a different modality may require materially different target density, internalization, tissue distribution, exposure, trafficking, or safety characteristics. The useful diligence question is which elements of the original validation transfer to the new modality, and which have to be established again.
The same problem appears after the science has worked. Clinical differentiation is allowed to convert directly into assumptions about adoption and reimbursement without sufficient attention to treatment burden, biomarker logistics, cost of goods, site of care, competing therapies, or payer skepticism about durability of benefit. A clinically differentiated product can still be commercially unattractive, and the evidence underlying those two propositions sits at different levels.
The most useful evidence package is rarely the largest; it is the one that most efficiently eliminates the uncertainty that would otherwise make the next decision irrational. For each major claim, five questions expose the bridge.
Broad propositions conceal gaps. Define the statement narrowly enough that it can be supported or falsified.
Identify evidence supporting this proposition, not an adjacent one.
Make the translational, modality, regulatory, manufacturing, or commercial bridge explicit.
If nothing could change the conclusion, the proposition is not being treated as a testable claim.
The structure of the decision should reflect what has been demonstrated and what still has to be earned.
The exercise is most valuable before a financing, a licensing process, or a major indication expansion. It often reveals that a proposed transaction is sound in principle but structured as though critical risks had already been retired. That difference is what deal terms are for.
A disciplined organization makes the evidence gap visible before it becomes expensive.
Build development plans around decision-critical uncertainties rather than a conventional checklist of studies. Each meaningful study should carry an explicit purpose: what it is intended to establish, and what outcome would redirect or stop the program. A study that cannot change a decision deserves scrutiny before it is funded, not after.
Underwrite the next value inflection rather than the endpoint implied by the deck. Observed data, literature-supported interpretation, and scenario analysis are all legitimate inputs, and they should not be treated as the same evidentiary category. The central question is what has to become true between today's evidence and tomorrow's valuation, and who is paying to find out.
Use transaction structure to reflect the evidence gap. Upfront consideration, milestones, option timing, and governance rights can distinguish risks that have been resolved from risks that remain contingent. A deal can make sense before the proposition has been fully validated; the economics should not pretend the validation has already occurred.
Know which claims in the narrative are direct observations, which rest on published evidence or external precedent, and which remain assumptions. The objective is not to remove the assumptions from the story. It is to know which one matters most, and what the shortest credible path is to turn it into evidence before the next raise.
Related Alacrita services
Alacrita conducts technical and commercial due diligence for investors, pharma BD teams, and companies preparing for incoming diligence. See Due Diligence and Investor Services.
Three regulatory items test how far agencies are prepared to let validated lower-level evidence carry.
Biotech development requires decisions before certainty exists, and there is nothing wrong with that. The mistake is letting the narrative conceal where the certainty ends.
The useful question is not how much evidence supports a proposition, but what decision that evidence legitimately supports and how far the claim has traveled beyond it.
The strongest biotech narratives do not make uncertainty disappear. They show why the next investment will convert the most important remaining in-deck assumption into direct evidence.
In deck describes a proposition assembled from synthesis rather than from direct testing by the program: published biology, competitor outcomes, market research, expert opinion, mechanistic analogy, regulatory precedent, and curated public data. The individual inputs may be sound. The limitation is that the combined proposition has not been tested in the relevant product, population, setting, or development context.
The largest published analysis, covering more than 21,000 compounds, estimates the Phase 1 to approval probability at 13.8% overall, with oncology substantially lower at 3.4%.
Yes. The proportion of drug mechanisms with direct human genetic support rises from 2.0% at the preclinical stage to 8.2% among mechanisms behind approved drugs. A 2024 analysis of 29,476 target-indication pairs estimated that genetically supported mechanisms carry 2.6 times the probability of success of unsupported mechanisms.
In the Amgen exercise reported by Begley and Ellis, findings from 53 landmark preclinical cancer studies were confirmed in six cases. An independent Bayer effort found published data consistent with in-house results in roughly a quarter of projects examined. Neither result shows preclinical research is unreliable in principle, but both indicate that positive in vivo data carries less weight than it is often given.
In defined circumstances. The FDA April 2025 roadmap set out a staged shift toward validated New Approach Methodologies in preclinical safety testing, and March 2026 draft guidance describes a validation framework built on context of use, human biological relevance, technical characterization, and fit-for-purpose. The bridge is established for a specific context of use rather than granted generally.
Platform validation establishes that an enabling technology works: a delivery system, a manufacturing process, an engineering principle, or a screening approach. Product validation establishes that a specific disease-directed application of that platform works. Evidence for the first does not transfer automatically to the second.
A target validated with an antibody may not be validated for a CAR-T or a radiopharmaceutical, because different modalities can require different target density, internalization, tissue distribution, exposure, trafficking, and safety characteristics. The diligence question is which elements of the original validation transfer and which have to be established again.
What precisely is the claim; what is the highest evidentiary level directly supporting it; what assumptions are required to extend it further; what experiment, dataset, or precedent could falsify it; and does the proposed transaction or capital commitment match the residual uncertainty.