Experimental Results Can Expire

Institutions often treat successful experiments as foundations for durable decisions. A pilot shows that an intervention improves an outcome, an evaluation identifies a promising programme, a trial demonstrates that a particular process works better than the alternative, and the organisation then scales, formalises or embeds what it has learned. This is a reasonable way to govern under uncertainty. Institutions cannot wait for perfect knowledge before acting, and experimental evidence provides a disciplined basis for moving from hypothesis to policy. Yet the authority of an experimental result is not timeless. Evidence that was once sufficient to justify a decision can become less reliable as the conditions that made it valid begin to change.

The distinction is easy to miss because the historical result does not itself deteriorate. A well-conducted experiment remains a well-conducted experiment. Its data do not become false merely because time passes, and the causal relationship it identified may have been entirely real under the circumstances in which it was observed. What can change is the relationship between those circumstances and the world in which the institution is now acting. The original finding may survive perfectly as a historical fact while becoming progressively weaker as a guide to present action.

This is why evidence preservation and evidence validity are different institutional problems. An organisation can remember exactly what was tested, which population was involved, what outcomes were observed and why the result originally justified a particular policy. Its institutional memory can function flawlessly. Yet if the surrounding environment has changed substantially, preserving the evidence does not answer the question that now matters: does this result still tell us enough about what will happen today?

Experimental evidence is always conditional, even when those conditions remain largely invisible. An intervention may work because of the technology available at the time, the behaviour of the population, the organisational capacity implementing it, the incentives surrounding the programme or the regulatory and economic environment in which the experiment occurred. If those conditions remain broadly stable, the original result may continue to travel well. If they change, the distance between the experimental world and the current world begins to grow.

Technology provides an obvious example. A process tested when citizens relied primarily on physical service channels may behave differently once most interactions become digital. An intervention evaluated before automated decision systems were introduced may encounter a different institutional workflow afterwards. A programme whose effectiveness depended on a particular information constraint may lose or gain value when new data sources become available. The experiment did not become wrong; the system to which its result is being applied became different.

Social conditions can change in the same way. Behavioural responses evolve, expectations shift, populations change and practices that once seemed unusual become normal. A programme evaluated in one demographic or cultural configuration may later operate in another. Incentives can also move. An intervention that worked while participation was voluntary may produce different effects once it becomes mandatory, just as a small pilot may behave differently after becoming part of a large administrative system. Scale changes relationships, not merely quantities.

Organisational conditions matter just as much. A successful pilot may have been delivered by an unusually experienced team, supported by exceptional leadership attention or protected from the constraints that affect routine operations. Once the intervention becomes standard practice, those conditions may disappear. The institution can continue citing the experimental result long after the implementation environment has ceased to resemble the environment in which that result was generated. What looked like evidence about the intervention may partly have been evidence about a temporary organisational configuration.

This creates a particular risk for institutions because successful evidence tends to become embedded. Once an experiment has justified a policy, the policy acquires procedures, budgets, staff, systems and constituencies. The original evidential question gradually disappears behind the operational reality that followed from it. The institution no longer asks whether the intervention works; it manages the fact that the intervention exists. Over time, a result that was once provisional support for action can acquire the status of an unquestioned historical foundation.

The problem is not that institutions should distrust old evidence simply because it is old. Age alone tells us very little. Some causal relationships remain remarkably stable, while others become fragile quickly. What matters is whether the conditions relevant to the original inference still hold. An experiment may remain informative decades later if its underlying mechanism and context have changed little, whereas evidence only a few years old may already be poorly transferable if the system has been transformed substantially.

This means that the useful question is not “How old is this evidence?” but “What had to remain true for this evidence to retain its authority?” That shift turns revalidation from a calendar exercise into a cognitive one. The institution needs to understand which assumptions connected the original result to the policy decision and whether those assumptions still deserve confidence.

Sometimes the relevant change is visible. A new technology is introduced, a population shifts or the legal framework changes. At other times, evidence can expire more quietly. Staff adapt their behaviour to the policy, citizens learn how to respond to it, complementary programmes accumulate or organisational routines drift. The intervention continues to operate, but the causal environment around it is no longer the one originally studied. Without deliberate attention, the institution may not notice that the basis of its confidence has weakened.

This is different from asking whether the policy currently performs well. A programme can continue producing apparently acceptable outcomes while the original causal justification becomes increasingly uncertain. Conversely, performance can deteriorate for reasons unrelated to the original evidence. Evaluation of current performance and revalidation of evidential authority therefore overlap, but they are not identical. One asks what is happening now; the other asks whether the historical evidence that continues to justify the intervention still transfers to present conditions.

The distinction becomes particularly important when policies are scaled. Experiments are often conducted under bounded conditions precisely so that institutions can learn before committing more resources. Yet scaling changes the system. Participants become less selected, implementation becomes more variable, administrative attention is diluted and interactions with other policies become more complex. Treating the original pilot result as permanently sufficient can therefore turn experimental success into a kind of evidential inertia.

This does not imply that institutions should continually rerun every experiment. That would be impractical and often unnecessary. Revalidation can take different forms. Sometimes monitoring key assumptions is enough. Sometimes new observational evidence can show whether the original mechanism still appears to be operating. In other cases, changing conditions may justify a new experiment or a more substantial reassessment. The important capability is not constant retesting, but knowing when the relationship between past evidence and present reality has become uncertain enough to require renewed inquiry.

Artificial intelligence and rapidly changing digital systems make this challenge more visible. Evidence about a model, workflow or human–machine interaction can become obsolete quickly when the underlying technology changes, data distributions shift or users adapt their behaviour. Yet the same logic applies far beyond AI. Any policy whose effects depend on a changing social, technological, organisational or environmental configuration carries some risk that yesterday’s evidence will gradually lose authority without ever being formally revoked.

Institutions therefore need to treat evidential authority as something that can require maintenance. The archive should preserve not only the finding but the conditions under which the finding was generated: who was studied, what assumptions mattered, which implementation capabilities were present and what surrounding systems shaped the result. Without that context, future decision-makers inherit a conclusion without inheriting the information needed to judge whether the conclusion still travels.

The deeper issue is temporal. Evidence enters an institution at a particular moment, but policies often persist across many moments. The world continues to change while the justification remains fixed in the past. If institutions do not revisit the relationship between the two, they can become highly disciplined about evidence at the moment of adoption and surprisingly undisciplined about evidence afterwards.

Experimental results can expire because evidential authority depends not only on what an experiment found, but on whether the conditions that made that finding transferable still exist. A historical result can remain perfectly valid as a description of what happened and yet become progressively weaker as a basis for present action. The challenge for an intelligent institution is therefore not simply to remember its evidence, but to know when the world has changed enough that yesterday’s justified confidence needs to become today’s question again.