Evidence Needs Maintenance Too

Institutions know that important assets require maintenance. Buildings deteriorate, software becomes obsolete, datasets accumulate errors, professional qualifications expire and infrastructure needs periodic inspection if it is expected to remain reliable. Yet evidence is often treated very differently. Once an experiment has been completed, an evaluation published or a causal claim accepted, the result can acquire a strangely permanent status. The institution stores the report, incorporates the finding into policy and moves on. The implicit assumption is that the evidence, unlike the systems around it, requires no further care.

That assumption is difficult to defend because evidential authority depends on conditions that can change. An experiment does not merely produce a conclusion; it produces a conclusion within a particular configuration of population, technology, institutions, incentives, implementation arrangements and surrounding systems. If those conditions remain sufficiently stable, the result may continue to provide a strong basis for action. If they change, the original evidence may become progressively less informative about the present, even though nothing about the historical study itself has become incorrect.

This creates a maintenance problem. The institution does not necessarily need to repeat the original experiment, but it does need some way of noticing whether the relationship between past evidence and present reality is weakening. Evidence maintenance is therefore not the same as archiving. A perfectly preserved evaluation can still become a poor guide to current action. Nor is maintenance synonymous with constant retesting. The relevant capability lies between those extremes: preserving enough knowledge about why the evidence was trusted, monitoring the conditions on which that trust depended and reopening inquiry when important assumptions no longer appear secure.

Consider how differently institutions often treat operational and evidential dependencies. If a digital service relies on a particular technical system, the organisation normally knows that the system must be updated, monitored and eventually replaced. If the same service was designed because a study conducted ten years earlier showed that citizens behaved in a particular way, the evidential dependency may receive no comparable attention. The technology is recognised as something that can age; the behavioural assumption can remain embedded indefinitely. Yet either dependency can fail.

Evidence maintenance begins by making those dependencies visible. A useful evidential record contains more than the final conclusion. It should preserve the population in which the result was observed, the institutional conditions under which the intervention operated, relevant implementation capabilities, important assumptions about behaviour and the environmental factors necessary for the causal interpretation to hold. Without that contextual information, future decision-makers inherit a result while losing the means to judge how far it can safely travel.

Once those conditions have been identified, maintenance can become selective rather than indiscriminate. Institutions do not need to monitor every contextual variable with equal intensity. They can ask which changes would most seriously weaken the inference that justifies the policy. An intervention may depend heavily on a particular incentive structure but only weakly on demographic composition. Another may be robust across populations but sensitive to the technology through which it is delivered. Maintenance effort should therefore follow evidential vulnerability rather than chronological age alone.

This is one reason why simple expiry dates are rarely sufficient. Time matters because change accumulates through time, but old evidence is not automatically bad evidence and recent evidence is not automatically reliable. A relationship grounded in stable mechanisms may remain useful for decades. A result from only two years ago may already require reconsideration if the institutional or technological environment has changed dramatically. Evidence maintenance is therefore fundamentally about conditions, not calendars.

Replication is one possible maintenance tool, but it is not the only one. A full replication can provide strong evidence where stakes are high or conditions have changed substantially, but institutions may also use lighter forms of revalidation. They can examine whether the relevant mechanism still appears to operate, compare outcomes across newer populations, inspect whether implementation has drifted, test key causal assumptions or use new observational data to check whether the expected relationship remains visible. Maintenance is best understood as a family of practices calibrated to evidential risk.

This also means that maintenance should begin before doubt becomes crisis. Institutions often revisit evidence only after a policy appears to fail. By that point, the organisation may be trying simultaneously to diagnose current performance, reconstruct the original rationale and determine which assumptions changed. A maintenance architecture creates a different temporal relationship with evidence. Instead of asking only whether the policy still works after something goes wrong, the institution keeps some awareness of whether the evidential conditions supporting the policy are becoming more or less similar to those originally studied.

The distinction becomes especially important when policies become deeply embedded. The longer an intervention exists, the more infrastructure, routines, budgets and expectations can accumulate around it. That institutionalisation can make the original evidence harder, rather than easier, to reconsider. A policy may come to look inevitable because the organisation has adapted around it, even if the evidence that once justified it has become increasingly remote from present conditions. Evidence maintenance counteracts this tendency by keeping the justificatory relationship visible.

There is also an important difference between maintaining evidence and maintaining confidence. Institutions should not preserve confidence merely because a decision has already been made. Proper maintenance can increase confidence when relevant conditions remain stable and new observations continue to support the original inference, but it can also reduce confidence when transferability weakens. The purpose is not to defend the historical decision. It is to maintain an honest estimate of how much authority the evidence still deserves.

This requires governance as much as analysis. Someone must be responsible for noticing when important conditions change, and there must be a route for that observation to influence the policy or programme that depends on the evidence. Without institutional ownership, maintenance easily becomes nobody’s task. Researchers may consider the original evaluation complete, operational teams may focus on delivery and senior decision-makers may assume that the evidential question was settled when the policy was approved. The gap between those responsibilities is exactly where evidential obsolescence can become invisible.

Maintenance can also fail through excessive abstraction. Institutions sometimes retain a headline conclusion while forgetting the conditions attached to it. A finding such as “the intervention increases participation” may circulate long after the organisation has forgotten which population participated, what alternatives were available, which implementation supports were present or how the outcome was measured. Over time, a conditional claim becomes an unconditional institutional belief. Evidence maintenance must therefore preserve not only conclusions but the boundaries of those conclusions.

Digital systems could make this easier by linking policies to the evidence on which they depend and by recording important contextual assumptions alongside the result. Changes in population, implementation model, technology or regulatory context could then trigger review rather than relying entirely on institutional memory. Yet technology alone cannot determine when evidence requires reconsideration. The difficult judgement remains epistemic: which changes matter enough to weaken the original inference?

This is why evidence maintenance is not simply another compliance process. A checklist can confirm that a review occurred without establishing whether the relevant assumptions were actually examined. The institution needs enough substantive understanding of the evidence to know what could make it stop travelling well. Maintenance therefore requires not only procedural regularity but cognitive ownership of the causal reasoning behind the original decision.

It also creates a healthier relationship between institutional learning and institutional memory. Memory allows the organisation to preserve what it learned. Maintenance asks whether what was learned still deserves the same authority. These capabilities reinforce one another but cannot substitute for one another. An institution without memory repeatedly loses its evidence; an institution without evidential maintenance can remember its evidence perfectly and continue applying it long after its relevance has weakened.

The practical objective is not to turn every policy into a perpetual experiment. Institutions need stability, and many decisions should continue without constant reconsideration. The objective is to avoid a more subtle failure: allowing evidential confidence to become permanent simply because organisational attention has moved elsewhere. When the cost of being wrong is significant, when the environment changes quickly or when the original result depended on fragile assumptions, stronger maintenance becomes justified. Where mechanisms are stable and risks low, lighter monitoring may be enough.

Evidence needs maintenance because the authority of a result depends on more than its original quality. It depends on whether the conditions that made the result informative continue to hold. A mature institution therefore does not choose between preserving evidence and repeatedly rerunning every experiment. It develops the ability to preserve the reasoning behind its evidence, monitor the assumptions on which that reasoning depends and reopen inquiry when those assumptions begin to weaken. In this sense, maintaining evidence is part of maintaining the institution’s capacity to remain justified over time.