Institutional Cognition Needs Stress Tests

An institution can appear highly capable when nothing unusual is happening. Experienced people are in their usual positions, information arrives through familiar channels, established routines work as expected and decisions can be made with enough time to consult the right colleagues. Memory seems available, coordination seems effective and judgement seems sound. From the evidence of everyday performance, we might reasonably conclude that the institution possesses all of these capabilities. Yet routine conditions can be surprisingly generous. Sometimes what looks like institutional capability is partly the result of the environment never asking the institution a sufficiently difficult question.

This is why institutional cognition needs stress tests.

The basic idea is familiar from other domains. A system that works under ordinary conditions may behave very differently when placed under pressure. The important question is not simply whether it performs well today, but whether the properties on which that performance depends remain available when some of the normal supports are weakened. For institutional cognition, this means asking whether capacities such as remembering, coordinating and judging remain functional when circumstances become more demanding.

Consider institutional memory. Under normal conditions, an organisation may appear to remember its history perfectly well because experienced employees are present, archives are accessible and familiar cases allow people to rely on established routines. But what happens if several key employees are unavailable at the same time? What if relevant records become harder to retrieve precisely when a decision must be made quickly? The purpose of asking such questions is not to predict every future crisis. It is to reveal what the apparent capability actually depends on.

Coordination presents a similar problem. A government organisation may coordinate effectively when responsibilities are clear, workloads are manageable and information follows predictable paths. That performance tells us something, but perhaps not enough. A more revealing diagnostic question would ask what happens when a problem crosses several organisational boundaries, instructions conflict or authority becomes temporarily ambiguous. If coordination collapses immediately, routine performance may have been concealing a fragile architecture.

Judgement can also look stronger than it is when conditions are comfortable. Decisions may be good when evidence is clear, objectives align and there is enough time for consultation. Institutional judgement becomes more visible when evidence is incomplete, objectives compete and decisions must still be made. The pressure exposes whether the institution possesses mechanisms for handling uncertainty or whether apparently good judgement depended on circumstances doing much of the cognitive work.

The same logic increasingly applies to Human–AI systems. An institution may appear to have developed an effective hybrid capability while human expertise and automated systems agree with one another. The more interesting diagnostic situations arise when the model produces a plausible but questionable output, when human and machine assessments conflict, or when the normal source of technical expertise is unavailable. A capability that exists only when every component agrees and every support remains intact may be much more fragile than ordinary performance suggests.

These examples point towards a general method. A cognitive stress test begins by identifying the capability we believe the institution possesses. We then identify the ordinary conditions that support its apparent performance: familiar staff, accessible information, stable interfaces, sufficient time, predictable workloads, established relationships or redundant expertise. The diagnostic task is to introduce enough pressure to make those dependencies visible and then observe whether the capability remains available.

This requires care. The purpose is not simply to make the institution fail. An arbitrarily difficult scenario proves very little. Almost any system can be overwhelmed if enough pressure is applied. A useful stress test therefore needs a capability-survival criterion: a clear idea of what would count as the capability continuing to function under demanding but meaningful conditions. The question is not whether performance remains perfect. It is whether the underlying cognitive capacity survives sufficiently well to remain institutionally usable.

This also means distinguishing stress testing from ordinary performance measurement. Routine indicators tell us how the institution performs under the conditions it normally encounters. Stress tests ask a different question: what becomes visible when some of the conditions supporting normal performance are no longer guaranteed? The distinction matters because ordinary metrics may repeatedly report success while important dependencies remain hidden.

A team may consistently meet its targets because one exceptionally experienced employee quietly resolves every difficult case. A cross-departmental process may appear reliable because the same small group of people maintain informal relationships that compensate for weak formal coordination. A decision system may appear robust because ambiguous cases are rare. A Human–AI workflow may produce good outcomes because disagreements between humans and models have not yet become consequential. Routine success does not tell us whether these arrangements would remain effective if their hidden supports changed.

Stress testing can therefore reveal something that measurement under ordinary conditions cannot: latent fragility. It can show where an apparent institutional capability depends excessively on particular people, relationships, information channels, technologies or favourable operating conditions. It can also reveal genuine resilience. A capability that remains available when familiar supports are weakened gives us stronger grounds for believing that the institution possesses something durable rather than merely benefiting from favourable circumstances.

There is, however, an important methodological boundary. A cognitive stress test is primarily diagnostic. Its purpose is to activate a capability under demanding conditions so that its dependencies and failure signals become observable. It is not automatically an experiment designed to establish causal relationships by systematically varying institutional architecture. Those activities can overlap in practice, but their questions are different. Stress testing asks whether a capability survives pressure; experimentation asks what we can learn by deliberately changing conditions. Keeping that distinction clear prevents diagnostic exercises from claiming more causal knowledge than they actually produce.

The value of stress testing therefore lies less in generating a single score than in making architecture visible. If memory fails, where did retrieval break? If coordination collapses, which boundary became decisive? If judgement deteriorates, which support had been carrying more of the cognitive burden than expected? If a Human–AI configuration fails, was the vulnerability located in the technology, human interpretation, oversight or the interface between them? The failure signal matters because it points towards the architecture beneath performance.

This changes what it means to say that an institution has a capability. Routine success is evidence, but it is not always sufficient evidence. An institution that remembers only while particular people remain present, coordinates only while informal relationships compensate for structural gaps, judges well only when uncertainty is low or uses AI effectively only when human and machine outputs agree may possess capabilities that are substantially more conditional than they appear.

Institutional cognition should therefore be assessed not only by watching institutions perform when their normal supports are intact, but by observing what remains when some of those supports are deliberately placed under pressure. Stress tests make hidden dependencies visible, distinguish robust capability from favourable circumstance and expose forms of cognitive fragility that routine performance can quietly conceal. Sometimes the clearest way to discover what an institution can really do is to stop asking how it performs when everything is working normally.