The dashboard of most development programmes shows how much was spent and what was achieved. It does not show whether anything remains once the support ends.
Two programmes that are in fact entirely different can therefore look the same in the reports for years. The difference becomes obvious only when there is no longer any opportunity to change course.
Catch-up and development programmes are today assessed almost exclusively with input and output indicators: how much funding was used, how many people were trained, how many jobs, patents and projects came into being. These are important indicators, they are precisely measurable, and they are collected competently.
One quantity is nevertheless missing from the dashboard: the capacity for local knowledge recirculation. That is, whether the knowledge created by the programme is able to produce new value on its own once the external support has ended.
This is not a data collection error, since it follows from the nature of the current indicators. Two programmes can present an entirely identical picture on every conventional measure while in one of them the knowledge concentrates at the external supplier and in the other it becomes self-sustaining within the local community.
The difference becomes visible when the support comes to an end and the resources accumulated earlier run out, because from that point on what shows is no longer what the programme built, but whether it is able to keep building without it.
Phantom measurement: the indicator measures something accurately, while the quantity that the decision depends on is not in it.
The catch-up research thread approaches the same gap from the other side: catching up isn't a question of import asks what remains once the imported actor leaves, while this page asks why the measurement does not show it in time.
Two programmes received identical support for ten years and produced identical values on every standard indicator. One of them still operates a decade after the support ended; the other stopped within three years. Choose, before reading on.
The data is not incomplete and the question is not a trick: these are the indicators on which the funding decision, the mid-term review and the closing report actually rest. The next section lets you try to separate them anyway.
Weight the five standard indicators however you judge best. The composite index is calculated for both programmes at year 10, and the third figure shows the distance between them.
No weighting separates them, because the two programmes carry the same value on all five indicators. A composite index cannot recover information that is absent from every one of its components, and this holds regardless of how sophisticated the weighting scheme becomes.
Support runs for ten years and stops at year 10. Note what the standard indicators do in years 11 and 12: they keep rising for both programmes, because the projects already funded finish and the patents already filed are granted. This lag is the reason the failure is usually noticed two years too late.
Nothing was hidden. The persistence indicators were recorded from the first year and were available to anyone who asked for them. They were not part of the assessment, so nobody looked.
The pattern is not specific to development policy. It appears in each of the frameworks that describe capability, and in each it wears a different name, which is part of why it is rarely recognised as one problem.
Absorptive capacity. Cohen and Levinthal define the construct as an unobservable organisational ability, while empirical work measures it with R&D intensity, doctoral headcount and patent counts. The construct is cumulative and relational; the measure is an input.
Dynamic capabilities. Here the problem surfaces as a tautology objection: a firm that performs well is said to have dynamic capabilities, and the evidence that it has them is that it performs well. In the absence of an independent measure, the capability is inferred backwards from the outcome.
Smart specialisation. The core of the framework is the entrepreneurial discovery process, which is relational by definition, while evaluation runs on funds allocated, projects launched and jobs created. The mechanism the framework rests on is the one part that is not measured.
Economic complexity. This framework partly escapes, since it infers capability from what is actually produced rather than from what was put in. In exchange it sees only what appears in trade data, so the institutional and relational layer stays outside its field of view.
Measurement theory has known the general version of this for decades under the heading of construct validity, and Marilyn Strathern's formulation of Goodhart's law describes its incentive consequence. The defensible claim is therefore narrower: this well-established methodological critique does not cross over into the practice of development policy, so something understood in one literature stays invisible in the place where the decisions are made.
The critique is old; the replacement is the part that can be new. A persistence indicator asks what continues without the support, rather than what was produced with it. Four of them are used in the model above, and all four can be applied retrospectively to closed programmes.
All four share a property that distinguishes them from the standard set: they measure what survives withdrawal, so they cannot be satisfied by spending. They remain exposed to Goodhart's law like any indicator, which is a limit worth stating rather than concealing.
Can persistence be measured while a programme is still running, without the measurement itself becoming the target?
This is a constructed model, not empirical data. The two programmes are defined to be identical on all five standard indicators through the support period, which is why no weighting can separate them. The impossibility shown in section 03 follows from that construction and is therefore a property of the model, not a finding about the world.
The empirical claim behind it is weaker and is stated separately: programmes that differ substantially in local recirculation frequently produce similar values on standard indicators, and that similarity is what makes early detection fail. Testing this claim requires closed programmes with post-withdrawal follow-up, which is the research extension of this piece rather than part of it.