Insight, Measurement & Cities
New evidence that upgrading informal settlements delivers heat resilience raises a harder question than whether it is true. Slum upgrading has been happening for fifty years. The benefit was not hidden. Nobody was collecting the data, because it was not in the results framework.
A study published in npj Urban Sustainability in August 2026 finds empirical evidence that upgrading informal settlements produces meaningful heat-resilience benefits even where climate adaptation was never the purpose of the intervention. Better infrastructure, better health outcomes during heatwaves. Drainage, housing, streets, shade and basic services functioning as adaptation technology without being labelled as such.
The finding is welcome and plausible. The more useful question is a different one.
Settlement upgrading has been carried out at scale for the better part of fifty years, across dozens of countries, with substantial evaluation attached to almost every programme. If the heat resilience effect is real, it has been occurring the entire time. It was not concealed. It was not too subtle to detect. It was simply never measured, because it was not in anybody's results framework.
That is not an accident, and understanding why is more valuable than the finding itself.
Robert Merton set out the problem in 1936, in a paper that remains the clearest statement of it. The unanticipated consequences of purposive social action are not a marginal case. They are a general feature of deliberate intervention in social systems, and Merton's taxonomy is explicitly symmetrical: an action can produce an unexpected drawback, a perverse result that defeats its own purpose, and an unexpected benefit.
The development sector has internalised two thirds of that taxonomy and almost entirely ignored the third.
There is elaborate institutional machinery for detecting unanticipated harm. Environmental and social impact assessment. Safeguard policies. Do no harm frameworks. Grievance mechanisms. Resettlement standards. All of this exists because harm creates liability, reputational exposure and, in the better cases, genuine moral concern.
There is almost nothing equivalent for detecting unanticipated benefit. No standard instrument, no budget line, no requirement in any major funder's reporting template. A programme that produced an enormous unmeasured good is not sanctioned for failing to notice, and so nobody looks.
The result is a systematic bias in what the sector knows about itself: well documented on the harms it causes, poorly documented on the goods it produces incidentally. That bias then feeds directly into what gets funded, because funding follows evidence and evidence follows measurement.
The mechanism is not mysterious. Indicators are specified in advance, before the intervention begins, in a results framework agreed with the funder. Data collection is then built to those indicators. A drainage and housing upgrade is financed from an urban development or sanitation budget line, so the framework contains drainage and housing indicators, so those are the data that exist. Heat-related morbidity was not in the framework, so no baseline was taken, so the question cannot be answered afterwards even in principle.
It is important to be fair about why the system works this way. Pre-specification of outcomes is a defence against a real problem. An evaluator free to choose which outcomes to report after seeing the data will find something positive in almost any programme, and the literature on outcome switching in clinical trials exists precisely because this is not a hypothetical failure. Pre-registration is methodologically sound and should not be abandoned.
But pre-specification answers one question well and forecloses another entirely. It is a confirmatory instrument. It tests whether the thing you expected to happen happened. It has no capacity whatsoever to discover what else occurred, and treating it as though it does is a category error that the sector makes constantly.
The fix is not to loosen the logframe. It is to stop asking a confirmatory instrument to do exploratory work, and to fund the second function separately.
Three of these quadrants have named professions attached to them. The fourth has almost nobody.
Open-ended qualitative fieldwork is the only method that reliably detects an outcome nobody specified in advance. This is not a claim about depth or richness or giving people a voice, all of which are true and all of which have made the argument sound softer than it is.
It is a claim about the structure of the instrument, and it is the reason societal readiness resists assessment through indicators alone. A survey can only return answers to questions someone thought to ask. An unstructured conversation starts from what the respondent considers significant, which means it can surface a change that appears in no indicator anywhere. Someone mentioning that they no longer sleep outside in April, or that a child's asthma settled, or that the clinic queue is shorter in the hot months, is producing information that no results framework would have requested and no closed instrument could have captured.
That is why this kind of work is not decoration on a quantitative study. It is the discovery function, and the quantitative work that follows is only as good as the questions it was given.
The sequencing matters and is usually reversed. Exploratory fieldwork identifies candidate effects, and confirmatory measurement then tests the ones worth testing. A programme that runs a household survey first and adds qualitative work at the end to explain the results has used the discovery instrument as a garnish. This is what field research is for, and it is a large part of why we argue that the social side has to lead rather than follow.
If ordinary settlement upgrading delivers material heat resilience, then adaptation finance has a tagging problem, and it is a serious one.
Adaptation spending is classified by stated intent. A project counts as adaptation if it is designed and described as adaptation. That produces two errors running in opposite directions, and both are large.
Effective adaptation that is not counted. Drainage, housing, street layout, shade, water supply and basic services in fast-growing cities are, on this evidence, adaptation infrastructure. Almost none of it is tagged as such, which means the sector systematically understates what is already being spent and misattributes the results.
Counted adaptation that may be less effective. Specialised interventions with adaptation in the title, from cool-roof coatings to sensor networks, are tagged in full. Some are excellent. The point is not that they do not work. The point is that the tagging system cannot distinguish between them and the alternatives, because it measures the label rather than the effect.
This is the mirror image of the reporting problem we described in One Month Is Not a Trend. There, a real number was asked to support a claim it could not carry. Here, a real effect produces no number at all. Any classification based on stated intent will reward projects that describe themselves in the funder's vocabulary. That is not corruption, it is what rational applicants do in response to the incentive presented to them. But it means adaptation finance totals are, in part, a measure of how programmes were written up.
This is a checkable claim rather than a rhetorical one. It would require retrospective outcome measurement on upgrading programmes that were never evaluated for climate effects, which is expensive but entirely feasible, and considerably cheaper than continuing to allocate adaptation budgets on the basis of project titles.
A finding of unanticipated benefit is not permission to stop looking. Merton's taxonomy runs in both directions, and settlement upgrading has a well-documented capacity to produce unanticipated harm alongside the good.
Upgrading changes land values. It changes tenure security, sometimes for the better and sometimes decisively for the worse. It can displace the poorest residents from the neighbourhood that was improved, so that the population enjoying the heat-resilience benefit at endline is not the population that was measured at baseline. A study finding better health outcomes in an upgraded settlement is measuring the people who are still there.
That is not a criticism of the npj work, which is doing something valuable. It is the reason the same argument applies to itself. If the sector's instruments cannot see unanticipated benefit, they are equally blind to unanticipated harm that falls outside the safeguard categories, and the honest response to this finding is to build the discovery function properly rather than to celebrate a result that happens to be favourable.
Fund discovery separately from verification. A small, ring-fenced budget for open-ended fieldwork, contracted and reported separately from the results framework evaluation, so neither contaminates the other. The confirmatory study keeps its pre-registration and its discipline. The discovery study is allowed to find whatever is there. Both are needed and they should not be the same contract.
Take wider baselines than the theory of change requires. The marginal cost of adding health, thermal comfort, tenure and displacement questions to a baseline that is being collected anyway is small. The cost of not having them is that a question can never be answered retrospectively. Baseline breadth is the cheapest insurance in evaluation and it is routinely cut first.
Go back and look at completed programmes. The largest available body of evidence about what upgrading does is sitting in fifty years of finished projects that were never examined for effects outside their original frameworks. Retrospective measurement is unglamorous and no funder's reporting cycle rewards it, which is precisely why it is where the undiscovered findings are.
The Lab works on this in climate and ecosystems and through monitoring and evaluation designed to detect what a programme actually did rather than only what it promised. The same reasoning underpins our argument about the distance between the work and the reward: what gets measured is decided by the financing structure, not by what matters.
If you are running or funding an upgrading programme and suspect it is doing more than its logframe records, that is a measurable question.
This is an independent insight piece by Transitions Lab. For the Lab's applied work, see Monitoring, Evaluation & Dissemination and Impact Measurement. See also Nobody Buys a Chiller on the same measurement-instrument failure inside an efficiency contract, The Right That Matters Is to the Tree on the version that appears in restoration-programme accounting, The Cheaper It Gets to Verify, the Less Anyone Visits on what falls out of view when the field visit is no longer required, A Warm House Is Not a Cheaper One on the same discovery-versus-verification split inside a Just Transition programme, Who Does It Fail For? on evaluation that only returns answers to the questions it was built around, The Trough Before the Dividend on the reversion rate no case-study synthesis measures, and Adoption Is Not the End of the Research on the abandonment finding that only a follow-up outside the delivery consortium can recover. To discuss a study, see Contact.