Skip to main content
Back to Research

Tesseract Foundational Research: Evaluation Methods

What government evaluates, and how openly it says so

Government publishes a great deal of evaluation, and asks suppliers to build on it. But the published record has never been mapped as a whole: what kinds of evaluation dominate, who commissions them, and how clearly they state their own methods. This is a first pass at that map, built from open data, and one number in it is uncomfortable.

1,770
evaluation publications
886
impact evaluations
11%
declare a method in metadata
1996–2026
years covered

The direction of travel

The centre of government keeps raising the evidential bar behind spending. The Evaluation Task Force now presses departments not just to evaluate but to let the evidence shape what is continued or stopped, with the Magenta Book as the standard. Suppliers are routinely asked to ground bids in the existing evaluation evidence base. Doing that well requires knowing the shape of that base, and it is exactly what no one publishes.

The method: classify what the catalogue shows

We harvested every GOV.UK publication of type research and independent report whose title declares it an evaluation, then classified each by evaluation type and by any method its metadata names. Classification is from the title and description only, the publication's own framing, by a locally hosted open-weights model: it measures how evaluations present themselves in the public catalogue, which is what a searcher or an automated evidence-synthesis tool sees first.

Impact
886
Process
391
Mixed
139
Feasibility
108
Economic
91
Evidence synthesis
67
Unclear
88

The uncomfortable number

Impact evaluations dominate the record, which is what the Magenta Book agenda would predict. But only 11% of these publications declare a recognisable method in their title or description. Where a method is named, qualitative and survey/monitoring designs lead; a randomised design is named in just nineteen. That is not a claim that the methods are absent from the reports, they are surely inside the PDFs, but it is a claim that the catalogue itself is largely method-blind: an automated reader, or a hurried human, cannot tell a robust impact study from a light-touch one without opening every file. The largest commissioners are DfID, the Department for Education and the Department for Work and Pensions.

Where this goes

The atlas is the corpus companion to our machine-checkable Theory-of-Change method: one measures how the evaluation estate presents itself, the other shows how a single evaluation's logic can be made auditable. Extended, it supports evidence-gap mapping by policy area and a method-transparency indicator per department, the raw material for a searchable evaluation evidence base.

Independent, self-initiated open research. Contains public sector information licensed under the Open Government Licence v3.0. Classification is automated and metadata-based; it describes how publications present themselves, not an audit of their contents.

Explore the atlas

1,770 classified publications as CSV, the summary breakdowns, and the reproducible pipeline.