Tesseract Foundational Research: Evaluation Methods
What government evaluates, and how openly it says so
Government publishes a great deal of evaluation, and asks suppliers to build on it. But the published record has never been mapped as a whole: what kinds of evaluation dominate, who commissions them, and how clearly they state their own methods. This is a first pass at that map, built from open data, and one number in it is uncomfortable.
The direction of travel
The centre of government keeps raising the evidential bar behind spending. The Evaluation Task Force now presses departments not just to evaluate but to let the evidence shape what is continued or stopped, with the Magenta Book as the standard. Suppliers are routinely asked to ground bids in the existing evaluation evidence base. Doing that well requires knowing the shape of that base, and it is exactly what no one publishes.
The method: classify what the catalogue shows
We harvested every GOV.UK publication of type research and independent report whose title declares it an evaluation, then classified each by evaluation type and by any method its metadata names. Classification is from the title and description only, the publication's own framing, by a locally hosted open-weights model: it measures how evaluations present themselves in the public catalogue, which is what a searcher or an automated evidence-synthesis tool sees first.
The uncomfortable number
Impact evaluations dominate the record, which is what the Magenta Book agenda would predict. But only 11% of these publications declare a recognisable method in their title or description. Where a method is named, qualitative and survey/monitoring designs lead; a randomised design is named in just nineteen. That is not a claim that the methods are absent from the reports, they are surely inside the PDFs, but it is a claim that the catalogue itself is largely method-blind: an automated reader, or a hurried human, cannot tell a robust impact study from a light-touch one without opening every file. The largest commissioners are DfID, the Department for Education and the Department for Work and Pensions.
Where this goes
The atlas is the corpus companion to our machine-checkable Theory-of-Change method: one measures how the evaluation estate presents itself, the other shows how a single evaluation's logic can be made auditable. Extended, it supports evidence-gap mapping by policy area and a method-transparency indicator per department, the raw material for a searchable evaluation evidence base.
Independent, self-initiated open research. Contains public sector information licensed under the Open Government Licence v3.0. Classification is automated and metadata-based; it describes how publications present themselves, not an audit of their contents.
Explore the atlas
1,770 classified publications as CSV, the summary breakdowns, and the reproducible pipeline.
