Skip to main content

Research-Backed Implementation

We do not just "build". We validate. Our delivery models are rooted in rigorous academic and industrial research methods to ensure efficacy and reduce waste.

Our Approach

"Research-backed implementation" means that every technical decision is preceded by evidence gathering. We apply mixed-methods research (quantitative data analysis + qualitative user research) to define the problem space before writing a single line of code. This aligns perfectly with the GDS Discovery phase but adds a layer of academic rigour to the validation process.

Tesseract Foundational Research

Our self-funded research programme: open standards, evidence bases and reference implementations built on public data, published in full for independent verification and reuse. Each project ships with a complete write-up covering challenge, intervention, assurance and reusable assets. Browse by topic below.

Defence, security & space5

The Challenge

The UK zero-emission-flight sector has strong policy signals and funding under the Jet Zero Strategy, ATI FlyZero and the DfT-funded, CPC-led Zero Emission Flight Infrastructure (ZEFI) programme, but the data on who is building what, who funds whom, and which technologies gate which pathway sits scattered across press releases, programme pages and reports. That fragmentation is the problem any coordination tool has to solve.

The Outcome

42 entities, 55 relationships and 272 RDF triples validated at zero SHACL violations, with an interactive network view and a machine-readable graph. The three primitives shown, a typed stakeholder graph, a controlled-vocabulary technology model with sourced maturity, and referential-integrity validation, are the data spine of any single-view coordination tool for a fragmented sector.

The Challenge

The UK Defence Investment Plan (June 2026) commits over £5bn to autonomous systems and £7.5bn to a Digital Backbone and Digital Targeting Web, all of which depend on heterogeneous systems and allies sharing data a machine can reason over. On the UK side that vocabulary is the Information Exchange Standard (IES); the built environment runs on HQDM and the National Digital Twin. Both are 4D and share a common heritage, yet no machine-readable crosswalk between them had ever been published.

The Outcome

A shared starting point the field lacked, released so suppliers building across defence and built-environment data can start from something concrete. The divergences record names the traps, including the ies:Event to hqdm:event false friend, that a naive label-based mapping would fall into.

The Challenge

PYRAMID (UK Defence Standard 00-134) is the MOD open reference architecture for avionics and mission systems. By design it has no shared data model: its Technical Standard states that components "do not share interface definitions", so a deployment "will use bridges to close the semantic gap", and the accompanying MOD assessment confirms "the PRA does not define a data architecture". The meaning of data across components is left to bridges that are, today, hand-built per deployment.

The Outcome

A shared meaning layer beneath an open avionics architecture, demonstrated on the public standard. No published work grounds PYRAMID, or the related FACE Shared Data Model, in any upper ontology; this occupies that gap with ontologies the UK government already owns.

The Challenge

The Information Exchange Standard (IES) is the 4D ontology behind UK defence and security data, but authoring IES-conformant RDF by hand is slow and demands scarce ontology expertise. General-purpose language models cannot help: asked to write IES Turtle, a strong 30B code model invents terms that do not exist in the ontology 94% of the time. That hallucination is worse than useless in a standards context, because invalid data that looks plausible is harder to catch than data that fails outright.

The Outcome

A working assistant for IES data authoring, released as a research prototype with its dataset and eval code, so the small community that works with IES has a concrete starting point rather than a closed demo. It pairs naturally with the IES-to-HQDM crosswalk and the open-ontologies validation engine: the model drafts, the symbolic layer verifies.

The Challenge

Policy on the security of connective products (IoT, operational technology, computing devices, networking equipment, and software and firmware) is set against headline incident numbers that cannot be reconciled: no shared denominator, non-additive definitions, and vendor telemetry that is never quantified. The field lacks an honest, machine-readable baseline.

The Outcome

A shared, honest baseline that maps real incidents to product class, vector and impact, and states its own limits. It gives policy work something concrete to build on instead of irreconcilable headline figures.

Government & public services4

The Challenge

Government is raising the evidential bar behind spending: the Evaluation Task Force is shifting from checking that evaluation happens to ensuring evidence shapes decisions, and the Victims and Prisoners Act 2024 and Victims' Funding Strategy push victim services towards outcome-based, accountable commissioning. Most of that evaluation rests on a Theory of Change, which as a prose diagram hides its own weakest links: a causal claim with no evidence behind it looks identical to one with strong evidence, so effort pours into the well-understood links while the untested load-bearing assumption sits unexamined.

The Outcome

A method that turns "name the weak links" from a hope into a mechanical result: the products a commissioner sees are still the familiar Theory of Change diagram, narrative and M&E framework, while the structured graph underneath makes them auditable, internally consistent, and reusable in a full evaluation rather than redrawn from scratch.

The Challenge

Government publishes a great deal of evaluation and asks suppliers to build on it, but the published record has never been mapped as a whole: what kinds of evaluation dominate, who commissions them, and how clearly they state their own methods. Grounding a bid in the existing evidence base requires knowing the shape of that base, which no one publishes.

The Outcome

1,770 classified publications spanning 1996 to 2026. Impact evaluations dominate (886), as the Magenta Book agenda predicts, but only 11% declare a recognisable method in their metadata, so the catalogue itself is largely method-blind: an automated reader cannot distinguish a robust study from a light-touch one without opening every file. The largest commissioners are DfID, DfE and DWP.

The Challenge

Government runs thousands of public consultations and publishes their outcomes, but the estate has never been assembled as a single, queryable corpus: who consults, on what, and whether they publish what they heard back. Consultation-response coding and automated summarisation both need a structured view of that estate to sample and benchmark against.

The Outcome

6,260 consultations with published outcomes, coded by policy area at 98% coverage. Every one has a published outcome page, yet only 77% attach response documents: nearly a quarter close with an outcome carrying no published response analysis. Defra, DfT and MHCLG are the most prolific consulting bodies.

The Challenge

The UK asks public bodies to publish how they use algorithmic and AI tools in decisions affecting people, through the Algorithmic Transparency Recording Standard, the nearest thing to a public register of government AI. It is published one record at a time, which makes each disclosure visible but the pattern across government invisible.

The Outcome

136 records spanning 73 distinct public bodies, led by DSIT, DWP, the Money and Pensions Service, DfE, DESNZ and DBT. The corpus makes answerable the questions no single record can: which parts of government are most transparent about their AI, which kinds of tools recur, and how disclosure is spreading.

Economy, finance & skills6

The Challenge

Skills England publishes the national occupational maps (1,269 occupational standards with their routes, pathways and clusters, SOC classifications, apprenticeship and technical education products, green-jobs themes and progression pathways) through a public API. The data is relational but delivered as documents, has no published schema for the connected object, and buries the SOC crosswalk (the bridge to official labour-market statistics) inside each record. Consumed the usual way, a flattened slice ends up in a spreadsheet and the relationships that make it valuable are lost.

The Outcome

An open, machine-readable ontology of the entire national occupational map (51,355 triples: 1,269 standards, 15 routes, 35 pathways, 172 clusters, a 278-code SOC 2020 crosswalk, 1,313 products and 2,717 progression edges), with a live explorer on the Tesseract Academy for the Public Sector. It is the operational companion to our open PIAAC research on skills and social mobility: PIAAC measures what adults can do, the occupational maps show where those skills lead, and SOC is the join.

The Challenge

Local authorities must repeatedly evidence their housing market for local plans, strategic housing market assessments, stock-condition programmes and viability studies. The raw material, HM Land Registry Price Paid Data, is open and national but arrives as fourteen unlabelled columns and roughly nine hundred thousand rows a year, with no comparable per-authority summary.

The Outcome

Indicators for all 318 local authorities in England and Wales from 760,607 standard 2025 transactions, with local median prices spanning £135,000 to £1.16 million. A comparable, defensible starting point for any local housing-market evidence base, built entirely from open data.

The Challenge

"Structured data beats PDF for AI-driven reporting" is received wisdom in regulatory and reporting circles, but the empirical premise is rarely measured: across the accounts companies actually file, what fraction expose each core figure as a structured fact a machine can read without OCR or layout guessing?

The Outcome

Across 8,856 filings and 899 distinct concepts (median 49 tagged facts per filing), the core balance sheet is genuinely structured: equity exposed in 97.2% of filings, net assets 86.6%, creditors 75.6%. The tail falls away fast and unevenly, with debtors present in only 28%, and that unevenness is the finding: systems must be built on the measured exposure profile, not an assumption of uniform availability.

The Challenge

ESCO is the reference occupation and skills classification across Europe; England's Skills England maintains a separate, nationally specific set of occupational standards behind apprenticeships and technical education. To read English skills data against the international vocabulary you need a crosswalk, and one that is honest about how far each link can be trusted.

The Outcome

114 exact, 281 close, 270 related and 604 unmatched matches (52% any-band, 31% strong). The 604 unmatched occupations are the headline: a measurement of how far England's occupational language sits from ESCO's, and precisely the surface a semantic crosswalk or national ESCO extension must cover.

The Challenge

The FCA warns the public about unauthorised and scam firms through its warning list, a useful safety tool with one structural blind spot: it shows only the firms flagged right now. There is no history, so the trend, the very thing that says whether the scam problem is growing, is invisible to anyone designing fraud-prevention services or detection tooling.

The Outcome

18,224 live warning pages, 15,437 dated, revealing an unambiguous signal: FCA warnings ran at a few hundred a year through the late 2010s, then climbed to 1,616 in 2022 and around 1,900 a year across 2023 to 2025, a near four-fold rise on 2019. Just under a fifth flag clone firms, among the hardest frauds for a consumer to spot.

The Challenge

Employability programmes, skills bootcamps and local growth plans must target the right places, and the question of where the labour market is tightest has a monthly answer in the claimant count. It is open through the ONS NOMIS service, but delivered through a query builder few outside statistics teams use, so it usually is not to hand.

The Outcome

A like-for-like reading for all 374 local authorities, refreshed monthly, with claimant rates spanning 1.1% to 10.1% against a 3.5% national mean, Birmingham highest at 10.1%. Exactly the spread a place-based programme needs before deciding where to act.

Environment & climate3

The Challenge

The UK is rolling out continuous water-quality monitoring on wastewater assets under the Environment Act 2021, and the Environment Agency is exploring whether that data can serve as a regulatory tool. The binding question is data quality: calibration drift, fouled probes, telemetry gaps and transcription errors all masquerade as real signals.

The Outcome

The data is complete and internally consistent (COD ≥ BOD holds on every one of 1,382 rows), yet the battery still surfaces multi-year baseline drift and dozens of statistical outliers per determinand. The SHACL rules reject every physically impossible record and pass every clean one, making the trust verdict auditable.

The Challenge

In January 2026 the UK Government published its National Security Assessment on global biodiversity loss, ecosystem collapse and national security, assessing with high confidence that six strategic ecosystems are on a pathway to collapse with security consequences. The assessment is a one-off written document: it ships no machine-readable dataset, no shared ontology and no standing monitoring mechanism, so its cascade chains cannot be queried, versioned or stress-tested.

The Outcome

A first, correctable occupation of a genuine whitespace: the framing is government-endorsed and the science exists (IPBES Nexus 2024), but no open, computation-ready ontology and evidence base of nature-to-security cascades existed. The three government functions map to concrete methods: assess (a nature-security exposure profile binding open indicators to transmission channels), monitor (early-warning statistics including critical slowing down), and mitigate (typed intervention points per cascade, mapped to National Risk Register, IPBES and Kunming-Montreal levers).

The Challenge

Stakeholder mapping for the environment sector is usually a slide: boxes and arrows, no provenance, out of date the day it is drawn. The relationships that matter, who sponsors whom, who advises government internationally, whose data flows into the national record, which duties are statutory, are exactly the ones a diagram cannot defend, and getting them wrong in a bid or plan is a real risk.

The Outcome

A queryable graph that records the accurate relationship, not the convenient one: it correctly captures that the Forestry Commission is a non-ministerial department not a sponsored body, that NatureScot, NRW and SEPA answer to the devolved governments, and that Local Nature Recovery Strategy authorities are designated under Environment Act 2021 section 104, each carried in a cited basis rather than smoothed over.

Culture, health & science6

The Challenge

Most of the UK's heritage is digitised but not computable. Aerial photography collections holding tens of millions of frames are scanned, catalogued and mapped, yet a researcher cannot query them at scale by space and time or cross-reference them to other archives. This is the gap the Towards a National Collection programme (AHRC / UKRI) and its N-RICH work set out to close.

The Outcome

Measured finding: the collection is far closer to computation-ready than assumed. 100 percent of records already carry a machine-readable footprint and an ISO-8601 date; the one genuine gap is machine-readable rights. The same standard was then applied unchanged to two further national collections (NAPL Canada and WHAIFinder USA), each lifting to the identical NAPH Baseline at zero SHACL violations, with each collection missing a different Baseline field. The records export cleanly to STAC 1.0, GeoJSON and IIIF, and NAPH publishes the previously unoccupied RiC-O to STAC crosswalk that binds archival provenance to spatiotemporal computability.

The Challenge

Agri-environment schemes are the largest source of government funding for the rural historic environment, yet the wider co-benefits of heritage actions, for nature and for people, are real but evidentially fragmented. As nature recovery is delivered faster and under tighter budgets, heritage actions risk being overlooked unless that value can be evidenced.

The Outcome

A published, replicable protocol, an evidence-map and gap framework, and a real spatial analysis showing 84% of monuments at risk are on farmland, where heritage actions act. Together they make the case that heritage actions deliver multi-objective value, for nature and for people, in a form that withstands scrutiny.

The Challenge

Scientific data is widely described as FAIR (Findable, Accessible, Interoperable, Reusable), but the claim is rarely tested at scale against a machine-checkable contract, and "AI-ready" is asserted more often than it is evidenced. Without a measurable baseline, funders and repositories cannot see where the real gap is.

The Outcome

The datasets are overwhelmingly findable and accessible, but 0% are interoperable or AI-ready: 100% lack a machine-readable schema, checksums and provenance. The finding names the real gap precisely, and the paired ontology models the AI-ready layer that closes it.

The Challenge

UK museums are opening their collections as data faster than the data can be made usable. The records come out flat and free-text: the same polymer written four different ways as unrelated strings, relationships between objects trapped inside caption prose, no links to anything. The gap between a collection being published and being computable is real, and until it is measured it stays invisible.

The Outcome

99.9% of 35,172 free-text material assertions resolved to a science-grounded concept (476 raw strings collapsed to 137); 100% of technique and use-domain assertions; 55 verified Getty AAT exact matches; a 485,000-triple graph over all 11,865 objects; and 289 object-to-object variant edges recovered from prose no keyword search could traverse.

The Challenge

When a council commissions a health-and-wellbeing survey it asks residents to rate their own health, the same question the Census already asks the whole population. A survey read on its own is a number in a vacuum; the full-population baseline for that authority exists but is not laid out ready to compare against.

The Outcome

A comparable baseline for all 331 local authorities. Self-reported bad or very-bad health runs from 2.7% to 9.6%, highest in Blaenau Gwent, Merthyr Tydfil and Blackpool, the geography of health inequality made measurable per authority alongside disability and unpaid-care prevalence.

The Challenge

DCMS publishes monthly visitor figures for every museum and gallery it sponsors, one of the cultural sector's best open data sets, yet it circulates as static spreadsheet tables: no baselines, no seasonality context, no way for a trustee or a marketing team to answer 'are we back, and compared to what?'

The Outcome

A working demonstration of the difference between reporting and insight on data every cultural organisation already knows: recovery is uneven across museum groups, and within a single group the site-level trajectories diverge further still, which is where the actionable signal lives.

Commissioned delivery case studies

Contracts commissioned by government and government-funded programmes, each with a full write-up: challenge, intervention, assurance and reusable assets.

See all delivery case studies

Selected Publications & Talks

Grouped by topic. Open a group to see the studies, papers and talks inside it.

Safe & verifiable AI14
  • The Open-World Hole: Why SHACL Cannot Catch a Hallucinated Ontology Term

    Open benchmark, 2026

    SHACL, the default RDF validator, is open-world: it silently passes any ontology term it has no shape for, which is exactly the failure mode of an LLM authoring RDF. Measured on three real vocabularies (schema.org, IES4, and the OBO Foundry's PATO and RO), open-world SHACL validated as conformant every one of 300 data graphs carrying a fabricated term, across 418 fakes. A closed-world vocabulary gate caught all 300 with zero false positives on clean data. The correctness layer for AI-generated knowledge graphs, released reproducible.

  • Marginal Coverage Is Not Joint Coverage: Copula-Coupled Uncertainty Certificates

    Open case study, 2026

    Scientific models predict correlated vectors, and calibrating their uncertainty one output at a time gives a per-target guarantee that looks right and a joint guarantee that is quietly wrong. On 1,600 real OQMD materials, a "90%" independent conformal certificate delivers only 79% joint coverage; a coupled certificate restores 90% (22% tighter than Bonferroni) and a Gaussian-copula region matches it at less than half the size. Distribution-free, machine-checkable, reproducible.

  • Fluency Is Saturated, Correctness Is Not: An Un-Game-Able Scientific-LLM Benchmark

    Open benchmark, 2026

    Most LLM benchmarks grade on fluency or an LLM judge, so a model that invents a nonexistent gene or ontology term scores as correct. verifiabench grades with a deterministic closed-world oracle instead. Across nine models, including Claude Opus, Sonnet, Haiku and local open checkpoints, raw fluency saturates at 1.00 while verified capability ranges from 0.00 to 1.00 on the same tasks: a local Qwen3-Coder-30B ties Claude Opus at perfect verified correctness, while some equally fluent models invent roughly half their terms. Un-game-able, reproducible against any endpoint.

  • The Gate You Can Trust: A Proof-Carrying-Action Gatekeeper, Exhaustively Verified

    Open reference implementation, 2026

    A safeguarded AI permits an action only if it carries a certificate a small trusted component can check. This is a working reference gatekeeper for a bounded multi-agent system, verified exhaustively rather than tested: over 96 reachable states and 1,176 transitions it blocks 672 of 672 unsafe actions and admits 504 of 504 safe ones, at 0.36 microseconds per check with a 31-line trusted core. Without the gate, 57% of actions violate the safety spec.

  • The Privacy Fix That Quietly Breaks Fairness: One Report Card for Health AI

    Open case study, 2026

    Privacy and fairness are audited separately, but they interact. On the real UCI Diabetes readmission dataset, one report card runs membership inference and per-subgroup equity on the same clinical model across a differential-privacy sweep. At the tightest privacy, differential privacy costs minority subgroups 2.6x more accuracy than the majority (drops ordered by subgroup size), while the membership leakage it targets is near zero, almost all equity cost for almost no privacy gain, visible only when both planes sit on one card.

  • Proving the Utility of Large Language Models in Cybersecurity Simulations

    Peer research paper with The Alan Turing Institute, arXiv 2608.16422, 2026

    Language models write the network itself. The pipeline synthesises YAML topologies for cybersecurity simulation, checks each candidate against the emulator before it is kept, and uses the resulting environments to train and test reinforcement-learning defence agents. Across multiple synthetic topologies, LLM-instantiated Python attackers reached a 94.5 per cent compromise rate at 0.02 to 0.06 seconds per assessment, four to five orders of magnitude faster than the Double Q-learning with prioritised experience replay baseline they are measured against. Co-authored by Stylianos Kampakis and Fabio Rovai with Marcos Charalambides, Theodosis Mourouzis and Chris Hicks of The Alan Turing Institute.

    Read the paper on arXiv (opens in new tab)
  • Silence Is Not Assent: What Catches a Wrong AI Classification in Orbit

    Open knowledge graph and measurement, 2026

    A public 833,403-triple knowledge graph of all 70,122 catalogued space objects, aligned to the Space Situational Awareness Ontology. The ontology declares exactly one exclusion axiom, between two coordinate formats, so 351 of its 353 classes can never reject a wrong AI classification. The instance data can: LogMap 4.0 on this pair reports zero repair conflicts while the catalogue refutes nine of its candidates, catches a category error in its final output and rescues one it wrongly discarded. The threshold sweep also caught a circular measurement in our own earlier version, which is reported rather than buried. Includes the first openly published language model for a space ontology: hallucinated ontology terms fall from 13.81 to 0.06 per output.

  • Grounded, Not Retrieved: An Ontology-Validated Biomedical Knowledge Graph

    Open case study, 2026

    A gene-disease knowledge graph built from 40 real Open Targets associations, every edge typed with the Biolink Model and checked by a closed-world vocabulary gate before it enters the graph. The grounded graph has zero SHACL and zero vocabulary violations across 284 triples; an ungrounded twin with one fabricated Biolink predicate passes SHACL but is caught by the gate. Ships with a provenance-carrying hypothesis triage and a one-click reproduction.

  • Construction Data Standards Cannot Check Your AI: IFC, COBie, Uniclass and BOT Measured

    Open crosswalks and measurement, 2026

    AI tools now assign Uniclass codes and fill COBie workbooks automatically, and the target standards assert zero disjointness: measured falsifiability 0.00%, so no wrong mapping into them can ever be machine-rejected. Open, argued crosswalks across IFC4, COBie, Uniclass 2015 and the W3C Building Topology Ontology: 49 correspondences, 7 recorded refusals including the bot:Zone false friend and the framed-structures part-whole trap, all 96 identifiers verified against pinned sources, every Uniclass code confirmed against the live NBS service. BOT reaches 80.95% checkability from 9 axioms; IFC4 reaches 11.45% from 2,443.

  • Is the Semantic Web Counterfactual-Ready? A Tractability Census of Public Ontologies (opens in new tab)

    Open dataset and preprint (forthcoming), 2026

    The first tractability map of the public semantic web. A compiler certifies whether each of 528 public ontologies can, as written, support a counterfactual ("what if") query. The finding: 56% declare no constraints and are counterfactual-blind, including both of the UK's 4D data standards, IES4 and HQDM, alongside schema.org, CIDOC-CRM, SAREF and GS1. On the 233 that can, a certified counterfactual is computed for each (median 0.19 ms), confirming that the cost is governed by ontology structure, not by the degree-based hardness that would suggest infeasibility. Open, reproducible, and released with a "counterfactual-ready standards" proposal.

  • Open Governance: Open-Source AI Governance Server (opens in new tab)

    Open-Source Tool, Ongoing

    An open-source AI governance platform that helps organisations discover, assess, and monitor AI systems against EU AI Act, NIST AI RMF, and ISO 42001 frameworks. Provides automated risk classification, compliance matrices, bias and hallucination monitoring, policy enforcement gates, and audit-ready reporting through 48 governance tools.

  • The closest thing to a public register of government AI, as a corpus

    Open research, 2026

    The full set of 136 published UK Algorithmic Transparency Recording Standard records, structured as an open corpus: which of 73 public bodies have disclosed which algorithmic and AI tools. Published on GitHub and Hugging Face under the Open Government Licence.

  • Teaching an open-source LLM a government information standard: from 94% hallucination to 1%

    Open research, 2026

    Tesseract Academy fine-tuned Qwen3-Coder-30B on IES4, the UK government Information Exchange Standard for defence and national security data sharing. Term confabulation fell from 93.7% to 1.0%, term conformance rose from 0% to 88.6%, every claim machine-verified against the published ontology. Published openly on Hugging Face.

  • FAIR Dataset Contracts for Scientific Data

    Open research, 2026

    An open study of 1,738 real public biomedical datasets across three repositories (EMBL-EBI BioStudies, Dryad, PRIDE): overwhelmingly findable and accessible, but 0% interoperable or AI-ready (100% lack a machine-readable schema, checksums and provenance). Paired with an open, tool-certified OWL ontology that models the AI-ready dataset layer. Open source and reproducible.

Environment & climate5
  • Can a Parametric Climate Insurance Product Prove It Paid?

    Open standard and measurement, 2026

    An open vocabulary and twelve machine-checkable rules that audit a parametric climate insurance promise at the three points where it fails: whether the promise is specifiable at all, whether a registered household is payable before a storm, and whether the payout arrived after one. The replay produces a documented false negative from public records. Severe Tropical Storm Nalgae came ashore at 110 km/h, eight short of the 118 km/h that defines a typhoon, and destroyed roughly 67,000 tonnes of mostly rice worth about PHP 1.3 billion. A trigger set at typhoon strength pays nothing for that. Every one of the twelve rules ships with a case built to trip it, and the build report lists what we could not obtain.

  • Half the Detail Dies on the Way to the Return: What UK Waste Reporting Throws Away

    Open crosswalks, measurement and reporting engine, 2026

    AI analytics classify waste at the belt in over a hundred categories and the industry sells that granularity as regulatory reporting. Every UK channel that consumes composition data accepts between 7 and 47 values. Of 5.907 bits of composition detail, the List of Waste retains 53.8%, pEPR 52.1%, RAM 2027 49.5% and Simpler Recycling 38.4%. The crosswalk cannot be a function in either direction: it collapses up to 23 classes onto one value, 7 classes fix no regulatory value at all, and the packaging schemes cannot represent 19 to 22 of them. Of ten operational questions, two are answerable in every channel and four in none, including the aluminium against steel split that sets the pEPR fee and for which the List of Waste offers a single code. Ships a reporting engine that emits returns as intervals rather than point estimates, and a release gate that caught a language-tag defect hiding the most-used municipal code in the upstream ontology.

  • From raw museum records to a knowledge graph

    Open research, 2026

    A reproducible pipeline turning the Museum of Design in Plastics' raw open catalogue (11,865 objects, CC BY 4.0 via the Museum Data Service) into a standards-based knowledge graph: a SKOS materials taxonomy grounded in polymer science, a CIDOC-CRM instance graph, verified Getty AAT alignments and a variant graph recovered from description prose. 99.9% of 35,172 material tags resolved; zero SHACL violations. Published in Open Ontologies under CC BY 4.0.

  • The Value of Agri-Environment Heritage Actions

    Open research, 2026

    A self-initiated, open-data demonstration of how we evidence the wider value of agri-environment scheme heritage actions for nature and for people: a scoping-review framework on Natural England\'s NEER001 method, with a real spatial analysis of the 2,206 Scheduled Monuments on the Historic England Heritage at Risk Register 2024.

  • Mapping the UK Zero-Emission Flight Ecosystem

    Open research, 2026

    An open, SHACL-validated knowledge graph of the UK hydrogen and electric aviation ecosystem: 42 real entities and 55 relationships across organisations, airports, programmes, projects, funders and technologies, validated at zero SHACL violations with enforced referential integrity. Open source via Open Ontologies.

Industrial and freight standards assurance4
  • An overview of ontology quality issues in air cargo data standards, tested against IATA ONE Record

    IATA ONE Record audit, dual-engine verified, English and Korean, 2026

    An independent, dual-engine audit of the IATA ONE Record data model ontology across six releases. Every one of the 496 properties in the 2022-12 release declares an rdfs:domain. None of the 522 properties in 2023-12 do, and none of the 534 in 2024-12 do. Classes with no rdfs:label rose from 1 to 58 over the same span, and total lint issues rose from 2 to 650. Verified two independent ways, with rdflib and with the open-source Open Ontologies engine, which agree exactly. Written for Korean air cargo operators building on ONE Record, in English and Korean, with a reproducible check over the MIT-licensed public repository.

  • An overview of ontology quality issues in air and ocean freight data standards, tested against IATA ONE Record and DCSA

    IATA ONE Record and DCSA audit, English and Traditional Chinese, 2026

    An independent, dual-engine audit of the IATA ONE Record data model ontology across six releases, written for Taiwan's air and ocean freight sector. Every one of the 496 properties in the 2022-12 release declares an rdfs:domain. None of the 522 properties in 2023-12 do, and none of the 534 in 2024-12 do. Classes with no rdfs:label rose from 1 to 58 and lint issues from 2 to 650. Includes a contrast with the DCSA ocean container OpenAPI specifications, 174 files under Apache 2.0. Verified two independent ways with rdflib and the Open Ontologies engine. In English and Traditional Chinese.

  • An overview of ontology and schema quality issues in machinery data standards, tested against MTConnect and the Asset Administration Shell

    EU Machinery Regulation readiness, MTConnect and AAS audit, English and Korean, 2026

    A readiness assessment for Korean machine tool and equipment exporters against EU Machinery Regulation 2023/1230, which applies from 14 January 2027, verified against the EUR-Lex primary text at Article 54 because several vendor summaries give the wrong date. Audits the two semantic carriers the sector relies on. MTConnect publishes each version as both XSD and JSON Schema, and the JSON AlarmStateEnum admits INSTANT while the XSD AlarmStateType does not, identically in 2.0, 2.1 and 2.2. Includes a retracted finding: an initial count of 111 divergences collapsed to one after 105 proved to be a deliberate design difference and three were our own matching artefact. Also tests and refutes the common claim that AAS submodel compliance requires a paid ECLASS subscription. In English and Korean.

  • An overview of ontology and schema quality issues in machine tool data standards, tested against MTConnect and the Asset Administration Shell

    EU Machinery Regulation readiness, MTConnect and AAS audit, English and Traditional Chinese, 2026

    A readiness assessment for Taiwan's machine tool cluster against EU Machinery Regulation 2023/1230, which applies from 14 January 2027, verified against the EUR-Lex primary text at Article 54. Audits MTConnect and the Asset Administration Shell. The MTConnect JSON Schema AlarmStateEnum admits INSTANT while the XSD AlarmStateType does not, identically across 2.0, 2.1 and 2.2. Includes a retracted finding, an initial 111 divergences reduced to one verified defect. Counts 1,529 IDTA IRI references against 18 ECLASS and 9 IEC CDD, refuting the claim that compliance requires a paid dictionary subscription. In English and Traditional Chinese.

Security, defence & space6
  • An Open Ontology for Space Object Catalogues, Tested Against CelesTrak and the General Catalog of Artificial Space Objects

    Open ontology, two-catalogue measurement, 2026

    An open OWL 2, SKOS and SHACL ontology and reproducible audit of the boundary between the two open catalogues of objects in Earth orbit, built from keyless downloads of CelesTrak's SATCAT of 70,292 objects and Jonathan McDowell's General Catalog of Artificial Space Objects across all four of its catalogues, both pinned to 18 August 2026. GCAT marks 22 entries as corresponding to no real object, giving its reasons in its own words including radar error, cataloging error and duplicate, and CelesTrak carries all 22; one of them, catalog number 11006, has no decay date and therefore counts as a tracked object still in orbit. 1,094 objects GCAT records as no longer tracked carry no CelesTrak data status code, which is a finding rather than a scope difference because CelesTrak maintains that field and applies it to 1,292 other objects. 20,198 of 34,814 on-orbit objects, 58.0 per cent, have no published radar cross section, so no size or mass can be derived from the open record. CelesTrak's owner code collapses the Soviet Union and the Russian Federation into a single value where GCAT separates 16,142 objects from 9,016, and 157 objects are attributed to the United States by one catalogue and to New Zealand by the other. All 70,292 COSPAR designators are well formed with no collisions, a hypothesis that died and is reported as such. Every headline computed two independent ways.

  • Space debris metrics that cannot be added up: an open composition checker

    Open research, 2026

    The two premier orbital debris engineering models disagree by a factor of 2.3 to 3.0 at the 1 cm collision-risk threshold, and by almost two orders of magnitude in the sub-millimetre regime, because they partition the same physical population along orthogonal axes: MASTER-8 by generative source, ORDEM 3.1 by material density. Every indicator built on either model silently inherits that partition and never declares it. An open, MIT-licensed composition checker built on QUDT, SOSA/SSN and PROV-O that refuses invalid combinations with a reason and states what may be done instead.

  • IES to HQDM: an open 4D ontology crosswalk for defence data

    Open research, 2026

    The first public crosswalk between the UK Information Exchange Standard (IES) and HQDM, two 4D upper ontologies. Open SSSOM and RDF correspondences, a curated divergences record, SHACL validation, and a worked SAPIENT-node safety case grounding autonomy in IES-typed world states. Supports the Defence Investment Plan interoperability and autonomy-assurance agenda.

  • Grounding a PYRAMID avionics bridge in IES and HQDM

    Open research, 2026

    PYRAMID (Def Stan 00-134) is an open avionics reference architecture with no shared data model; it pushes interoperability into inter-component bridges. A worked, open, SHACL-validated example grounds those bridges in the IES and HQDM 4D ontologies, so three PRA components (Geography, Tactical Objects, Data Fusion) that model the same object resolve to one referent. No prior work grounds PYRAMID in an upper ontology.

  • Nature-Related Security Risk: an open evidence base and systems ontology

    Open research, 2026

    A self-initiated, open, machine-readable evidence base and systems ontology that operationalises the UK National Security Assessment on biodiversity loss and ecosystem collapse (2026): 14 documented nature-to-security cascades graded with the assessment\'s own confidence ratings, an OWL ontology (NSRO), a SKOS taxonomy cross-walked to IPBES and SHACL shapes, published on GitHub under CC-BY-4.0.

  • Cyber Incidents affecting Connective Products: an open evidence base

    Open research, 2026

    A self-initiated, open, machine-readable evidence base of cyber incidents affecting connective products (IoT, operational technology, computing devices, networking equipment and software/firmware): 16 documented incidents (2022 to 2025), a SKOS taxonomy and a source-quality rubric aligned to the six Government Data Quality Framework dimensions, published on GitHub under CC-BY-4.0.

Register assurance6
  • An Open Ontology for Italian Public Registers, Tested Against IPA, ANAC, OpenCUP and ISTAT

    Open ontology, five-register measurement and upstream SHACL contribution, 2026

    An open OWL 2, SKOS and SHACL ontology and reproducible audit of Italy's public register fabric, built from hands-on downloads of the IPA register of 23,735 public administrations, the ANAC contracts register, the OpenCUP register of 11.9 million public investment projects, both flavours of the ISTAT register of administrative units and the national semantic vocabulary, taken on 17 August 2026. The national cities vocabulary and the ISTAT CSV export both predate the 2026 Sardinian territorial reform: 380 dead municipality codes are published as current and 378 current codes are missing, so joins against the semantic layer silently drop every Sardinian municipality, while the operational IPA register already uses the new codes with zero pair mismatches across all 23,735 entities. All 23,735 IPA codici fiscali pass their checksum, a clean null result stated as a finding. In a seeded two-month sample of 2025 contracts, 140 well-formed CUPs referenced by 255 contracts are absent from the published OpenCUP exports, and 45 values in the CUP field fail the CUP grammar, including ESENTE and ATT000NON000CUP. The findings are delivered upstream as a contribute-first SHACL package answering the vocabulary maintainers' own request in issue #225, including the currency test that would have caught the drift in CI. Every headline computed two independent ways, reproducible from public data.

  • An Open Ontology for UK Public Registers, Tested Against Companies House, the Charity Commission and the Global LEI System

    Open ontology, four-register measurement, 2026

    An open OWL 2, SKOS and SHACL ontology and reproducible audit of the boundary between the UK's public registers, built from complete keyless downloads of the Companies House bulk file of 5,695,465 live companies, the full Charity Commission register, the 12.9 GB PSC snapshot and the GLEIF LEI golden copy, pinned to snapshots of 1 and 17 August 2026. 1,268 of 31,612 Registered charitable companies (4.01 per cent) cite a company number that resolves to nothing on the live register, including 18 charities registered under the placeholder 00000000, 9 under 01234567 and 5 under 12345678, all of which pass the format rule and none of which resolve. 1,082 charities disagree with Companies House about their own company's name, and 662 charitable companies are in a distress state at Companies House, including 223 in liquidation, while the charity remains Registered. Of 117,324 GLEIF records naming UK Companies House as registration authority, 36 live issued LEIs cite format-valid, in-range company numbers absent from the live register, 2,666 disagree with Companies House on the name, and one in three has lapsed certification. Companies House's own bulk file carries 5,912 live companies whose numbers violate the documented format, and its PSC snapshot misspells its own enum value in 621,705 production records. Every headline computed two independent ways, reproducible from public data.

  • An Open Census of UK Operational Surveillance Reporting, and What Its Identifiers Actually Denote

    Open ontology, complete publisher census, 2026

    A companion to our census of the FSA alert register, asking the prior question: can the report that carries a surveillance finding be cited at all? Every publication held by APHA, UKHSA and Cefas on GOV.UK was censused on 18 August 2026 through the open Search and Content APIs, giving 6,897 publications indexed and 2,704 report pages carrying 12,769 published files. Two of our own hypotheses died and are reported: every page does carry a persistent content_id, and the report-level unique_reference field is populated on 38.62 per cent of files rather than never. What survives is that 10,582 files, 82.9 per cent, share their only persistent identifier with every other file on the same page, with up to 165 under one identifier, so a specific edition cannot be durably cited. Of 8,758 dated editions, 5,212 carry no report-level identifier and 242 recurring series carry none on any edition. The governance finding is that the UKHSA weekly national flu and COVID-19 surveillance series ran at 99.2 per cent identifier coverage in 2023 to 2024 and 0.7 per cent in 2025 to 2026, a range of 98.5 points across consecutive seasons of one series, which makes the practice a habit rather than a rule. A field named unique_reference is non-unique on 204 values after excluding 362 legitimate format variants. The dual-computation gate caught a 14-page duplication in our own harvest before publication.

  • The Missing Pathogen Crosswalk: An Open Ontology for One Health Biosurveillance Registers

    Open ontology, complete register census, 2026

    One Health is an integration claim, and it can be tested by asking whether a record in one surveillance register can actually be joined to a record in another. Measured across the Food Standards Agency's complete open food alert register, all 1,348 alerts published from 9 January 2018 to 10 August 2026 under the Open Government Licence, the answer for most pathogen alerts is no. The register maintains 27 allergen concepts, granular enough to distinguish macadamia from pecan and gluten-free oats from oats, and 5 pathogen concepts whose scheme declares its last modification as 4 September 2017. Of 234 alerts naming a pathogen, 57 name it in prose with no concept to join on. Pathogen coding was zero across all 188 alerts issued in 2018, including the register's first record about Salmonella in pulled pork, and has never exceeded 19.7 per cent in any year, while allergen coding has run between 47.7 and 64.6 per cent every year without exception. Four organisms appear with no concept available to code them: Bacillus cereus, hepatitis A, Cronobacter sakazakii and norovirus. No pathogen concept carries any external mapping, and all five denote at genus or species level while genomic food chain surveillance identifies isolates at serovar and sequence type, so a record coded "Salmonella spp" cannot be joined to one about Salmonella Typhimurium ST34. The study separates three failure modes with three different remedies, ships the NCBI Taxonomy crosswalk the register lacks with every taxon resolved two independent ways and zero disagreements, and computes every headline twice with a build that fails if the two paths differ.

  • The Missing CIK-to-LEI Crosswalk: An Open Ontology for US Securities Entity Registers

    Open ontology, four-register measurement, 2026

    An open OWL 2, SKOS and SHACL ontology and reproducible audit of the boundary between the SEC's EDGAR entity register and the Global LEI System, built from complete downloads of the SEC's entire entity register and every open register that embeds or is embedded by it, all taken on 17 August 2026. The lei field the register carries on every one of its 981,355 records is populated in 773 of them, and only 667 of those are valid LEIs: the remainder includes telephone numbers, IRS employer identification numbers, entity names typed into the identifier field and thirty-five literal "N/A" strings. In a single quarter 1,973 fund registrants stated their LEI to the SEC on Forms N-PORT and N-CEN, and the entity register surfaces 12 of them, while for CIK 892538 it carries the investment adviser's identity on the fund, the same defect class found at the FDIC a week earlier. GLEIF publishes the reverse crosswalk for 27,718 EDGAR-registered LEIs daily under CC0; the SEC publishes no mapping in either direction. Reconciling the two halves across 12,604 fund series finds 100 carrying two different LEIs on the two sides of the register boundary, including a checksum-failing character substitution and three ETF series holding each other's LEIs in a cycle, and 56.6 per cent of all US LEI records are LAPSED. An open OWL 2 ontology, SKOS scheme registry and SHACL governance layer, 526,098 triples, every headline computed two independent ways, reproducible from public data.

  • Register Assurance: Why Every Public Register Fails at Its Boundary

    Research programme pillar, 2026

    Public registers increasingly assure their own records, and GLEIF is the standing proof that it can be done: 99.99 average data-quality score across 3.39 million LEI records. But assurance stops at the register boundary. Nothing checks conformance when one register embeds another's identifiers, nothing verifies that embedded identifiers still resolve, and nothing reconciles what two registers claim about the same entity. Measured across six domains with one open method, the boundary fails in the same four ways every time: embedded-scheme non-conformance, resolution failure, cross-register disagreement, and missing governance metadata. All 2,252 FDIC LEIs are truncated and invalid, 19.5 per cent of active EEA insurers carry no LEI, the largest index fund is missing from the open identifier map, the registers of retraction agree on 72.42 per cent, all 67,141 dereferenced US academic-standards identifiers return 404, and zero of 300 sampled GOV.UK documents carry any maintenance commitment. The pillar names the discipline, indexes the six open ontologies as one program, and announces the Register Integrity Index.

Economy, finance & skills20
  • What 13 Million Triples Reveal About the Quality of US Federal Vocabularies

    Open register, reproducible assessment, 2026

    Consultancies sell ontology assessment, and almost none has published an assessment of a single named real-world ontology. This register does it in public against 28 vocabularies from 10 US federal publishers, retrieved and hashed on 16 August 2026 and put through 26 checks. The estate holds up better than a gotcha audit would claim: every asset retrieved, HTTPS compliance was complete, and nothing was logically inconsistent. What makes the ten real failures actionable is that every check names the authority it derives from and declares whether failing it violates a specification or departs from a convention the publisher never agreed to. The Library of Congress MADS/RDF ontology does not parse, using rdf:resource on a node element 23 times where the RDF/XML grammar forbids it, and the Thesaurus for Graphic Materials types its 7,782 concepts against it. All 507 ISO 639-2 concepts violate SKOS integrity condition S14, because the vocabulary of language codes carries three preferred labels each and no language tag on any of them. The strain in the estate is governance rather than logic: 21 of 28 assets declare no licence, 20 name no publisher, 14 carry no version. Findings are W3C EARL assertions anchored to dated hashed snapshots, publishers have a documented right of reply, and two flaws in the register's own method were caught and discarded before publication.

  • An Open Ontology for US Bank Registers, Tested Against the FDIC, the Federal Reserve, and the Global LEI System

    Open ontology, full-register measurement, 2026

    An open OWL 2, SKOS and SHACL ontology for the entity and regulatory-reporting fabric of US banking, measured against the FDIC BankFind register, the Federal Reserve's MDRM dictionary and the complete GLEIF golden copy of 3,403,760 records. Every LEI value in the FDIC BankFind register is exactly 16 characters, truncated from the 20 that ISO 17442 requires, discarding both check digits, so no value can be validated and none matches the Global LEI System as published. Resolving all of them against the golden copy shows why arithmetic cannot repair it: 16 characters puts 6.37 per cent of the entire global LEI population into a collision, and nine FDIC values denote up to six unrelated companies each, across four countries. Two values, confirmed by GLEIF's own consolidation records, identify the wrong legal entity: Associated Bank carries the LEI of Associated Banc-Corp, its parent holding company, while the bank's own active LEI appears nowhere in the register. The Federal Reserve's MDRM dictionary shows the mirror-image defect in the concept fabric: item names and definitions are pure functions of the item code across forms with different consolidation bases, and 27,445 of 47,305 item codes carry no definition at all. The artefact comprises 1,123,634 triples, with every headline computed twice, reproducible from public data.

  • Your Search Index and Your AI Pipeline Are Reading Different Corpora

    Open ontology, corpus measurement, 2026

    Every organisation is wiring its documents into an AI assistant and almost none can say whether the knowledge is any good, because the field answers that question with adjectives. We built an open OWL ontology defining authority as a property of a maintenance commitment rather than of content, ten SKOS schemes, a SHACL publish gate that exits non-zero in CI, and the Corpus Readiness Index, then ran the instrument against 54,222 GOV.UK guidance documents. Across 300 randomly sampled documents and 133 distinct schema keys, not one field expresses a review date, an owner, a maintainer or a verification date. GOV.UK search correctly returns zero withdrawn documents, while its public sitemap advertises roughly 55,000 of them still serving body text at a median 5.87 years since withdrawal, so any pipeline that crawls rather than consuming the curated index ingests exactly what that index was careful to exclude. 4,810 documents are owned only by organisations that no longer exist, including a 1955 treaty attributed to a department merged away in 2020. 63.5 per cent are unchanged in over two years and 13,595 are still tagged to the 2010 to 2015 coalition government, more than are tagged to the current one. 29.7 per cent carry under 500 characters of indexable text because the answer sits inside an attachment. Reproducible from public data, with a build report listing four defects we found in our own code.

  • Science Has No Shepard's: Measuring How Far the Registers of Retraction Disagree

    Open ontology, four-register measurement, 2026

    Legal research has known since 1873 whether an authority still stands, because Frank Shepard began printing citation labels marking which decisions had been overruled. Science has registers of retraction instead, and they do not agree. We built an open OWL ontology, SKOS status vocabulary and three-layer SHACL suite for the integrity status of scholarly works, then measured it against four public registers. Crossref and Retraction Watch, both published by Crossref since it acquired the database in September 2023, agree on only 72.42 per cent of the retracted DOIs between them. OpenAlex flags 94.5 per cent of retraction notices as retracted research, a category error confirmed independently at 95.95 per cent against Europe PMC's MEDLINE publication types, while Crossref does this for 0.91 per cent of the same notices and Europe PMC for 0.32 per cent, which shows the distinction is achievable rather than hard. The central metadata field is declared a closed 12-value enumeration on a required attribute, yet the live index holds 34 values, 22 of them invalid, including two different misspellings of retraction, a bare integer and a test string reading this_is_some_update_23. Only 19.24 per cent of a 137,243 DOI union is agreed by all four registers. 43,683 citations to the most-cited retracted works post-date their retraction, including 1,171 to the Wakefield paper retracted in 2010 for falsification, none carrying any machine-readable warning. 3.19 million triples, reproducible from public data.

  • The Content Registry Television Measurement Runs On Covers 1.84 Per Cent of Television

    Open ontology, full-corpus identifier measurement, 2026

    EIDR is the audiovisual industry's own content registry, and it models properly the distinction measurement depends on: a title is not a cut, and a cut is not an encoding, because nobody watches an abstraction and duration is what completion and frequency are computed from. In the open graph it reaches 53.44 per cent of films and 1.84 per cent of television series. We harvested 224,710 EIDR identifier assertions across 224,182 works and resolved 1,179 of them against EIDR's public registry to learn what each actually denotes. Of the 522 works carrying more than one EIDR identifier, 285 hold identifiers at more than one level of the abstraction hierarchy at once, pairing a title-level record with a specific cut as though they were interchangeable, and a further 133 identifiers are each claimed by two distinct works. There is a trap for implementers underneath all of it: EIDR runs ISO 7064 MOD 37,36 but maps the supplementary value to the digit 0 where the standard convention is an asterisk, so a validator built from the standard alone rejects 6,252 perfectly valid identifiers, one in thirty-six, as corrupt, while across all 224,577 distinct identifiers exactly zero genuinely fail their arithmetic. IMDb mints no season-level identifier at all and covers 0.78 per cent of seasons. The IAB Tech Lab taxonomies, the advertising industry's shared vocabularies for content, audience and ad product, contain no RDF, SKOS or OWL of any kind, and are published here as SKOS for the first time, 2,845 concepts, with the 10 structural defects found in the source reported rather than repaired. schema.org carries 2,987 labelled terms and not one expresses viewership, impressions, exposure or reach. 2,702,154 triples, findings computed twice and agreeing exactly, reproducible from public data.

  • Every Identifier Underneath American Academic Standards Is Dead, and the Redirect Still Works

    Open ontology, three-population census, 2026

    The Achievement Standards Network minted the persistent identifiers that United States K-12 academic standards were published under, and that the LRMI educationalAlignment pattern was designed around, so they sit inside learning-resource metadata across the open education web. We dereferenced 67,141 of them across three populations, two of which are complete censuses rather than samples. Every single one returns HTTP 404. The vocabulary that describes them still returns 200, served from a static object store, so a system checking whether the scheme is alive gets a reassuring and false answer. The dependency reaches the standards bodies: 1EdTech, publisher of the successor specification CASE, ships a QTI v3 example package whose curriculum references are dead ASN URIs, and so do DCMI's own worked LRMI examples. Alongside the census we measured what the surviving artefacts can reject. The CEDS Ontology v14 from the US Department of Education declares zero owl:ObjectProperty and zero rdfs:domain across its 2,336 properties, types 965 terms as both an owl:Class and a skos:ConceptScheme, and publishes 19,546 concepts with no broader relations at all; given six specific mis-statements it detects none, five of them being inexpressible in it, while the abandoned 465-triple ASN schema detects one. The corpus fails in both directions at once: 433 identifiers name materially different standards, one attaching four statements about controlled investigations and one about cellular respiration to the same name, while the single most replicated statement text carries 718 distinct identifiers. In the live Georgia CASE package, 2,453 of 2,483 associations are document structure and cross-framework associations number zero. 1,931,913 standard statements across 771 jurisdictions, 21,404,069 triples, reproducible from public data.

  • An Open Ontology for Insurance and Reinsurance, Tested Against Every Insurer in Europe

    Open ontology, full-register measurement, 2026

    There was no open, usable ontology for insurance and reinsurance entity data, so we built one and proved it on the hardest available evidence. An open OWL 2 ontology, two SKOS registries and a three-layer SHACL governance suite, run against the complete EIOPA Register of Insurance Undertakings (33,924 rows), joined to a same-day GLEIF harvest of all 3,630 identifiers it files, then cross-checked against BaFin's German register. 643 of 3,304 active insurers and reinsurers carry no LEI at all. Four filed values are arithmetically impossible, including a letter O where the real identifier carries a zero, which quietly disconnects a Danish insurer from the global system. 118 identifiers have lapsed, 42 name entities GLEIF says no longer exist, 283 cross-border passports outlive the authorisation they depend on, and one identifier is filed for both SCOR Global Reinsurance France and SCOR Global Reinsurance Ireland, two distinct reinsurance legal entities. The German cross-register test recovers 14 valid identifiers that the national regulator publishes and the European register does not, while confirming that where both populate the field they agree on all 344 matched undertakings: a coverage problem, not a contradiction problem. Reproducible from public data with five commands, no licensed sources.

  • The Largest Index Fund Is Missing From the Open Identifier Map: the US Fund Register as a Governance Graph

    Open ontology, full-universe measurement, 2026

    Fund data is an identifier problem before it is anything else. An open OWL 2 ontology, twenty-scheme SKOS identifier registry and SHACL governance layer, built and run against the whole public US fund universe: 2,316 registrants, 14,841 funds and 43,344 share classes joined to four quarters of Form N-CEN and all 9,119,948 GLEIF ISIN-LEI pairs, every one check-digit validated with zero failures. SEC filings are not as clean: 19 of 14,960 self-reported LEIs (0.13%) fail the same check digit. Only 12.3% of self-reported ETF fund LEIs (497 of 4,053) carry any ISIN in the open mapping, including the Vanguard 500 Index Fund, which has a valid ISIN in commercial data that GLEIF's open file does not carry, and 96.8% of register quotations (29,258 of 30,238) carry no venue field, mostly because the field is ETF-only in the source schema. Class-level ISIN resolution is enclosed behind licensed CUSIP data: 259 of 19,803 funds (1.3%) resolve from public data alone.

  • Provenance Beats Plausibility: Catching Wrong Financial Answers Without a Gold Key

    Open measurement and benchmark audit, 2026

    The failure that matters when a model reads a filing is a confidently wrong number wearing plausible provenance. FinanceBench ships 2,400 model answers carrying human correctness labels, an unused supervision set for verification. A deterministic check that never sees the gold answer recovers 57.4% of labelled errors and lifts served accuracy from 68.8% to 78.0% at 60.5% coverage, with positive recall in all sixteen configurations. Two negative results matter more: the same check against the whole filing recovers 3.6%, because a filing carries a median of 1,270 numbers against 69 on the cited page, and excusing answers reconstructible by arithmetic excuses 98.6% of them while the excused are more often wrong than the rest. The audit also finds evidence page numbers are 0-indexed and that the two shipped gold files disagree on 15 of 51 numeric cases, every one by exactly 100x.

  • FCA Consultation: Stablecoins and UK Crypto Regulation (opens in new tab)

    Financial Conduct Authority Regulatory Consultation, 2025

    Contributed expert analysis to the FCA's consultation on stablecoin regulation and the future of crypto asset oversight in the UK. Provided evidence-based commentary on regulatory frameworks, consumer protection mechanisms, and systemic risk considerations for digital assets.

  • Is Blockchain Part of the Future of Art? (opens in new tab)

    Journal of the British Blockchain Association (JBBA), Peer-Reviewed

    Peer-reviewed research exploring the intersection of distributed ledger technology and the creative industries. Examined provenance tracking, digital ownership, and the implications of blockchain for cultural asset management and intellectual property governance.

  • Welsh Government Land Valuation Research Report (opens in new tab)

    Commissioned by Welsh Government, 2025–2026

    Independent research into the feasibility of land value tax models for Wales. Combined statistical analysis of land registry data with international comparator evidence and stakeholder interviews across Welsh local authorities. Findings presented to Welsh Government officials and cited in Senedd committee proceedings.

  • What Adult Skills Reveal About Social Mobility That Qualifications Hide (opens in new tab)

    Open research, OECD Survey of Adult Skills (PIAAC) public-use data, 2026

    A reproducible analysis of educational and skills mobility in England using open OECD PIAAC data. Among adults with the same qualification, those from a higher-educated background score about 37 points higher in numeracy, so qualifications understate the advantage of social origin. The data harmonisation is published as an open, machine-readable scheme.

  • AI Skills for the UK Workforce - Skills England (opens in new tab)

    Skills England / UK Government Publication, 2025

    Tesseract Academy is cited as an AI training provider and consultancy in Skills England's official research into AI skills for the UK workforce. The publication's methodology included stakeholder workshops with 43 organisations, with Tesseract Academy contributing alongside institutions including The Alan Turing Institute and the Surrey AI Centre.

  • A live list with no memory: reconstructing the FCA scam-warning signal

    Open research, 2026

    The FCA warning list of unauthorised firms shows only the current set with no history. This open dataset reconstructs it as a monthly time series from 18,224 live warning pages: FCA warnings rose nearly four-fold from around 500 a year in 2019 to roughly 1,900 across 2022 to 2025, 18 percent flagging clone firms. Published on GitHub and Hugging Face.

  • Does structured data actually help AI read company accounts? A controlled pilot

    Open research, 2026

    A self-funded controlled pilot comparing how well an open-weights language model reads the same UK annual reports as structured XBRL facts, HTML text and PDF text. Three FY2026 filed reports, 21 tasks, 189 scored model calls. Financial extraction accuracy: XBRL 88.9%, PDF 86.7%, HTML 80.0%; only XBRL ever retrieved a bracketed negative equity figure with the correct sign, at 2.9x the token cost.

  • Every government consultation that reached an outcome, as one corpus

    Open research, 2026

    An open, structured corpus of 6,260 UK government consultations with a published outcome, harvested from the GOV.UK APIs and coded by policy area. 98 percent carry policy-area coding; only 77 percent attach response documents to their outcome. Published on GitHub and Hugging Face under the Open Government Licence.

  • What government evaluates, and how openly it says so

    Open research, 2026

    An open atlas of 1,770 UK government evaluation publications, harvested from the GOV.UK Search API and classified by evaluation type and declared method. Impact evaluations dominate, but only 11 percent declare a recognisable method in their metadata. Published on GitHub and Hugging Face under the Open Government Licence.

  • The most timely local labour-market signal, tidied for every authority

    Open research, 2026

    An open dataset of the latest monthly claimant count for all 374 local authorities in Great Britain, from the ONS via NOMIS. Claimant rate ranges from 1.1 to 10.1 percent against a 3.5 percent mean. Published on GitHub and Hugging Face under the Open Government Licence.

  • The baseline a local health survey should be read against

    Open research, 2026

    An open small-area health profile for all 331 local authorities in England and Wales, combining Census 2021 self-reported general health, disability and unpaid care. Self-reported bad health ranges from 2.7 to 9.6 percent. Published on GitHub and Hugging Face under the Open Government Licence.

Talks & community3
  • UK Government Business Academy - AI Webinar Series

    Department for Business and Trade, Business Academy, 2025 and 2026

    Delivered a series of three official UK Government Business Academy webinars on AI adoption for growing businesses, led by Dr Stylianos Kampakis. Topics covered: designing AI roadmaps, choosing the right AI tools using the OCT (Objectives-Capabilities-Tools) methodology, and building internal AI capability including skills-gap analysis and organisational models for long-term success.

    The Business Academy recommissioned the series for autumn 2026, taking the programme to six official webinars. The three 2026 sessions run free and online at 11:00 on 24 August, 24 September and 6 October 2026, covering roadmap design, tool-agnostic automation of existing workflows, and what has to change inside a team for AI adoption to hold once the pilot ends.

    2025 series

    2026 series, booking open

  • London Data Week 2026: AI Tools for Everyone, Advancing Disability Inclusion

    London Data Week, Shaw Library, LSE, 8 July 2026

    The speaker panel at Tesseract Academy's AI Tools for Everyone: Advancing Disability Inclusion event in the Shaw Library, LSE, with King's Institute for Artificial Intelligence, London Data Week and LSE Data Science Institute banners

    Hosted at the London School of Economics on Wednesday 8 July 2026, Tesseract Academy convened a public session bringing together speakers from technology, disability, research and policy to explore how AI can advance accessibility and support for people with diverse disabilities. Held in the Shaw Library, LSE, and hosted by the LSE Data Science Institute with the King's Institute for Artificial Intelligence, with speakers from Vision Ability CIC, Imperial College London, King's College London and UCL.

  • London Data Week 2025: AI Tools for the Visually Impaired (opens in new tab)

    London Data Week, co-organised with Vision Ability CIC, 2025

    Co-organised a public workshop and demonstration on making AI tools accessible to people with visual impairments. Delivered at Chabad Islington Community Centre as part of London Data Week 2025, in partnership with Vision Ability CIC.