Research programme, August 2026
Register assurance: why every public register fails at its boundary
Public registers are the reference layer of the modern economy: who is a bank, who may write insurance, which paper stands retracted, which standard a curriculum aligns to, which guidance is current. The operators of the best registers now assure their own records seriously, and GLEIF is the standing proof that it can be done. But assurance stops at the register boundary. Nothing checks conformance when one register embeds another register's identifiers, nothing verifies that the embedded identifiers still resolve, and nothing reconciles what two registers claim about the same entity. We measured that boundary with one open method across six domains, and it fails in the same four ways every time. This page names the discipline, states the four defect classes, presents the evidence, and announces the instrument.
The thesis: assurance stops at the boundary
Start with the positive example, because it defines the standard. GLEIF, the Global Legal Entity Identifier Foundation, runs a genuine data-quality programme over the Global LEI System: monthly public reports, defined quality criteria, per-issuer scoring. Its July 2026 report gives an average Total Data Quality Score of 99.99 across 3,390,204 records, and when we validated its open ISIN-to-LEI mapping file in full rather than sampling it, all 9,119,948 pairs passed check-digit arithmetic with zero failures. Inside its boundary, the Global LEI System is the best-governed identifier register we have measured anywhere.
Now look one step past the boundary. The FDIC embeds LEIs in its register of insured banks: all 2,252 of them are truncated and invalid. EIOPA embeds LEIs in its register of insurers: four are arithmetically impossible and 42 point at entities GLEIF says have ceased to exist. The SEC receives LEIs in structured filings: 19 fail the same one-line arithmetic that GLEIF passes nine million times in a row. None of this is GLEIF's fault, and that is precisely the point. GLEIF's conformance checks are scoped, correctly, to what LEI issuers submit into the system. No check anywhere in the world measures what a register publishes back out when it embeds somebody else's scheme.
Each register assures its own records. No one assures the seams. The seams are where registers cite each other, embed each other's identifiers, and make overlapping claims about the same entities, and the seams are exactly where automated systems, retrieval pipelines and AI agents now read. A defect inside a register gets caught by its operator eventually. A defect at the boundary has no operator, so it survives indefinitely, producing values that still look like identifiers and claims that still look like facts.
The four boundary defect classes
Six domains, measured with the same method, produce the same taxonomy. Every boundary defect we have found falls into one of four classes, and every class appeared in more than one domain, which is what justifies stating them as classes rather than anecdotes.
1. Embedded-scheme non-conformance
A register publishes values under another register’s scheme that violate that scheme’s own structural rules, and no check on either side notices.
Measured instances: All 2,252 FDIC LEIs truncated to 16 characters. Four EIOPA LEI values that fail the ISO 7064 check digits, one of them twenty zeros. 19 SEC N-CEN LEIs failing the same arithmetic. Five BaFin values, three of them 19 characters.
2. Resolution failure
An embedded identifier refers to nothing: the register still carries it, the target scheme no longer answers for it, and every consumer stores a dead pointer.
Measured instances: All 67,141 dereferenced ASN standards identifiers return HTTP 404 while their describing vocabulary still returns 200. 16 FDIC values compatible with no LEI on earth. The Vanguard 500 Index Fund absent from GLEIF’s open ISIN-to-LEI map.
3. Cross-register disagreement
Two registers make incompatible claims about the same entity or work, and no system reconciles them, so downstream consumers inherit whichever register they happened to read.
Measured instances: Crossref and Retraction Watch, both published by Crossref, agree on 72.42 per cent of retracted DOIs. 42 EEA insurers authorised in EIOPA while GLEIF marks the entity inactive. Associated Bank carrying its parent holding company’s LEI, proven wrong by GLEIF’s own consolidation records.
4. Missing governance metadata
The register carries no machine-readable statement of who maintains a record, on what cadence, or when it was last verified, so staleness is undetectable by construction.
Measured instances: Zero of 300 sampled GOV.UK documents, across 133 distinct schema keys, carry any maintenance-commitment field. 27,445 of 47,305 MDRM item codes carry no definition at all. 643 of 3,304 active EEA insurers carry no LEI, twelve years after guidelines asked for one.
The evidence, one domain at a time
Each study below is an open OWL 2 ontology with SKOS registries and a SHACL governance layer, built against the complete public register fabric of its domain, with every headline computed twice by independent implementations and a build report that records what could not be obtained. Each links its full write-up and its repository.
Banking, United States
The FDIC publishes 2,252 LEIs and not one is a valid LEI
Every LEI value in the FDIC BankFind register is truncated to 16 of the 20 characters ISO 17442 requires, discarding both check digits. Measured against the complete GLEIF golden copy of 3,403,760 records, 16 characters puts 6.37 per cent of the global LEI population into a collision, nine FDIC values are ambiguous across up to six unrelated companies, and two resolve to the wrong legal entity, including Associated Bank carrying the LEI of its parent holding company.
Insurance, European Union
643 of 3,304 active EEA insurers carry no LEI, and four carry impossible ones
The EIOPA Register of Insurance Undertakings, joined to a same-day GLEIF harvest and cross-checked against the German national register: 19.5 per cent of active insurers have no global identifier, four filed values fail the check-digit arithmetic, 42 name entities GLEIF says no longer exist, one identifier is shared by two distinct SCOR reinsurance companies, and 283 cross-border passports outlive the authorisation they depend on. Where both registers populate the field they agree perfectly, so the problem is coverage, not contradiction.
Investment funds, United States
GLEIF is checksum-clean across 9.1 million pairs; the register boundary is not
All 9,119,948 pairs in GLEIF’s open ISIN-to-LEI file pass check-digit validation with zero failures, while hand-keyed SEC N-CEN filings carry 19 LEIs that fail the same arithmetic. Only 12.3 per cent of self-reported ETF fund LEIs have any ISIN in the open mapping, and the Vanguard 500 Index Fund, the largest index fund, is among the missing: it holds a valid ISIN in commercial data that the open file simply does not carry.
Scholarly record
The registers of retraction agree on 72.42 per cent, and both are published by one organisation
Crossref and the Retraction Watch database, both published by Crossref since its 2023 acquisition, agree on only 72.42 per cent of the retracted DOIs between them. OpenAlex flags 94.5 per cent of retraction notices as retracted research, a category error confirmed independently against Europe PMC. Only 19.24 per cent of a 137,243 DOI union is agreed by all four registers measured.
Academic standards, United States
67,141 identifiers dereferenced, every single one returns 404
The Achievement Standards Network was the identifier layer for US K-12 academic standards, embedded in learning-resource metadata across the open education web. Three censuses, two of them complete: every one of 67,141 dereferenced identifiers returns HTTP 404, while the vocabulary describing them still returns 200 from a static object store, so a liveness check on the scheme gets a reassuring and false answer. The successor bodies’ own example packages still ship the dead references.
Enterprise knowledge, United Kingdom
Zero of 300 sampled documents carry any maintenance commitment
Across 54,222 GOV.UK guidance documents, a random sample of 300 fetched in full exposes 133 distinct schema keys, and not one field expresses who maintains the content, on what cadence, or when it was last verified. The public sitemap advertises roughly 55,000 withdrawn pages that search correctly excludes but that still serve body text, so a retrieval pipeline that crawls ingests exactly what the curated index was careful to exclude.
Credit where it is due, and what is actually new here
Three attributions need making plainly, because a category is only worth naming if its parts are honestly sourced.
GLEIF owns the positive example. Its data-quality programme is the existence proof that a register operator can assure its records rigorously and publicly, and everything on this page treats it as the standard the rest of the fabric should be held to, not as a target.
Jodi Schneider and colleagues own the result that the registers of scientific retraction disagree. Salami, McCumber and Schneider (2026) compare eleven sources across a union of 83,317 items and find that 92.64 per cent carry some disagreement. Our 72.42 per cent Crossref against Retraction Watch figure is a two-source special case of their programme, and anyone citing our scholarly study should cite theirs.
K-AI, a document knowledge platform, named corpus readiness first, in a May 2026 piece arguing that AI readiness frameworks omit unstructured documents. Our enterprise knowledge study credits them as such; what we added is the scoring method, the implementation and the measured study.
What is ours is the cross-domain generalisation, the boundary framing, and the instrument. Individual defects in individual registers have been reported before, and prior ontologies exist in several of these domains, credited in each study. What we have not found anywhere is the observation that the same four defect classes recur at the boundary of every public register fabric measured, banking, insurance, funds, scholarship, education and government knowledge alike, nor a method that makes the boundary checkable: identifiers as reified assertions carrying scheme, source and computed validation state, disagreement recorded as data rather than resolved silently, and every check shipped as executable SHACL.
Announcing the Register Integrity Index
The six studies were built one register fabric at a time. The Register Integrity Index generalises them into a single open instrument: a reproducible scoring framework that measures any public register against the four boundary defect classes, embedded-scheme conformance, resolution, cross-register agreement, and governance metadata, and returns graded findings with the queries attached. It is being stood up now, building directly on the ontologies, SHACL suites and pipelines already published, and it will be developed in the open like everything else in this programme.
The ontology family: one program, six instruments
The six ontologies are one program, not six projects. They share a design stance: identifiers are reified assertions rather than string attributes, scope is declared as data, lifecycles are nodes, disagreement is recorded rather than adjudicated, and arithmetic is computed in code while policy is enforced in shapes. Four publish their terms under the https://gov.tesseract.academy/def/ namespace root, and two publish under sibling Tesseract Academy hosts as recorded in their repositories. Each namespace below is taken from the repository's own Turtle source, not from documentation.
| Ontology | Namespace | What it models |
|---|---|---|
| Bank Register Ontology (BRO) | https://gov.tesseract.academy/def/banking#https://gov.tesseract.academy/def/banking/scheme# | The entity fabric and regulatory-reporting concept fabric of US banking: FDIC BankFind, the GLEIF golden copy and the Federal Reserve MDRM, with identifiers as reified, validated assertions. |
| Insurance Register Ontology (IRO) | https://gov.tesseract.academy/def/insurance#https://gov.tesseract.academy/def/insurance/scheme# | Insurance and reinsurance undertakings, their authorisations, cross-border passports and identifier fabric, instantiated against the EIOPA register joined with GLEIF. |
| Enterprise Knowledge Ontology (EKO) | https://gov.tesseract.academy/def/knowledge#https://gov.tesseract.academy/def/knowledge/scheme# | The distinction between a working document and a maintained knowledge asset, with the SHACL publish gate and the Corpus Readiness Index that make it testable. |
| Investment Fund Ontology (IFO) | https://gov.tesseract.academy/def/fund#https://gov.tesseract.academy/def/fund/scheme# | The registered fund product hierarchy, registrant to fund to share class to listing, with a twenty-scheme identifier registry, instantiated against the full public US fund universe. |
| Scholarly Record Ontology (SRO) | https://ontology.tesseract.academy/sro/ | The integrity status of scholarly works, retractions, corrections and expressions of concern, as claims made by named registers that can and do disagree. |
| Learning Standards Ontology (LSO) | https://learning.tesseract.academy/lso#https://learning.tesseract.academy/lso/scheme/ | Academic standards, the identifiers that name them and the alignment claims made about them, built against 1.93 million standard statements and the census of the dead ASN identifier layer. |
All six are open: ontologies and shapes under CC BY 4.0, code under MIT or equivalent, every headline reproducible from public data via the pipelines in each repository.
Common questions
What is register assurance?
Register assurance is the discipline of verifying public registers at their boundaries rather than only inside them. Registers increasingly assure their own records: GLEIF, the best example, runs a monthly data-quality programme over the Global LEI System. What no register operator currently does is check conformance where registers meet: whether the identifiers a register embeds from another scheme are structurally valid, whether they still resolve, whether two registers agree about the same entity, and whether the register declares who maintains each record and when it was last verified. Register assurance names that gap, defines the four defect classes found at the boundary, and provides an open instrument for measuring them.
Which public registers publish invalid identifiers?
Measured with the same open method in 2026: the FDIC BankFind register publishes 2,252 LEI values and every one is truncated to 16 of the required 20 characters, so none is valid under ISO 17442. The EIOPA Register of Insurance Undertakings carries four LEI values that fail the ISO 7064 check digits and so cannot exist, and the German BaFin register carries five more, three of them 19 characters long. SEC Form N-CEN filings carry 19 LEIs that fail the same arithmetic. In education, all 67,141 dereferenced Achievement Standards Network identifiers, the identifier layer of US K-12 academic standards, return HTTP 404. Each measurement is reproducible from public data via the linked repositories.
Does GLEIF check the quality of LEI data?
Yes, and it does so well. GLEIF publishes monthly data-quality reports over the Global LEI System, and its July 2026 report gives an average Total Data Quality Score of 99.99 across 3,390,204 records. Checked in full rather than sampled, all 9,119,948 pairs in GLEIF’s open ISIN-to-LEI mapping file pass check-digit validation with zero failures. The limitation is scope, not rigour: GLEIF’s conformance checks apply to what LEI issuers submit into the system. Nothing in the Global LEI System constrains what a downstream register, such as a national banking or insurance regulator, publishes back out, which is exactly where the defects measured across this programme sit.
What are the four boundary defect classes?
First, embedded-scheme non-conformance: a register publishes values under another scheme that violate that scheme’s own rules, such as the FDIC’s 2,252 truncated LEIs or EIOPA’s four arithmetically impossible ones. Second, resolution failure: an embedded identifier refers to nothing, such as the 67,141 dead ASN identifiers or the 16 FDIC values matching no LEI on earth. Third, cross-register disagreement: two registers make incompatible claims about the same entity, such as Crossref and Retraction Watch agreeing on only 72.42 per cent of retracted DOIs, or 42 EEA insurers shown as authorised while GLEIF marks the entity inactive. Fourth, missing governance metadata: the register declares no owner, review cadence or verification date for its records, such as zero of 300 sampled GOV.UK documents carrying any maintenance-commitment field.
What is the Register Integrity Index?
The Register Integrity Index is the instrument that generalises these six studies: an open, reproducible scoring framework that measures any public register against the four boundary defect classes, embedded-scheme conformance, resolution, cross-register agreement, and governance metadata, and returns graded findings rather than adjectives. It is being stood up at github.com/fabio-rovai/register-integrity-index, building on the ontologies, SHACL suites and pipelines already published for the banking, insurance, fund, scholarly, learning-standards and enterprise-knowledge domains.
How is register assurance different from ordinary data quality management?
Data quality management operates inside one system’s boundary: completeness, validity and consistency of the records that system owns. Register assurance operates at the seams between systems, where no single operator has authority and therefore nobody measures. The defects it finds are invisible to each register’s own quality programme: the FDIC truncation survived indefinitely because every value still looked like an identifier, GLEIF scores 99.99 because its checks correctly stop at what issuers submit, and both are right by their own scope. The practical differences are that identifiers are treated as assertions to be validated rather than as attributes to be trusted, that disagreement between sources is recorded as data rather than silently resolved, and that the checks are shipped as executable SHACL rather than as a procedures manual.
Working with us
Everything on this page is a worked example you can run today. Each of the six repositories contains the ontology, the shapes, the harvesting pipeline and the queries that produced its headlines, and each study page walks through the method in enough detail to apply it without us. If you operate a register, the constructive reading of this page is that the boundary checks are cheap: check-digit validation is one modulo operation, resolution is a lookup, and a SHACL shape enforcing your embedded schemes runs on every ingest.
If you run entity data, reference data, or a corpus that embeds identifiers from registers you do not control, a typical first engagement is bounded and diagnostic: we build the identifier fabric over your data as it is, run the four defect classes against it, and return a graded findings report with the queries attached so your team can re-run it. That report is useful whether or not the work continues. Write to fabio@thetesseractacademy.com with a description of the registers your systems depend on, and we will tell you which of the four defect classes we would expect to find, before any commitment.
Related work: the methodological foundations are in our machine-validated open ontologies study and the open-world hole benchmark, which measure when an ontology can actually reject a wrong statement.
