Open research, August 2026
The FDIC publishes 2,252 Legal Entity Identifiers. Not one of them is a valid LEI.
Every figure on this page is from a build of 16 August 2026 and is reproducible from the open repository. We pulled all 27,836 institution records from the FDIC's BankFind register and measured the length of every LEI in it. The histogram has exactly one bucket: 16 characters, 2,252 times, without a single exception, truncated by four characters from the twenty the standard requires. We then resolved every value against the complete Global LEI System, found nine that denote up to six unrelated companies each, and found two, confirmed by GLEIF's own consolidation records, that identify the wrong legal entity entirely. This page explains what was measured, why the four missing characters are the four that matter, and what the same census says about the global identifier system itself.
The short version
- The finding: all 2,252 LEI values in the FDIC BankFind register are truncated to 16 characters, discarding both ISO 7064 check digits and two entity-identifying characters. JPMorgan Chase Bank appears as
7H6GLXDRUGQFU57Ragainst a real7H6GLXDRUGQFU57RNE97. - The census: truncating all 3,403,760 LEIs in the GLEIF golden copy to 16 characters puts 216,965 of them, 6.37 per cent of the global population, into a collision. Nine FDIC values fall into exactly that hole.
- The case that matters: Associated Bank, National Association carries an identifier that completes to exactly one real LEI, and that LEI belongs to Associated Banc-Corp, its parent holding company. GLEIF's Level 2 relationship file confirms the parent-child link independently.
- The mirror image: the Federal Reserve's MDRM dictionary attaches one definition to each item code across forms with different consolidation bases, and 58 per cent of item codes carry no definition at all. Scope is never in the identifier, in either fabric.
- The artefact: an open OWL 2 ontology, SKOS registries and a SHACL governance layer, 1,123,634 triples, code MIT, ontology and shapes CC BY 4.0, reproducible from public data.
Why this becomes a live problem on 1 October 2026
The Financial Data Transparency Act Joint Data Standards final rule (91 FR 38246, published 25 June 2026, effective 1 October 2026) commits nine federal financial regulators, the FDIC and the Federal Reserve among them, to a single answer for how a legal entity is identified. In its codified text, the legal entity identifier is established to be ISO 17442, the LEI.
The LEI is twenty characters. Four identify the issuing operating unit, two are reserved, twelve identify the entity, and the last two are check digits computed under ISO 7064 MOD 97-10. Those last two are the reason the LEI exists in the form it does. They let any system, anywhere, decide whether a value it has just been handed could possibly be real, without asking anyone.
Precision matters about what this is and is not. The FDTA rule binds the agencies, not banks, and its own DATES section says it will not change any reporting requirements without further action by the agencies. The agency-specific rules that will have teeth are due around mid-2028 and, as of today, not one has been proposed. So nobody is in breach of anything. What follows is a baseline measurement taken before the implementing rules are written. That framing matters, because the rule's preamble expressly preserves the practices that let identifier quality rot: a reporting entity need only report an LEI if it has one, lapsed LEIs may still be reported, and agencies may adopt other standards entirely. Those tolerances are the mechanism by which today's defects survive into the post-2028 regime unless somebody measures them first.
There is also a sharp irony sitting in existing law. Under 12 CFR 1003.4(a)(1)(i), every HMDA Universal Loan Identifier begins with the reporting institution's LEI and ends with a check digit computed under ISO 7064 MOD 97-10. HMDA cannot function on a truncated LEI. The same federal government requires twenty characters in one collection and publishes sixteen in another.
Why the missing four characters are the four that matter
Two of the discarded characters are check digits, and two are part of the entity portion. Losing the check digits means the value can no longer be validated in isolation: the cheapest control in the entire identifier system, one modulo operation with no network call, no licence and no lookup table, has been removed from every record. Losing two entity characters means the damage cannot be undone by arithmetic. Given sixteen characters there are thirty-six possibilities for character seventeen and thirty-six for character eighteen, so 1,296 candidates, and only then do the check digits become determined. You cannot repair a truncated LEI by recomputing anything. You can only look it up.
So we looked all of them up, against the entire Global LEI System. We took the GLEIF golden copy published at 08:00 on 16 August 2026, all 3,403,760 LEI records, and truncated every one to sixteen characters to see how often the result still identifies something unique. It collapses 3,403,760 identifiers into 3,220,636 distinct prefixes. 33,841 of those prefixes are shared by more than one legal entity, putting 216,965 LEIs, 6.37 per cent of the entire global population, into a collision. The worst single prefix is shared by one hundred distinct entities. That is a census, not a projection: six per cent of the world's legal entity identifiers are not uniquely recoverable from sixteen characters.
Nine of the FDIC's own values fall into exactly that hole. Every one was issued under the LOU prefix 894500, which allocates sequentially, so consecutive registrations differ only in their final characters. The result is that 894500C8YTS0IB1B, as published in a United States federal banking register, denotes any of Waterfall Bank in the United States, a Bulgarian agricultural company, Societe du Passage Agard in France, and three Indian companies. Madison County Bank shares its published identifier with a Canadian bond fund and a Guernsey limited partnership. The community banks that used the cheapest registrar are precisely the ones whose federally published identifier now points at half a dozen unrelated companies.
The findings, graded
| Finding | Number | Class | Operational consequence |
|---|---|---|---|
| LEI values in the FDIC register truncated to 16 of the 20 characters ISO 17442 requires, discarding both check digits. The length histogram has exactly one bucket | all 2,252 (100%) | defect | No value the FDIC publishes can be validated in isolation, and none matches the Global LEI System as published. Every downstream join keyed on the LEI fails or guesses. |
| LEIs across the entire Global LEI System that stop being uniquely recoverable when truncated to 16 characters: 33,841 prefixes shared by more than one legal entity, worst prefix shared by 100 | 216,965 of 3,403,760 (6.37%) | signal | A census, not a projection. Sixteen characters is demonstrably insufficient as a global entity key, so the FDIC defect is not repairable by arithmetic, only by lookup. |
| FDIC values that are consequently ambiguous, compatible with more than one real LEI. All were issued under LOU prefix 894500, which allocates sequentially | 9 | defect | A value published in a US federal banking register denotes up to six unrelated companies across four countries. The community banks that used the cheapest registrar are hit precisely. |
| Values matching no LEI in the global system at all after case normalisation, and values containing lowercase characters that ISO 17442 forbids | 16 unresolvable, 12 lowercase | defect | The 16 refer to nothing. The 12 all resolve once uppercased, which makes them the cheapest fix in this study, and until then any case-sensitive join fails silently. |
| Institutions whose published identifier resolves to a different legal entity, confirmed against GLEIF’s own Level 2 consolidation records | 2 confirmed, 15 parent-shaped | defect | Associated Bank carries the LEI of Associated Banc-Corp, its parent. Consolidated and bank-only reporting then attribute to one legal person whatever trusts the register. |
| Active insured institutions holding a lapsed, retired or duplicate LEI, including Santander Bank, N.A. GLEIF’s July 2026 report puts 35.0% of all LEIs worldwide in a lapsed state, so 16.5% among FDIC-insured banks is materially better than the global population | 283 | signal | Lapsed records stop being refreshed, so attributes read from them age silently while the bank trades normally. |
| MDRM item codes carrying no definition anywhere in the dictionary. Keyed on the eight-character MDRM identifier instead, the figure is 33,281 of 75,264; both denominators are stated because they move the headline | 27,445 of 47,305 (58.0%) | gap | More than half the official glossary of US bank regulatory reporting is a label with nothing behind it. |
| Item codes confidential on one form and public on another; identifiers whose confidentiality changes across their own date ranges | 1,816 / 1,342 | signal | Disclosure status is scoped by form and period and is not derivable from the item code, which is exactly what a metadata layer keyed on the item code would assume. |
All figures are computed from the 16 August 2026 FDIC fetch, the GLEIF golden copy published at 08:00 that day, and the MDRM file rebuilt at 04:00 that day. All three are living systems, so a later run produces different totals while the method reproduces exactly.
The one that made us stop
Once each value resolves to a real LEI you can ask a better question than whether it is well formed. You can ask whether it identifies the right company. Mostly it does: on 2,173 institutions the resolved name matches the bank. On two, GLEIF's own consolidation records prove it does not, and one of those two is worth the whole exercise.
FDIC certificate 5296 is Associated Bank, National Association, of Green Bay, Wisconsin. Its record carries the LEI value 549300N3CIN473IW, which completes to exactly one real identifier, 549300N3CIN473IW5094. That identifier belongs to Associated Banc-Corp, of Madison. The holding company. The parent. The bank's own LEI is ZF85QS7OXKPBG52R7N18, registered in Green Bay, status ACTIVE. It appears nowhere in the FDIC's register. GLEIF's Level 2 relationship file confirms the parent-child link independently, which is how this was detected rather than guessed.
Fifteen further institutions resolve to names ending in BANCORP, BANCSHARES or HOLDING, including SoFi Bank resolving to Golden Pacific Bancorp. GLEIF publishes no active consolidation edge for those fifteen, so we do not count them as confirmed. They are the same shape.
This is not a clerical curiosity. A bank and its holding company file different reports on different consolidation bases: FR Y-9C is the holding company consolidated, FFIEC 041 is the bank alone. If the identifier meant to distinguish them points at the same entity, then consolidated and unconsolidated positions are attributed to one legal person by any system that trusts the register. The same record shows a second kind of drift. The FDIC reports Associated Bank at $45.5bn with a reporting date of 31 March 2026. Associated Banc-Corp's acquisition of American National Corporation closed on 1 April 2026, the next day, and the company now reports roughly $52 billion. One record, carrying the parent's identifier and a balance sheet one acquisition out of date.
One further case needs no name judgement at all. CBI Bank & Trust, an Iowa institution, resolves to Herausgebergemeinschaft Wertpapier-Mitteilungen, a German company. The country code settles it.
Fairness requires one more number. 283 active insured institutions hold a lapsed, retired or duplicate LEI, including Santander Bank, N.A. That sounds worse than it is: GLEIF's July 2026 report puts 35.0 per cent of all LEIs worldwide in a lapsed state, so 16.5 per cent among FDIC-insured banks is materially better than the global population.
Where the responsibility actually sits, and where it does not
It would be easy and wrong to make this a story about GLEIF failing. GLEIF's July 2026 data quality report gives an average Total Data Quality Score of 99.99 across 3,390,204 records, and its conformance check on ISO 17442 structure is scoped to what LEI issuers submit. That is the whole point. GLEIF polices what issuers put in. Nothing in the Global LEI System constrains what a regulator publishes back out. No check anywhere measures the length of an LEI as republished downstream. That gap is not anyone's fault in particular, and it is exactly where 2,252 truncated identifiers have been sitting.
Which makes one open GitHub issue worth reading. On 25 January 2022, Pete Rivett filed issue #218 against GLEIF's own RDF repository, proposing SHACL for constraints rather than OWL restrictions, on the grounds that closed-world constraint execution suits the GLEIF dataset. It has zero comments, four and a half years later. Issue #91, asking for invalid sample data for validation testing, has been open since January 2019. Had either been actioned, a shape enforcing twenty characters and the check digits would have caught this on any ingest.
The same failure, one layer up: the MDRM
While the registers were open we pulled the Federal Reserve's MDRM, the published dictionary of regulatory reporting items: 87,702 rows, 47,305 item codes, 181 reporting forms. It is the closest thing American banking has to an official business glossary, and its mnemonics are the element-naming backbone of the FFIEC Call Report XBRL taxonomy, so whatever is true of it is inherited by every Call Report filed.
We expected the classic glossary pathology, one code carrying different definitions on different forms. We were wrong, and the truth is more interesting. The Item Name and the Description are pure functions of the four-digit item code. Byte for byte identical, on every form, in every period, with zero divergence across all 47,305 codes. Item code 2170, TOTAL ASSETS, carries one definition across 117 mnemonics and 64 reporting forms.
The dictionary achieves perfect consistency by being unable to represent the difference. Total assets on FR Y-9C is a consolidated holding company figure. On FFIEC 041 it is the bank alone. On FR Y-14A it is a projection under stress. The MDRM says the same sentence for all three, and where it does acknowledge that forms differ it does so in free text inside the definition, under a heading marked COMPARABILITY. 1,278 concepts carry their scope that way, as prose. Meanwhile 27,445 of the 47,305 item codes carry no definition anywhere at all; they are labels. That denominator is the four-digit item code. Keyed instead on the eight-character MDRM identifier, mnemonic plus code, the figure is 33,281 of 75,264, or 44.2 per cent. Both are worth stating, because the choice of key moves the headline and a reader is entitled to know which one is in use. Keyed instead on the full eight-character MDRM identifier the figure is 33,281 of 75,264, and both are worth stating, because the denominator moves the headline.
And where meaning cannot vary, the facets do. 1,816 item codes are confidential on one form and public on another. More pointedly, 1,342 MDRM identifiers change their confidentiality across their own date ranges. Disclosure status is scoped by form and by period, and is not derivable from the item code, which is exactly what a metadata layer keyed on the item code would assume.
One problem, twice: what the ontology does about it
Put the two halves together and they are the same sentence. In the entity fabric, scope was thrown away when the identifier was truncated, then attached to the wrong level of the corporate hierarchy. In the concept fabric, scope was never encoded, because the dictionary attaches meaning to a code spanning sixty-four forms with different consolidation bases. Scope is never in the identifier. Both public registers embody it.
So the ontology is shaped by that rather than around it. Identifiers are reified assertions, not designators, which is the only way to state that a published value is truncated, ambiguous, or belongs to somebody else. Resolution is data: an assertion records what it resolved to and how many candidates it resolved to. The insured institution and the legal entity are separate classes, which is what makes the Associated Bank case statable rather than merely wrong. On the reporting side, confidentiality and item type hang off a reified usage of a concept on a form in a period, because that is where the source data shows them varying.
The result is 1,123,634 triples gated by SHACL. The shapes are the interesting part: each encodes a defect class found in the sources, so the shape file doubles as the specification of what went wrong. Running pyshacl over the graph rediscovers, from the recorded state alone, all 1,733 truncated values carried by active institutions, 283 lapsed registrations, 9 ambiguities, 3 unresolvable values and 2 wrong-entity assertions. The governance report computes every headline twice, once set-based from source and once through the shipped SPARQL, and refuses to write itself if the two disagree.
What we are not claiming
We are not the first to put US bank regulatory data into OWL. Jürgen Ziemer's FinRegOnt has published the FFIEC 031 Call Report transliterated into OWL, in FIBO namespaces with real filings keyed on FDIC certificate number, since 2017. The Office of Financial Research modelled NIC ownership and control in OWL 2 in 2018 (Fan and Flood, Staff Discussion Paper 18-01), though no artifact was ever released. Nor does FIBO ignore this territory: it declares the RSSD identifier, the FDIC certificate number, the bank holding company, the LEI, registry lifecycle states, and the complete 43-code NIC entity-type vocabulary. Anyone claiming otherwise has not looked.
What none of them do is validate anything. Measured against FIBO's current master: zero SHACL shapes across 295 files, zero xsd:pattern on the LEI class, zero SKOS concept schemes. FIBO asserts ISO 17442 conformance in prose and carries no length or check-digit constraint, so a FIBO graph cannot detect that a sixteen-character string is not an LEI. Neither can FinRegOnt, whose own documentation notes that all joins are simple text label comparisons as in the original source.
That is the narrow, defensible contribution: the first SHACL validation in this domain, the first SKOS registries carrying scheme rules as data, the first RDF rendering of MDRM and of BankFind, and the first computed disagreement between a US banking register and the Global LEI System. On novelty of the findings themselves we will say only what we can defend: no prior report of either the truncation or the collision census was found in a targeted search of GLEIF's publications, the LEI Regulatory Oversight Committee, FSB progress reports, OFR working papers, arXiv and GitHub. SSRN and Google Scholar were not searched.
Two gaps are worth naming. The FFIEC's National Information Center bulk download, which holds the Federal Reserve's own view of who owns whom, sits behind a CAPTCHA and could not be fetched, so the comparison we most wanted to make is not in this build. The HMDA panel file, a genuine published crosswalk between LEIs and RSSDs, sits on a host that is edge-blocked. Both are recorded in the build report rather than papered over.
None of this is evidence of incompetence
A sixteen-character column is an ordinary engineering decision that became wrong when an external standard specified twenty, and it survived because nothing downstream ever failed loudly. That is how identifier defects always survive: they produce values that still look like identifiers. The fixes are correspondingly ordinary. Widen the field to twenty characters. Uppercase on ingest, which repairs twelve records immediately. Validate the check digits at the point of entry, one modulo operation. Record the identifier of the insured institution rather than of whichever entity in the group was to hand.
We are sending these findings to the FDIC's Chief Data Officer and to GLEIF, because a coverage result about the Global LEI System belongs with the people who run it, and a defect in a federal register belongs with its publisher. The repository, pipeline, shapes and queries are open, so anyone can re-run the census on tomorrow's golden copy and get tomorrow's numbers.
Common questions
Does the FDIC publish valid Legal Entity Identifiers?
No. As of the 16 August 2026 build measured in this study, all 2,252 LEI values in the FDIC BankFind register are exactly 16 characters long, truncated from the 20 characters ISO 17442 requires. The four discarded characters include both ISO 7064 MOD 97-10 check digits, so no value the FDIC publishes can be validated as an LEI, and none matches any record in the Global LEI System as published. JPMorgan Chase Bank appears as 7H6GLXDRUGQFU57R where its actual LEI is 7H6GLXDRUGQFU57RNE97. This is a field-width defect, uniform across the register, and it is repairable by widening the field and re-ingesting the full identifiers.
Can a truncated 16-character LEI be recovered by computation?
Not by arithmetic, and not always by lookup. Two of the four missing characters are entity-identifying, so a 16-character value has 1,296 arithmetically possible completions before the check digits are determined. Measured against the complete GLEIF golden copy of 3,403,760 LEI records, truncating every LEI to 16 characters leaves 33,841 prefixes shared by more than one legal entity, putting 216,965 LEIs, 6.37 per cent of the global population, into a collision. Nine values published by the FDIC fall into exactly that hole: one of them, 894500C8YTS0IB1B, is compatible with a US bank, a Bulgarian agricultural firm, a French property company and three companies in India.
Is the FDIC in breach of the Financial Data Transparency Act?
No. The FDTA Joint Data Standards final rule (91 FR 38246, effective 1 October 2026) establishes ISO 17442, the LEI, as the legal entity identifier standard for nine federal financial regulators, but its own DATES section states that it changes no reporting requirements without further agency action, and the agency-specific implementing rules expected around mid-2028 have not yet been proposed. Nobody is in breach of anything. This study is a baseline measurement taken before the implementing rules are written, which is exactly when such a measurement is most useful.
Is there an open ontology for US bank registers?
The Bank Register Ontology published at github.com/fabio-rovai/bank-register-ontology is an open OWL 2, SKOS and SHACL artefact for the entity fabric and regulatory-reporting concept fabric of US banking, built against the FDIC BankFind register, the GLEIF golden copy and the Federal Reserve MDRM dictionary. Prior art exists and is credited: FinRegOnt has published the FFIEC 031 Call Report in OWL since 2017, and the Office of Financial Research modelled NIC ownership in OWL 2 in 2018. What none of the prior artefacts do is validate anything: FIBO carries zero SHACL shapes across 295 files and no length or check-digit constraint on its LEI class, so a FIBO graph cannot detect that a 16-character string is not an LEI.
How do you validate a Legal Entity Identifier?
An LEI is 20 characters under ISO 17442, ending in two check digits computed under ISO 7064 MOD 97-10. Expand each character to its base-36 numeric value, concatenate, and confirm the resulting integer modulo 97 equals 1. The check is one modulo operation, needs no network call, no licence and no lookup table, and it is the cheapest control in the entire identifier system. Truncating an LEI to 16 characters removes both check digits, so every value the FDIC publishes has had that control amputated.
What is the MDRM and what does this study find in it?
The MDRM is the Federal Reserve’s published dictionary of regulatory reporting items: 87,702 rows, 47,305 item codes, 181 reporting forms, and the element-naming backbone of the FFIEC Call Report XBRL taxonomy. The study finds that its Item Name and Description are pure functions of the four-digit item code, byte for byte identical on every form and period, so the dictionary cannot express that total assets on FR Y-9C (consolidated holding company) has different scope from FFIEC 041 (bank only) or FR Y-14A (projected under stress). 27,445 of 47,305 item codes carry no definition at all, 1,816 are confidential on one form and public on another, and 1,342 identifiers change confidentiality across their own date ranges.
Open, reproducible, and free to use
The ontology, the SKOS registries, the SHACL suite, the harvesters, the resolution pipeline, the collision census and the query library are public: code under MIT, ontology and shapes under CC BY 4.0. A worked example modelling the Associated Bank case end to end ships in the repository, so you can see exactly how a wrong-entity assertion is stated, validated and queried. The build report lists what could not be obtained as carefully as what could.
Working with us
If your organisation runs on bank entity data, regulatory reporting lineage, counterparty resolution, holding-company hierarchies, or a business glossary that has to reconcile to what the regulator actually publishes, this repository is the open baseline of that discipline. The public version took days against three sources; the private version is the same method applied to the systems inside your firm that disagree with each other in ways nobody has measured yet.
A typical first engagement is bounded and diagnostic: take your entity and reference data as it is, build the identifier fabric, run the validation layers, and return a graded findings report that separates the impossible from the missing from the merely stale, with the queries attached so your team can re-run it. For the applied version on your own data, write to fabio@thetesseractacademy.com.
Related work: this is the same register-boundary discipline proved on the US fund universe and on every insurer in Europe, and the category it belongs to is set out in Register assurance: why every public register fails at its boundary.
