Skip to main content
Back to research

Medicinal product registers, measured for joinability

28 August 2026

ISO IDMP exists so that a medicine, its substance and its holder can be identified the same way everywhere. The pharmaceutical industry is spending heavily to comply, and regulators run master data services built for the purpose. On 28 August 2026 we measured what the public can actually join today, across the three open surfaces: the EMA medicines report, the openFDA NDC directory, and the GLEIF legal entity register. The EU register publishes its holders as free text. A quarter of them join to GLEIF by exact name. Fourteen of the joined identities land on an LEI that is not in good standing. The EU and US registers join to each other for one holder in twelve. And SPOR, the service that should carry the join, answers 401 to the public. Every number is computed two independent ways, and the pipeline fails its own build if they disagree.

The short version

  • The EMA medicines report publishes no organisation identifier for marketing authorisation holders. 392 distinct free-text names cover the 1,567 authorised human medicines; 13 of them are spelling variants of each other.
  • Joined to GLEIF by exact legal name (trivial normalisation allowed, no fuzzy matching), 102 of 392 holders resolve, 26.0 per cent. The mechanism of failure is spelling: AbbVie Deutschland GmbH & Co. KG misses solely because GLEIF stores "GmbH And Co. KG".
  • Of the 102 that join, 11 LEIs are LAPSED (AbbVie Limited among them), 2 DUPLICATE and 1 RETIRED: 14 of 102 joined identities are not in good standing.
  • EU to US: 31 of 379 EMA holder names equal an openFDA labeler name after normalisation, 8.2 per cent.
  • Register-internal hygiene is high and its defects are precise: one malformed product number in 2,732 (EMEA/H/C005201), one truncated ATC code (L01XE3), four prose strings in the ATC column, zero date-sanity failures. openFDA: zero NDC scheme violations in 137,521 records, 1,904 NDCs listed more than once, 1,103 labeler-name variant groups.
  • EMA SPOR, the IDMP master data service, has no anonymous read surface: 401 and a login redirect on every route tried. The public join has to run on the XLSX shadow, and the shadow is free text.

Why this register pair, and why now

Regulatory affairs departments across the industry are staffing data standards and governance roles whose job descriptions name the same list: EMA standards published through SPOR, the FDA Data Standards Catalog, WHO ATC codes, ISO IDMP, and ontology services that compare internal dictionaries to external standards. AbbVie advertised exactly such a director role in August 2026 with a posted range of 160,500 to 305,000 dollars. The budgets are real because the obligation is real: EU implementation of IDMP has been moving for a decade, xEVMPD submission is mandatory today, and every company holds product master data that must reconcile with what the regulators publish.

All of that work rests on an assumption: that the public registers on the other side of the reconciliation can be joined to. That assumption is measurable from outside, with no credentials, in a day. This study measures it.

What we measured, and from where

Three public surfaces, all open and keyless, pinned to 28 August 2026. The EMA medicines report, the agency's own XLSX export of centrally authorised products: 2,732 rows, 2,339 human and 393 veterinary, of which 1,567 are authorised human medicines. The openFDA NDC directory bulk file: 137,521 product records in the same-day export. And the GLEIF API, queried twice per holder name, once with the exact legal-name filter and once fulltext, into a resumable cache.

One surface refused. The SPOR referentials and organisations APIs, the master data services that IDMP implementation in Europe runs through, answered HTTP 401 and a login redirect to every anonymous request. We record that rather than work around it, and it frames everything else: the join we are about to measure is the one actually available to the public, through the XLSX shadow of a database we cannot read.

The joins are deliberately conservative. A holder name matches a GLEIF record only if the legal name is equal verbatim or after trivial normalisation: whitespace, case, trailing punctuation. No fuzzy matching contributes to any headline. A separately reported near-miss class, equal after transliterating an ampersand to "and", isolates one mechanism without inflating the match rate. WHO's ATC index is licensed, so ATC values are validated structurally against the code grammar only, and the index is not republished.

Finding one: the holder join runs on spelling, and loses

The EMA register names the organisation holding each marketing authorisation in one free-text column. There is no LEI, no SPOR organisation identifier, no company number. For the 392 distinct holder names on authorised human medicines, the GLEIF join resolves 102 and fails 290.

GLEIF join, 392 EMA holder namesCount
Exact or trivially-normalised legal name match102
Matched LEI in good standing (ISSUED)88
Matched LEI LAPSED / DUPLICATE / RETIRED11 / 2 / 1
No exact match290
Of which near miss, ampersand transliteration only1

The near miss is the mechanism in miniature. AbbVie Deutschland GmbH & Co. KG, as the EMA register spells it, has a perfectly good LEI, 549300FI0P3XDOCBXJ78. The join fails because GLEIF stores the legal name as "AbbVie Deutschland GmbH And Co. KG". Same entity, same jurisdiction, one character class apart. We nearly filed this as an encoding bug in our own query; verification showed the query was fine and the registers genuinely disagree on the spelling of a legal name. Multiply that by every umlaut, comma and legal-form abbreviation across 290 misses and you have the state of the join.

The 290 misses are statements about the join, not about the companies. Bayer AG appears in GLEIF as Bayer Aktiengesellschaft; the entity exists, the identifier exists, and the free-text bridge between the two registers still fails. That distinction is held explicitly in the ontology, where a failed observation records no-exact-match-in-register and never asserts that the entity lacks an LEI.

And the joins that succeed are not all safe to build on. Eleven of the 102 matched LEIs are LAPSED, including AbbVie Limited, the UK entity; two are DUPLICATE records; one is RETIRED. An organisation column that did carry LEIs would surface this immediately. A free-text column hides it.

Finding two: the EU and US registers barely speak

The same conservative join between the EMA holder column and the openFDA labeler column resolves 31 of 379 normalised names, 8.2 per cent. The corporate groups overlap far more than that; the registers describe them through different local subsidiaries, different spellings and no shared identifier. A regulatory data office reconciling a global product portfolio against both registers is doing, by hand or by vendor, exactly the join the public infrastructure does not provide.

Finding three: inside each register, precise defects in a clean field

Fairness requires saying that both registers are internally disciplined. In 2,732 EMA product numbers there is exactly one scheme violation, EMEA/H/C005201, missing the slash before its six-digit block. Every date-sanity check passed: zero authorisations before decisions, zero withdrawals before authorisations, zero status contradictions. Those nulls are results and we report them as such.

The ATC column on authorised human medicines: 1,399 full seventh-level codes, 149 partial-level codes, 17 blanks, four occurrences of the prose string "Not yet assigned" sitting in a code field, and one truncated code, L01XE3 on Vargatef, where valid codes end with two digits. Prose in a code slot and a truncated code are exactly the class of defect a scheme-aware shape catches at submission time and a spreadsheet column never will.

On the US side, the NDC identifier column is spotless: zero scheme violations in 137,521 records. The directory nonetheless lists 1,904 product NDCs more than once, each under distinct structured product labeling documents, 1,151 of the duplicate groups sharing the same finished flag, and its labeler column holds 1,103 variant groups: 9,771 spellings for 8,343 names. Identifier hygiene and name hygiene are different disciplines, and the registers prove it in both directions.

The model: identity as a claim, not a property

The Medicinal Product Register Ontology (MPRO) is a small OWL 2 ontology with the same load-bearing decision as our other register work: identity and coding are dated, reified assertions by a named register. An IdentifierAssertion keeps the published value verbatim, prose included, with conformance to the scheme's own declared rules recorded on the assertion. Schemes declare their rules as data in a SKOS registry, and a test pins the SHACL shapes to the registry so the two cannot drift. A CrossRegisterObservation records what two registers say about the same entity, and the absence of an exact match is a statement about the join, never about the entity.

Three SHACL layers turn the validation report into the findings table: zero structural violations, 155 scheme violations (the one product number plus 154 non-full ATC values), and 1,760 defect-class results (4 prose, 2 truncated or malformed, 1,116 name-variant groups, 638 join failures), reconciling exactly with the dual-computed counts. The graph is 81,008 triples, emitted as text and parse-verified.

What would move the numbers

For EMA, one column: publish an organisation identifier, SPOR ORG-ID or LEI, alongside the holder name in the public medicines report. The agency already maintains those identifiers inside OMS; the public export simply does not carry them. A second, cheaper change: open an anonymous read surface on RMS and OMS, which are reference data services whose entire purpose is to be referenced.

For GLEIF, the ampersand case is worth a data-quality rule: legal names transliterated at registration time diverge from the names other registers publish, and the divergence defeats exactly the join the LEI exists to enable. For industry data offices, the practical reading is that name-based reconciliation against the public registers has a measured ceiling of about a quarter, and any RIM programme assuming better is assuming, not measuring.

Prior art, method, and the offer

UNICOM, the Horizon 2020 project with 41 partners including 21 medicines agencies, built and piloted IDMP infrastructure across Europe and analysed identification data from the inside. EMA and the Heads of Medicines Agencies published the Data Quality Framework for EU medicines regulation in December 2023. This study claims none of that ground: it is an outside-in measurement of what the public surfaces actually join today, reproducible by anyone with a laptop.

Everything is public: the repository at github.com/fabio-rovai/medicinal-product-register-ontology holds the pipeline, ontology, shapes, queries, tests and a build report that includes the errors we made, one of which became the mechanism finding. The method transfers to any register pair and has run on UK health data linkage, EU health dataset catalogues, fund registers, bank registers and the scholarly record.

If your organisation maintains regulatory master data against SPOR, IDMP or an internal RIM and wants its holder, substance and code columns to survive these shapes, write to fabio@thetesseractacademy.com. The first engagement is a fixed-scope review of your dictionaries against external standards, with a findings table your data governance board can act on.