Open research, July 2026
Why industrial data crosswalks fail, measured across seven standards
Open crosswalks between four pairs of industrial data standards, and the measurement that explains why this kind of work goes wrong. Four of the seven standards examined cannot reject a mis-mapping at all, so the reasoner check most projects rely on could never have failed.
Correspondences
51 + 31
mappings plus asserted non-mappings, across 7 standards
Cannot be contradicted
4 / 7
zero disjointness axioms, so zero checkability
Reasoner-certified
21 / 0
bridge axioms; zero new incoherences
The Challenge
A process plant is designed in an ISO 15926 world and handed to an operator whose asset system runs on IFC or CFIHOS. A factory is specified in ISA-95 and expected to expose itself through an Asset Administration Shell or OPC UA. Every party holds valid, standards-conformant data, and the join still fails. This is the capital-facilities handover problem, and it recurs in every infrastructure, energy and defence estate programme that has to carry information from design into operation.
The usual diagnosis is that the file formats differ. That diagnosis is wrong, and expensively so. Measured across the published ontologies, the exact normalised term overlap between CFIHOS V2.0 and IFC4 is one class name, across 1,397 and 1,286 classes respectively. The failure is semantic: the same word denotes different things, and the same thing carries different words.
What we built
| Deliverable | What it is |
|---|---|
| Four crosswalks | ISO 15926-14 to IFC4, ISA-95 to the Asset Administration Shell, CFIHOS to IFC4, and SAREF4INMA to OPC UA. 51 SSSOM correspondences in total, each carrying predicate, confidence, justification and provenance. Every one of the 120 distinct identifiers used is checked to exist in the source it claims to come from. |
| Asserted non-mappings | The pairs that look alignable and are not, plus the concepts with no counterpart at all, published as machine-readable denials rather than left as silent omissions. 31 across the four pairs, and on two of them the denials outnumber the mappings. That imbalance is the result, not a gap in the work. |
| Axiomatic Strength Index | A reproducible count, per standard, of the axioms actually capable of producing a contradiction. |
| Falsifiability rate | The fraction of possible mis-groundings a vocabulary can reject. The metric that decides whether any automated check of an alignment can work at all. |
| Documented schema lifts | ISA-95 ships XML Schemas and OPC UA ships an address-space NodeSet, so both need transforming before alignment. The transformations are first-class artefacts that count and report what they discard. |
| Reasoner-certified bridge | The flagship crosswalk promoted to OWL and certified with the HermiT reasoner: 21 of 24 candidate axioms survive, with zero new unsatisfiable classes and zero conservativity violations. |
| Independent audit | A published third-party CFIHOS alignment audited for how much of it a reasoner is in a position to check. |
The finding: most of these standards cannot tell you that you are wrong
An ontology can only prove a mapping impossible if it contains an axiom capable of producing a contradiction. Count, for each standard, the fraction of class pairs that provably cannot share an instance. That is its falsifiability rate: the ceiling on what any automated check of an alignment into it can ever catch.
| Standard | Classes | Disjointness axioms | Falsifiability |
|---|---|---|---|
| ISO 15926-14 (IDO) | 49 | 15 | 75.94% |
| ISO 15926-2:2003 | 201 | 781 | 46.81% |
| IFC4 ADD2 | 1,286 | 2,443 | 11.45% |
| SAREF core | 95 | 0 | 0.00% |
| Asset Administration Shell | 64 | 0 | 0.00% |
| CFIHOS V2.0 (IDO-aligned) | 1,397 | 0 | 0.00% |
| OPC UA Device Information | 53 | 0 | 0.00% |
Four of the seven score exactly zero. No mis-mapping into CFIHOS, SAREF, the Asset Administration Shell or OPC UA Device Information can ever be rejected by a reasoner, because none of them asserts a single disjointness axiom. Running a reasoner over such an alignment and reporting that it is consistent measures the vocabulary, not the alignment. It would return the same answer for a deliberately absurd mapping.
The second result is the one that changes practice. IFC4 carries 163 times more disjointness axioms than ISO 15926-14 and is 6.6 times less checkable. Axiom count is a poor proxy for rigour; axiom placement decides it. The 15 axioms in ISO 15926-14 sit at the top of a 49-class hierarchy and propagate down to three quarters of all class pairs. The 2,443 in IFC4 sit between leaf siblings, distinguishing a wall from a beam while saying nothing across branches. For anyone choosing a hub vocabulary for an estate or a digital twin, that is the number to procure against.
Why reviewing mappings one at a time cannot work
Both flagship ontologies are individually sound: reasoned alone, ISO 15926-14 and IFC4 each have zero unsatisfiable classes, and merging them with no bridge changes nothing. So any damage is attributable to the crosswalk alone.
Each of the 24 correspondences was then added on its own and measured. Every single one is harmless in isolation: zero unsatisfiable classes, zero invented subsumptions, 24 times out of 24. Assert the same 24 together and 29 IFC classes lose all possible instances, including IfcSite, IfcBuilding, IfcBuildingStorey and IfcSpace, while the merge invents 91 subsumptions that neither standard states. One of them entails that every process stream is a manufactured article.
The failure is emergent, and it needs at least two correspondences to interact. That has a direct consequence for assurance: reviewing a crosswalk correspondence by correspondence, which is what mapping review and most alignment tooling do, cannot detect this class of fault by construction. The smallest unit of failure is a pair of mappings, not a mapping.
Named divergences, not silent omissions
ido:PhysicalQuantity vs ifc:IfcPhysicalQuantityThe pair a lexical matcher ranks first, and it is wrong. In ISO 15926-14 a PhysicalQuantity is a Quality borne by the pipe. In IFC it is a recorded measurement value. The correct target is ido:QualityDatum, which sits in a branch the first model declares disjoint from qualities.
ido:System vs ifc:IfcSystemIdentical names concealing a genuine disagreement about the world. ISO 15926-14 permits one pump to be both a functional object and a physical object. IFC places IfcSystem under IfcGroup, which it declares disjoint from IfcProduct. Assert both mappings and any functional-and-physical item becomes impossible.
ido:Site vs ifc:IfcSiteProcess industry treats space as a frame of reference; construction treats it as a physical object with geometry. The two sit on opposite sides of ISO 15926-14 own disjointness, and the reasoner rejects this pair in every direction.
ido:Stream vs ifc:IfcFlowSegmentContents versus container. The first is the fluid in motion, the second is the pipe. Tooling that equates them attaches composition and temperature to pipework, and diameter and insulation to the fluid, with nothing in either schema to stop it.
ido:Actual / ido:Specified vs ifc:IfcObject / ifc:IfcTypeObjectThe strongest agreement in the set, and one no label matcher would find. Two committees, one from process industry and one from construction, working decades apart, independently drew the same line between the thing as designed and the thing as built, and independently declared it exclusive. A crosswalk between industrial standards should start at the matching disjointness, not the matching labels.
ido:Function and ido:Capability vs IFCIFC has no class for what a thing is for. It can name a pump as a pump through a type enumeration, but cannot say that a valve is capable of isolating a line without asserting that it currently is. Recorded as an asserted absence, because no counterpart exists and no mapping found are different claims.
Each of these is published as a machine-readable asserted non-mapping, so a tool consuming the crosswalk is told not to make the join rather than left to discover the problem in production.
What this means for a buyer
Three procurement-relevant conclusions follow, and none of them requires an ontologist to act on.
- Do not accept a clean reasoner report as evidence of a good alignment. Ask for the falsifiability rate of the target vocabulary alongside it. If that rate is zero, the report is uninformative by construction.
- Do not accept per-mapping review as assurance. It is structurally blind to the most common failure mode. Ask what pairwise or set-level checking was done.
- Ask which rendering was aligned. ISA-95 and the Asset Administration Shell are each published in two independent machine-readable renderings, and the two ISA-95 renderings agree on only 5.4% of their object-model concepts. A crosswalk against one does not transfer to the other.
Learn the method: free 15-lesson course
The whole method above is taught as a free course on the Tesseract Academy platform, worked end to end on the same standards and the same measurements. Every figure quoted in the lessons is one produced by the repository, so the course and the crosswalks check each other. Fifteen lessons, each with a structural diagram and a graded quiz.
| Lessons | Module | What it covers |
|---|---|---|
| 1 to 4 | Foundations | The handover problem and why it is semantic, not a file-format issue. What a crosswalk is: SSSOM, the five SKOS mapping predicates, and why every correspondence needs reified provenance. Fetching standards by IRI with a checksum lockfile. Reading an unfamiliar ontology by its disjointness rather than its labels. |
| 5 to 7 | Measuring a standard | The Axiomatic Strength Index: counting the axioms that can actually produce a contradiction. Refutation-inert ontologies and why a clean reasoner report against them proves nothing. The falsifiability rate, and why axiom placement beats axiom count. |
| 8 to 12 | Building the crosswalk | Lifting standards that ship schemas rather than ontologies, and what you must refuse to lift. Candidate generation once lexical matching has failed. Authoring correspondences, confidence and asserted non-mappings. A field guide to false friends. SHACL shapes that reject a lazy crosswalk. |
| 13 to 15 | Proving it holds | Consistency is not coherence: what happens when SKOS becomes OWL. The emergent-failure result and why per-mapping review cannot work. Orientation search: certifying a bridge and knowing when to drop a mapping. |
Open, and reproducible
Every figure on this page is produced by scripts in the repository, run against artefacts fetched by IRI and pinned by checksum. No standards body material is redistributed. Two commands reproduce the measurements, and a third re-runs the reasoner certification.
industrial-ontology-crosswalks on GitHubReleased CC BY 4.0. The CFIHOS audit examines a third-party ontology by Abad-Navarro, Fernandez-Breis and Garcia-Castro at Universidad de Murcia, and is offered as an independent measurement rather than a competing alignment. Related work: the IES to HQDM crosswalk applies the same reasoner-certification method to defence data, and the method itself is taught in the free Running Open Ontologies course.
