Ontology engineering
51 open studies in OWL 2, SKOS and SHACL, across 29 public repositories. Every one is reproducible from public data, every headline is computed at least two independent ways, and the hypotheses that died during the work are reported rather than dropped. This page is the index. If you are here to commission work rather than read it, the service page is the shorter route.
What an ontology is, and what it is for
An ontology is a formal, machine-readable model of a domain: what kinds of thing exist, what properties they have, what relationships can hold between them, and what constraints must hold. It is normally written in OWL 2, with SKOS carrying the controlled vocabulary and SHACL carrying the constraints. Writing it formally buys you one thing that a data dictionary or a Confluence page cannot: software can read it and reject a statement that contradicts it.
That capability is why ontologies came back. A retrieval pipeline can hand a language model relevant text, and the model will produce something fluent. Nothing in that loop can tell you the entity it just asserted does not exist. An ontology plus a closed-world check can, at the term level, cheaply, and without a human in the loop. We published the measurement behind that claim: open-world SHACL accepted all 300 data graphs we seeded with fabricated terms, and a closed-world vocabulary gate caught all 300 with no false positives on clean data.
The distinction people most often get wrong is between the three layers. A taxonomy is a hierarchy of concepts. An ontology adds the relationships and the constraints. A knowledge graph is the instance data underneath: the actual entities and the actual claims about them. The ontology says what can be said. The graph says what is being claimed. Our glossary covers the rest of the vocabulary, and the free 15-lesson ontology engineering course teaches the modelling itself, from RDF through SPARQL, SHACL and GraphRAG.
The published work
Grouped by what the study does. Open a group to see the studies inside it, each with its repository where one exists.
Register integrity ontologies16
Public registers are the identity layer that finance, government and science all join against. Each of these studies models one register in OWL 2, SKOS and SHACL, then measures where the published record contradicts itself.
- Register assurance: why every public register fails at its boundary
The programme note that sets out the method the studies below all share.
fabio-rovai/register-integrity-index on GitHub (opens in new tab) - UK public registers: Companies House, the Charity Commission and the Global LEI System
What the UK register fabric can and cannot verify about a legal entity.
fabio-rovai/uk-register-ontology on GitHub (opens in new tab) - Italian public registers: IPA, ANAC, OpenCUP and ISTAT
380 dead municipality codes published as current, delivered upstream as a SHACL package.
fabio-rovai/italy-register-ontology on GitHub (opens in new tab) - US bank registers: the FDIC, the Federal Reserve and the Global LEI System
Three registers of the same banks, and the identifiers that fail to join them.
fabio-rovai/bank-register-ontology on GitHub (opens in new tab) - US securities registers, and the missing CIK to LEI crosswalk
The crosswalk regulators assume exists, built and measured.
fabio-rovai/securities-register-ontology on GitHub (opens in new tab) - Insurance and reinsurance: the EU register of 3,304 insurers
What EIOPA publishes about who is authorised, and where it is wrong.
fabio-rovai/insurance-register-ontology on GitHub (opens in new tab) - The US fund register as a governance graph
The largest index fund in the world is missing from the open identifier map.
fabio-rovai/investment-fund-ontology on GitHub (opens in new tab) - Land register integrity across all 33 Scottish registration counties
Registers of Scotland publishes an area field that is unpopulated on every parcel.
fabio-rovai/scotland-land-register-ontology on GitHub (opens in new tab) - Space object catalogues: CelesTrak and the General Catalog of Artificial Space Objects
GCAT marks 22 entries as corresponding to no real object. CelesTrak carries all 22.
fabio-rovai/space-object-register-ontology on GitHub (opens in new tab) - On-chain control across 6,316 contracts on seven EVM chains
Control modelled as a dated assertion with an evidence class, never as a property of a contract.
fabio-rovai/chain-control-ontology on GitHub (opens in new tab) - Mass spectrometry reference libraries: the register that publishes its own contradiction
GNPS derives structure twice per record, so its disagreement with itself is countable.
fabio-rovai/spectral-library-ontology on GitHub (opens in new tab) - One Health biosurveillance registers and the missing pathogen crosswalk
Human, animal and environmental pathogen registers that name the same organism differently.
fabio-rovai/biosurveillance-ontology on GitHub (opens in new tab) - Retraction registers: how far Crossref, PubMed and Retraction Watch disagree
Science has no Shepard’s citator, and the registers that stand in for one do not agree.
fabio-rovai/scholarly-record-ontology on GitHub (opens in new tab) - American K-12 academic standards: 67,141 identifiers dereferenced, none resolve
An identifier scheme in national use where every URI is dead.
fabio-rovai/learning-standards-ontology on GitHub (opens in new tab) - The content registry television measurement runs on covers 1.84 per cent of television
What the audience measurement chain can actually identify.
fabio-rovai/media-attention-ontology on GitHub (opens in new tab) - US federal vocabularies: what 13 million triples reveal about their quality
A census of the vocabularies US federal agencies publish for reuse.
fabio-rovai/semantic-asset-register on GitHub (opens in new tab)
Crosswalks and standards audits13
Interoperability projects usually assume a crosswalk between two standards exists and is lossless. These studies build the crosswalk and measure what it loses, which is often the thing the programme depended on.
- IES to HQDM: an open 4D ontology crosswalk for defence data
The two UK defence 4D data standards, aligned and SHACL-validated.
fabio-rovai/ies-hqdm-crosswalk on GitHub (opens in new tab) - Grounding a PYRAMID avionics bridge in IES and HQDM
Def Stan 00-134 has no shared data model. This grounds three components in one referent.
- Why industrial data crosswalks fail, measured across seven standards
Seven industrial standards, and the information that dies between them.
fabio-rovai/industrial-ontology-crosswalks on GitHub (opens in new tab) - Construction data standards cannot check your AI: IFC, COBie, Uniclass and BOT measured
Four construction standards tested for whether they can reject a wrong classification.
fabio-rovai/construction-standards-crosswalks on GitHub (opens in new tab) - Ontology quality in air cargo standards, tested against IATA ONE Record
All 496 properties in the 2022-12 release declare a domain. None of the 534 in 2024-12 do.
fabio-rovai/cargo-semantics on GitHub (opens in new tab) - Air and ocean freight standards, tested against IATA ONE Record and DCSA
The same audit with a DCSA ocean container contrast, in English and Traditional Chinese.
fabio-rovai/cargo-semantics on GitHub (opens in new tab) - Machinery data standards, tested against MTConnect and the Asset Administration Shell
Whether the machinery standards can carry what the EU Machinery Regulation asks for.
fabio-rovai/machinery-semantics on GitHub (opens in new tab) - Machine tool data standards, tested against MTConnect and the Asset Administration Shell
The machine tool variant, written for Taiwan’s sector.
fabio-rovai/machinery-semantics on GitHub (opens in new tab) - The Skills England occupational maps as an ontology
51,355 triples covering 1,269 standards, 15 routes, 35 pathways and a 278-code SOC 2020 crosswalk.
- Where England’s occupational standards meet Europe’s skills vocabulary
An open SKOS crosswalk from 1,269 Skills England standards to ESCO.
- What UK waste reporting throws away, measured in bits
Of 5.907 bits of composition detail, Simpler Recycling retains 38.4 per cent.
fabio-rovai/waste-vocab-crosswalk on GitHub (opens in new tab) - Food loss and waste data taxonomy, commissioned by WRAP
A commissioned taxonomy, delivered and in use.
- Space debris metrics that cannot be added up
An open composition checker for metrics that look additive and are not.
Verification: proving an ontology is right9
An ontology that validates is not the same as an ontology that is correct. This is our research line on the gap, and it is why every engagement ships a check that can fail.
- The open-world hole: why SHACL cannot catch a hallucinated ontology term
Open-world SHACL passed all 300 data graphs carrying a fabricated term. A closed-world gate caught all 300.
fabio-rovai/open-ontologies on GitHub (opens in new tab) - SHACL validates shapes, not vocabulary
The case for a closed-world companion to the standard RDF validator.
- Semantic-integrity defects evade static analysis and shape validation
A cross-library detection study of the defects no existing tool reports.
- Foundry-grade guarantees for machine-authored ontologies
What has to hold before you let a model write your schema.
fabio-rovai/open-ontologies on GitHub (opens in new tab) - Where symbol-existence checking belongs in a neuro-symbolic pipeline
The missing box, and what goes wrong without it.
- Beyond existence: certified denotation, the next gate
Checking a term exists is not checking it means what the sentence needs it to mean.
- Fluency is saturated, correctness is not
Across nine models fluency sits at 1.00 while verified capability ranges from 0.00 to 1.00.
- Checking the knowledge graph your pipeline just built
The practical version, for teams already running an extraction pipeline.
- Publishing machine-validated open data structures for the public sector
The publication standard we hold our own releases to.
Grounding language models in an ontology8
The commercial reason ontologies came back is that agents and retrieval pipelines need something that can tell them they are wrong. These studies fine-tune and constrain models against a published vocabulary, and measure the result.
- Teaching an open-source LLM a government information standard
On IES4, term confabulation fell from 93.7 per cent to 1.0 per cent and conformance rose from 0 to 88.6 per cent.
fabio-rovai/open-ontologies on GitHub (opens in new tab) - An open language model for IES4 data
The published checkpoint, on Hugging Face.
fabio-rovai/open-ontologies on GitHub (opens in new tab) - An open, conformant language model for biomedical knowledge graphs
The same method carried into the OBO Foundry vocabularies.
fabio-rovai/open-ontologies on GitHub (opens in new tab) - Grounded, not retrieved: an ontology-validated biomedical knowledge graph
What changes when the graph is checked against the ontology instead of trusted.
- Silence is not assent: what catches a wrong AI classification in orbit
833,403 triples over 70,122 catalogued space objects, where 351 of 353 ontology classes can never reject anything.
fabio-rovai/neurosymbolic-space-kg on GitHub (opens in new tab) - Your search index and your AI pipeline are reading different corpora
The enterprise knowledge case, which is where most buyers first meet this problem.
fabio-rovai/enterprise-knowledge-ontology on GitHub (opens in new tab) - From raw museum records to a knowledge graph
A reproducible pipeline over the Museum of Design in Plastics catalogue of 11,865 objects.
- The UK nature-governance landscape, as a graph you can cite
A provenance-first reference graph of UK nature and environment governance.
Metadata and catalogue conformance5
Data catalogues are ontologies people forget are ontologies. These studies measure real national and federal catalogues against the profiles they claim to follow.
- Nordic health dataset catalogues measured against HealthDCAT-AP Release 7
2,811 Nordic health datasets, and not one conformant to the profile.
fabio-rovai/health-dataset-catalogue-ontology on GitHub (opens in new tab) - Fifty-nine federal agencies are validated against the schema that exempts them
What US federal open data validation actually checks.
fabio-rovai/open-data-catalog-ontology on GitHub (opens in new tab) - We probed every .gov domain for DKAN and found three portals
All three publish the same wrong identity.
fabio-rovai/dkan-portal-profile on GitHub (opens in new tab) - FAIR dataset contracts for scientific data
1,738 biomedical datasets: findable and accessible, zero per cent interoperable.
fabio-rovai/fair-scientific-data on GitHub (opens in new tab) - Nature-related security risk: an open evidence base and systems ontology
An OWL ontology and SKOS taxonomy cross-walked to IPBES, published under CC-BY-4.0.
Questions we get asked
- What is an ontology?
- An ontology is a formal, machine-readable model of a domain: the kinds of thing that exist in it, the properties those things have, the relationships that can hold between them, and the constraints that must hold. In practice it is written in OWL 2, usually with SKOS for the controlled vocabulary and SHACL for the constraints. The point of writing it formally is that software can then reason over it and reject statements that contradict it, which prose documentation cannot do.
- What is the difference between an ontology, a taxonomy and a knowledge graph?
- A taxonomy arranges concepts in a hierarchy: broader and narrower, and little else. An ontology adds the relationships and the constraints, so it can express that a bank holding company controls a subsidiary and that control has to be dated and sourced. A knowledge graph is the instance layer: the actual banks, the actual subsidiaries, the actual dates. The ontology says what can be said; the knowledge graph says what is claimed. You can build a graph without an ontology, but then nothing can tell you the graph is wrong.
- What is the difference between OWL and SHACL, and do I need both?
- OWL 2 describes what is true in the domain and lets a reasoner infer more of it. SHACL checks whether a particular data graph conforms to a set of shapes. They answer different questions, so most real projects use both. The trap is assuming SHACL is a sufficient check. SHACL is open-world: if your data uses a property the shapes say nothing about, SHACL passes it. We measured this across three vocabularies and open-world SHACL accepted every one of 300 graphs carrying a fabricated term. If a language model is writing your RDF, you need a closed-world vocabulary gate as well.
- Is the semantic web the same thing as a knowledge graph?
- No, though the technologies overlap almost completely. The semantic web is the W3C programme and its stack: RDF for the data model, RDFS and OWL for the schema, SKOS for vocabularies, SHACL for validation, SPARQL for querying, and dereferenceable URIs so that data published by different people can be joined. Linked data is the publishing discipline that goes with it. A knowledge graph is the artefact you get: a graph of entities and relationships, usually stored in a triple store, that a system actually queries. Most enterprise knowledge graphs are built on the semantic web stack precisely because they need identifiers that survive being joined against somebody else’s data.
- What is GraphRAG, and does it need an ontology?
- GraphRAG retrieves over a knowledge graph rather than over a flat vector index, so the model gets structured neighbours and relationship paths instead of loose passages. It works better than plain retrieval on multi-hop questions. It does not need an ontology to run, and this is where projects go wrong: without a schema and a term-level check, the graph the pipeline builds inherits every entity the extraction step invented, and GraphRAG then retrieves that error confidently. The ontology is what lets you reject the bad node before it enters the graph.
- Do AI agents and RAG pipelines actually need an ontology?
- They need something that can tell them they are wrong, and an ontology is the cheapest thing that does that job at the term level. Retrieval gives a model relevant text; it does not give it a way to detect that the entity it just asserted does not exist. When we fine-tuned an open model on the UK IES4 standard, confabulated ontology terms fell from 93.7 per cent of outputs to 1.0 per cent, but the number that made that measurable at all was the closed-world check against the published vocabulary. Without it, fluent and wrong looks the same as fluent and right.
- How long does an ontology project take, and what does it cost?
- A scoped audit of an ontology or knowledge graph you already have takes about two weeks and sits below the £10,000 direct award threshold for public bodies. A domain ontology with a validated instance graph over real source data is typically six to twelve weeks depending on how many source registers have to be reconciled. Public sector buyers can commission through CCS RM6200 for build work or RM6126 for research and audit work.
- Who builds ontologies in the UK?
- The established UK and European names include Semantic Partners, Ontotext, DNV and the Ontology Engineering Group at Universidad Politécnica de Madrid, alongside in-house teams at large publishers and banks. Tesseract Academy works in the same space with a specific emphasis: everything we publish is reproducible from public data, every headline is computed at least two independent ways, and we report the hypotheses that died. Our work for the National Digital Twin Programme is open source under Apache 2.0.
Send us the model you are working against
Give us the ontology, standard or schema and we will tell you what it can and cannot verify, at no cost, within five working days.
