Skip to main content
Back to research

Nordic health dataset catalogues measured against HealthDCAT-AP Release 7

25 August 2026

The European Health Data Space obliges health data access bodies to publish a machine readable national dataset catalogue, and the European Commission publishes HealthDCAT-AP as the profile those descriptions should follow. On 25 August 2026 we asked what the Nordic region actually publishes today. Across the eleven Nordic national catalogues harvested by data.europa.eu there are 2,811 dataset descriptions carrying the EU health theme, and not one of them satisfies all eight properties Release 7 makes mandatory. Three of the eight are present on exactly zero. Everything below comes from keyless public sources, every headline is computed two independent ways, and the pipeline fails its own build if the two ever disagree.

The short version

  • 2,811 health themed dataset descriptions across eleven Nordic national catalogues. Zero satisfy all eight mandatory HealthDCAT-AP Release 7 properties on dcat:Dataset.
  • The three health specific mandatory properties, healthdcatap:hdab, healthdcatap:healthCategory and healthdcatap:hasStructuredData, are present on exactly none of the 2,811.
  • The generic layer is in good order. dct:title and dct:description are on all 2,811 and dct:publisher on 2,785. The gap is confined to the layer the Regulation adds.
  • Finland contributes 1,146 themed descriptions to data.europa.eu and not one uses the EU data theme authority vocabulary, so a European health filter returns zero Finnish datasets while avoindata.fi returns 57.
  • 36 distinct IRIs are in use inside the EU data theme authority namespace and the authority defines 14. 1,238 datasets carry one of the other 22, including 806 on the literal string data-theme/undefined.
  • The authority host answers HTTP 200 for any IRI in its namespace and returns an empty 170 byte graph for terms it has never defined, so a status code check cannot detect any of this.
  • Findata's Aineistokatalogi holds 89,368 variable descriptions, the most detailed health metadata in the region, of which 84.74 per cent carry no English label. It publishes nothing into the DCAT layer.

Why now, and why the Nordics

The Nordic Council of Ministers funds VALO, a project coordinated by Sitra on value from Nordic health data. Its full members are the Danish Health Data Authority, Sitra with THL, the Finnish Ministry of Social Affairs and Health and Findata, the Icelandic Ministry of Health with the Directorate of Health and the University of Iceland, the Norwegian Institute of Public Health, and the Swedish Ministry of Health and Social Affairs with the National Board of Health and Welfare and the Swedish eHealth Agency. Estonia and Lithuania are observers. Phase one ran from February 2024 to October 2025 and VALO2 runs to October 2026. Two of its stated goals are preparing jointly for the European Health Data Space and maintaining Nordic leadership in the secondary use of health data.

That is a well resourced, well governed, five country programme with exactly the right people in the room. It is also the reason the finding matters. If the region that leads on secondary use of health data publishes nothing today that would pass the profile, the gap is not a Nordic failing. It is a measure of how far the profile sits from the installed base everywhere, and the Nordics are simply the place where someone is far enough along to measure it against.

On timing. The Regulation entered into force on 26 March 2025. The European Commission's own page states that the rules on secondary use start to apply for most data categories in March 2029, with the remaining categories including genomic data in March 2031, and that the Commission's deadline for adopting key implementing acts is March 2027. We could not verify those dates against the Official Journal text, because EUR-Lex answered HTTP 202 with a zero length body on every route we tried. Anyone citing chapter or article numbers should read the primary text first. We are not.

What we measured, and from where

Two public surfaces, both keyless, both pinned to 25 August 2026. The data.europa.eu SPARQL endpoint, which harvests the national open data portals of the Member States and held 1,908,938 instances ofdcat:Dataset at the time of the run. And the Findata Aineistokatalogi public API, which returns the whole Finnish social and health data catalogue, instance variables inlined, in a single response of roughly 116 MB.

The requirement set is not typed by hand anywhere in this work. The pipeline fetches the official HealthDCAT-AP Release 7 SHACL shapes from the European Commission's GitLab at code.europa.eu, parses them, and generates the requirement registry the measurement runs against. A test fails the build if the committed registry drifts from what the generator produces, so a future release of the profile shows up as a diff rather than as a silent divergence between our claims and the standard.

That matters more than it sounds, because the profile has moved and its own signposting has not kept up. HealthDCAT-AP used to live at healthdcat-ap.github.io. That page was decommissioned on 22 September 2025 and its README still directs readers to Release 5 as the current and authoritative version. Release 5 is not current. Release 6 was deprecated on 24 April 2026 and Release 7 is the live one. Several national and institutional implementations on GitHub still point at the decommissioned location.

Finding one: the gap is entirely in the health layer

HealthDCAT-AP Release 7 makes eight properties mandatory ondcat:Dataset in its public layer. Here is what the Nordic region publishes against them.

Mandatory propertyPresentShare of 2,811
dcat:theme2,811100.00%
dcat:distribution2,58291.85%
dct:accessRights2,33883.17%
dct:identifier1,76162.65%
dcatap:applicableLegislation210.75%
healthdcatap:hdab00.00%
healthdcatap:healthCategory00.00%
healthdcatap:hasStructuredData00.00%

One line in that table is not a finding and we would rather say so than let a reader take it for one. Thedcat:theme figure is 100 per cent by construction, because descriptions were selected by carrying the health theme in the first place. It is in the table so the table is complete.

What the Nordic region publishes against each mandatory property

2,811 health themed dataset descriptions, eleven Nordic national catalogues, 25 August 2026

dcat:theme100.00%dcat:distribution91.85%dct:accessRights83.17%dct:identifier62.65%dcatap:applicableLegislation0.75%healthdcatap:hdab0healthdcatap:healthCategory0healthdcatap:hasStructuredData00%25%50%75%100%

Share of the 2,811 descriptions carrying each property. The dcat:theme row is 100 per cent by construction, because descriptions were selected by carrying the health theme; it is shown so the set is complete, not as a finding. The three properties at zero are the ones HealthDCAT-AP adds to DCAT-AP for health, and no catalogue in the region can currently express any of them.

The useful shape here is the contrast with the generic layer. dct:title is on all 2,811. dct:description is on all 2,811.dct:publisher is on 2,785. These are catalogues run by people who are good at cataloguing. Nothing in the numbers suggests carelessness, and reading them as a scorecard of national diligence would be reading them wrong.

The operational consequence is narrower and more useful than a scorecard. Three properties sit at zero across five countries and eleven catalogues, which means no Nordic catalogue can currently express who the responsible health data access body is, what category of health data a dataset holds, or whether it has structured data at variable level. Those are not fields anyone forgot to fill in. They are fields that do not exist in the systems doing the publishing, so closing that gap is a procurement and system change question rather than a data entry question. That is worth knowing four years before the obligation bites rather than one.

Finding two: Finland is absent from European health discovery

Finland's national portal contributes 2,259 dataset descriptions to data.europa.eu. 1,146 of them carry a dcat:theme. Not one of those values is drawn from the European Union data theme authority table. Every one is a local CKAN group UUID underavoindata.suomi.fi/data/group/.

Finland is alone in this among the eleven catalogues we measured. data.norge.no binds 1,547 of 1,591 themed descriptions to the authority. Datavejviser binds 3,460 of 3,461. dataportal.se binds 18,635 of 18,701. The Swedish INSPIRE node binds all 178 of its 178. The Icelandic geoportal binds 383 of 422. Even the weakest of the others, Geonorge at 112 of 347, is not at zero.

Share of themed descriptions bound to the EU data theme authority

Share of each catalogue's themed descriptions carrying at least one value from the EU authority table, across all themes and not only health

Swedish INSPIRE node100.00%Datavejviser (DK)99.97%Norwegian INSPIRE node99.66%dataportal.se99.65%data.norge.no97.23%Icelandic geoportal90.76%Paikkatietohakemisto (FI)80.50%Geodata.se69.57%Geonorge32.28%avoindata.fi0 of 1,1460%25%50%75%100%

Finland's national portal is the only one of the eleven at zero. All 1,146 of its themed descriptions carry a local CKAN group identifier instead of a value from the European authority table, so a European filter on the health theme returns no Finnish datasets at all.

The consequence is immediate and it is the whole point of the exercise. A European health filter returns zero Finnish datasets. Finland has them, and avoindata.fi's own search returns 57 results for health against the same corpus. The data is published, it is indexed at home, and it is invisible on the route the European Health Data Space is being built on. This is a one property fix in a harvest mapping and it is the highest leverage single change available to any organisation in this study.

A fair objection deserves answering here, because the well informed view in Finland is that the country is about as close to ready for the European Health Data Space as anyone can be at this stage, and on most axes that view is correct. Kanta aggregates primary and secondary care and social care data, every care provider public and private is mandated to deposit into it, Kanta PHR lets individuals contribute their own measurements, and the European electronic health record exchange format is on its roadmap. Findata has run a working permit process and a secure processing environment for years. Measured on infrastructure, on legal machinery and on the willingness of providers to participate, Finland is ahead of most Member States and the comparison is not close.

Our finding does not contradict any of that, and it would be a misreading to take it as a scorecard on Finnish readiness. It is about one specific obligation on one specific axis: the machine readable dataset catalogue through which someone in another Member State discovers that a Finnish dataset exists at all. Having the data, having the permit process and having the exchange format are necessary and none of them makes a dataset findable. On that axis the measured value today is zero, and the distance between zero and a good score is a theme mapping rather than a programme. Being nearly ready everywhere else is precisely what makes the gap worth naming: it is the cheapest remaining item on the list.

Finding three: the EU authority namespace accepts terms it never defined

DCAT-AP requires dcat:theme values to come from the EU Data Theme authority table. We asked the endpoint which IRIs are actually in use inside that namespace, and then dereferenced every one of them. 36 distinct IRIs are in use. The authority defines 14. The other 22 were minted by publishers inside the European Union's own namespace, and 1,238 datasets carry one.

Datasets carrying a data theme IRI the EU authority never defined

Of 36 distinct IRIs in use inside the authority namespace, the authority defines 14

undefined806ENV237UKLF34VERWALTUNG29BEVOELKERUNG26TRANSPORT-VERKEHR23the other 16 values830200400600800

1,238 datasets in total. The largest value is the literal string undefined, published by the Moldova government portal on 647 datasets and the Zagreb city portal on 157. ENV is used by the London Datastore on 237 datasets where the authority defines ENVI. The authority host answers HTTP 200 for every one of these and returns an empty 170 byte RDF document, so a check on the status code sees success.

806 datasets carry the literal string undefined as a theme. That is a serialiser writing a JavaScript value into a public authority namespace, and it has been sitting there long enough to be harvested and republished.

The reason nobody has caught this is worth stating carefully, because it generalises well beyond themes. The authority host answers HTTP 200 for any IRI in the namespace. For a term it has never defined it returns a well formed RDF document of 170 bytes containing no concept. For a real term such asHEAL it returns 16,063 bytes. A validator checking that a theme value resolves sees success on a term that does not exist. Membership can only be decided by parsing the body and looking for a concept, which is what our pipeline does and what a test in the repository pins so that a future rewrite cannot quietly reintroduce the status code check.

Being fair about scope: the health slice of this is three datasets. It is a small finding for health and a large one for the endpoint, and we report it that way rather than inflating it. What it does mean for health is that a theme based count of health datasets is a floor and not a census, which is why the 2,811 figure above is described as one.

Finding four: the richest health metadata in the region is the least reachable

Findata is the Finnish social and health data permit authority. Its Aineistokatalogi holds 2,835 dataset descriptions and 89,368 instance variable descriptions, each with a distinct identifier. That is finer grained than anything in the Nordic DCAT layer by a wide margin, and none of it is in that layer, so none of it appears anywhere in finding one.

The metadata is substantively complete. Only 2.01 per cent of the 89,368 variables lack a description. People have done a great deal of careful work here. It is also almost entirely closed to anyone who does not read Finnish: 84.74 per cent of the variables carry no English label, and only 7.51 per cent of the dataset descriptions do. Cross border discovery is the stated purpose of the Regulation, and a description that is complete domestically and monolingual is not a complete description for that purpose however good it is.

English coverage in Findata's Aineistokatalogi

The most detailed health metadata in the Nordic region, and the share of it a researcher who does not read Finnish can use

Dataset descriptions with an English label7.51%

213 of 2,835

Dataset descriptions with an English description7.20%

204 of 2,835

Instance variables with an English label15.26%

13,639 of 89,368

Instance variables with a description in any language97.99%

87,569 of 89,368

The bottom row is the point of the top three. The catalogue is substantively complete, with only 2.01 per cent of its 89,368 variable descriptions missing a description in any language, and it is almost entirely closed to a reader outside Finland. Cross border discovery is the stated purpose of the European Health Data Space, so a record that is complete domestically and monolingual is not yet a complete record for the purpose the Regulation sets.

A second gap compounds it. All 2,948 concept tags in the catalogue, spread across 2,059 descriptions, report a null concept scheme. There are 508 distinct concepts and 477 of them have English labels. They are good tags bound to nothing, so they cannot be mapped to the EU data theme vocabulary or to any other published vocabulary, and the English labels that do exist cannot be reached through them.

The asymmetry is the interesting part. Mapped onto the profile, Findata could source six of the eight mandatory properties from fields it already populates, includinghealthdcatap:hasStructuredData, which its 89,368 variable descriptions answer better than any other catalogue in the region could. Exactly two have no source field at all: dcatap:applicableLegislation andhealthdcatap:hdab, the identifier of the health data access body, which in this case is Findata itself. The organisation best placed in the region to satisfy the hardest requirement in the profile is currently absent from the measurement entirely.

The layer the Finnish national proposal does not name

In February 2026 Sitra published study 255, Terveystiedon tulevaisuus tekoälyn aikakaudella, by Olli Kallioniemi and Kimmo Porkka, proposing a national Finnish Health Data Space on a one, one, one model: FHDS-Tieto as a single national data infrastructure, FHDS-Lupa as a single coordinating authorisation authority, and FHDS-TKI as a single national collaboration body. It is a serious document and its diagnosis is blunt. In its own English summary, “Finland's fragmented data production and governance, along with regional data silos, hinder the full-scale utilisation of AI, the formation of a comprehensive national overview, and compatibility with the EHDS.”

Our finding two is a measured instance of exactly that sentence. Where the report asserts a compatibility gap, the endpoint shows one: 1,146 themed Finnish descriptions, zero bound to the vocabulary a European health filter reads.

There is a second observation worth making carefully, because it is about an absence and absences are easy to overstate. The report is thorough on data content standards. It names HL7 FHIR, openEHR and the OMOP Common Data Model repeatedly, it puts the FAIR principles at the centre of what FHDS-Tieto is for, and it cites Findata's own binding regulation 1/2021 on the content, concepts and structures of dataset descriptions, noting that standardisation of metadata is essential for findability. Those are the right instincts and the right instruments.

The extracted text of the report mentions DCAT zero times and HealthDCAT-AP zero times. FHIR, openEHR and OMOP describe what is inside a dataset. HealthDCAT-AP describes the dataset so that someone in another Member State can find it in the first place, and it is the profile the European Commission publishes for that purpose. A national programme can get every content model right and still be undiscoverable, which is the position finding two measures. We state this as an observation about the document's text and not about the authors' knowledge: figures in the report are images, we read the extracted text, and the omission may well be deliberate scoping rather than oversight.

This is not an abstract gap, and Finland is already paying to close it. On 20 May 2026 Business Finland announced funding for Roadmap to Finnish Health Data Space, a joint project with a budget of 7.7 million euros of which Business Finland funds 4.6 million. It is the first joint project of the Finnish wellbeing regions to be funded. The participants are HUS, the Pirkanmaa wellbeing region, the University of Helsinki, Orion, GE Healthcare Finland, Productivity Leap and Biocomputing Platforms, and the announcement is explicit that the work is driven by what the European Health Data Space Regulation requires.

Seven named organisations, a four year public investment and a regulation that starts to bite in 2029 are the reason a measurement like this one is worth taking now rather than in 2028. The number that matters to that project is not our headline. It is the per record list underneath it, which says exactly which descriptions fail which requirement, and which of those requirements the current schema cannot source at all. The first is a data entry problem. The second is a system design decision, and it is much cheaper to take at the start of a roadmap than at the end of one.

The constructive version is short. Finland already has a binding national regulation on dataset descriptions, a permit authority that operates a catalogue, and, in that catalogue, 89,368 variable descriptions that would answer the single hardest requirement in the European profile. What is missing between those assets and European discoverability is a mapping and a theme binding, both of which are small pieces of work compared with everything else the FHDS proposal contemplates.

The method, and why it is worth borrowing

The transferable part of this is not the Nordic numbers. It is the recording discipline, and it applies to any register that republishes another register's identifiers or another profile's properties.

Record absence, do not infer it. When a catalogue does not publish a required property, that silence is a position the catalogue has taken. We emit it as a dated observation with a verdict and a reason drawn from a controlled scheme, because absent, present but bound to no vocabulary, and present but injected by an aggregator are three different failures with three different remedies. A bare count of missing fields collapses them and destroys the remedy.

Separate what the publisher said from what the aggregator computed. This one cost us a finding. An early run measured dqv:hasQualityAnnotation at 98.40 per cent and we were briefly delighted. The property is present, in named graphs underdata.europa.eu/88u/metrics/, which is where the portal stores the output of its own metadata quality assessment. Those triples were computed downstream and no publisher supplied any of them. Excluding the metrics graphs moved the figure to 0.96 per cent, a factor of a hundred. Every measurement here now excludes them, and the ontology carries a provenance observation type that exists solely because of that mistake.

Compute every headline twice. Seventeen figures in this study are computed set based in Python and again by SPARQL over the emitted graph, sharing no code path, and the build fails if any pair disagrees. All seventeen agree in the published run. They did not on the first attempt: one cross check summed across two distinct defect classes and reported 4,186 where the true figure is 1,238. Doing the work twice is the only reason we know the number we published is the right one.

Do not trust a status code as evidence of meaning. Finding three exists because we parsed bodies. An earlier version of the same check read status codes, got 200 for every term including the invented ones, and would have reported zero squatted IRIs with complete confidence.

We also got the requirement set itself wrong on the first pass. A regular expression over the official SHACL file reported eight mandatory properties, five of which are not mandatory at all, because SHACL property shapes are nested blank nodes that a line oriented regex cannot bracket. An entire conformance table was computed against the wrong list before parsing the shapes properly gave the right one. All four of these mistakes are written up in the repository's build report rather than quietly fixed, because a method is only worth borrowing if you can see where it broke.

Prior art

HealthDCAT-AP was designed under the HealthData@EU pilot, and the design account is published by its authors: Pascal Derycke, Beatriz J. Barros, Nienke M. Schutte, Charles-Andrew Vande Catsyne and Martina Bargeman Fonseca of Sciensano, “Designing DCAT-AP Extensions for Common European Data Spaces: The EHDS HealthDCAT-AP Case Study”, presented at the NeXt-generation Data Governance workshop at SEMANTiCS 2025 in Vienna and published in CEUR-WS Volume 4064. It sets out the requirements gathering, the stakeholder working groups, the multi country sandbox validation against real world examples and the public consultation. It is the right citation for what the profile is and why it has the shape it has, and anyone reading our numbers should read it first.

Our claim is narrower and does not overlap theirs. They designed the profile and validated it against curated examples. We measured the installed base of live national catalogues against the shipped Release 7. Neither piece of work answers the other's question.

A second piece of prior art sits alongside this one and shares an author with the Finnish proposal. Ole A. Andreassen and twenty one colleagues, with Olli Kallioniemi among them, published “An AI-Health infrastructure for the Nordic region: technical foundations, data assets, and a roadmap for deployment” in Nature Medicine in 2026 (doi 10.1038/s41591-026-04575-4). It proposes a Nordic AI-Health platform over the region's linked registries and biobanks, a Nordic Common Data Model built on OMOP, federated learning across national trusted research environments, and an integrated high performance computing layer. It cites the Sitra report as reference 15. We read the SSRN preprint rather than the published version, so the two may differ in detail.

Its roadmap is the most useful thing here, because it argues our timing better than we would. Step five says the project must “treat EHDS compliance as a primary design constraint rather than a retrofit”. Step two commits to a common data model “supported by rigorous schema auditing and FAIR-compliant metadata”, and the body promises that datasets will undergo “schema auditing, ontology mapping, and controlled-access onboarding”. We agree with all of it.

The observation to make, carefully, is the same one the Finnish proposal invites. The paper names FAIR, whose first letter is findability, and it names OMOP, ICD, imaging and pathology standards, all of which describe what is inside a dataset. It does not name DCAT, DCAT-AP or HealthDCAT-AP, which is the profile that decides whether anyone outside your own country can find the dataset at all. Two of the most considered documents produced in this region this year both treat data standardisation as a content problem, and the discovery layer is where the measurement above finds zero. If EHDS compliance is to be a primary design constraint rather than a retrofit, the catalogue metadata is part of what that constrains, and it is the part currently at zero across all five countries.

The European Commission also publishes an official HealthDCAT-AP validator, built on the Interoperability Test Bed. It is the right tool for checking a single record before submission and we have not reimplemented it. The shapes in our repository do the opposite job: they read a graph of dated observations about many catalogues at once and produce a report that can be handed to the body able to fix it. If you are a publisher preparing one record, use the Commission's validator.

The artifact

Everything above is reproducible from a public repository: an OWL 2 core, a requirement registry generated from the official Release 7 shapes, a SKOS registry in which each scheme declares its own conformance rules as data, three gated SHACL layers with one shape per defect class, a resumable pipeline, six verified SPARQL queries and offline tests with CI. The graph is 221,430 triples, emitted as Turtle text in 0.05 seconds and parsed back in 3.35 seconds to prove it is well formed.

No harvested payload is redistributed. Findata publishes no reuse licence for the Aineistokatalogi metadata, so redistributing it would assume a permission nobody granted. The harvester regenerates it in one call.

github.com/fabio-rovai/health-dataset-catalogue-ontology

If you run one of these catalogues

Every aggregate in this study resolves to named records. We can send you the individual dataset descriptions behind any figure above for your own catalogue, as a list you can open, along with the query that produced it so you can rerun it yourself whenever you like.

If it is useful, a bounded first engagement is a two week conformance baseline for a single catalogue: your descriptions measured against HealthDCAT-AP Release 7 property by property, the per record evidence rows, a mapping of which mandatory properties your current schema can source and which have no source field at all, and a written note on what closing each gap actually requires. Fixed scope, fixed price, and the pipeline handed over so the baseline can be rerun without us.

Corrections are welcome and are applied on this page rather than silently. If a number here is wrong, tell us, and the correction and its cause will stay visible alongside it.

fabio@thetesseractacademy.com