None of the 2,811 health-themed dataset descriptions reviewed from 11 Nordic national catalogues met all eight mandatory properties under HealthDCAT-AP Release 7. The result comes from a dated snapshot of machine-readable records, so it describes what the portal exposed at the time rather than a permanent property of every description.
The shortfall was concentrated in health-specific fields. Three mandatory properties appeared on exactly zero descriptions across five countries and all 11 catalogues. A fourth field, for applicable legislation, appeared in only 0.75 per cent of the sample.
The missing information was not the basic description
The absent fields were the health data access body, the health category and a field for structured data. By contrast, every selected record had a title and a description, while 2,785 had a publisher listed. The figures show that complete title and description fields did not amount to full profile conformance.
The study treated each verdict as an observation made against a stated requirement on a particular date. It recorded an absence when a field was not present, rather than inferring that a publisher had supplied information elsewhere.
A deliberately narrow count
The primary census used records carrying the EU authority term HEAL in the dcat:theme field. That is a conservative way to identify health-themed descriptions, but it leaves out datasets with no theme and datasets using a local vocabulary. The 2,811 records should therefore be read as a floor for the visible sample, not as a count of every health dataset in the Nordic countries.
To reduce the risk of measurement errors, the researchers derived the mandatory requirement set from the published shapes rather than entering it by hand. They recomputed all 18 headline figures through both set-based and SPARQL methods, and the two routes agreed in the reported run. Validation also used three gated SHACL layers, a rule system for checking whether data follows a specification: the recording layer had no violations, the defect layer reported 21,431, including 12,975 mandatory-property and 8,433 health-specific absences, and the third layer reported 11.
The wider catalogue picture was uneven too
A separate Metadata Quality Assessment, or MQA, of the broader DCAT-AP standard pointed in the same direction, although it was not a second measurement of the HealthDCAT-AP result. The strongest Nordic catalogue score was 300 out of 405, below the 351-point threshold for the portal’s Excellent band. Five catalogues were reported at zero per cent DCAT-AP compliance.
The review also showed how theme binding can affect what users see in a portal. Finland contributed 2,259 descriptions to the European portal, and 1,146 of them carried a theme. None used the EU data-theme authority vocabulary, however, so the portal’s European health filter returned zero Finnish datasets from that catalogue.
Across the wider portal, 36 theme identifiers were in use. Fourteen were defined, while 22 were undefined across 1,238 datasets. Requests for those undefined terms returned an HTTP 200 response with a well-formed but empty RDF document measuring 170 bytes. In other words, those terms produced a successful response with no RDF records in it.
Findata had detail, but not easy reach
The researchers separately harvested Findata rather than treating it as part of the primary 2,811-record census. Among 2,835 Findata records, six of the eight mandatory properties had source fields, one only thinly, and exactly two had no source field. Having a place from which to draw a value is not the same as already conforming to the HealthDCAT-AP profile.
The direct harvest showed substantial metadata detail alongside limited reach for people searching across languages. Findata contained 89,368 instance-variable descriptions, but 75,729 variables had no English label. Only 1,799 variables lacked a description, while all 2,948 concept tags had a null concept scheme.
A snapshot of discoverability, not health-data quality
The study measures catalogue metadata and how portals expose it, not the quality, accessibility or reuse of the underlying health data. The conservative theme filter means the main result may miss records that use local or no themes, so its scope is limited to machine-readable descriptions visible through the European portal.
The comparison with the portal’s MQA needs the same caution: it covered whole catalogues rather than the health subset, used its own scoring scheme and was retrieved two days after the primary census. It does not independently establish HealthDCAT-AP conformance.
Future reruns can test whether catalogues add the missing legal and health-specific fields, whether the EU authority service handles undefined terms clearly, and whether Findata fills the two source gaps while improving English labels and concept-scheme coverage. The snapshot identifies gaps, but it does not test whether fixing them would improve discovery.
Paper data and sources
Original title: Measuring the Installed Base: Nordic Health Dataset Catalogues Against HealthDCAT-AP Release 7
Authors: Fabio Rovai
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-27
DOI: Not available
Original paper · Full text