Large language models can generate DCAT-compatible metadata for data catalogs at quality close to human annotations, though the strongest evidence is for simple extraction tasks.
The W3C Data Catalog Vocabulary, Version 2: Rationale, Design Principles, and Uptake
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
DCAT is an RDF vocabulary designed to facilitate interoperability between data catalogs published on the Web. Since its first release in 2014 as a W3C Recommendation, DCAT has seen a wide adoption across communities and domains, particularly in conjunction with implementing the FAIR data principles (for findable, accessible, interoperable and reusable data). These implementation experiences, besides demonstrating the fitness of DCAT to meet its intended purpose, helped identify existing issues and gaps. Moreover, over the last few years, additional requirements emerged in data catalogs, given the increasing practice of documenting not only datasets but also data services and APIs. This paper illustrates the new version of DCAT, explaining the rationale behind its main revisions and extensions, based on the collected use cases and requirements, and outlines the issues yet to be addressed in future versions of DCAT.
fields
cs.IR 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Exploring LLM Capabilities in Extracting DCAT-Compatible Metadata for Data Cataloging
Large language models can generate DCAT-compatible metadata for data catalogs at quality close to human annotations, though the strongest evidence is for simple extraction tasks.