Pith. sign in

REVIEW 4 cited by

Ontology Generation using Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.05388 v1 pith:6MM5LTPU submitted 2025-03-07 cs.AI

classification cs.AI
keywords ontologyengineersevaluationllmsontologiesqualityengineeringlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The ontology engineering process is complex, time-consuming, and error-prone, even for experienced ontology engineers. In this work, we investigate the potential of Large Language Models (LLMs) to provide effective OWL ontology drafts directly from ontological requirements described using user stories and competency questions. Our main contribution is the presentation and evaluation of two new prompting techniques for automated ontology development: Memoryless CQbyCQ and Ontogenia. We also emphasize the importance of three structural criteria for ontology assessment, alongside expert qualitative evaluation, highlighting the need for a multi-dimensional evaluation in order to capture the quality and usability of the generated ontologies. Our experiments, conducted on a benchmark dataset of ten ontologies with 100 distinct CQs and 29 different user stories, compare the performance of three LLMs using the two prompting techniques. The results demonstrate improvements over the current state-of-the-art in LLM-supported ontology engineering. More specifically, the model OpenAI o1-preview with Ontogenia produces ontologies of sufficient quality to meet the requirements of ontology engineers, significantly outperforming novice ontology engineers in modelling ability. However, we still note some common mistakes and variability of result quality, which is important to take into account when using LLMs for ontology authoring support. We discuss these limitations and propose directions for future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Leveraging Large Language Models for Generating Research Topic Ontologies: A Multi-Disciplinary Study

    cs.DL 2025-08 conditional novelty 6.0 of 10

    Fine-tuned open-weight LLMs classify research-topic relationships with up to 93.5% F1 on a new multi-disciplinary benchmark, and cross-domain transfer loses only about 5 points.

  2. Retrieval-Augmented Generation of Ontologies from Relational Databases

    cs.DB 2025-06 conditional novelty 6.0 of 10

    An iterative RAG-LLM pipeline converts relational schemas into OWL ontology fragments, achieving LLM-judged quality scores of 4.2 to 4.6 out of 5 on two medical databases.

  3. Benchmarking Resource-Efficient LLMs for Research Topic Ontology Generation in the Biomedical Field

    cs.DL 2026-07 conditional novelty 5.0 of 10

    Fine-tuning small open-source LLMs on a new MeSH-derived benchmark (MeSH-Rel-4K) raises semantic-relation classification F1 by 34.1 points on average, reaching 91.6% for gemma-2-9b.

  4. Streamlining Knowledge Graph Creation with PyRML

    cs.DB 2025-05 conditional novelty 5.0 of 10

    PyRML is a Python-native RML engine that passes 306 of 323 RML-Core tests and shows lower runtimes than RMLMapper on the small test suite.

Pith tools