Pith. sign in

REVIEW 3 cited by

Assessing SPARQL capabilities of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.05925 v2 pith:PLPM5AEL submitted 2024-09-09 cs.DB cs.AIcs.CLcs.IR

classification cs.DBcs.AIcs.CLcs.IR
keywords sparqlllmscapabilitiesmodelsqueriesselectsemantictasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The integration of Large Language Models (LLMs) with Knowledge Graphs (KGs) offers significant synergistic potential for knowledge-driven applications. One possible integration is the interpretation and generation of formal languages, such as those used in the Semantic Web, with SPARQL being a core technology for accessing KGs. In this paper, we focus on measuring out-of-the box capabilities of LLMs to work with SPARQL and more specifically with SPARQL SELECT queries applying a quantitative approach. We implemented various benchmarking tasks in the LLM-KG-Bench framework for automated execution and evaluation with several LLMs. The tasks assess capabilities along the dimensions of syntax, semantic read, semantic create, and the role of knowledge graph prompt inclusion. With this new benchmarking tasks, we evaluated a selection of GPT, Gemini, and Claude models. Our findings indicate that working with SPARQL SELECT queries is still challenging for LLMs and heavily depends on the specific LLM as well as the complexity of the task. While fixing basic syntax errors seems to pose no problems for the best of the current LLMs evaluated, creating semantically correct SPARQL SELECT queries is difficult in several cases.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MetaboT: An LLM-based Multi-Agent Frameworkfor Interactive Analysis of Mass SpectrometryMetabolomics Knowledge Graphs

    cs.AI 2025-10 conditional novelty 6.0 of 10

    A multi-agent LLM system converts natural-language metabolomics questions into SPARQL queries over the ENPKG knowledge graph, reaching 83.67% accuracy with GPT-4o versus 8.16% for the standalone model.

  2. SPARQL Query Generation with LLMs: Measuring the Impact of Training Data Memorization and Knowledge Injection

    cs.IR 2025-07 conditional novelty 6.0 of 10

    A controlled prompting protocol shows open-weight LLMs rely heavily on memorized Wikidata entities when writing SPARQL queries, with accuracy collapsing on a rarely used benchmark.

  3. Knowledge Conceptualization Impacts RAG Efficacy

    cs.AI 2025-07 conditional novelty 6.0 of 10

    An empirical study showing that both schema complexity and representation format affect how well GPT-4o generates SPARQL queries from competency questions, with mixed results across two knowledge graph families.

Pith tools