Pith. sign in

REVIEW 3 cited by

SPARQL Generation: an analysis on fine-tuning OpenLLaMA for Question Answering over a Life Science Knowledge Graph

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.04627 v1 pith:ILATG7HM submitted 2024-02-07 cs.AI cs.CLcs.DBcs.IR

classification cs.AIcs.CLcs.DBcs.IR
keywords knowledgeansweringfine-tuninggraphqueriesquestionapproachclues
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent success of Large Language Models (LLM) in a wide range of Natural Language Processing applications opens the path towards novel Question Answering Systems over Knowledge Graphs leveraging LLMs. However, one of the main obstacles preventing their implementation is the scarcity of training data for the task of translating questions into corresponding SPARQL queries, particularly in the case of domain-specific KGs. To overcome this challenge, in this study, we evaluate several strategies for fine-tuning the OpenLlama LLM for question answering over life science knowledge graphs. In particular, we propose an end-to-end data augmentation approach for extending a set of existing queries over a given knowledge graph towards a larger dataset of semantically enriched question-to-SPARQL query pairs, enabling fine-tuning even for datasets where these pairs are scarce. In this context, we also investigate the role of semantic "clues" in the queries, such as meaningful variable names and inline comments. Finally, we evaluate our approach over the real-world Bgee gene expression knowledge graph and we show that semantic clues can improve model performance by up to 33% compared to a baseline with random variable names and no comments included.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Role of Visualization in LLM-Assisted Knowledge Graph Systems: Effects on User Trust, Exploration, and Workflows

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Visualizations designed to increase transparency in an LLM-based knowledge graph query tool actually led users, including experts, to overtrust incorrect outputs.

  2. KERL: Knowledge-Enhanced Personalized Recipe Recommendation using Large Language Models

    cs.LG 2025-05 conditional novelty 5.0 of 10

    KERL uses a food knowledge graph and three LoRA adapters on one LLM to recommend constrained recipes, generate cooking instructions, and produce micro-nutrition details.

  3. Enhancing Manufacturing Knowledge Access with LLMs and Context-aware Prompting

    cs.AI 2025-07 conditional novelty 4.0 of 10

    Feeding an LLM a reduced, question-relevant slice of a manufacturing ontology improves SPARQL query accuracy by roughly 20 to 30 percent relative to the full ontology.

Pith tools