REVIEW 6 cited by
Leveraging small language models for Text2SPARQL tasks to improve the resilience of AI assistance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this work we will show that language models with less than one billion parameters can be used to translate natural language to SPARQL queries after fine-tuning. Using three different datasets ranging from academic to real world, we identify prerequisites that the training data must fulfill in order for the training to be successful. The goal is to empower users of semantic web technology to use AI assistance with affordable commodity hardware, making them more resilient against external factors.
Forward citations
Cited by 6 Pith papers
-
SPARQL Query Generation with LLMs: Measuring the Impact of Training Data Memorization and Knowledge Injection
A controlled prompting protocol shows open-weight LLMs rely heavily on memorized Wikidata entities when writing SPARQL queries, with accuracy collapsing on a rarely used benchmark.
-
Conversational Lexicography: Querying Lexicographic Data on Knowledge Graphs with SPARQL through Natural Language
This paper introduces a taxonomy and 1.27M-example template dataset for text-to-SPARQL on Wikidata lexicographic data, and finds that only GPT-3.5-Turbo generalizes to novel query types.
-
QUPID: Quantified Understanding for Enhanced Performance, Insights, and Decisions in Korean Search Engines
A fine-tuned ensemble of a generative small language model and an embedding model outperformed zero-shot LLMs on Korean search relevance labeling, with reported Cohen's kappa of 0.646 versus 0.387 and 60x lower latency.
-
Ontology-grounded Automatic Knowledge Graph Construction by LLM under Wikidata schema
An LLM pipeline that generates competency questions from documents, aligns extracted relations to Wikidata properties, and outputs RDF triples grounded in the resulting ontology achieves competitive partial-F1 on Wiki...
-
Enhancing Text2Cypher with Schema Filtering
Schema filtering, especially exact-match pruning, reduces prompt length and cost for Text2Cypher and improves accuracy for smaller models, though larger models gain less.
-
Thinking with Knowledge Graphs: Enhancing LLM Reasoning Through Structured Data
Representing knowledge graph triples as Python code improved LLM multi-hop reasoning accuracy over text and JSON in this study, though the effect is modest and possibly due to explicit inference steps in the code.
Discussion (0). Continue with ORCID to comment.