An 8B model trained by supervised imitation of a SPARQL-guided teacher plus GRPO reinforcement learning outperforms all compared frozen frontier-LLM systems on WebQSP, CWQ, and GrailQA.
C omp KBQA : Component-wise Task Decomposition for Knowledge Base Question Answering
1 Pith paper cite this work, alongside 1 external citations. Polarity classification is still indexing.
1
Pith paper citing it
1
external citations · OpenAlex
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Search-on-Graph-R1: Training Large Language Models to Search Knowledge Graphs with Reinforcement Learning
An 8B model trained by supervised imitation of a SPARQL-guided teacher plus GRPO reinforcement learning outperforms all compared frozen frontier-LLM systems on WebQSP, CWQ, and GrailQA.