A small LLM trained with GRPO and LLM-judge rewards rewrites simple prompts into more effective ones, improving question-answering and arithmetic accuracy over base prompts while giving mixed, often negligible gains on summarization.
Title resolution pending
1 Pith paper cite this work, alongside 30 external citations. Polarity classification is still indexing.
1
Pith paper citing it
30
external citations · OpenAlex
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter
A small LLM trained with GRPO and LLM-judge rewards rewrites simple prompts into more effective ones, improving question-answering and arithmetic accuracy over base prompts while giving mixed, often negligible gains on summarization.