RLAP uses a trained Q-value estimator to adaptively order LLM subtask execution and reports accuracy gains on MRC, IE, and text completion benchmarks.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
RLAP: A Reinforcement Learning Enhanced Adaptive Planning Framework for Multi-step NLP Task Solving
RLAP uses a trained Q-value estimator to adaptively order LLM subtask execution and reports accuracy gains on MRC, IE, and text completion benchmarks.