A 0.5-1.5B parameter proxy with three cooperating agents, trained by reinforcement learning with tree-structured rollouts, improves RAG question answering by 8-13 points on multi-hop datasets without changing the retriever or LLM.
Consider the specificity, complexity, and clarity of the question
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented Generation
A 0.5-1.5B parameter proxy with three cooperating agents, trained by reinforcement learning with tree-structured rollouts, improves RAG question answering by 8-13 points on multi-hop datasets without changing the retriever or LLM.