ARBOR introduces a reusable rubric buffer that consolidates contrastive trajectory drafts into cross-query rubrics for online process rewards, outperforming GRPO and DAPO on multi-hop QA benchmarks.
Hy- brid reward normalization for process-supervised non-verifiable agentic tasks
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
MultiSearch uses parallel multi-query retrieval plus explicit merging inside a reinforcement-learning loop to improve retrieval-augmented reasoning, outperforming baselines on seven QA benchmarks.
citing papers explorer
-
ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents
ARBOR introduces a reusable rubric buffer that consolidates contrastive trajectory drafts into cross-query rubrics for online process rewards, outperforming GRPO and DAPO on multi-hop QA benchmarks.
-
Scaling Retrieval-Augmented Reasoning with Parallel Search and Explicit Merging
MultiSearch uses parallel multi-query retrieval plus explicit merging inside a reinforcement-learning loop to improve retrieval-augmented reasoning, outperforming baselines on seven QA benchmarks.