This survey introduces the Generate-Filter-Control-Replay (GFCR) taxonomy to structure rollout pipelines for RL-based post-training of reasoning LLMs.
T ree RL : LLM reinforcement learning with on-policy tree search
4 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.
years
2026 4verdicts
UNVERDICTED 4representative citing papers
HarnessForge co-evolves harness-policy pairs in LLM agents via fault-guided tailoring and alignment, reporting up to 12% gains over single-component baselines on five benchmarks.
Bidirectional Evolutionary Search augments autoregressive expansion with evolutionary recombination operators and dense backward subgoal feedback to produce better candidates than standard best-of-N or tree search for language model self-improvement.
LEAF recovers tree structure from rollout batches for span-level credit assignment in GRPO-style speech LLM post-training, improving over baselines on QA and translation tasks.
citing papers explorer
-
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning
This survey introduces the Generate-Filter-Control-Replay (GFCR) taxonomy to structure rollout pipelines for RL-based post-training of reasoning LLMs.
-
HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems
HarnessForge co-evolves harness-policy pairs in LLM agents via fault-guided tailoring and alignment, reporting up to 12% gains over single-component baselines on five benchmarks.
-
Self-Improving Language Models with Bidirectional Evolutionary Search
Bidirectional Evolutionary Search augments autoregressive expansion with evolutionary recombination operators and dense backward subgoal feedback to produce better candidates than standard best-of-N or tree search for language model self-improvement.
-
LEAF: Growing Trees Without Branching for Speech-Aware Large Language Model Post-Training
LEAF recovers tree structure from rollout batches for span-level credit assignment in GRPO-style speech LLM post-training, improving over baselines on QA and translation tasks.