Pith. sign in

hub

Tree search for llm agent reinforcement learning

23 Pith papers cite this work. Polarity classification is still indexing.

23 Pith papers citing it

hub tools

citation-role summary

background 3 baseline 1

citation-polarity summary

years

2026 21 2025 2

representative citing papers

Uno-Orchestra: Parsimonious Agent Routing via Selective Delegation

cs.AI · 2026-05-06 · unverdicted · novelty 6.0

A learned orchestration policy for LLM agents that jointly optimizes task decomposition and selective routing to (model, primitive) pairs, delivering 77% macro pass@1 at 10x lower cost than strong baselines across 13 benchmarks.

APPO: Agentic Procedural Policy Optimization

cs.LG · 2026-06-10 · conditional · novelty 5.0

APPO improves LLM agent training by branching at tokens selected for both uncertainty and future impact, then scaling credit for consequential reasoning procedures.

citing papers explorer

Showing 23 of 23 citing papers.