Pith. sign in

hub Canonical reference

arXiv preprint arXiv:2503.23383 , year=

Canonical reference. 70% of citing Pith papers cite this work as background.

30 Pith papers citing it
1 external citations · external index
Background 70% of classified citations

hub tools

citation-role summary

background 8 baseline 1 other 1

citation-polarity summary

years

2026 23 2025 7

representative citing papers

Agent Explorative Policy Optimization for Multimodal Agentic Reasoning

cs.CL · 2026-05-27 · unverdicted · novelty 6.0 · 2 refs

AXPO addresses the Thinking-Acting Gap in agentic RL training by targeted resampling of tool calls in all-wrong subgroups, delivering +1.8pp gains over GRPO on nine multimodal benchmarks with an 8B model beating a 32B baseline on Pass@4.

Harnessing LLM Agents with Skill Programs

cs.AI · 2026-05-18 · conditional · novelty 6.0

HASP upgrades textual skills into executable Program Functions that intervene in LLM agent loops at inference, post-training, or self-evolution, delivering 25% gains over ReAct and 30.4% over Search-R1 on reasoning benchmarks.

APPO: Agentic Procedural Policy Optimization

cs.LG · 2026-06-10 · conditional · novelty 5.0

APPO improves LLM agent training by branching at tokens selected for both uncertainty and future impact, then scaling credit for consequential reasoning procedures.

citing papers explorer

Showing 30 of 30 citing papers.