Pith. sign in

hub Mixed citations

Agentgym-rl: Training llm agents for long-horizon decision making through multi-turn reinforcement learning

Mixed citation behavior. Most common role is background (50%).

17 Pith papers citing it
Background 50% of classified citations

hub tools

citation-role summary

background 3 baseline 2 method 1

citation-polarity summary

years

2026 16 2025 1

representative citing papers

SkillGrad: Optimizing Agent Skills Like Gradient Descent

cs.AI · 2026-05-26 · unverdicted · novelty 6.0

SkillGrad applies a gradient-descent-inspired optimization to structured agent skill packages using trajectory-level losses, diagnostic text gradients, momentum accumulation, and LLM-based patching, yielding 6.7 percentage point average gains over training-based baselines on SpreadsheetBench Verifie

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning

cs.AI · 2026-05-13 · unverdicted · novelty 6.0

ICRL uses joint RL training of solver and critic with distribution-calibration re-weighting and role-wise advantage estimation to internalize critique into unassisted LLM performance, yielding 6.4-point gains on agentic tasks and 7.0 on math reasoning with Qwen3 models.

Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Control

cs.LG · 2026-05-12 · unverdicted · novelty 6.0 · 2 refs

Entropy polarity is a signed token-level quantity derived from a first-order approximation of entropy change that predicts whether RL updates expand or contract policy entropy in LLM fine-tuning, revealing an asymmetry between high- and low-probability tokens.

GRAFT: Graph-Tokenized LLMs for Tool Planning

cs.LG · 2026-05-12 · unverdicted · novelty 6.0

GRAFT internalizes tool dependency graphs via dedicated special tokens in LLMs and applies on-policy context distillation to achieve higher exact sequence matching and dependency legality than prior external-graph methods.

citing papers explorer

Showing 17 of 17 citing papers.