Pith. sign in

Canonical reference

Arpo: End-to-end policy optimization for gui agents with experience replay

Canonical reference. 80% of citing Pith papers cite this work as background.

12 Pith papers citing it
Background 80% of classified citations

citation-role summary

background 5

citation-polarity summary

years

2026 11 2025 1

roles

background 5

polarities

background 4 unclear 1

representative citing papers

Faithful Mobile GUI Agents with Guided Advantage Estimator

cs.AI · 2026-05-02 · unverdicted · novelty 7.0

Faithful-Agent raises Trap SR in GUI agents from 13.88% to 80.21% via faithfulness-oriented SFT and GuAE-enhanced RFT with consistency rewards while retaining general performance.

Learning with a Single Rollout via Monte Carlo Pass@k Critic

cs.LG · 2026-06-24 · unverdicted · novelty 6.0

SR-PPO trains a Pass@k critic from single-rollout Monte Carlo outcomes to enable token-level advantage estimation in language model RL, yielding stable training and Pass@128 gains on math benchmarks.

Agentic Reasoning for Large Language Models

cs.AI · 2026-01-18 · unverdicted · novelty 4.0

The survey structures agentic reasoning for LLMs into foundational, self-evolving, and collective multi-agent layers while distinguishing in-context orchestration from post-training optimization and reviewing applications across domains.

citing papers explorer

Showing 12 of 12 citing papers.