Pith. sign in

Paper Citation Record · LEDGER

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

As of 28 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2505.16282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16282 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-27T06:30:09.085275+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T12:50:16.625077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.408321Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation efa7606d-90a1-4c89-a24d-f755e76324f5 · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.047730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:e7be357aed3222d87dc33030a10d3fe78bf62ff54b19ce3023e0935ce8160999

Observation f7a6fa83-59c1-42f8-bd82-ecc76dcd26e5 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.866371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:0258a38c50b572d19a8a65f8d165a084b1de6ddd2049c1de53b45b1441d93afb

Observation a00a8d54-94c6-4eff-8a4b-afedab8048e3 · inbound

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization cites this paper.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:36:35.212903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:3338766e50970b4186ca0089e881e92198e9517f4478bc971fb892aa02c1feed

Observation 32ddfcb8-f0f4-4f19-ad06-b0543b4f3018 · inbound

Faithful Mobile GUI Agents with Guided Advantage Estimator cites this paper.

Faithful Mobile GUI Agents with Guided Advantage Estimator ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:24.147115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-09T15:17:05.477466Z digest=sha256:5de502a821420bbbd8452290904a521239cca4365efa070af464c55250f54598

Observation 3ab69e72-872e-4417-b40f-66d6b72ad9eb · inbound

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length cites this paper.

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:40:40.641959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=arxiv_source observed=2026-05-08T18:13:25.735085Z digest=sha256:4dc7f6c8cab7ad8db3f2c45289bcbc9902a198ba66783d9524c862338dee4a9d

Observation 61a0ea97-88b7-4880-83e2-222fca09fa1a · inbound

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization cites this paper.

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:17.873771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-12T03:00:01.549385Z digest=sha256:b530b8faffefe2236099f3c2805c1155bbe9eb198d6852bd6558e4b1f96204b6

Observation e43a52a5-801e-44c0-a4e7-1f00e7d6cdf0 · inbound

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization cites this paper.

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:07.683646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-13T00:58:25.205386Z digest=sha256:ba7bfad6afa98a1510dbafc43d689b95bfd9f2dc82c7ed774e24c2a8d0d9ce8c

Observation 494b1bdb-c798-4aa8-b384-8e948dd3bd4f · inbound

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents cites this paper.

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.365420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-05-13T03:51:54.401200Z digest=sha256:f1ad23df7042a0dfeee99e5ac23666181c2249384ee150442a27454f0233a82c

Observation ad7335ec-cd17-4d1b-85cc-f1e71fcfc475 · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.369302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:e29fa33581f9c4f3acac1180ef63b49ee7ac777d78768c6fa9ec99521d602857

Observation 44daa594-25cc-420c-a599-f689908c2cdb · inbound

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents cites this paper.

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:47:09.594291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-06-27T22:24:41.353303Z digest=sha256:34c3d846ade96ce83b119524043188bd940b6bbc6ae5d738c09765575a7f418e

Observation 46e7f937-ef14-4065-9ae5-ae9c9c4aa368 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.731617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:4c82f7c425dcb04b6b506699bf8fc310150dcf61e07293a541d89beb60800ea1

Observation 3450fa5e-24d1-467e-99a6-dbf7f0c44049 · inbound

Learning with a Single Rollout via Monte Carlo Pass@k Critic cites this paper.

Learning with a Single Rollout via Monte Carlo Pass@k Critic ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:00:08.409789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=arxiv_source observed=2026-06-25T20:53:10.016447Z digest=sha256:bb5f1c563a5fdc197cc9a611d08402fdea1764976b23032b730a1946650567d3

Observation 490492e7-74a4-411e-afa8-b66bcd8c2a8d · inbound

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents cites this paper.

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:53:26.532661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.

source=pdf_text observed=2026-06-29T12:50:16.625077Z digest=sha256:6682261d30151dc605dc4eabba47ba12208879118c450f092daa928c9a2b55cc