Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 28 July 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2505.16282.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-07-27T06:30:09.085275+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T12:50:16.625077Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:00:08.408321Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation efa7606d-90a1-4c89-a24d-f755e76324f5 · inbound
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.
Observation f7a6fa83-59c1-42f8-bd82-ecc76dcd26e5 · inbound
Agentic Reasoning for Large Language Models ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.
Observation a00a8d54-94c6-4eff-8a4b-afedab8048e3 · inbound
Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.
Observation 32ddfcb8-f0f4-4f19-ad06-b0543b4f3018 · inbound
Faithful Mobile GUI Agents with Guided Advantage Estimator ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.
Observation 3ab69e72-872e-4417-b40f-66d6b72ad9eb · inbound
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.
Observation 61a0ea97-88b7-4880-83e2-222fca09fa1a · inbound
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.
Observation e43a52a5-801e-44c0-a4e7-1f00e7d6cdf0 · inbound
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.
Observation 494b1bdb-c798-4aa8-b384-8e948dd3bd4f · inbound
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.
Observation ad7335ec-cd17-4d1b-85cc-f1e71fcfc475 · inbound
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.
Observation 44daa594-25cc-420c-a599-f689908c2cdb · inbound
StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.
Observation 46e7f937-ef14-4065-9ae5-ae9c9c4aa368 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.
Observation 3450fa5e-24d1-467e-99a6-dbf7f0c44049 · inbound
Learning with a Single Rollout via Monte Carlo Pass@k Critic ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.
Observation 490492e7-74a4-411e-afa8-b66bcd8c2a8d · inbound
DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-07-27T06:30:09.085275+00:00.