Pith. sign in

Paper Citation Record · LEDGER

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 17 inbound Pith citation observations for arXiv:2505.16282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16282 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:10.796768Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:19:15.434119Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.408321Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a97a4621-6c2e-4838-8b44-aac68c14df56 · outbound

This paper cites Qwen2.5-VL Technical Report.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:08.732949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:08.732949Z digest=sha256:464506ef6eacbb9e5d5f1052edd428cbd5c115c1fdfb2a2d2181f85e9e640d7c

Observation a5544cf9-7703-4840-aa66-c648472154de · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:08.818285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:08.818285Z digest=sha256:09e5953512ed5d5465be728b35dc919ba54e73851a4b5920bd63a95723f9867f

Observation 06a674f1-8a7e-48ab-8c65-6f05654887e2 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay KTO: Model Alignment as Prospect Theoretic Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:08.943210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:08.943210Z digest=sha256:46202c8b45641ee06932661ffa37d96602defc4d22f6b4e5ee9f8467afbf8c3b

Observation a204acfe-dfc4-4120-a7ae-98cec497d88e · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.054423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.054423Z digest=sha256:ee20de0bfab614facb9e9ee00f5b554018dcf5105d69a68fc1141ba6154637c5

Observation 96a71639-ea78-4543-8235-922a76b611cc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.148499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.148499Z digest=sha256:23aefa474795e41b4dfa69343c2352a8ca7c2d45b2ed5450437a5535f0161468

Observation f2923ec7-65ca-4e66-a110-9ab98fa6804c · outbound

This paper cites Cogagent: A visual language model for gui agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Cogagent: A visual language model for gui agents

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.440915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:07:09.257984Z digest=sha256:64c2c097d29664fef36101bc464cb49b89b8ee615fec415f2c8c725543446e07

Observation 08626187-7624-473a-848e-7420015aef08 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.358618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.358618Z digest=sha256:12e89f8a39c8369f78b6dea9c478b702f2621384e4ed3c4e5c69a1e35e44efa7

Observation 48d13b54-9277-4d20-9f3b-7d8563c4eae4 · outbound

This paper cites OpenAI o1 System Card.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay OpenAI o1 System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.452790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.452790Z digest=sha256:5bde93a820afd196150303953233b9f3a334fa6abf447452d937c1e0a9d74933

Observation 4b5481a0-74d7-401d-881c-9dac23cee2e4 · outbound

This paper cites OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.526745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.526745Z digest=sha256:d7128773bfd850616ad53c253835f468b03768c3bdc505bc713d7ac23172f89e

Observation e0a72d0c-1f2b-439e-b76a-0815a1f77501 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Efficient memory management for large language model serving with pagedattention

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.374189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:07:09.645013Z digest=sha256:022bab10b087da7b6f8199c096ad9be6fc197d16e40b4e16d553da00071969ab

Observation 723a7e63-be07-4902-9360-6ab11888e498 · outbound

This paper cites Code-r1: Reproducing r1 for code with reliable rewards.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Code-r1: Reproducing r1 for code with reliable rewards

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.737013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.737013Z digest=sha256:304abca0e1cfbe187e19705ee724d71d58f5a5cfc6230fd3ff97ca6776652566

Observation 8ea43a1d-8027-47f5-b596-6acbc00c0eb8 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.853262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.853262Z digest=sha256:28d34a4b68b66b01da7dda264d83390a33e7c9b409af59d85ff8e49e0c1f9c53

Observation bc7d8539-4150-48cb-832c-6b4548dd2e82 · outbound

This paper cites Decoupled Weight Decay Regularization.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Decoupled Weight Decay Regularization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.981894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.981894Z digest=sha256:e9372a91e69f035afca0d3055f3c944ccd15aa15236f2e5b9c0b21f941f4edbb

Observation 2a897a41-6265-4c2c-abd7-a4289024e73f · outbound

This paper cites ScreenAgent: A Vision Language Model-driven Computer Control Agent.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay ScreenAgent: A Vision Language Model-driven Computer Control Agent

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.028219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.028219Z digest=sha256:050096494825b05146e4b18e5d573665b50ada40bc2ccb2454799aefa2ae1488

Observation 53b21184-ee19-4e15-a7de-b8d3d4fa34d7 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay ToolRL: Reward is All Tool Learning Needs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.067180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.067180Z digest=sha256:90a8cdd0271907a11872c4c8a8310e1ff0dd73b9a4f96d416aa96b3dff664442

Observation 202e5032-e30c-4814-905a-6ed48af0b141 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.120332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.120332Z digest=sha256:f2e5f0155c4e18e7742181d6f718100c34ca70f98e1203ffddfb5d0a30cff03b

Observation 15b11a46-da31-4f47-aea5-c4c68fc8cadd · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Direct preference optimization: Your language model is secretly a reward model

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.288367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:07:10.174192Z digest=sha256:4ec7761824b249dbb8869f8f8b0c11454d5adfb8aed3e60da462f7fb9b8fc5a5

Observation 1f8adde8-ef64-4dfe-a2b9-59b519a037ef · outbound

This paper cites Proximal Policy Optimization Algorithms.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Proximal Policy Optimization Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.211575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.211575Z digest=sha256:8aecebbcc271f4e20f1294ebe69c95b46af2e5327abab5f7ecf2279a69bf8ae2

Observation 1d28fa66-0dad-4959-9ec7-efc2a223b9c2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.279353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.279353Z digest=sha256:d2ea137b70ca6da567ec5c6cd30c8513fcb0131250eac04706a8ab1ff411d856

Observation 440ba8ab-3650-4f73-bf52-0b10af21267c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay HybridFlow: A Flexible and Efficient RLHF Framework

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.332778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.332778Z digest=sha256:33f7f4af80557568763220dcacd00f782c980332100906ea2f7e4a9d960e1d91

Observation 6ababca5-5736-47a2-924b-fd00f9d5b887 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.363216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.363216Z digest=sha256:bcb52b017601f5de87a7a29f70ddb3b68d504a9a35aa452524816825ecea9845

Observation eda3ee73-fe85-4711-a4cc-bec481c35515 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Chain-of-thought prompting elicits reasoning in large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.194077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:07:10.403817Z digest=sha256:fe5945671161d492ab87d7ac84e7dc8b1eadbe0051b98b7f2ad93a9e9c3916e6

Observation 8f180272-3fd7-41f7-af45-ed5731c23050 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.444576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.444576Z digest=sha256:1ad9887d3728b57ee1094f3220490fb8865a3dd2e626577a1f60ef2f4b82ef46

Observation 58f4df80-56d4-438d-8f34-5af8fc91a67a · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.491309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.491309Z digest=sha256:7aa6a16ec712f248a94249ca41f934c4b69cff9c6ad54ac8bbe403b72950c029

Observation f4d6f9ea-76b2-4537-a06c-9ecfa4747bd3 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.128405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T15:07:10.537390Z digest=sha256:ef33afcff6298a16bbd109bd9b9e1ccc6b8efc637ceec035f550253518ef29ef

Observation ca674936-63b4-4926-aa11-f4c6420fc5b6 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.588099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.588099Z digest=sha256:f9b8c51c2f7f7cbd070ed4864d43b7355e911aec18e26681591683fa321442a9

Observation a363f308-e927-4e0c-9fbf-e202d39c717e · outbound

This paper cites Aria-UI: Visual Grounding for GUI Instructions.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Aria-UI: Visual Grounding for GUI Instructions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.649086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.649086Z digest=sha256:8038fea0313ceb472db27d4d9d7ca0f8bb278e2a20b3853bc5a863073bf69a93

Observation f9423a99-3bc6-4f52-a8cb-ee85a07db547 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.758735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.758735Z digest=sha256:3b41ca4a0833dd1c4841ff429b8e8ca350e2aba0f6124e9de3480cf4024e61d8

Observation b169f929-fb11-4c83-8be7-a284473d8032 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.796768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.796768Z digest=sha256:2cfce8af0e2376937dba3d7132500d58afc56086bca344c5e6b1afc00e190e9d

Pith citing papers

Observation c7df2476-8c0a-48d9-8a9f-2cf60b09acd7 · inbound

MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment cites this paper.

MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:14.276099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:14.276099Z digest=sha256:1f593f765a2581e223b8d9b62d9a91d57663d0ae1d0ecea34d4c6c685eb60d62

Observation efa7606d-90a1-4c89-a24d-f755e76324f5 · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.047730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:75e04e9703867d4a95907fd99d49bf97bd57246994a09e392708602e0c7a46cd

Observation 0acd32bd-456b-4de0-87eb-b50696880523 · inbound

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning cites this paper.

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:52:27.271355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:52:27.271355Z digest=sha256:b59430b7c002d14a285281c8bc72cd7401be52768ff30e4f2affe48f1866c161

Observation 61c88db6-3fd5-4294-852b-ee3cc6395ca1 · inbound

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning cites this paper.

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.322149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:30:39.322149Z digest=sha256:be9b9ff84e33526730afade59e80c797223e03a1e7a0682ba6166bd3fe4db5f1

Observation f7a6fa83-59c1-42f8-bd82-ecc76dcd26e5 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.866371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:d96321aedff4fd303f78830a5680cda5d0e91b1f9e668db88de5c646fe36ad84

Observation a00a8d54-94c6-4eff-8a4b-afedab8048e3 · inbound

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization cites this paper.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:36:35.212903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:18e5c0a087c7052d8f2631b72237285304d2c95d6b61b7663c22b92b3d178d66

Observation 32ddfcb8-f0f4-4f19-ad06-b0543b4f3018 · inbound

Faithful Mobile GUI Agents with Guided Advantage Estimator cites this paper.

Faithful Mobile GUI Agents with Guided Advantage Estimator ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:24.147115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T15:17:05.477466Z digest=sha256:cff94f96971726f00dda758668c9c863f41f566477ed423e2b23add28a5fbd24

Observation 3ab69e72-872e-4417-b40f-66d6b72ad9eb · inbound

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length cites this paper.

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:40:40.641959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T18:13:25.735085Z digest=sha256:726c75bc053f481e93ef10b7e525a7e4967e914c92f1579fcf33a52518a3e21b

Observation 61a0ea97-88b7-4880-83e2-222fca09fa1a · inbound

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization cites this paper.

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:17.873771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T03:00:01.549385Z digest=sha256:3d17be90600fecca34c7b1f6a427e3c0a288c433c003f18d815535169acb296d

Observation e43a52a5-801e-44c0-a4e7-1f00e7d6cdf0 · inbound

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization cites this paper.

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:07.683646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T00:58:25.205386Z digest=sha256:cc17dd99629f055c8581ea3d86d6075bd9c446392e9a004a2251cbba5a4f8f6e

Observation 494b1bdb-c798-4aa8-b384-8e948dd3bd4f · inbound

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents cites this paper.

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.365420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T03:51:54.401200Z digest=sha256:dcd9f28852c2428ac602ba023ee73a719d420e8f06049b4fc00bc3fd2bcfe856

Observation ad7335ec-cd17-4d1b-85cc-f1e71fcfc475 · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.369302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:9a5e5a677b3ad512e4187f881ee56a8f1adf4ae5c2340f2237db1a04be0fce5f

Observation 44daa594-25cc-420c-a599-f689908c2cdb · inbound

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents cites this paper.

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:47:09.594291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T22:24:41.353303Z digest=sha256:91a9d6ef1a527758d22d187baf4c623c98ce46f609c937adab920be1cf697959

Observation 46e7f937-ef14-4065-9ae5-ae9c9c4aa368 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.731617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:d9ad0300d6889a053ec427871f08c01bef442c9cbc798bf70ddc588a811586c8

Observation 3450fa5e-24d1-467e-99a6-dbf7f0c44049 · inbound

Learning with a Single Rollout via Monte Carlo Pass@k Critic cites this paper.

Learning with a Single Rollout via Monte Carlo Pass@k Critic ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:00:08.409789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-25T20:53:10.016447Z digest=sha256:8f55e8974e21d1cc0d3b375020f8d02b48bd7f9541e879fdfd42e16880ae7707

Observation 490492e7-74a4-411e-afa8-b66bcd8c2a8d · inbound

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents cites this paper.

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:53:26.532661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T12:50:16.625077Z digest=sha256:5d516fa83f8896c426c75d16a36e58c0c26ec632c890936da0438c0792fbd454

Observation 6632b7f6-e440-45f9-bf76-c7f0b764d1da · inbound

Software Engineering for and with GUI Agent cites this paper.

Software Engineering for and with GUI Agent ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-11T20:19:15.434119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:19:15.434119Z digest=sha256:540fc78760029d5895a0797ecc343a7cb19ba71eed10d7e4cc1788a56cbfe806