Pith. sign in

Paper Citation Record · LEDGER

ProgRM: Build Better GUI Agents with Progress Rewards

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 5 inbound Pith citation observations for arXiv:2505.18121.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18121 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:49.267563Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T19:26:37.281563Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T16:47:09.687885Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aec7cd57-76d8-4c11-a7ff-73119782bc71 · outbound

This paper cites Digi-q: Learning vlm q-value functions for training device-control agents.

ProgRM: Build Better GUI Agents with Progress Rewards Digi-q: Learning vlm q-value functions for training device-control agents

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.949072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:45.129305Z digest=sha256:701dbc09cadccfa8b02da130b5034a4992f40fa5a058d3676dc8c620aba811cc

Observation f7254977-e58d-4cd0-abab-2582825f887b · outbound

This paper cites Digirl: Training in-the-wild device-control agents with autonomous re- inforcement learning.

ProgRM: Build Better GUI Agents with Progress Rewards Digirl: Training in-the-wild device-control agents with autonomous re- inforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.762133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:45.198300Z digest=sha256:dd2db9e823e3dd8d3cc739253f21080c17dc6f7a01b1e64e4fef63036095e897

Observation a5895103-8609-41bb-847b-5452b5004258 · outbound

This paper cites Learning about progress from experts.

ProgRM: Build Better GUI Agents with Progress Rewards Learning about progress from experts

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.581616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:45.251487Z digest=sha256:b851a80b869151ef286b0311e7b03bff88af23d1dcea6db6a37543146f5ad964

Observation 879f22e9-00ab-40d8-9a23-badacf93344a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ProgRM: Build Better GUI Agents with Progress Rewards Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.337260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.337260Z digest=sha256:d22693e24a02a419d8318e374f781e614ea6a885f73c051f9c28e9555d5b0162

Observation 53605823-cbb5-4d82-80f2-6a76092eb65e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.419221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.419221Z digest=sha256:77dfa7372c38a1f0d11a0b881fd21953e225b57d1118bd1cfbf403f0e0ad59c0

Observation 9bb8ec98-2eac-4dc7-8c12-73a5d9dd3a60 · outbound

This paper cites Dungeons and data: A large-scale nethack dataset.

ProgRM: Build Better GUI Agents with Progress Rewards Dungeons and data: A large-scale nethack dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.400251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:45.508048Z digest=sha256:db73443a3b398988eba23e4e7ebe07b2f2043284cb0673dc7727eb4c6119a73f

Observation e7dc26f3-39d4-42cb-91c6-261d8a91c241 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

ProgRM: Build Better GUI Agents with Progress Rewards REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.597283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.597283Z digest=sha256:a88757f03ea9b631a6dfdd6b95eb202d0a3ddf91565faf544ef5798a3546ee52

Observation 58e481f7-a907-4cbb-9051-e260b61b71d4 · outbound

This paper cites Exploring Expert Failures Improves LLM Agent Tuning.

ProgRM: Build Better GUI Agents with Progress Rewards Exploring Expert Failures Improves LLM Agent Tuning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:37:50.068075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:45.644348Z digest=sha256:6580b39b7e8de753cb9b903f5d99dbba95715ad0ec3d7e1c5defe76109a08b93

Observation a4ae53f1-e8c2-4bcc-ac12-def7e240e590 · outbound

This paper cites Making language models better reasoners with step-aware verifier.

ProgRM: Build Better GUI Agents with Progress Rewards Making language models better reasoners with step-aware verifier

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.740939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.740939Z digest=sha256:44b8c83cdce6f9378259c8019da3764be1aa5b2b0f018fe03d22b6aadf419ef2

Observation 8b426dbb-a6d4-449f-b9d5-68468ac898a7 · outbound

This paper cites Let’s verify step by step.

ProgRM: Build Better GUI Agents with Progress Rewards Let’s verify step by step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.786546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.786546Z digest=sha256:b91e453a601bb464c04b90491b96354fa38845e2512314e98157686b9249891a

Observation 4a6e6515-a8a4-4cac-a2c8-9cca2eecd7b2 · outbound

This paper cites UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.862230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.862230Z digest=sha256:4c98eca9f785d4b164294025be5d8381b8ed3892e31ddec62c8d9a308a06e3a1

Observation 50f6f860-4632-4c42-8826-d133a6404fdb · outbound

This paper cites NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild.

ProgRM: Build Better GUI Agents with Progress Rewards NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.943894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.943894Z digest=sha256:1ebc9f6e2a0795a4b234cc448c722695fc042beb2ff6d4b76dca1bfa5dacc193

Observation 47146464-9f0a-4fd5-8d7f-bafd72db849c · outbound

This paper cites Xu, Aman Madaan, Jiarui Liu, Robert Lo, Abishek Sridhar, Sudipta Sengupta, Dan Roth, Graham Neubig, and Shuyan Zhou.

ProgRM: Build Better GUI Agents with Progress Rewards Xu, Aman Madaan, Jiarui Liu, Robert Lo, Abishek Sridhar, Sudipta Sengupta, Dan Roth, Graham Neubig, and Shuyan Zhou

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.202532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:46.013705Z digest=sha256:18ea23aed815705fdce0d9cec697507f32bd5eef740ccfefbb2119782fc91dfd

Observation f3f932b9-fe28-43b6-9684-dfa389c944ac · outbound

This paper cites Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents.

ProgRM: Build Better GUI Agents with Progress Rewards Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.119179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.119179Z digest=sha256:bfb6859044b7c57defe9597cb830631a569ac1295e98c613e85665945caaba3f

Observation b5b3b549-039f-4284-99eb-97810c3fe91c · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

ProgRM: Build Better GUI Agents with Progress Rewards Autonomous Evaluation and Refinement of Digital Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.228152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.228152Z digest=sha256:636d00271191d7278e3148d38083a272ab971d89bad21fd3b868022432ca5c96

Observation 4efd62e6-4159-4d0d-b049-b694982913da · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.329762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.329762Z digest=sha256:3dc87a45676c19c02a82c2bddffd156a72e06100a72bccd39eb31520242cbbed

Observation 8e89bd61-c05a-4608-9851-7de29189cb15 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

ProgRM: Build Better GUI Agents with Progress Rewards UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.431629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.431629Z digest=sha256:4d8ff3cf81a75570cfc4d7ba8387c761ee522121d9150f0d5be42e4861ce07c0

Observation 18d80b30-3761-4e5e-b55e-c5e36827d8ce · outbound

This paper cites Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning.

ProgRM: Build Better GUI Agents with Progress Rewards Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.502834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.502834Z digest=sha256:3b6ac0bde7c293407234f28dc9756ab933660d697a1cd9fc838575e350dec6e0

Observation d4bdfdd7-bfd0-451a-a2a5-ffbd21fb45d1 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert- networks.

ProgRM: Build Better GUI Agents with Progress Rewards Sentence-bert: Sentence embeddings using siamese bert- networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.001358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:46.642165Z digest=sha256:cc484cc1f1965b2bb428afe2d27bd24820dddfc23700921da908ffbda31a240c

Observation c464d765-6783-4c58-ba4f-a10440cdcbe1 · outbound

This paper cites Step: Stacked llm policies for web actions.

ProgRM: Build Better GUI Agents with Progress Rewards Step: Stacked llm policies for web actions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.788218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:46.921764Z digest=sha256:b38e7ec341af47015d13bdd4fad08b5093d813a26ae2594a8eddbaf0d5d542c1

Observation c9cf3a05-12da-493e-98f2-dce10c1ab90d · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.033881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.033881Z digest=sha256:a51e8d206c86290d1d8d534c0e64d8aa7d16ddc68ccbec93c94375045748fe41

Observation 3eb32102-56d0-490f-a9e5-92e377d9949e · outbound

This paper cites Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments.

ProgRM: Build Better GUI Agents with Progress Rewards Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.108403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.108403Z digest=sha256:2b15a7515d2afbdef8b5e823b278eeb02ae0eb6e256e0f6885a0fb66af546c00

Observation 486f8fcf-9f3c-4375-be1f-f7656a1a18bf · outbound

This paper cites OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis.

ProgRM: Build Better GUI Agents with Progress Rewards OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.185783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.185783Z digest=sha256:5a0b18a2828f18218daf32baabdc5363b29f32b5e0d10aa42a04350a0930f863

Observation c6421f46-7dfa-4f7a-8464-a8c8d1c8b002 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

ProgRM: Build Better GUI Agents with Progress Rewards Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.238146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.238146Z digest=sha256:1c19888612f66d9cec25511fed90ea3c4be846f2953e041016056c2516f7652c

Observation 19a62709-4c66-4c0c-b607-d22e1aa6c55a · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

ProgRM: Build Better GUI Agents with Progress Rewards Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.380315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.380315Z digest=sha256:69ad85ec8a2583a709c91f3ea499bc6254a8bf54bbc42d8026311c017ea54983

Observation 053bc850-e181-465d-8960-9813e0b5f08f · outbound

This paper cites DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents.

ProgRM: Build Better GUI Agents with Progress Rewards DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.485240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.485240Z digest=sha256:ddda84e3b906fdc632026b0685e9715d794d9bc1623b9760cde53a5b9d449ace

Observation 8029fb17-4eaa-44da-a0b9-37a8b23ba8c9 · outbound

This paper cites Reinforcing Language Agents via Policy Optimization with Action Decomposition.

ProgRM: Build Better GUI Agents with Progress Rewards Reinforcing Language Agents via Policy Optimization with Action Decomposition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.590688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.590688Z digest=sha256:cd414a2f7e94eb659b4bbeebf0713915ed658c3bad89cb7672ce89ab974dc179

Observation a3e49663-2092-4537-8873-2260e3312495 · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

ProgRM: Build Better GUI Agents with Progress Rewards GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.700415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.700415Z digest=sha256:854173c475580458c6d13b343d3dfee3b8ad3c88936aa413aa77b8ba3ad1f346

Observation 19c05f48-6249-4353-acce-6092cd98e096 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.804784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.804784Z digest=sha256:4b2473f4e01cb9b86db1fcb52438c82f9533d144dbeb464cee3e3c47d7077e0b

Observation 5a6a36ef-b5cf-4007-84c6-d943abba95e4 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

ProgRM: Build Better GUI Agents with Progress Rewards Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.937341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.937341Z digest=sha256:0debdf51bf7e52e5edfe541a0e506c1359bdae9cb8c35cf0ab234c4743eed97d

Observation 835c7d55-3d95-4219-9486-da9488b80267 · outbound

This paper cites Agenttrek: Agent trajectory synthesis via guiding replay with web tutorials.

ProgRM: Build Better GUI Agents with Progress Rewards Agenttrek: Agent trajectory synthesis via guiding replay with web tutorials

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.672358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:48.039904Z digest=sha256:48cfa9ed39f944868f74da6e61f9ddb4bc0a2d89849d73de3e3ddefc0dbf1a34

Observation 3450faf6-7e13-4dc7-8f4e-c5518ff85d63 · outbound

This paper cites Qwen2.5 Technical Report.

ProgRM: Build Better GUI Agents with Progress Rewards Qwen2.5 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.119240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.119240Z digest=sha256:68d08772faa65c4b90058799542791ef2fc67f50899559be99ec5d9847f194fb

Observation bd0674c3-a3c1-4f74-af92-f26678f2827c · outbound

This paper cites OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning.

ProgRM: Build Better GUI Agents with Progress Rewards OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.213972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.213972Z digest=sha256:ed9f66d1454b9f538fb510de41b5c52a0f7a2545959915821abc6c053de8967e

Observation 750b3d33-83c5-4606-8ed4-199bcbfa0d59 · outbound

This paper cites Free Process Rewards without Process Labels.

ProgRM: Build Better GUI Agents with Progress Rewards Free Process Rewards without Process Labels

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.320653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.320653Z digest=sha256:9fba3fc655032a737bbee674c8f1b3318b56d9c9663105c43798cd59e6f68288

Observation 5dc867a9-8a39-4ef1-ae06-9424c5cdf5d8 · outbound

This paper cites UFO: A UI-Focused Agent for Windows OS Interaction.

ProgRM: Build Better GUI Agents with Progress Rewards UFO: A UI-Focused Agent for Windows OS Interaction

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.529789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.529789Z digest=sha256:eebed8c09936b1659878783a875ccce362e871a8ea827eab006da8fee883f6ed

Observation 8db64b3d-63b0-47ef-b6ef-d92909160a9a · outbound

This paper cites Appagent: Multimodal agents as smartphone users.

ProgRM: Build Better GUI Agents with Progress Rewards Appagent: Multimodal agents as smartphone users

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.664282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.664282Z digest=sha256:2c1f4e3b53e0d52d53135e6f5d59413c11984409482a6837804a91d49d55532e

Observation d601e41c-ad3e-4be6-92fe-9e5ad8d03326 · outbound

This paper cites Large language models are semi-parametric reinforcement learning agents.

ProgRM: Build Better GUI Agents with Progress Rewards Large language models are semi-parametric reinforcement learning agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.549721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:48.771877Z digest=sha256:b20d3855cb3894d23cf39d38805d39454f4b966f986491ce7abb4370aec5bc1e

Observation 32ad9c9e-7ec1-40c1-99e7-bdf899b9763d · outbound

This paper cites Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction.

ProgRM: Build Better GUI Agents with Progress Rewards Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.872226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.872226Z digest=sha256:d2a5575cfd498bcc0c0291e368cadcdded48f016186f5719194948462701859b

Observation 129c17e0-bb24-4cf3-ad7e-bbdcd6ef5525 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

ProgRM: Build Better GUI Agents with Progress Rewards The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.959454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.959454Z digest=sha256:72b37acdc285e0fd8dc9de9a60d67b3d89e6ec0fd8d0f1d37741ddc03b46a641

Observation a0dee5fd-9a2b-44d8-a18e-b0cb28957dc1 · outbound

This paper cites Gpt-4v(ision) is a generalist web agent, if grounded.

ProgRM: Build Better GUI Agents with Progress Rewards Gpt-4v(ision) is a generalist web agent, if grounded

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.386446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:49.041108Z digest=sha256:9d9f6229b41700976da5a1c4415602ccd024d83db533e24f5acd23e40ec2e376

Observation 54802348-9e11-49bd-a93e-166b08e21376 · outbound

This paper cites VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model.

ProgRM: Build Better GUI Agents with Progress Rewards VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.160588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.160588Z digest=sha256:7faa7cd3bd3969edc369a0c2eb67a8afa3f381d6411361b14dbac8c1f29eb4d3

Observation 901c9048-5ede-4c28-ae38-b51dac1b22fd · outbound

This paper cites Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents.

ProgRM: Build Better GUI Agents with Progress Rewards Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.207033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.207033Z digest=sha256:6b17554964025b1d88f563746ba1a4f2bfe7f94d0c430c6ff461d126a08a4b09

Observation b897fa9b-2afd-4675-9485-e29701257712 · outbound

This paper cites prototype.

ProgRM: Build Better GUI Agents with Progress Rewards prototype

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.229956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:37:49.267563Z digest=sha256:c1a3cb9dd60a727ae36487bcf7df5b9b1bc6d2398a2ca17a3242c65f1f2bdc2f

Observation 853b27f3-2c39-475c-a86a-ae09ebfd1b8f · outbound

This paper cites URL https://doi.org/10.18653/v1/D19-1410.

ProgRM: Build Better GUI Agents with Progress Rewards URL https://doi.org/10.18653/v1/D19-1410

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.769331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.769331Z digest=sha256:bac60ae40aabd29a6f158ffb7f204d302ca8f797a22954bc94d63b6fd7069a8a

Observation 9d04d547-2c34-45ac-a980-11f7e7716d76 · outbound

This paper cites Free Process Rewards without Process Labels.

ProgRM: Build Better GUI Agents with Progress Rewards Free Process Rewards without Process Labels

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.423228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.423228Z digest=sha256:ed70f52ca6db1ae6a3b6fe575c7675d29a3a72124f33dfcf47b4107361e64fa7

Pith citing papers

Observation 32164952-9169-4ee7-af71-03c15a4fad7f · inbound

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics cites this paper.

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics ProgRM: Build Better GUI Agents with Progress Rewards

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:00:31.666131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T02:59:27.807789Z digest=sha256:ce7ead9ddb00a77302d87d72dd9c6b6588d5ecbcc1ce50ee27273a2cb2ffb542

Observation 81223177-eb8d-47d6-92ed-1976515667bf · inbound

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants cites this paper.

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants ProgRM: Build Better GUI Agents with Progress Rewards

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:29.615649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T05:48:00.486572Z digest=sha256:df528f9361a9bc92e1938ecfee3d6a5bdbabb92589d994950e19d5657d6ea1c3

Observation 331eab4f-af11-4384-8355-7588e38095c9 · inbound

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability cites this paper.

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability ProgRM: Build Better GUI Agents with Progress Rewards

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:56.866360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:15:19.239355Z digest=sha256:df450bc33bc8fa2c9404ff17ef0580b18f1bc24cd15b4d0941a661fb9952162e

Observation 106cf729-2ff8-4143-ace8-1a0164978cba · inbound

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding cites this paper.

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding ProgRM: Build Better GUI Agents with Progress Rewards

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:32:35.010485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T19:26:37.281563Z digest=sha256:1273cc5fe9e3275797acf3072e9859353e5e8f5edd497c1060641565089b76a9

Observation 6c8feefd-d674-4c20-a654-fa417d286531 · inbound

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents cites this paper.

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents ProgRM: Build Better GUI Agents with Progress Rewards

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:47:09.689614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T22:24:41.353303Z digest=sha256:dd852dfb990cfd33593c55a169acf702c8e701f05deac8be51cee8cde5ccff0f