Pith. sign in

Paper Citation Record · LEDGER

ProgRM: Build Better GUI Agents with Progress Rewards

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 5 inbound Pith citation observations for arXiv:2505.18121.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18121 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:49.267563Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T19:26:37.281563Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T16:47:09.687885Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aec7cd57-76d8-4c11-a7ff-73119782bc71 · outbound

This paper cites Digi-q: Learning vlm q-value functions for training device-control agents.

ProgRM: Build Better GUI Agents with Progress Rewards Digi-q: Learning vlm q-value functions for training device-control agents

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.949072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:45.129305Z digest=sha256:74b0c7f046cf3efbca6cebd2b0ece18062cdc229e216936475fb6b9add94d7f6

Observation f7254977-e58d-4cd0-abab-2582825f887b · outbound

This paper cites Digirl: Training in-the-wild device-control agents with autonomous re- inforcement learning.

ProgRM: Build Better GUI Agents with Progress Rewards Digirl: Training in-the-wild device-control agents with autonomous re- inforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.762133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:45.198300Z digest=sha256:beaa9cac4e62f34bdbbad50537c67373898bc0f6bdeb1a48e2e40e7e44612347

Observation a5895103-8609-41bb-847b-5452b5004258 · outbound

This paper cites Learning about progress from experts.

ProgRM: Build Better GUI Agents with Progress Rewards Learning about progress from experts

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.581616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:45.251487Z digest=sha256:f4b1877f8defdacf0a05055876335c8aa027f6deb09d4ad923522872180d5c1a

Observation 879f22e9-00ab-40d8-9a23-badacf93344a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ProgRM: Build Better GUI Agents with Progress Rewards Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.337260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.337260Z digest=sha256:a92a5eaf11da15af46128e8d42146f8bbd97e8b584deac7bc0ad5284ed9d52a3

Observation 53605823-cbb5-4d82-80f2-6a76092eb65e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.419221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.419221Z digest=sha256:e95b57a654a9c5fdc4532a95ab8b3a7bdec94436681f4d80387e7229eec9339b

Observation 9bb8ec98-2eac-4dc7-8c12-73a5d9dd3a60 · outbound

This paper cites Dungeons and data: A large-scale nethack dataset.

ProgRM: Build Better GUI Agents with Progress Rewards Dungeons and data: A large-scale nethack dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.400251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:45.508048Z digest=sha256:31a5b07bccbbc018f13665546b8ad3fe515a31a49ec1ae9f1a44b7d39852a962

Observation e7dc26f3-39d4-42cb-91c6-261d8a91c241 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

ProgRM: Build Better GUI Agents with Progress Rewards REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.597283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.597283Z digest=sha256:465f778fdff0cc0bb46a352d735c7f7fb9daa15270b0a6c2a20a0f49d78523e7

Observation 58e481f7-a907-4cbb-9051-e260b61b71d4 · outbound

This paper cites Exploring Expert Failures Improves LLM Agent Tuning.

ProgRM: Build Better GUI Agents with Progress Rewards Exploring Expert Failures Improves LLM Agent Tuning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:37:50.068075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:45.644348Z digest=sha256:8eae692d7c6c9159edd3a7abae9f7a581e521e6c43850f592dbbdf382c738207

Observation a4ae53f1-e8c2-4bcc-ac12-def7e240e590 · outbound

This paper cites Making language models better reasoners with step-aware verifier.

ProgRM: Build Better GUI Agents with Progress Rewards Making language models better reasoners with step-aware verifier

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.740939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.740939Z digest=sha256:72a08b36fff6b90a58675f83e6546f458659699afd4c53801d18c3b98c0b9093

Observation 8b426dbb-a6d4-449f-b9d5-68468ac898a7 · outbound

This paper cites Let’s verify step by step.

ProgRM: Build Better GUI Agents with Progress Rewards Let’s verify step by step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.786546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.786546Z digest=sha256:fa133e1434173a3655a73a02c5de0406f1d04be0707d3b57f3335145150f53b2

Observation 4a6e6515-a8a4-4cac-a2c8-9cca2eecd7b2 · outbound

This paper cites UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.862230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.862230Z digest=sha256:d8beb554792c833d6a14532143c9a1a6dbd639bf50421cd1e77b714664493318

Observation 50f6f860-4632-4c42-8826-d133a6404fdb · outbound

This paper cites NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild.

ProgRM: Build Better GUI Agents with Progress Rewards NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.943894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.943894Z digest=sha256:e28862e1a4fcc67407aafbd1af82ade39d30f5ac6507e81078e035438f9b4da1

Observation 47146464-9f0a-4fd5-8d7f-bafd72db849c · outbound

This paper cites Xu, Aman Madaan, Jiarui Liu, Robert Lo, Abishek Sridhar, Sudipta Sengupta, Dan Roth, Graham Neubig, and Shuyan Zhou.

ProgRM: Build Better GUI Agents with Progress Rewards Xu, Aman Madaan, Jiarui Liu, Robert Lo, Abishek Sridhar, Sudipta Sengupta, Dan Roth, Graham Neubig, and Shuyan Zhou

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.202532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:46.013705Z digest=sha256:f8e0325dc9943e709f6cb6e87b5c6c4a6b6622ab09c3c19e0cddc8e9707a7cc5

Observation f3f932b9-fe28-43b6-9684-dfa389c944ac · outbound

This paper cites Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents.

ProgRM: Build Better GUI Agents with Progress Rewards Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.119179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.119179Z digest=sha256:3b72767a4af3d95a097147816787fef5dd54223b122f431aa6681402e2b60f69

Observation b5b3b549-039f-4284-99eb-97810c3fe91c · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

ProgRM: Build Better GUI Agents with Progress Rewards Autonomous Evaluation and Refinement of Digital Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.228152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.228152Z digest=sha256:8b0d65c4100328d324fba423abe9ac5967ade489f20cafb12efe48517edb376d

Observation 4efd62e6-4159-4d0d-b049-b694982913da · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.329762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.329762Z digest=sha256:c6ad2727c93ab7c3ba0a1184f91b994fccca1c28ca0d5079a6f74cb3874d2a6c

Observation 8e89bd61-c05a-4608-9851-7de29189cb15 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

ProgRM: Build Better GUI Agents with Progress Rewards UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.431629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.431629Z digest=sha256:3a3455dd8d87515151c33caed8300b07a7c0d0cdf7a1714371840508d04604cb

Observation 18d80b30-3761-4e5e-b55e-c5e36827d8ce · outbound

This paper cites Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning.

ProgRM: Build Better GUI Agents with Progress Rewards Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.502834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.502834Z digest=sha256:6dce20724086139ad3a70eb24a00724bd31f17c37e6a7b191756752ad1dcf92d

Observation d4bdfdd7-bfd0-451a-a2a5-ffbd21fb45d1 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert- networks.

ProgRM: Build Better GUI Agents with Progress Rewards Sentence-bert: Sentence embeddings using siamese bert- networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.001358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:46.642165Z digest=sha256:0cad86d1b98c1d8b8901be15c4fa89d8bfa0011edb24e3375baf8ce4551fe645

Observation c464d765-6783-4c58-ba4f-a10440cdcbe1 · outbound

This paper cites Step: Stacked llm policies for web actions.

ProgRM: Build Better GUI Agents with Progress Rewards Step: Stacked llm policies for web actions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.788218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:46.921764Z digest=sha256:477e7ade135e8c2c3a1888dda8a5cc100bf9d506951d83bb369383c0d1a6abec

Observation c9cf3a05-12da-493e-98f2-dce10c1ab90d · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.033881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.033881Z digest=sha256:969d6329daec3d1ca5fd8530a857516aa8ba92f6a7dd888d7e6b93277ac1e6be

Observation 3eb32102-56d0-490f-a9e5-92e377d9949e · outbound

This paper cites Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments.

ProgRM: Build Better GUI Agents with Progress Rewards Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.108403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.108403Z digest=sha256:ec1604356cc427cc96c98f6dfd9ce1734e4e13c84c75518a77331f3dce93fe6c

Observation 486f8fcf-9f3c-4375-be1f-f7656a1a18bf · outbound

This paper cites OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis.

ProgRM: Build Better GUI Agents with Progress Rewards OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.185783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.185783Z digest=sha256:d7fb4d2ced708c8fe52ab87e2dd2c5cde2cf06f46b2df474c0afd62cf6770dc6

Observation c6421f46-7dfa-4f7a-8464-a8c8d1c8b002 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

ProgRM: Build Better GUI Agents with Progress Rewards Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.238146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.238146Z digest=sha256:8e5515a5bb18d2af700daadeed5f8aead4fa9e8e18cc55e40360ec0018813931

Observation 19a62709-4c66-4c0c-b607-d22e1aa6c55a · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

ProgRM: Build Better GUI Agents with Progress Rewards Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.380315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.380315Z digest=sha256:05e82c0e04eba70ed1d5ed43c021af8dccf33cb9e819e0a924b55f45709cce14

Observation 053bc850-e181-465d-8960-9813e0b5f08f · outbound

This paper cites DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents.

ProgRM: Build Better GUI Agents with Progress Rewards DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.485240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.485240Z digest=sha256:f9587a4f9751e91ca8407677f114a063f7ea0df7e0d196b42778f1f8d55c11d0

Observation 8029fb17-4eaa-44da-a0b9-37a8b23ba8c9 · outbound

This paper cites Reinforcing Language Agents via Policy Optimization with Action Decomposition.

ProgRM: Build Better GUI Agents with Progress Rewards Reinforcing Language Agents via Policy Optimization with Action Decomposition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.590688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.590688Z digest=sha256:38ccb217bb49fcc5e420f850e7a6a534d0dbc55ba32e388b6373114438569bee

Observation a3e49663-2092-4537-8873-2260e3312495 · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

ProgRM: Build Better GUI Agents with Progress Rewards GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.700415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.700415Z digest=sha256:8db1869118563e1f0466e00418c6e5991db0823e8b4b0a1e45e3e6c1eed4b3ba

Observation 19c05f48-6249-4353-acce-6092cd98e096 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.804784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.804784Z digest=sha256:a865425b72eb34b1bab58fff5ab1eeb286fdc792d1419ab2babc54f48b3352f4

Observation 5a6a36ef-b5cf-4007-84c6-d943abba95e4 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

ProgRM: Build Better GUI Agents with Progress Rewards Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.937341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.937341Z digest=sha256:d9d560333861598804366d4f3890339afb67df67703a99699e0afd63da0a67b4

Observation 835c7d55-3d95-4219-9486-da9488b80267 · outbound

This paper cites Agenttrek: Agent trajectory synthesis via guiding replay with web tutorials.

ProgRM: Build Better GUI Agents with Progress Rewards Agenttrek: Agent trajectory synthesis via guiding replay with web tutorials

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.672358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:48.039904Z digest=sha256:907d331cf6779901af59a7100e808d09c117943c2c8a452d38329ead1c69e324

Observation 3450faf6-7e13-4dc7-8f4e-c5518ff85d63 · outbound

This paper cites Qwen2.5 Technical Report.

ProgRM: Build Better GUI Agents with Progress Rewards Qwen2.5 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.119240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.119240Z digest=sha256:3fb2c08e3da51bdb2ef08ea2b0f91d3caf405c67c3b9be28c02aae238c8a6534

Observation bd0674c3-a3c1-4f74-af92-f26678f2827c · outbound

This paper cites OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning.

ProgRM: Build Better GUI Agents with Progress Rewards OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.213972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.213972Z digest=sha256:efff264c4e1c4e387adce3be9e5a6336dc6a8ce8889d86f27a233cdd8e64f9d6

Observation 750b3d33-83c5-4606-8ed4-199bcbfa0d59 · outbound

This paper cites Free Process Rewards without Process Labels.

ProgRM: Build Better GUI Agents with Progress Rewards Free Process Rewards without Process Labels

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.320653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.320653Z digest=sha256:dd6cf1e0e8a427177688c4fee1cbd5aadb08e10310d29d3cd4d8df38b3439528

Observation 5dc867a9-8a39-4ef1-ae06-9424c5cdf5d8 · outbound

This paper cites UFO: A UI-Focused Agent for Windows OS Interaction.

ProgRM: Build Better GUI Agents with Progress Rewards UFO: A UI-Focused Agent for Windows OS Interaction

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.529789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.529789Z digest=sha256:fc1969bc61dd64951137dba551f352ad4dbd925d1f82aee4e111abc8d8a54a94

Observation 8db64b3d-63b0-47ef-b6ef-d92909160a9a · outbound

This paper cites Appagent: Multimodal agents as smartphone users.

ProgRM: Build Better GUI Agents with Progress Rewards Appagent: Multimodal agents as smartphone users

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.664282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.664282Z digest=sha256:2e38f6e7328f242df38c6ed940e570339163feaf1bafd1ded5db4ac7ef953f50

Observation d601e41c-ad3e-4be6-92fe-9e5ad8d03326 · outbound

This paper cites Large language models are semi-parametric reinforcement learning agents.

ProgRM: Build Better GUI Agents with Progress Rewards Large language models are semi-parametric reinforcement learning agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.549721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:48.771877Z digest=sha256:48e00ef548258240a8dda72556df36268dcfd176d63eb20c6e3b8ebaf1cf9eed

Observation 32ad9c9e-7ec1-40c1-99e7-bdf899b9763d · outbound

This paper cites Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction.

ProgRM: Build Better GUI Agents with Progress Rewards Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.872226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.872226Z digest=sha256:3700de73bd00b58b4fa44b94e5669cbf83b4573caf69d0d3b4daaffde6743b12

Observation 129c17e0-bb24-4cf3-ad7e-bbdcd6ef5525 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

ProgRM: Build Better GUI Agents with Progress Rewards The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.959454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.959454Z digest=sha256:33babfbefade73851782dfeeca3b66d3cb063fdee5b76cf0824ad911ee211760

Observation a0dee5fd-9a2b-44d8-a18e-b0cb28957dc1 · outbound

This paper cites Gpt-4v(ision) is a generalist web agent, if grounded.

ProgRM: Build Better GUI Agents with Progress Rewards Gpt-4v(ision) is a generalist web agent, if grounded

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.386446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:49.041108Z digest=sha256:9d32e02ae5a36390ee999f4f6c4783f0c9c3871e1bc3f966a535cf3cd3dc3c6c

Observation 54802348-9e11-49bd-a93e-166b08e21376 · outbound

This paper cites VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model.

ProgRM: Build Better GUI Agents with Progress Rewards VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.160588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.160588Z digest=sha256:3ab78a3ad61f625bdb39489d5c15d6b52c3013586320e11c5d3a7e69aa36e90d

Observation 901c9048-5ede-4c28-ae38-b51dac1b22fd · outbound

This paper cites Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents.

ProgRM: Build Better GUI Agents with Progress Rewards Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.207033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.207033Z digest=sha256:f252d51d78e611dbc0be35f12a82bff59cf7e0289df42f24500ff1e8089f24df

Observation b897fa9b-2afd-4675-9485-e29701257712 · outbound

This paper cites prototype.

ProgRM: Build Better GUI Agents with Progress Rewards prototype

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.229956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:37:49.267563Z digest=sha256:c32d63a7024898a27e7c67785bc13e6afb830c4397e65c962bf2ade60146da1c

Observation 853b27f3-2c39-475c-a86a-ae09ebfd1b8f · outbound

This paper cites URL https://doi.org/10.18653/v1/D19-1410.

ProgRM: Build Better GUI Agents with Progress Rewards URL https://doi.org/10.18653/v1/D19-1410

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.769331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.769331Z digest=sha256:9a70339cc1bd485fa33c36e63363f2f9a5eb51231c21fc5f887921316b8fa0c5

Observation 9d04d547-2c34-45ac-a980-11f7e7716d76 · outbound

This paper cites Free Process Rewards without Process Labels.

ProgRM: Build Better GUI Agents with Progress Rewards Free Process Rewards without Process Labels

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.423228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.423228Z digest=sha256:6bb57da3b24164c948d3804632de74a7f6b69a76fb4131ffb0ff57cfa2b7f917

Pith citing papers

Observation 32164952-9169-4ee7-af71-03c15a4fad7f · inbound

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics cites this paper.

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics ProgRM: Build Better GUI Agents with Progress Rewards

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:00:31.666131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T02:59:27.807789Z digest=sha256:28c6025a23e279d12cd030033de0cf2f60abbde7b7d2b72b353e69d303f0afd8

Observation 81223177-eb8d-47d6-92ed-1976515667bf · inbound

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants cites this paper.

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants ProgRM: Build Better GUI Agents with Progress Rewards

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:29.615649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T05:48:00.486572Z digest=sha256:aa338528ad06669979739ac44ccbcea67be75a829e82db827c732ed6269202de

Observation 331eab4f-af11-4384-8355-7588e38095c9 · inbound

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability cites this paper.

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability ProgRM: Build Better GUI Agents with Progress Rewards

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:56.866360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:15:19.239355Z digest=sha256:4f1fd995243e84fc4135a324685fc4534d8523dbd9d1a5e338a3285a8b9abc74

Observation 106cf729-2ff8-4143-ace8-1a0164978cba · inbound

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding cites this paper.

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding ProgRM: Build Better GUI Agents with Progress Rewards

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:32:35.010485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T19:26:37.281563Z digest=sha256:2f38b55a2d861007f994fe660e2c22c6918775bcfbce18b1be8eaea9d7b32bb0

Observation 6c8feefd-d674-4c20-a654-fa417d286531 · inbound

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents cites this paper.

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents ProgRM: Build Better GUI Agents with Progress Rewards

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:47:09.689614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:24:41.353303Z digest=sha256:6afd04b984f8600fd99e43db1060de6781fce756047335f21f8471def2631a98