Pith. sign in

Paper Citation Record · LEDGER

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

As of 4 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2607.28609.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.28609 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T02:18:02.002869Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved81
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f0ece2c-cfe6-4b21-867c-df1a2639c695 · outbound

This paper cites Introducing Claude Haiku 4.5.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Introducing Claude Haiku 4.5

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.559074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.559074Z digest=sha256:25a89dffef53cc267f6876a134a5fd60e5d966f722bb614df9a24f34bda282a6

Observation 86a9152b-3e6b-4a1c-af7f-bea75ae96674 · outbound

This paper cites Introducing Claude Opus 4.6.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Introducing Claude Opus 4.6

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.565401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.565401Z digest=sha256:5365a772c445265766bb86c42de55f91e2516d8bbaa10db30d550254121c8e6d

Observation be37175b-0cae-4f68-bd30-485ee20ce7dd · outbound

This paper cites Introducing Claude Opus 4.8.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Introducing Claude Opus 4.8

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.570451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.570451Z digest=sha256:b1f119d3641375d10ae585050ca9d0304e032a0f3c4557eab214b2494c9ac58b

Observation 7a2945e3-38c9-49de-9c5e-9df9d711e777 · outbound

This paper cites Introducing Claude Sonnet 4.6.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Introducing Claude Sonnet 4.6

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.575664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.575664Z digest=sha256:e6b1d4d5bed946b5fd89ac40e172d846303e928a58d4a8cd8410cdea8b52de70

Observation becb788d-5be3-4eb4-b251-72025671a35b · outbound

This paper cites DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.580579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.580579Z digest=sha256:fcbde495fe03bc5a93e74922e9b488929bbd01152767fcf0ffbdfde9de22d2d1

Observation a39e32f3-cf99-4948-87b7-1cf2397d94fd · outbound

This paper cites Qwen3-VL Technical Report.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Qwen3-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.586349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.586349Z digest=sha256:8d79caaa30b8c3d3fe2c9855c9904d26c336f57e08ea880fdf5a82003e19c929

Observation a8e28b60-e354-4d99-84b9-588fabf998a6 · outbound

This paper cites Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.593694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.593694Z digest=sha256:a53b9eb4c495e4857cba5cd4f291b0a0536c7f3246ee7dae8a0b2a33a1ec4257

Observation 44906056-6099-4c43-9fec-8c20240e35b6 · outbound

This paper cites Seed2.0 model card: Towards intelligence frontier for real-world complexity.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Seed2.0 model card: Towards intelligence frontier for real-world complexity

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.599021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.599021Z digest=sha256:cda49b591806af975d2dc3ac1b7314d48b661c831d8ee4268e40d467348637a0

Observation 122a9ba3-117f-4d40-a272-4ddddee64a8b · outbound

This paper cites Web-shepherd: Advancing PRMs for reinforcing web agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Web-shepherd: Advancing PRMs for reinforcing web agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.604132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.604132Z digest=sha256:c2d559f0c4b32a19b1b4b2f109a9950b408df9830a5fce200e9a3a6c7051f3d9

Observation 2b1dc412-a8c4-4c38-81fd-de27687ab0d7 · outbound

This paper cites Gui-shepherd: Reliableprocessrewardand verification for long-sequence gui tasks, 2025a.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Gui-shepherd: Reliableprocessrewardand verification for long-sequence gui tasks, 2025a

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.609714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.609714Z digest=sha256:7dee1596a7158bf19b02d7540be4d6bee80838a72f91ba16a65f3be75a56f208

Observation 50983d72-a1f0-41fc-9126-68fad673c8f0 · outbound

This paper cites MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.617281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.617281Z digest=sha256:91362690fb47eb93c18221cd409418566c7d50683b4f5f9b93801c399733be57

Observation 1535705b-2d8b-45b2-9893-b62b1256c7c5 · outbound

This paper cites OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.622960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.622960Z digest=sha256:1f0c3ceff0b078b78a9c624318aff26fe54cc9b0ab240ed684ce3cfcf2a64130

Observation d6bb6f28-dd1c-4612-8a17-ab45032375c9 · outbound

This paper cites SeeClick: Harnessing GUI grounding for advanced visual GUI agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models SeeClick: Harnessing GUI grounding for advanced visual GUI agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.629392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.629392Z digest=sha256:ab0daf950c349a46fa36c9005179a7b8087b07a9445195382250454b4e800e19

Observation d84f5190-b757-43af-a4f7-56f8cd635c9c · outbound

This paper cites OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.634294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.634294Z digest=sha256:2a39c2cbcdfa6bddef552300aa4dd790f263f7005d3003e38f64dde6c57dda42

Observation dadbfdf6-b7d9-4902-81d8-6210156fcf77 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Training Verifiers to Solve Math Word Problems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.639292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.639292Z digest=sha256:39c1abf1f7660d6a6c04040ada5ee4b53f531dbebe620fe5d558a5e2e1930653

Observation a180f979-82fa-42f8-84fa-d9bdc7fbafaa · outbound

This paper cites A coefficient of agreement for nominal scales.Educational and Psychological Measure- ment, 20(1):37–46, 1960.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models A coefficient of agreement for nominal scales.Educational and Psychological Measure- ment, 20(1):37–46, 1960

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.643927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.643927Z digest=sha256:a8d643804a32ef0dcaeb39a74f4dfd7cc2cdadb6c1c5dbaaaf803935b56c3aff

Observation 39f6ae8c-87b7-4b0c-b087-26f7ba5d86ba · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.649199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.649199Z digest=sha256:32b802ea4077c47121948ce3c8a6c6aa80818bc3a09b441fb89ac62fdeb36e4d

Observation 5df6b5e5-ba55-4476-9855-200a6c82ce69 · outbound

This paper cites Gemini 3 pro model card, November 2025.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Gemini 3 pro model card, November 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.654732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.654732Z digest=sha256:44c51d6c20d6d6a507a33b9465c50bd5905b3c87895069976709307f914cbaab

Observation c38733fd-ba89-42bd-a48f-2022a4e2693a · outbound

This paper cites Navigating the digital world as humans do: Universal visual grounding for GUI agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Navigating the digital world as humans do: Universal visual grounding for GUI agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.659486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.659486Z digest=sha256:3e585cf5bac3b25d4c7a7a2e48435619583bf74fbc160297fe24600c14095b53

Observation 97edba44-382a-4689-83b4-f3ca99232e22 · outbound

This paper cites GPT-4o System Card.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models GPT-4o System Card

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.664032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.664032Z digest=sha256:edb6bda8466b0da77e332b61e96fad31850951661608a46e9b6c43687c0625bc

Observation 0def549d-f8df-41a6-a0fc-8b055e01c4cd · outbound

This paper cites AgentStore: Scalable integration of heterogeneous agents as specialized generalist computer assistant.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models AgentStore: Scalable integration of heterogeneous agents as specialized generalist computer assistant

Reference 21

Resolution
verified exact
doi, observed 2026-07-31T02:21:00.870350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-31T02:18:01.668582Z digest=sha256:ce9bdd2dee3e555d9f573f1380b5f4e9261e030827efd137794758c4088eeb90

Observation 0f075693-9c1c-44d6-ad9d-83017b7704c5 · outbound

This paper cites TreeCUA: Efficiently scaling GUI automation with tree-structured verifiable evolution.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models TreeCUA: Efficiently scaling GUI automation with tree-structured verifiable evolution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.675609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.675609Z digest=sha256:10baedff0e16184aae24b939e314e404bbb814a1ef7a1d20bf7849a7cbd171da

Observation b1c09d96-f1fa-49f5-b354-354ba75bc02d · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Kimi K2.5: Visual Agentic Intelligence

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.680543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.680543Z digest=sha256:8f183757d3086fbf57957d7b2db2df2b31900098159e17dee260173a75901bce

Observation 1e795199-04f4-425b-912b-7f58b460ebf8 · outbound

This paper cites Mobileworld: Benchmarking autonomous mobile agents in agent-user interactive and mcp-augmented environments.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Mobileworld: Benchmarking autonomous mobile agents in agent-user interactive and mcp-augmented environments

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.688571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.688571Z digest=sha256:6937d486a4f7bf368afb55d73d2add366c1b3fd6dfd75514b76dc3beae881cd9

Observation 6596f939-5842-46d7-92c7-a871d155d462 · outbound

This paper cites Computing Krippendorff’s alpha-reliability.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Computing Krippendorff’s alpha-reliability

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.696580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.696580Z digest=sha256:b49d39091bc688a134770c1ba3308f75ef7495c71edbe89dabde4f4e930d0710

Observation 43abe647-55ec-4006-8221-de383503935b · outbound

This paper cites Vl-rewardbench: a challenging benchmark for vision-language generative reward models.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Vl-rewardbench: a challenging benchmark for vision-language generative reward models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.704054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.704054Z digest=sha256:b234b0f854c92953c6e600d23a4a4706848a0a34204ff40b475ff1f448e46ad9

Observation 7ab18965-3338-461f-a5d7-d4fa8b553275 · outbound

This paper cites Os- themis: A scalable critic framework for generalist gui rewards, 2026.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Os- themis: A scalable critic framework for generalist gui rewards, 2026

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.709379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.709379Z digest=sha256:686863787cdcc90c24baa7623d1c4fa3ab306d6666aa5d8cb28f5ea1761690d9

Observation 53522066-f390-47fe-ae4e-86ed83598176 · outbound

This paper cites Cuarewardbench: A benchmark for evaluating reward models on computer-using agent, 2025.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Cuarewardbench: A benchmark for evaluating reward models on computer-using agent, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.714047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.714047Z digest=sha256:ed2afa79d4b5a9e41907a0e199914db50d89bf16f3adffa1c1700ef5012a28f7

Observation bd9d6844-5bc6-4bbc-b48c-a79602639f15 · outbound

This paper cites ScaleCUA: Scaling open- source computer use agents with cross-platform data.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models ScaleCUA: Scaling open- source computer use agents with cross-platform data

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.718753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.718753Z digest=sha256:4d86790a9c2bac49eff0a07811e9eb7d726752b3e5265145fb7120e44b692c33

Observation 9ef05bf4-ca94-4e20-a70c-5f39c166c972 · outbound

This paper cites Pal, and Siva Reddy.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Pal, and Siva Reddy

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.724504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.724504Z digest=sha256:3063c51261e5ddfb668cf4097c15c3689b35dbe82d3b02080f6541c5ae19e572

Observation a0eddbdf-16e5-4e1d-8023-0e22724882a0 · outbound

This paper cites Computer-using agent: Introducing a universal interface for ai to interact with the digital world, 2025a.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Computer-using agent: Introducing a universal interface for ai to interact with the digital world, 2025a

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.729104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.729104Z digest=sha256:76fa285da6463ad58fbf18e8d590b810c1540a7d48327c17bd12a58b44e234ba

Observation 7ca4923f-a0f6-4b92-9b04-969e31d5c3b3 · outbound

This paper cites GPT-5 System Card.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models GPT-5 System Card

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.735104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.735104Z digest=sha256:3ee23d1d9176009cdac52c58ac00079bc3744edff5c6b7d37a6887d942614df1

Observation 2268efc2-f51c-4442-a390-d4966aabe0ef · outbound

This paper cites GPT-5.4 Thinking System Card.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models GPT-5.4 Thinking System Card

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.739977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.739977Z digest=sha256:ea4ec7394e996b4d0978b33a21c633757f2f6d7545e38e050245c927b373301a

Observation 72f5b3c5-cd05-45d9-b64f-3750870cbec4 · outbound

This paper cites GPT-5.5 System Card.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models GPT-5.5 System Card

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.751636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.751636Z digest=sha256:2e90294e31d8955ecf6c6c2213f15e757734b943a3fc38a41a262db00a083962

Observation 54a84a0f-6312-4ffb-9af5-d3af6d168d01 · outbound

This paper cites Autonomous evaluation and refinement of digital agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Autonomous evaluation and refinement of digital agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.756864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.756864Z digest=sha256:598adf8989a979c179459daf2c9363bfe7a5f4adcdd732fda7ee1e91f753abbc

Observation 258eb65e-660f-489a-ac9d-68331475c65b · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.764481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.764481Z digest=sha256:069a57cc612da9eedf56b1e59ffaba174b9a484c75ee9d42575a3413bd390dde

Observation 4b9082c2-ae6b-46b9-ba30-b77e8c7cc08a · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.769432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.769432Z digest=sha256:8ca0dba3541e777b94bb93f7eedea236cac59485d2816df64e78f49f2671c82d

Observation 7c6c2432-e0fc-4ff6-a181-a8c2e7672d75 · outbound

This paper cites Qwen3.5: Towards native multimodal agents, February 2026.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Qwen3.5: Towards native multimodal agents, February 2026

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.775221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.775221Z digest=sha256:8433e0ee5530c64d19e5e5a9d47116dd92b76e6ad58e3c0041a8143ea4973ddb

Observation 06082ed4-d196-45e0-baa1-d74569d6ecf8 · outbound

This paper cites Androidworld: A dynamic benchmarking environment for autonomous agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Androidworld: A dynamic benchmarking environment for autonomous agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.782849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.782849Z digest=sha256:dd2e61e486475048d2b52695cf6803e0239509d0627c0398e61d3f908a16afa6

Observation 23e7573e-c538-4c57-bb3e-ec4926b5fa4c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.789891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.789891Z digest=sha256:89932946d75d93035635789ef3c645f64002e2ffb9ca39a424ccf8ffde6755e6

Observation 3bbdcff4-df2c-4b7b-ae41-c080f3c3758d · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models HybridFlow: A Flexible and Efficient RLHF Framework

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.794382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.794382Z digest=sha256:a4207d275de80b9f7d694a3b1223f414776dc1780d207a1995ed80b3f6ed38d5

Observation 5056fae7-ba0e-470f-af70-e3c0863311c1 · outbound

This paper cites A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models A Survey of Neural Code Intelligence: Paradigms, Advances and Beyond

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.800198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.800198Z digest=sha256:3518f19d6a33964184c8b3aa88c7e8d8062117473bdfa14508c4a95b5bb5b9b3

Observation 52a0cabb-29e5-48f3-a5db-b731a8d93ff4 · outbound

This paper cites Os-genesis: Automating gui agent trajectory construction via reverse task synthesis.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Os-genesis: Automating gui agent trajectory construction via reverse task synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.804558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.804558Z digest=sha256:64dfc8a23ed96e1e5b0f014fa2e0a7ba5f8b2e3f4e95801356aa88f36d3c9c06

Observation c256251f-d902-424e-9f2f-a7ba998c57f7 · outbound

This paper cites OS- sentinel: Towards safety-enhanced mobile GUI agents via hybrid validation in realistic workflows.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models OS- sentinel: Towards safety-enhanced mobile GUI agents via hybrid validation in realistic workflows

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.809067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.809067Z digest=sha256:94986f8ed7b4424935d0c227e729d0875ee96249f35a29049120e7f9e889b2b1

Observation 0cf55a1f-3454-4825-a8d3-cdb34e56c5a0 · outbound

This paper cites Scienceboard: Evaluating multimodal autonomous agents in realistic scientific workflows.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Scienceboard: Evaluating multimodal autonomous agents in realistic scientific workflows

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.813520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.813520Z digest=sha256:3d0cb0e7e8241918dac326dc8f7e0b425d43eed6c9491bc80099721eedc0e2b3

Observation 9d79eac9-d5b5-4c6e-bdd7-ca9363ed2204 · outbound

This paper cites CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.817666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.817666Z digest=sha256:9297742c50b9ccdb97161a92b782faf5151f41f97a87fb25da3e4bb8d117ba15

Observation ae9c1a53-c75c-4d84-82f0-0a7ebaf1b595 · outbound

This paper cites InSTA: Towards Internet-Scale Training For Agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models InSTA: Towards Internet-Scale Training For Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.822251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.822251Z digest=sha256:718d47ddf62a228588980f5e37292363bbd600acc2929217a5caa8c062cbbfd8

Observation a344d516-3ad7-478a-9ec2-ba53f7da9601 · outbound

This paper cites Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.826853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.826853Z digest=sha256:c264bff4bf0df6d95c1c9f1bddbb8f0640afde3fdc716c2f73c758b0e974fa27

Observation 382433c4-c952-4f06-8636-369b4aa22cf1 · outbound

This paper cites Charles, Zhilin Yang, and Tao Yu.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Charles, Zhilin Yang, and Tao Yu

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.831507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.831507Z digest=sha256:ffcd62e38e8961bff2cffb310c7481677b092171ab2e1f0cf206a0c54b80e085

Observation 8ba75244-b467-48c5-a019-eb17f4854ad6 · outbound

This paper cites SynthAgent: Adapting web agents with synthetic supervision.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models SynthAgent: Adapting web agents with synthetic supervision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.836647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.836647Z digest=sha256:938fdc4d154ba7f7bcb3281b2ef5fa749644bb99a39867589dc9a2ec5f38c250

Observation 2fb1c925-27ce-4cfa-9784-d2c3c57326e4 · outbound

This paper cites GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.841745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.841745Z digest=sha256:b1c1f675680af0b9844c453fffa48ae8a32625223683e90240c6e6a72267be4a

Observation 8de53c12-3cf2-4433-8b59-ddcb3b99b2f6 · outbound

This paper cites Os-oracle: A comprehensive framework for cross-platform gui critic models.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Os-oracle: A comprehensive framework for cross-platform gui critic models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.846403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.846403Z digest=sha256:049b6b041d57b455f39dabb008e5f20fd6082ccb6b7b6bb89607d3543d793618

Observation e325db33-46bf-447a-8482-e35b1fd0b48e · outbound

This paper cites Os-copilot: Towards generalist computer agents with self-improvement,.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Os-copilot: Towards generalist computer agents with self-improvement,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.851384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.851384Z digest=sha256:65f29fdfba7ee3faefbc9a3296cdadfa046ba65be23edba3c099ecaa81957339

Observation 53fbdd16-afea-40f8-9d1b-451ea7e1672b · outbound

This paper cites OS-ATLAS: Foundation action model for generalist GUI agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models OS-ATLAS: Foundation action model for generalist GUI agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.862476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.862476Z digest=sha256:b3a0ca7c26ee4b8550e8cc7c4c66535ca6e6b15abac7b6c2f1522bb28ad7c3dc

Observation 7c64305a-28dd-4fa5-a81f-0a21ec06e366 · outbound

This paper cites UI- genie: A self-improving approach for iteratively boosting MLLM-based mobile GUI agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models UI- genie: A self-improving approach for iteratively boosting MLLM-based mobile GUI agents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.868149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.868149Z digest=sha256:e547780687764dbe6851026f8a54fd82c96a6d5c7be987456d9f2033e3b4edd6

Observation f54037f1-612c-44b4-a6d2-6a4e82dda62e · outbound

This paper cites OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.874190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.874190Z digest=sha256:ba6edf62999a7bc4fe04eebe7f115c435a8f0ec0d2321d6d0b42a3849c56e832

Observation 5ef1ba9d-0674-405a-8456-d7756185cea0 · outbound

This paper cites Introducing osworld-verified.xlang.ai, July 2025.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Introducing osworld-verified.xlang.ai, July 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.880834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.880834Z digest=sha256:d9f4b16b6930e862315443717dffbfd20016082353cb72c7ba3b89ea25725c42

Observation e53c8bbd-d6b2-4c87-a4c7-57212f0f7113 · outbound

This paper cites OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.888233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.888233Z digest=sha256:7dc0f0ab6be964d70fd2c7e0a98ed813edce172195f44643ff460417d5cbb710

Observation ebdfee6d-fda7-4ea4-812a-5e48bfefec8e · outbound

This paper cites Agenttrek: Agenttrajectorysynthesisviaguidingreplaywithwebtutorials.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Agenttrek: Agenttrajectorysynthesisviaguidingreplaywithwebtutorials

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.893840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.893840Z digest=sha256:063e00b028e818e6232545ce23f96b2ea06ec7c450555b3b462624d14c7c6132

Observation 7b6bec37-da01-43a7-8c67-8644bae97af3 · outbound

This paper cites Aguvis: Unified pure vision agents for autonomous GUI interaction.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Aguvis: Unified pure vision agents for autonomous GUI interaction

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.899968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.899968Z digest=sha256:b85e808b763ebf38ada4903b033328992cd8a10b68275e33385010419bac1e5e

Observation dfbe6fd2-2ad7-4bc1-827e-b61fd9eaff40 · outbound

This paper cites EvoCUA: Evolving computer use agents via learning from scalable synthetic experience.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models EvoCUA: Evolving computer use agents via learning from scalable synthetic experience

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.905245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.905245Z digest=sha256:67640bfc2356e82d4d63eaf321ccf72f91ce8d03c47f72d59abe5994b2d4aadd

Observation 45fd7341-0114-40c8-9992-91d3a5d21d60 · outbound

This paper cites An illusion of progress? assessing the current state of web agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models An illusion of progress? assessing the current state of web agents

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.909883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.909883Z digest=sha256:6359adaf3c12b36ea9a4aa623011191a00e10330e311d47d6eac6eef2ae034c9

Observation 1c59e694-eefd-46d3-856c-637e8e3fa841 · outbound

This paper cites Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.915161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.915161Z digest=sha256:efe1e1cb662024f26a57c1cc013811de1ea197a9571cc65d22e1abda87ec75d5

Observation 515506ea-fced-47ac-8251-51fcb02a2fad · outbound

This paper cites OS-symphony: A holistic framework for robust and generalist computer-using agents.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models OS-symphony: A holistic framework for robust and generalist computer-using agents

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.920231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.920231Z digest=sha256:9fb3cc4c8b9c3f9cd10c08dec62776d921e3889c965ad04b1a1837bfc73ea676

Observation 59b4aa22-7ead-4508-a002-6c176b145308 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.925929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.925929Z digest=sha256:0ab59c4ae231747d583078495f2f54c25e22ee887389f863474e121ef3f7f435

Observation 60050239-3626-401e-8617-c643f1b35367 · outbound

This paper cites Breaking the data barrier – building GUI agents through task generalization.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Breaking the data barrier – building GUI agents through task generalization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.931168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.931168Z digest=sha256:0d8d1324832e8ac51a20c2dd8dc5e11d6d9ec90d2209826381dc978826d95caf

Observation 722de82b-810a-4df9-ae7b-196d301cc430 · outbound

This paper cites Agent learning via early experience.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Agent learning via early experience

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.935775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.935775Z digest=sha256:01e7d07e8b568483f626dde9e23bcdf5bc8775cefd7c0f3042c4c00ccbb6bf42

Observation 9a6c3e94-d649-4be3-9a8d-dfd27e49b1f1 · outbound

This paper cites Gpt-4v(ision) is a generalist web agent, if grounded.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Gpt-4v(ision) is a generalist web agent, if grounded

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.940319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.940319Z digest=sha256:99173adec93aadb9c0402bff31f68e895208011386d100e7602e299eeedaf6f9

Observation 6d9cf97e-1a11-4e3e-9660-34cabeac9e2b · outbound

This paper cites Xing, Hao Zhang, Joseph E.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Xing, Hao Zhang, Joseph E

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.944619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.944619Z digest=sha256:995c816af8a1281ec9ec83b3584203e2153ae246f023a6f9392646c9836c95e4

Observation 5d8ce686-afd0-4d44-91a6-837c3db0be87 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models SGLang: Efficient Execution of Structured Language Model Programs

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.949050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.949050Z digest=sha256:fbfdd64d335192644a95ece0ec56cafceb1a57c93f97087dddcaf760279a58f3

Observation 4334b6a6-f6ac-45bf-b3c8-71cd76f6f721 · outbound

This paper cites RMB: Comprehensively benchmarking reward models in LLM alignment.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models RMB: Comprehensively benchmarking reward models in LLM alignment

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.953939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.953939Z digest=sha256:5d586a5bcf34705478b4bc44fc9a41bfa9c91478e467cf2c1e17393a6ccff870

Observation 870583bd-5863-476d-b880-c2ca9d5b31e6 · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.958291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.958291Z digest=sha256:0eb9671a86e9b13776d424b4d925e54975d7e7f53a38082376fb8aea85e0a833

Observation 3b2cee03-64b1-4456-af73-6f29faa3c7a3 · outbound

This paper cites Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, and Jürgen Schmidhuber.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Ashley, Wenyi Wang, Dmitrii Khizbullin, Yunyang Xiong, Zechun Liu, Ernie Chang, Raghuraman Krishnamoorthi, Yuandong Tian, Yangyang Shi, Vikas Chandra, and Jürgen Schmidhuber

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.962720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.962720Z digest=sha256:d90dc1cc49ab5da673070fe797631ea6b49f39a1c4dc0e7a80f28dffeda84e88

Observation 76ee57d2-1d50-41cd-a2d2-cd65237e00a6 · outbound

This paper cites Intern-s1-pro: Scientific multimodal foundation model at trillion scale.arXiv preprint arXiv:2603.25040, 2026.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Intern-s1-pro: Scientific multimodal foundation model at trillion scale.arXiv preprint arXiv:2603.25040, 2026

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.967005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.967005Z digest=sha256:3b75ae312be527b53bccc0cf6ef35af35cbbaa9944e208cf34f65c9fd2ad627f

Observation ac6c3393-3e62-4f2a-bd6f-798aabcc1911 · outbound

This paper cites an unresolved cited work.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.973135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.973135Z digest=sha256:137a91a4c9d98f4d8f03a5dcfcff1284ea78b3ec2f7c8988cf3b969e7e7bb98a

Observation 7c080dcb-e7c2-427f-b0a8-fce226eade0a · outbound

This paper cites For click-like actions, the action point may be highlighted with a red circle on the screenshot.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models For click-like actions, the action point may be highlighted with a red circle on the screenshot

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.977691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.977691Z digest=sha256:b6781606c487ed44c35e463e59aa45f3cc7be71d2cd29ece911ceba0596e8f43

Observation 170dc844-e7e2-4626-b93b-1481eddf78b9 · outbound

This paper cites [EVALUATION GOAL] Your task is to synthesize all the evidence above and determine whether the agent reasonably completed the task according to the user’s instruction.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models [EVALUATION GOAL] Your task is to synthesize all the evidence above and determine whether the agent reasonably completed the task according to the user’s instruction

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.982279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.982279Z digest=sha256:3c22646ca048b77c1220daab1fd550bdaeecdf7b3459c76a3fb0ddce059072f6

Observation e5a70814-dea6-49c3-b94b-57bc23dc8484 · outbound

This paper cites an unresolved cited work.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.986907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.986907Z digest=sha256:b00b1c6e7ee018038aa0391df64f2ffe78eae7fb385f9274c52e3b94c3b8c324

Observation 587bb7b2-d8ea-4d5b-aa2e-109c97fea10b · outbound

This paper cites an unresolved cited work.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.992781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.992781Z digest=sha256:e55682244ae6b1e1f82fe33e4b276a6ba77d4e024d3e1db14e2df2b66bddf812

Observation fb4a70b5-e1b6-4b5a-a475-9e36db622242 · outbound

This paper cites - Examples include system or OS barriers, login walls, CAPTCHAs, paywalls, region restrictions, network failures, and unavailable pages or apps.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models - Examples include system or OS barriers, login walls, CAPTCHAs, paywalls, region restrictions, network failures, and unavailable pages or apps

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.997718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.997718Z digest=sha256:8d2819f0312f68d5842cb5f469603aa5f2085c73841194fb7627eb97e9d57ba8

Observation 10c3efbc-d33b-4617-ad09-84cd573180ff · outbound

This paper cites recently popular.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models recently popular

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:02.002869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:02.002869Z digest=sha256:7568a85b075d872b10ddb14cff0bc845cc430f1dcb974109f0b71a20feb8466d

Observation 6a7ab6ba-fdce-46df-ab5b-df00bf530e3a · outbound

This paper cites OS-Copilot: Towards Generalist Computer Agents with Self-Improvement.

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models OS-Copilot: Towards Generalist Computer Agents with Self-Improvement

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-07-31T02:18:01.857270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T02:18:01.857270Z digest=sha256:91a503829a47d629febd650a4c3f36ccf418168e7241ba3fd5b4ebda5fd4a216

Pith citing papers

No inbound Pith citation observations are available.