Pith. sign in

Paper Citation Record · LEDGER

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

As of 4 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2606.11042.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.11042 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T13:28:05.021411Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T08:15:49.428124Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T10:59:46.438679Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact14
  • verified fuzzy0
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 739c0fd9-0b1d-48c4-9a17-d6a92b49e40d · outbound

This paper cites Scuba: Salesforce computer use benchmark.arXiv preprint arXiv:2509.26506.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Scuba: Salesforce computer use benchmark.arXiv preprint arXiv:2509.26506

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.678521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:9b276065760248ebb1d042c20cbd6c3f1963e950e64d3cf4dc2e1761e034cede

Observation 1dea3782-7b77-4dd0-b23f-f44813dccae6 · outbound

This paper cites Mobile-bench: An evaluation benchmark for llm-based mobile agents.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Mobile-bench: An evaluation benchmark for llm-based mobile agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T13:28:05.021411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:2b46166db57380a789382f60be394cf80fcd25ecc9ae13348e7231455e68c9a3

Observation d584761e-f27d-4849-b843-09e1ad3769ea · outbound

This paper cites Nl2repo-bench: Towards long-horizon repository generation evaluation of coding agents.CoRR, abs/2512.12730.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Nl2repo-bench: Towards long-horizon repository generation evaluation of coding agents.CoRR, abs/2512.12730

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.665109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:fa49d267ab0bc26443ba78681fb961b3733a105a98252a6e77dd6aff79625183

Observation 0a99caf9-2e69-451f-a7fd-a6cbb348ad72 · outbound

This paper cites Gemini 3.1 Pro Model Card, 4 2026.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Gemini 3.1 Pro Model Card, 4 2026

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T13:28:05.021411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:c87e061f53b95531ee8b6a960e2f9628ad24c8eaf796da63557ddaadda3e3c3d

Observation 5d21d009-ed9c-48a7-a507-f62fada26e51 · outbound

This paper cites Gemini 3 Flash Model Card, 4 2026.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Gemini 3 Flash Model Card, 4 2026

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T13:28:05.021411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:aba112fd9793c5fd8ccd077cd606c59bcc096fa6310fd84f1fc52395d434f0dd

Observation d241ce8f-b2d6-4327-a520-42ac33d38aee · outbound

This paper cites PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.659003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:f6ea2ffa8ea32df26b8d496e5c6e6b7ad153199c9befd8243f573e8ad9048f8f

Observation b667ee27-817f-4460-9754-f6cdcc575b4e · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:07:38.670710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:6e2e22e0ac90cb6193269236406ead0f69882dcdacaa92bb6b5ddf15a7d55cad

Observation 8875a492-36e8-47e7-8540-6f62fa8b3cbe · outbound

This paper cites Kimi k2.6: Advancing open-source coding.https://www.kimi.com/blog/kimi-k2-6, 2026.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Kimi k2.6: Advancing open-source coding.https://www.kimi.com/blog/kimi-k2-6, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T13:28:05.021411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:98dde1b9ae8a74210b8fb2dc402f5bc115c79d06549d6bb590eaceaa6b3c23a7

Observation bcc1f9c1-1d84-4617-a312-be721c649a08 · outbound

This paper cites Gui-360: A comprehensive dataset and benchmark for computer-using agents.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Gui-360: A comprehensive dataset and benchmark for computer-using agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T13:28:05.021411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:d431cb04c7d89551eac5d686ab62675984f4f01d543895bef565b3d48d13acea

Observation 75598cf3-99c0-46cb-adfd-fea6bd5dea9a · outbound

This paper cites Gpt-5.4 thinking system card.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Gpt-5.4 thinking system card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T13:28:05.021411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:d35a28c2a983300e38a6a78f9c7005a6c736535af718982b908d4b6017b1df4b

Observation 6e9fa475-88ba-4233-8b68-c303dd04aadc · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:07:38.674949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:bbcf4df2d62eb8b8a6197412779b67e42191cb5b49351e66ca3a79537541433d

Observation 17706fb0-47d6-490a-b0fd-599b1311df14 · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:07:38.667134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:bd24b163d888fa9e3e73a19ebb95e5c1484fc6b2d99c4e37911b314e6febf183

Observation 6c1ddbea-2477-4852-9c3b-567f4b9b491c · outbound

This paper cites Seed1.8 Model Card: Towards Generalized Real-World Agency.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Seed1.8 Model Card: Towards Generalized Real-World Agency

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:07:38.675829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:0bfa316a7a0ac74b4b1588f24a92e9b83997b8af0821b5dccf09eca28c6f52e0

Observation 3f38fb24-7d95-40ea-ab79-7ef0765cada2 · outbound

This paper cites an unresolved cited work.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T13:28:05.021411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:4b09e401fc54368bcf9cc26f7086e874318d6fd376e04fb3b5ec497945c44871

Observation a6c52a60-b540-43ea-9d65-26348d409fb4 · outbound

This paper cites Gui knowledge bench: Revealing the knowledge gap behind vlm failures in gui tasks.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Gui knowledge bench: Revealing the knowledge gap behind vlm failures in gui tasks

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.667708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:8f3601901267c1a57ee48086e4cd9464ef5f37c83dc6454e39148396675732c5

Observation e83153bb-2699-4bd1-93d9-dde3ae3fa131 · outbound

This paper cites Ambibench: Benchmarking mobile gui agents beyond one-shot instructions in the wild.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Ambibench: Benchmarking mobile gui agents beyond one-shot instructions in the wild

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.640426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:9faa7f8956f097522466ee62430ff29ea9b3871ed918ad112a9a43e4d5321e0e

Observation cd5a18d6-3e34-4632-9bdc-dca0ecd1c6e2 · outbound

This paper cites ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:07:38.650508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:24dba371bc1802eff3d1d069ef6e3c42d4e7073a069fcb9e25a41d03a96fce48

Observation 4b4a3584-ce2e-4e57-af1c-679cc85e6fd8 · outbound

This paper cites UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:07:38.654126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:5a90dd68ca9aabc7746e5adce0b89c08442b43ad4d1f6143c7efa00975fd03a1

Observation 1a31f9f4-15b1-4723-88a6-67440f298c3f · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:07:38.673400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:af518aeca7bebab06ddd0a7180a62ad8eee5eb05d9f52ff03cbc98670b51b6d0

Observation 526cc51e-cb90-40dd-bb92-f1c803c8ec40 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.Advancesin Neural Information Processing Systems, 37:52040–52094, 2024.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.Advancesin Neural Information Processing Systems, 37:52040–52094, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T13:28:05.021411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:35fd9dec5fef7b7c90e83b9c739d09346d767bfd09f3e7c52354e8032ea85829

Observation 517d22f5-d31f-4a81-a0ab-5a6bcad9c604 · outbound

This paper cites Step-gui technical report.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Step-gui technical report

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.661888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:3e198ef0675d120df3f9d828ceb083d6fadbf71fe9c724249bd1d36ec6e54635

Observation d49dfc8a-2439-4240-93be-1f2fca21552c · outbound

This paper cites $OneMillion-Bench: How far are language agents from human experts?arXiv preprint arXiv:2603.07980.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields $OneMillion-Bench: How far are language agents from human experts?arXiv preprint arXiv:2603.07980

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.656946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:e52b077d2bf3b1efd7b4cb3e698aa8d3a5ca7f8026a74a78fda89b01f37a8b62

Observation 97b7b588-17be-4d09-a2d8-1e819bc140a1 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields React: Synergizing reasoning and acting in language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-27T13:28:05.021411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T13:28:05.021411Z digest=sha256:0632c84d94661015b0ffc86e178c6ddb0c117b2814a65133ad0493328bb1e8e8

Pith citing papers

Observation e706fb87-e870-4490-9019-da08b19cd982 · inbound

PhoneBuddy: Training Open Models for Agentic Phone Use cites this paper.

PhoneBuddy: Training Open Models for Agentic Phone Use Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T10:59:46.439972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-26T08:15:49.428124Z digest=sha256:6064dfb618baaca33214dec0eb36989dcaedefeafafc85fa8802ce77905191d2