Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T13:28:05.021411Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2606.11042.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T13:28:05.021411Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T08:15:49.428124Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T10:59:46.438679Z
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 739c0fd9-0b1d-48c4-9a17-d6a92b49e40d · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Scuba: Salesforce computer use benchmark.arXiv preprint arXiv:2509.26506
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1dea3782-7b77-4dd0-b23f-f44813dccae6 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Mobile-bench: An evaluation benchmark for llm-based mobile agents
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d584761e-f27d-4849-b843-09e1ad3769ea · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Nl2repo-bench: Towards long-horizon repository generation evaluation of coding agents.CoRR, abs/2512.12730
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0a99caf9-2e69-451f-a7fd-a6cbb348ad72 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Gemini 3.1 Pro Model Card, 4 2026
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d21d009-ed9c-48a7-a507-f62fada26e51 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Gemini 3 Flash Model Card, 4 2026
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d241ce8f-b2d6-4327-a520-42ac33d38aee · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields PC-Agent: A Hierarchical Multi-Agent Collaboration Framework for Complex Task Automation on PC
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b667ee27-817f-4460-9754-f6cdcc575b4e · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8875a492-36e8-47e7-8540-6f62fa8b3cbe · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Kimi k2.6: Advancing open-source coding.https://www.kimi.com/blog/kimi-k2-6, 2026
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcc1f9c1-1d84-4617-a312-be721c649a08 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Gui-360: A comprehensive dataset and benchmark for computer-using agents
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75598cf3-99c0-46cb-adfd-fea6bd5dea9a · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Gpt-5.4 thinking system card
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e9fa475-88ba-4233-8b68-c303dd04aadc · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 17706fb0-47d6-490a-b0fd-599b1311df14 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6c1ddbea-2477-4852-9c3b-567f4b9b491c · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Seed1.8 Model Card: Towards Generalized Real-World Agency
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f38fb24-7d95-40ea-ab79-7ef0765cada2 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c52a60-b540-43ea-9d65-26348d409fb4 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Gui knowledge bench: Revealing the knowledge gap behind vlm failures in gui tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e83153bb-2699-4bd1-93d9-dde3ae3fa131 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Ambibench: Benchmarking mobile gui agents beyond one-shot instructions in the wild
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cd5a18d6-3e34-4632-9bdc-dca0ecd1c6e2 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific Workflows
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4b4a3584-ce2e-4e57-af1c-679cc85e6fd8 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1a31f9f4-15b1-4723-88a6-67440f298c3f · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields OpenHands: An Open Platform for AI Software Developers as Generalist Agents
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 526cc51e-cb90-40dd-bb92-f1c803c8ec40 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.Advancesin Neural Information Processing Systems, 37:52040–52094, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 517d22f5-d31f-4a81-a0ab-5a6bcad9c604 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Step-gui technical report
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d49dfc8a-2439-4240-93be-1f2fca21552c · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields $OneMillion-Bench: How far are language agents from human experts?arXiv preprint arXiv:2603.07980
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 97b7b588-17be-4d09-a2d8-1e819bc140a1 · outbound
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields React: Synergizing reasoning and acting in language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e706fb87-e870-4490-9019-da08b19cd982 · inbound
PhoneBuddy: Training Open Models for Agentic Phone Use Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.