Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:43:27.037355Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 3 inbound Pith citation observations for arXiv:2604.10866.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:43:27.037355Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:13:59.086704Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T17:50:00.070007Z
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5c445c85-7ae2-432b-b043-7b93b33415ec · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 802a4a1f-9dd6-4778-8cf7-4247fae0ace8 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Internalizing world models via self-play finetuning for agentic RL.arXiv preprint arXiv:2510.15047
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c44ffc03-2ed3-4aea-bab0-c0f43cefca63 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 94a5adaf-1242-4b0c-af22-d1c020b95e45 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f90b6c4-3ccf-42b9-87d7-f15b37d8cc2b · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Cl-bench: A benchmark for context learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af91293d-5720-4698-8270-78209db53825 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation GLM-5: from Vibe Coding to Agentic Engineering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc736636-c5c8-4df0-9680-a9466798b0bc · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c9ef5b4-240d-48e6-b590-7ff9a4f97561 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Kimi K2.5: Visual Agentic Intelligence
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 070a4869-ea72-4541-a805-2f59d9baadab · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Simulating environments with reasoning models for agent training
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ef63fb1-351e-476c-b9e9-1c0ad813fb96 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation ViMo: A Generative Visual GUI World Model for App Agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1cb947c-175e-4db4-9988-9380e13e9586 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1d17204-c528-4417-86f1-e567da618b05 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation GAIA: a benchmark for General AI Assistants
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6055ae1-9489-40bc-b083-a8db4e8816d9 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2be0ed0e-7423-401e-9fc9-da6612ca8508 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation OpenAI GPT-5 System Card
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 737c3d97-8bf7-4b8f-acd5-3aebad835fb5 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Generative Agents: Interactive Simulacra of Human Behavior
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c34966c-72ce-4aeb-b1fa-b9388f0d54d1 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfdf2bd0-6e26-4321-9851-8e8e7855c463 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 813d494a-e6e8-4ce6-bc8f-18e8949d23f8 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 742b2b78-558f-453f-98c2-83c6bc9d7a6d · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Mcpmark: A benchmark for stress-testing realistic and comprehensive mcp use
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a4bcc152-078d-48fe-9d69-598b237cfe14 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Webworld: A large-scale world model for web agent training
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83543edd-fa4c-4576-998b-f8fcd193ece6 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation $OneMillion-Bench: How far are language agents from human experts?arXiv preprint arXiv:2603.07980
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5aab5c21-e55e-4c8f-a900-81e7e61dd8ee · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8bd3fc51-c1ea-4313-8db6-6a3b6e6948e7 · outbound
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ef7c2d8-170f-4967-9f36-10c905273466 · inbound
LaGO: Latent Action Guidance for Online Reinforcement Learning OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93ed0ce3-6533-48a8-851a-9aa3c8f22b85 · inbound
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad3f85ea-6b51-4094-a9a1-7b753821c9f5 · inbound
Quo Vadis, World Modeling? OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.