Pith. sign in

Paper Citation Record · LEDGER

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation

As of 7 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 3 inbound Pith citation observations for arXiv:2604.10866.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.10866 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:43:27.037355Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:13:59.086704Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T17:50:00.070007Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact22
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c445c85-7ae2-432b-b043-7b93b33415ec · outbound

This paper cites MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:57.821340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:9f1da43aa308e17e0ca811db64dd6b3fdcbfaecf3c6b45d25f20703866acb5c9

Observation 802a4a1f-9dd6-4778-8cf7-4247fae0ace8 · outbound

This paper cites Internalizing world models via self-play finetuning for agentic RL.arXiv preprint arXiv:2510.15047.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Internalizing world models via self-play finetuning for agentic RL.arXiv preprint arXiv:2510.15047

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:57.915436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:dd89b67ad346666d5bf20a8dff1a698e0b73342ca4043c3940b9e6f1076e7c95

Observation c44ffc03-2ed3-4aea-bab0-c0f43cefca63 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:57.872101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:c2f38c140aa508c805db288c0a749cf416ea1124aef18b15f29a949fd8312a5b

Observation 94a5adaf-1242-4b0c-af22-d1c020b95e45 · outbound

This paper cites Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:57.851105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:43f3fffefb7bbc5ed029461c67d779f23e00d69b68cb3837f3530d5f479dcd44

Observation 2f90b6c4-3ccf-42b9-87d7-f15b37d8cc2b · outbound

This paper cites Cl-bench: A benchmark for context learning.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Cl-bench: A benchmark for context learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:57.920830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:77eac9782a7f3e290dcc961cc069217f8eadc33990c3fbf4e1a42772bcd49f4c

Observation af91293d-5720-4698-8270-78209db53825 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation GLM-5: from Vibe Coding to Agentic Engineering

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:57.903585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:1a3bfc58d3eec40a35142d740a07e23000172f14a796b230223d15301ec29520

Observation fc736636-c5c8-4df0-9680-a9466798b0bc · outbound

This paper cites Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Is Your LLM Secretly a World Model of the Internet? Model-Based Planning for Web Agents

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:57.887889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:05cfb6654e34e0ca536e6fcaeee52a16132beb0d52b326badc4600247d2456d8

Observation 5c9ef5b4-240d-48e6-b590-7ff9a4f97561 · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Kimi K2.5: Visual Agentic Intelligence

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:57.834346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:1e23f7c63b9ed1c8b325d2a492d9de5185a3806183b1f32c9b3aff1959b73da8

Observation 070a4869-ea72-4541-a805-2f59d9baadab · outbound

This paper cites Simulating environments with reasoning models for agent training.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Simulating environments with reasoning models for agent training

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:57.860037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:f212ad0c1faaf7faeaa7df0428315f19513d1fd7a6db5b4b5a938c8ea4cf8fca

Observation 9ef63fb1-351e-476c-b9e9-1c0ad813fb96 · outbound

This paper cites ViMo: A Generative Visual GUI World Model for App Agents.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation ViMo: A Generative Visual GUI World Model for App Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:57.898393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:4b65c918e2cf26e989811376df8045c2047518cc8e7de8d418b2046246858408

Observation d1cb947c-175e-4db4-9988-9380e13e9586 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:57.808333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:e37693c15d5313ce8e5db92d74a2d6069313e9faffdddf58c49c0b52a0ed8309

Observation e1d17204-c528-4417-86f1-e567da618b05 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation GAIA: a benchmark for General AI Assistants

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:46:03.881709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:47f99c2f11f643ea17d483de74151f8b785e02680ada6f46f31c34f7416cd85a

Observation f6055ae1-9489-40bc-b083-a8db4e8816d9 · outbound

This paper cites SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.016292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:ed41d62467a5ebc0848debe5091c8ea6a9e7fc55f2a99586f051fe2128fd2c88

Observation 2be0ed0e-7423-401e-9fc9-da6612ca8508 · outbound

This paper cites OpenAI GPT-5 System Card.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation OpenAI GPT-5 System Card

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:57.984885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:fd93cacab9d5e37329a417e5ba2f8449ad77a19ea75c1dbeafd71620f3a8590e

Observation 737c3d97-8bf7-4b8f-acd5-3aebad835fb5 · outbound

This paper cites Generative Agents: Interactive Simulacra of Human Behavior.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Generative Agents: Interactive Simulacra of Human Behavior

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:05:14.047699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:63a7998ecf3b28328a92829216bec3e8bee45a05cca5330ac461dcb7de5e74f9

Observation 0c34966c-72ce-4aeb-b1fa-b9388f0d54d1 · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:11:04.750341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:4a9310f24f878fdf4eaea68431766f93cdc73645a3292397da734b4439dad03f

Observation cfdf2bd0-6e26-4321-9851-8e8e7855c463 · outbound

This paper cites MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:20:58.062227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:fad34a1a5f4fddf139e36f1d6b961eb84c5b6b11c9f88644b7c47646482e72dd

Observation 813d494a-e6e8-4ce6-bc8f-18e8949d23f8 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:44:32.193354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:573e4ead613a51f32455e00e61ac676fba8c464559f74cae0bfbfc6c700916e0

Observation 742b2b78-558f-453f-98c2-83c6bc9d7a6d · outbound

This paper cites Mcpmark: A benchmark for stress-testing realistic and comprehensive mcp use.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Mcpmark: A benchmark for stress-testing realistic and comprehensive mcp use

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.111675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:0c6f214f0730d0263a2af4b894009cce398e64a784bca803f8a9a078a3c84b40

Observation a4bcc152-078d-48fe-9d69-598b237cfe14 · outbound

This paper cites Webworld: A large-scale world model for web agent training.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Webworld: A large-scale world model for web agent training

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.008293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:8d50007168cdbe6eed74627ab1f14c6ca09427e46377a030293e26a0bee2ba92

Observation 83543edd-fa4c-4576-998b-f8fcd193ece6 · outbound

This paper cites $OneMillion-Bench: How far are language agents from human experts?arXiv preprint arXiv:2603.07980.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation $OneMillion-Bench: How far are language agents from human experts?arXiv preprint arXiv:2603.07980

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:20:58.126776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:d6f9cada37296b63b9b5244680e114db95ccbe4cb1bb4e89d141cf1828e91e5b

Observation 5aab5c21-e55e-4c8f-a900-81e7e61dd8ee · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:58.037323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:bd2f588cba1cbbbb3d4f0d02ee7b2ee315ec5e3f765eb476d73a0d5689c763b6

Observation 8bd3fc51-c1ea-4313-8db6-6a3b6e6948e7 · outbound

This paper cites Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents.

OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:57.970625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:43:27.037355Z digest=sha256:f459ce0d9d869ed30bd0fe843df6e5e278fbddd0f83fe7a031f55d17a7350382

Pith citing papers

Observation 1ef7c2d8-170f-4967-9f36-10c905273466 · inbound

LaGO: Latent Action Guidance for Online Reinforcement Learning cites this paper.

LaGO: Latent Action Guidance for Online Reinforcement Learning OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:50:00.071339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T23:26:55.307583Z digest=sha256:27440254eff5a950c64839df54f473d51df50ed4a08741afb5f26ba8b883d859

Observation 93ed0ce3-6533-48a8-851a-9aa3c8f22b85 · inbound

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL cites this paper.

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T14:50:11.936926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:50:11.936926Z digest=sha256:edc6120420696627b1db97f09e11208228bc7b80bf3eb32cb3ea9ea08766b734

Observation ad3f85ea-6b51-4094-a9a1-7b753821c9f5 · inbound

Quo Vadis, World Modeling? cites this paper.

Quo Vadis, World Modeling? OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language Environment Simulation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:59.086704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:59.086704Z digest=sha256:2ca60a9b6ad533263c4dfe96757aeb99b08ff4b66dacd306042c6a2dc09b7d1c