Pith. sign in

Paper Citation Record · LEDGER

OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2407.19056.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.19056 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:06:29.156277Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 328356e1-0374-400b-99f2-c3c82e635529 · inbound

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games cites this paper.

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T16:21:08.404854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:21:08.404854Z digest=sha256:146be32c68cade0a6105160e7c76f50bbb152a3a86fc902342d2c7c70f766e8f

Observation c58cb6eb-8d61-4432-b146-05acd5fc2ac8 · inbound

A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions cites this paper.

A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:02:35.183609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T04:59:36.994758Z digest=sha256:26f3dd174e2d4eb350495efa7e9f172c516aaa9220942feaa145fa7970b34059

Observation 7c9775cf-fff2-41b9-8e65-cec65a483072 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.324265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.324265Z digest=sha256:49a9528993b3a92b2a1b4e8f4f172b5dc62832ab533284e28621bb902150a4ef

Observation fa6b4586-a837-491a-b778-338618f93109 · inbound

ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution cites this paper.

ASPERA: A Simulated Environment to Evaluate Planning for Complex Action Execution OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:30.639518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:30.639518Z digest=sha256:3c29e7077925b71be50a80f3b4b89b71189cc8acac61d7e15312be95e5be33a2

Observation 89a0d352-e7fd-453a-84f5-9c2eb5443b92 · inbound

DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards cites this paper.

DashboardQA: Benchmarking Multimodal Agents for Question Answering on Interactive Dashboards OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:06:29.156277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:06:29.156277Z digest=sha256:abcd67ae6d91d5d93d5a5d6ac00f65b514a463473fb3d979b03cd14879205d23

Observation 190d4e78-2aa6-45b2-b793-a914e9469cc3 · inbound

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows cites this paper.

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:48:38.345661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T22:43:48.618334Z digest=sha256:2a6d1f638256ae67692e6e50d9d2fba48223f9e4aac7f233bb216a3d5182466c

Observation c6641c63-42b8-4ca7-94be-ee4dc21fd541 · inbound

ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces cites this paper.

ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:55:51.736435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:47:11.883682Z digest=sha256:9d38e53e3cb90aad9f2be329022ad986771518d1b995848e8b1ab81426917aee

Observation eceda6fe-72ac-4586-9f37-1665576155b5 · inbound

Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents cites this paper.

Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:10:36.935279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T18:52:16.162336Z digest=sha256:1ca0ec60a1b368aeef5520fb9e637e3f8b041eee42e9853d866220edcf2e9897

Observation 2193da31-5f56-4a81-8702-da43c9d03f2c · inbound

Agentic Performance at the Edge: Insights from Benchmarking cites this paper.

Agentic Performance at the Edge: Insights from Benchmarking OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:26:25.051461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T05:22:39.323345Z digest=sha256:120ac6b23d2baa3d9926311924f1218535b1a68867af6142a898aca21a198969

Observation 013aa40d-19f6-45a2-91d6-66f5db277ad9 · inbound

PREPING: Building Agent Memory without Tasks cites this paper.

PREPING: Building Agent Memory without Tasks OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:15:06.136812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T06:14:14.385586Z digest=sha256:f64b16dff4c56638da047700847446729cd431fc298758e5a848b5296b7229e0

Observation 6b947164-96fc-4a0e-a60e-1ac774deaf1d · inbound

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets cites this paper.

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:06:12.357210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T21:35:45.197423Z digest=sha256:9d0ad2b2cffddb78513bc6cd1021864f321b84ec915b8fff512fe86300bf8fee

Observation 4b235737-0405-4a4c-a78e-2228eec695fd · inbound

Signal-Driven Observation for Long-Horizon Web Agents cites this paper.

Signal-Driven Observation for Long-Horizon Web Agents OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:31:29.229037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T01:25:11.149977Z digest=sha256:ee1cbc9b81f9a581c2566708c87fd693aaf36ea4349018f3107e98d7f3df2a83

Observation 3518003a-bc10-4f58-b364-081854468251 · inbound

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents cites this paper.

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:35:48.605245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-30T01:19:54.081755Z digest=sha256:4441b03dcd67111f4941aa0fd8268955908b290a8fc2ce08ca04694ae45b03ad

Observation 0dd78781-e5e2-428c-b398-8a4df4fabe55 · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:14:21.123636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T07:10:38.909339Z digest=sha256:29990c0fca721ffe3b1775b0f3de2ff845e69d1351b683264710aca0831c40f4

Observation 4943d705-0fc6-4132-9ee3-cdf48f480e6d · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-15T10:24:53.345620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:24:53.345620Z digest=sha256:3c253eab1587217dd4ae638ce52e13e652781495a7c6dae76b45b3262cf58798

Observation d4f85837-c667-493b-bddf-ea639727d18b · inbound

PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows cites this paper.

PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-08T20:05:34.099480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T20:03:47.212663Z digest=sha256:e73d59e12ae8c67fbe599d3bc95c51703a4b4e9e79a0deffaca170109d1dae7b

Observation ac55d9ce-f3a7-4083-b841-78885a022dff · inbound

PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows cites this paper.

PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:37:42.726491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-11T01:37:23.646537Z digest=sha256:f87a88232e15dbd2b4120f91fcc489ea2cfecbb276228f8ae87fbb5d56d2e35f

Observation afff5545-726f-45fd-b48d-709ba2dcfc7c · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:17.347721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:17.347721Z digest=sha256:bbb75c35399c0eba98ae3b0d2c53679e0764a2ae6448849c54ca9985643aa8a4

Observation d5a4e0de-8ef9-4a4f-a39a-c9280cd26447 · inbound

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding cites this paper.

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-30T10:50:09.743229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T10:50:09.743229Z digest=sha256:2a304eed2efac400eb82a8058029d53555923fa35c013b314fb973210b7d0a15

Observation 3ee02c37-d69b-40b5-9bfe-df6e80b46eaf · inbound

Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability cites this paper.

Toward Reliable Context Compression for Long-Horizon Agents: An Empirical Study of Execution Instability OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T04:23:51.431618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:23:51.431618Z digest=sha256:781cf8d1af31571c72d35c30b43864652b246b27ae3627c2e8aedcb8bde59572

Observation f8c82d9a-2dc6-44fd-8cc6-cd146fffc92f · inbound

Software Engineering for and with GUI Agent cites this paper.

Software Engineering for and with GUI Agent OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 246

Resolution
unresolved
no resolver link, observed 2026-08-11T20:19:15.916578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:19:15.916578Z digest=sha256:792758bfe5ebfe5c93d60bbda56182125f6cbca61c81515c17364d17b73c2e1e

Observation cee6cd0a-078d-406f-a1ab-5e8e38fe15af · inbound

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? cites this paper.

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T14:25:46.473505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:25:46.473505Z digest=sha256:35bf35f044f04c4f74d545f585c1fbc58b9609c37dd0199d3a6fbcb9a62f78f8