Pith. sign in

Paper Citation Record · LEDGER

OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2407.19056.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.19056 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T03:38:17.347721Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c58cb6eb-8d61-4432-b146-05acd5fc2ac8 · inbound

A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions cites this paper.

A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:02:35.183609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T04:59:36.994758Z digest=sha256:abd1816ec484937e2c2439e7249908b06caec620c307872cb83105959431ca1f

Observation 190d4e78-2aa6-45b2-b793-a914e9469cc3 · inbound

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows cites this paper.

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:48:38.345661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T22:43:48.618334Z digest=sha256:aa70b5ce931e025921a9cffb3c8f059ea2ab168410788f1effd3bdd9442cb0ab

Observation c6641c63-42b8-4ca7-94be-ee4dc21fd541 · inbound

ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces cites this paper.

ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:55:51.736435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:47:11.883682Z digest=sha256:4cbf4a6a0d2fd50d70b48c6d8b222ec01eb958f2547cec914ebfc95edde060bb

Observation eceda6fe-72ac-4586-9f37-1665576155b5 · inbound

Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents cites this paper.

Augmenting Interface Usability Heuristics for Reliable Computer-Use Agents OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:10:36.935279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T18:52:16.162336Z digest=sha256:fe7cd9672f9df7df4f951592f808ba3582e9dacb3e089303c4a7676ea9d54194

Observation 2193da31-5f56-4a81-8702-da43c9d03f2c · inbound

Agentic Performance at the Edge: Insights from Benchmarking cites this paper.

Agentic Performance at the Edge: Insights from Benchmarking OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:26:25.051461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T05:22:39.323345Z digest=sha256:b77f294a269926d4a8e5ae65c63b5fbaa35c1d91ef70645650e65e19fab083ef

Observation 013aa40d-19f6-45a2-91d6-66f5db277ad9 · inbound

PREPING: Building Agent Memory without Tasks cites this paper.

PREPING: Building Agent Memory without Tasks OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:15:06.136812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T06:14:14.385586Z digest=sha256:e0c3cc263b3bc6ad82acca54a36c15e000eb4c2de040055e55382667d9caafaa

Observation 6b947164-96fc-4a0e-a60e-1ac774deaf1d · inbound

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets cites this paper.

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:06:12.357210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T21:35:45.197423Z digest=sha256:9346908da8738830a6eec79a5ee395db9430da985616bc852ad4f0ce5c269450

Observation 4b235737-0405-4a4c-a78e-2228eec695fd · inbound

Signal-Driven Observation for Long-Horizon Web Agents cites this paper.

Signal-Driven Observation for Long-Horizon Web Agents OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-28T01:31:29.229037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T01:25:11.149977Z digest=sha256:963808291d37fa02a1b58429f6b972d41ef052863e0f35737c5114c3c22b2dc8

Observation 3518003a-bc10-4f58-b364-081854468251 · inbound

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents cites this paper.

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:35:48.605245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T01:19:54.081755Z digest=sha256:7007d51d76630845cd1dc88709556b8a8b730511fb6d0f45a6a2e174aa12189b

Observation 0dd78781-e5e2-428c-b398-8a4df4fabe55 · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:14:21.123636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T07:10:38.909339Z digest=sha256:78c28a8833bfc47f93d74380647d6c81804eefd103d917644f9a4f571e6d231c

Observation 4943d705-0fc6-4132-9ee3-cdf48f480e6d · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-15T10:24:53.345620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:24:53.345620Z digest=sha256:164441be117eefd7b77feab35afa816fef0560e2cd94fe63310965894b37eb82

Observation d4f85837-c667-493b-bddf-ea639727d18b · inbound

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents cites this paper.

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-08T20:05:34.099480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T20:03:47.212663Z digest=sha256:22a2c93f3b2a34efcbfc4ac50ab7dd2986dcf85058b80e3e3aa0ef6add3cd027

Observation ac55d9ce-f3a7-4083-b841-78885a022dff · inbound

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents cites this paper.

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:37:42.726491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T01:37:23.646537Z digest=sha256:9f08f054b758be196f3992f252be914b03191e285fae8d313f90185e9d5bb899

Observation afff5545-726f-45fd-b48d-709ba2dcfc7c · inbound

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications cites this paper.

AgentOmnia: Scaling Agentic Models for Full-Scenario Applications OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T03:38:17.347721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:38:17.347721Z digest=sha256:180577550108f79cd8b5ff7de604c7a21a812ec1a0e3a4bc721ac06ce6001e30

Observation d5a4e0de-8ef9-4a4f-a39a-c9280cd26447 · inbound

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding cites this paper.

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-30T10:50:09.743229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T10:50:09.743229Z digest=sha256:8f6cae08073915a6dcc844b4975aba2be593b1ba52aa0d6c20b8f6f71ca181c5