Pith. sign in

Paper Citation Record · LEDGER

OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2508.09124.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09124 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:13:41.995563Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:37:42.801023Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 65d346f4-7218-450b-a8d3-1038b655014f · inbound

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows cites this paper.

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:48:38.315116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T22:43:48.618334Z digest=sha256:cb16eb08e72026a22431c3bb81f1299feb3fba3475011ec2a619e0c4ed0e768c

Observation 5072d6c0-9b4a-4c27-8feb-cd934da8508e · inbound

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions cites this paper.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.299526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.299526Z digest=sha256:80c3b912516dcdfb63e5aefba23f80adcb1a07ee76dff836fe53b5ec071f94ae

Observation c62324e5-15e4-4783-944e-7ea13ed7d6c7 · inbound

RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments cites this paper.

RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T23:45:24.809791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:45:24.809791Z digest=sha256:d1eb29b246d31452aa013defc2ee25f6eb100f00e361bff5022331ec228c445a

Observation 24c38b31-9cb7-4baa-9c58-338a028a99f0 · inbound

Opal: Private Memory for Personal AI cites this paper.

Opal: Private Memory for Personal AI OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:13:16.664821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T21:09:06.320543Z digest=sha256:5689ad23305968193d8f5ee493f429293912b3f11b7175b22a6f316637c3fca2

Observation 6b0e752a-c0bd-4e4b-ae6e-4ebeb9b0e139 · inbound

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows cites this paper.

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:43:01.095541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T08:42:57.035226Z digest=sha256:650627464b2e25134ea62543e9a854adf62fd211c271b21b7e346722521e86e3

Observation 1bea0ef8-cb73-4380-9fa3-c3f46052a504 · inbound

CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend cites this paper.

CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:13.490188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T07:47:51.195882Z digest=sha256:0370b82664f855b318b9597d46662c99d9695834966606eebdf53444e41a5152

Observation 4fa2cec6-8f8b-4a61-b35c-6e1a2595bf5d · inbound

Tools as Continuous Flow for Evolving Agentic Reasoning cites this paper.

Tools as Continuous Flow for Evolving Agentic Reasoning OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:51.700933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T01:30:19.859374Z digest=sha256:ab7141ac53f17b4bc027e221e35af6e4f088a6cca1961ee488de0ebaa19ec7b9

Observation 4d5e23ff-a394-4099-9135-f654742b73de · inbound

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation cites this paper.

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:23.957989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T03:40:00.725327Z digest=sha256:75e158c7501119be2d39ea55990a17c3449faaca3e0da3a74af80c4c0f9ba30d

Observation f5a64f86-eeb9-4c68-ba41-61067e440367 · inbound

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data cites this paper.

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:47:42.076382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-19T17:43:11.038874Z digest=sha256:22433f1be63a64e782060edf161875174bd27f71f56e8ffaaadc51e162d7ccce

Observation 6e0420ee-eb3c-4c33-8f04-924fbe58b065 · inbound

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data cites this paper.

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T05:13:41.995563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:13:41.995563Z digest=sha256:25fa8d3d5a1cae8f5d766f3212fe1be46788c02313616e1c87bf4260283129ad

Observation 4d295bdd-2f3e-46ed-8708-8acc3d463185 · inbound

SentinelBench: A Benchmark for Long-Running Monitoring Agents cites this paper.

SentinelBench: A Benchmark for Long-Running Monitoring Agents OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:16:48.301325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-28T06:09:38.698353Z digest=sha256:66a7ea7d0e1745e8e6beabda58366a44d0d23994ca46f762a16f78a42fe82c23

Observation 2fb866c7-3830-4bb4-b3fb-a9ecb1dd1d5a · inbound

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents cites this paper.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:26.489813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:db3d26ddccb6fa24791f83270fd097ddcfe0eaa1a0a6fe7bf2ce78f7dffb4698

Observation a4eb0f97-5ceb-46e3-8957-5d90ae170888 · inbound

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments cites this paper.

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:08:43.394326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-27T04:35:38.527140Z digest=sha256:7565d44f2f0d38bfa8adddb7091b029855deac8e6130f5fdf917725d969a8dda

Observation 3ab79eeb-82cd-4ef1-80d5-1a74c4fe99e4 · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:38:58.671790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T00:19:47.396643Z digest=sha256:c590d4a8d1f8e83ac8372de903369ef60a871227dacce9066b1e0ecadeb07b00

Observation f8e0eeee-4f78-4a5a-94ab-21a3e66e5da5 · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T11:01:20.504958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:01:20.504958Z digest=sha256:a51f76b5da740a8a2e4fc3a43418a205567233b25cb51a040844069f2a70e281

Observation ad13d3d4-f532-4450-96d8-0ce87595a72a · inbound

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks cites this paper.

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:38.855626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T14:05:28.423719Z digest=sha256:89add4b7393d21a1f8af25a95e49c027a3f74aa1f0b0fc77f1224583ce207313

Observation 561a7d5e-68da-4aaa-9250-13b69716b9d0 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:56.687129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:2e462f63fd52c999bc5cb9e24feaf2a544e0017b39171166bff9b719039fa479

Observation 0adcedff-cf5d-4e65-ac83-237afa52fb7e · inbound

Office Comprehension Benchmark cites this paper.

Office Comprehension Benchmark OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:29:15.017566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-04T00:20:42.974208Z digest=sha256:1b8d5140e73d879dc295252d68b03d7565ae4770efe86238972630ca975e9bb8

Observation 662aa882-8fc7-40db-979c-6a923384bff0 · inbound

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents cites this paper.

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-08T20:05:34.092723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T20:03:47.212663Z digest=sha256:381c48509805a99ebe8070ec0719980635141a475208f75140cae60cfbd9b4dd

Observation 3d6e37f4-3f48-4ee5-9232-60d5934d89f5 · inbound

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents cites this paper.

PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:37:42.831140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T01:37:23.646537Z digest=sha256:68202b0787ecc47936d001ea33aa3c526cdd27885161bc228f482d447f3acedd

Observation 35539d15-0a15-41c9-9662-f7ecf4bb764c · inbound

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding cites this paper.

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T10:50:09.643083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T10:50:09.643083Z digest=sha256:066f791caa00c19b4487b596f18b8fe61a9a77b5786051c0485f7ff3d64bf8d2