Pith. sign in

Paper Citation Record · LEDGER

OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2508.09124.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09124 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:36:02.661415Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:37:42.801023Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 65d346f4-7218-450b-a8d3-1038b655014f · inbound

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows cites this paper.

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:48:38.315116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T22:43:48.618334Z digest=sha256:a4ab88bb4760dc9d1c3c4d63cfb74de48794550caefa2d0f8b89861cd871e30a

Observation 5072d6c0-9b4a-4c27-8feb-cd934da8508e · inbound

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions cites this paper.

OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:08:12.299526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:08:12.299526Z digest=sha256:0fb56d6369c8a3a0835ff4025e81419a9ea58b9642d85c7263d3f8ba5ff21555

Observation c62324e5-15e4-4783-944e-7ea13ed7d6c7 · inbound

RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments cites this paper.

RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T23:45:24.809791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:45:24.809791Z digest=sha256:8e5b9caba92bf87f64d705ff3deb973d6da484c2455611661365e65bfbcd5e8e

Observation 24c38b31-9cb7-4baa-9c58-338a028a99f0 · inbound

Opal: Private Memory for Personal AI cites this paper.

Opal: Private Memory for Personal AI OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:13:16.664821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T21:09:06.320543Z digest=sha256:bee415da39a1eb556818869a56cd064478867c6893be8579d02e91314e451533

Observation 6b0e752a-c0bd-4e4b-ae6e-4ebeb9b0e139 · inbound

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows cites this paper.

GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:43:01.095541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T08:42:57.035226Z digest=sha256:6f6671da52b2401057392518239ee6bcf7db87b3b308b8cfd20a781a8970b3ed

Observation 1bea0ef8-cb73-4380-9fa3-c3f46052a504 · inbound

CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend cites this paper.

CUJBench: Benchmarking LLM-Agent on Cross-Modal Failure Diagnosis from Browser to Backend OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:51:13.490188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T07:47:51.195882Z digest=sha256:c48e8b00157ea5530a5f33a1958362ac31b22be16bc32b5409b9a876ee535813

Observation 4fa2cec6-8f8b-4a61-b35c-6e1a2595bf5d · inbound

Tools as Continuous Flow for Evolving Agentic Reasoning cites this paper.

Tools as Continuous Flow for Evolving Agentic Reasoning OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:51.700933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T01:30:19.859374Z digest=sha256:012862acf73f36e23e31c969cf4db7f4b9230050f3c07c7046f27f200199f713

Observation 4d5e23ff-a394-4099-9135-f654742b73de · inbound

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation cites this paper.

WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:23.957989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T03:40:00.725327Z digest=sha256:bdaf1e478e7fea1e7b1843789bc170a384407fc3d78086ccb615aa435b3f07a8

Observation f5a64f86-eeb9-4c68-ba41-61067e440367 · inbound

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data cites this paper.

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:47:42.076382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T17:43:11.038874Z digest=sha256:e019b59bbde689d804eed65527edd8e46081e116d1712dccf428e83473529c3a

Observation 6e0420ee-eb3c-4c33-8f04-924fbe58b065 · inbound

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data cites this paper.

EnergyAgentBench: Benchmarking LLM Agents on Live Energy Infrastructure Data OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T05:13:41.995563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:13:41.995563Z digest=sha256:d3c38cf678d635240a2a0a1e3014560c8aecf27d608e157cdc13c1f4ccd2e611

Observation 4d295bdd-2f3e-46ed-8708-8acc3d463185 · inbound

SentinelBench: A Benchmark for Long-Running Monitoring Agents cites this paper.

SentinelBench: A Benchmark for Long-Running Monitoring Agents OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:16:48.301325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T06:09:38.698353Z digest=sha256:f51ffaa75f2acacbfdefa18217ec163a73cd300f7efae59a3d37d57d9d4a1c0c

Observation 2fb866c7-3830-4bb4-b3fb-a9ecb1dd1d5a · inbound

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents cites this paper.

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:26.489813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T19:18:56.039587Z digest=sha256:98516c3d49a4cfc0581a66f4cc6b891fcb569cbbb749495400d8fb7f6040cbbf

Observation a4eb0f97-5ceb-46e3-8957-5d90ae170888 · inbound

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments cites this paper.

Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:08:43.394326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T04:35:38.527140Z digest=sha256:d95aec646df26437fefb000554a951808808bd9ee806805369ed9593fb2ff85b

Observation 3ab79eeb-82cd-4ef1-80d5-1a74c4fe99e4 · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T21:38:58.671790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T00:19:47.396643Z digest=sha256:1a02d4c7cbebca011f24f1447a66d95f20ce72ac72b8c1cdfbe58aa6a30ce3a0

Observation f8e0eeee-4f78-4a5a-94ab-21a3e66e5da5 · inbound

CEO-Bench: Can Agents Play the Long Game? cites this paper.

CEO-Bench: Can Agents Play the Long Game? OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T11:01:20.504958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T11:01:20.504958Z digest=sha256:c7b745434451283d4fc0d314386a6880ab2241df6eb9f5d99b38b84cc985a0f0

Observation ad13d3d4-f532-4450-96d8-0ce87595a72a · inbound

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks cites this paper.

ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:38.855626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T14:05:28.423719Z digest=sha256:eaec984a8e6619685441331a3d62c55ca5af4892f31dc069ec44b3056176c80e

Observation 561a7d5e-68da-4aaa-9250-13b69716b9d0 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:56.687129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:0fab42b838bebc05607d0fd6a255224a98d1c08e52ad93af182e7951a6bca366

Observation 0adcedff-cf5d-4e65-ac83-237afa52fb7e · inbound

Office Comprehension Benchmark cites this paper.

Office Comprehension Benchmark OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:29:15.017566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-04T00:20:42.974208Z digest=sha256:40a7cb22a02e2385784e6561fd77e8f80406bdeeff5221675c339a534052aa9f

Observation 662aa882-8fc7-40db-979c-6a923384bff0 · inbound

PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows cites this paper.

PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-08T20:05:34.092723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T20:03:47.212663Z digest=sha256:2ee10f3bb4ea8ec8b45a2aeadbd9b65ef879f3e2004202af3d65491cd2f30614

Observation 3d6e37f4-3f48-4ee5-9232-60d5934d89f5 · inbound

PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows cites this paper.

PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:37:42.831140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-11T01:37:23.646537Z digest=sha256:f77f2f5244939cb1afd2f0810f715c5bac8f473de5f2e28f24a0ddf118b4bc19

Observation ab5bce11-45c1-4c3c-bf0e-d09550334045 · inbound

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction cites this paper.

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:36:02.661415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:36:02.661415Z digest=sha256:6701b451ff8ecc2db84adea9ae47d16cfd51bf84c93a3ac77390ecd7b9336c7d

Observation 35539d15-0a15-41c9-9662-f7ecf4bb764c · inbound

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding cites this paper.

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T10:50:09.643083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T10:50:09.643083Z digest=sha256:dd8e5a3dd7a795161aadeace404cd3ef2c49bd8bd9aedb5b18e749d3c29a079f

Observation 9eac2e05-e348-4e38-86bb-fc182ea509c7 · inbound

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning cites this paper.

Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:44.200213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:44.200213Z digest=sha256:9ce3b4d5da40faefde25843476d6cf300b3562cb1bb19fa69ebf1a9e958d84ac

Observation ba1c042e-869e-4601-8b0d-3df27d192725 · inbound

OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality cites this paper.

OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T16:55:25.073757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:55:25.073757Z digest=sha256:f2465724081c1083c1f4d62927f2a221a558b9590fae4055a9db59254f916184