{"work":{"id":"061edd26-53a3-4802-a464-5ce90e859239","openalex_id":null,"doi":null,"arxiv_id":"2508.09124","raw_key":null,"title":"OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows","authors":null,"authors_text":"Weixuan Wang, Dongge Han, Daniel Madrigal Díaz, Jin Xu, Victor Rühle, and Saravan Rajmohan","year":2025,"venue":"cs.CL","abstract":"Autonomous agents powered by large language models (LLMs) are increasingly deployed in real-world applications requiring complex, long-horizon workflows. However, existing benchmarks predominantly focus on atomic tasks that are self-contained and independent, failing to capture the long-term contextual dependencies and multi-interaction coordination required in realistic scenarios. To address this gap, we introduce OdysseyBench, a comprehensive benchmark for evaluating LLM agents on long-horizon workflows across diverse office applications including Word, Excel, PDF, Email, and Calendar. Our benchmark comprises two complementary splits: OdysseyBench+ with 300 tasks derived from real-world use cases, and OdysseyBench-Neo with 302 newly synthesized complex tasks. Each task requires agent to identify essential information from long-horizon interaction histories and perform multi-step reasoning across various applications. To enable scalable benchmark creation, we propose HomerAgents, a multi-agent framework that automates the generation of long-horizon workflow benchmarks through systematic environment exploration, task generation, and dialogue synthesis. Our extensive evaluation demonstrates that OdysseyBench effectively challenges state-of-the-art LLM agents, providing more accurate assessment of their capabilities in complex, real-world contexts compared to existing atomic task benchmarks. We believe that OdysseyBench will serve as a valuable resource for advancing the development and evaluation of LLM agents in real-world productivity scenarios. In addition, we release OdysseyBench and HomerAgents to foster research along this line.","external_url":"https://arxiv.org/abs/2508.09124","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-11T01:37:42.801023+00:00","pith_arxiv_id":"2508.09124","created_at":"2026-05-10T08:43:01.093399+00:00","updated_at":"2026-07-11T01:37:42.801023+00:00","title_quality_ok":true,"display_title":"","render_title":""},"hub":{"state":{"work_id":"061edd26-53a3-4802-a464-5ce90e859239","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":15,"external_cited_by_count":null,"distinct_field_count":6,"first_pith_cited_at":"2025-12-15T10:28:45+00:00","last_pith_cited_at":"2026-07-07T08:50:09+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T00:39:55.398664+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":2},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"background","n":2},{"context_polarity":"use_dataset","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}