REVIEW 3 major objections 6 minor 47 references
On 100 long office tasks, today’s LLM agents are far cheaper and faster than people but still trail human deliverable quality.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 10:48 UTC pith:6DMB5TES
load-bearing objection Useful open office-agent suite with real per-task economics; the human–LLM quality gap is directionally right but inflated by best-of humans vs single-run models. the 3 major comments →
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Current frontier LLM agents on OmegaUse-OfficeVal remain substantially below a junior human baseline in deliverable quality (best model overall score 17.91 versus human 27.79) even though they are much cheaper and faster per task; value-weighted rankings further show that average quality and capture of high-labor or high-price work can diverge.
What carries the argument
OmegaUse-OfficeVal: 100 practitioner-derived long-horizon office tasks, each paired with human labor time and a task price proxy, and scored by code-based verifiers from usability plus weighted task-completion rubrics that zero out unusable files.
Load-bearing premise
That measuring agents only through a fixed programmatic (non-GUI) scaffold and automated rubric code, against best human submissions, fairly states the real quality and cost gap for office work.
What would settle it
Re-run the same 100 tasks with GUI-capable agents or alternate scaffolds and human-judged deliverables: if a model then matches or beats the human score of 27.79 at still-lower cost, or if code-verifier rankings reverse under expert review, the claimed quality gap and economic comparison fail.
If this is right
- Office-agent progress can be tracked on final file quality and dollars/hours saved, not only on GUI step success.
- Value-weighted scores will matter for deployment: winning on average may not mean winning on the highest-priced work.
- Longer human-labor tasks are a clearer difficulty signal; agents that hold quality as horizon grows are the ones that close the gap.
- Open release of tasks, rubrics, verifiers, and economic labels lets others reproduce cost–quality trade-offs as models change.
Where Pith is reading between the lines
- If quality catches up while cost stays low, routine multi-hour assistant and intern office work is the first large white-collar slice under real automation pressure.
- The human ceiling at 27.79 under strict rubrics implies many “done” office files still fail usability or damage checks—agents may need repair-aware objectives, not only completion.
- Programmatic file agents and GUI computer-use agents may need joint leaderboards on the same deliverables before claiming parity with junior staff.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces OmegaUse-OfficeVal, a fully open benchmark of 100 long-horizon office-suite tasks (DOCX/PPTX/XLSX/PDF and multimodal inputs) derived from practitioner requests via a privacy-preserving funnel. Each task carries task-level economic annotations—recorded human labor time (mean 2.32 h) and a hybrid task price proxy—and is scored by code-based verifiers built from fine-grained usability and completion rubrics (Table 2; §4.3). Several frontier LLMs are evaluated under a shared programmatic (non-GUI) scaffold against a human baseline. The headline empirical claim (Abstract; Table 3) is that LLMs are much cheaper and faster than humans but lag in deliverable quality (best model score 17.91 vs human 27.79), with value-weighted rankings that can diverge from unweighted averages.
Significance. If the construction and evaluation hold, this is a useful and timely contribution to agent evaluation. Relative to GDPVal/RLI/ALE and existing office or GUI benchmarks, the combination of (i) long-horizon office deliverables, (ii) per-task economic signals enabling cost and value-weighted analysis, (iii) deterministic code verifiers rather than LLM-as-judge, and (iv) a complete public release of instructions, files, rubrics, verifiers, and annotations is distinctive and practically valuable. The open assets and reproducible scoring protocol are real strengths and should support cumulative progress on “vibe working.” The significance of the human–LLM quality-gap claim is more conditional on how the human baseline and scaffold are interpreted.
major comments (3)
- [§5.1.1, Table 3, Abstract] §5.1.1 and Table 3: The central quality claim (human 27.79 vs best LLM 17.91; “have not yet approached human-level deliverable quality”) rests on an asymmetric aggregation. For each task the paper takes the highest-scoring human submission among ≥2 annotators, while each LLM is scored once under a fixed scaffold. The text itself notes that human quality “can vary substantially.” Best-of-N systematically raises the human bar relative to single-run models and is not the same object as the labor-time estimate (mean of the two shortest valid times; App. C / Algorithm 1) or the task price proxy used for cost comparisons. This is load-bearing for the Abstract and §5.2 narrative. Please report mean (and ideally median / per-annotator) human scores alongside best-of, and either (a) restate the claim as a gap to a best junior deliverable, or (b) primary-compare against mean human quality. Without
- [§5.1, Appendix E.1, Abstract] §5.1 / App. E.1: All LLM results use an in-house programmatic scaffold (tools, shell, file APIs) with “GUI-level computer-use capabilities … not included.” Office-suite work is often GUI-native (layout, animations, print areas, slide design; Fig. 2 examples). The paper’s claim is about LLM agents on office-suite workflows, not only scriptable file edits. A non-GUI-only protocol can understate attainable quality and distort model rankings for tasks where visual/GUI affordances matter (PPT-heavy and beautify/restructure intents in App. G). This need not invalidate the benchmark, but the Abstract/Conclusion should scope the negative quality result to programmatic agents, and the paper should either add a GUI/CUA condition for a subset of tasks or provide a clear limitation analysis of which task types are most scaffold-sensitive.
- [§4.3.3, Table 3, Figure 7] §4.3.3 scoring and Table 3: Even the best human scores only 27.79/100 under the two-stage usability gate and weighted completion rubric (negative items, clip at zero, usability fail ⇒ 0). That low ceiling is consistent with a harsh, repair-burden-aware protocol, but it also means absolute scores are hard to interpret as “deliverable quality” in a user sense, and zero-score mass is large even for humans (Fig. 7: 29%). For the economic story (cheaper/faster yet not approaching human quality), please add calibration evidence that code scores track expert usability judgments beyond the discrepancy-resolution process (e.g., correlation or agreement rates on held-out artifacts; sensitivity to ±1/±3/±5 weights). Otherwise it remains unclear whether the gap is mainly missing requirements, usability failures, or verifier strictness.
minor comments (6)
- [§4.2.2] §4.2.2: Task price proxy uses explicit prices for ~20% of tasks and three-expert aggregation otherwise, with a simple outlier rule (A/B ≥ 2). Report inter-expert dispersion, agreement with the explicit-price subset where both exist, and sensitivity of price-weighted rankings (Table 3) to alternative aggregations (median, trimmed mean).
- [§3.1] Exchange rate is given as 1 CNY = 0.14609 USD averaged Jan 19–Jul 16, 2026. For reproducibility, state the source series and keep CNY primary in tables with USD as secondary.
- [Figure 2] Fig. 2 task cards appear to reuse the same instruction text across unrelated examples (e.g., Corporate Operations Statement language under Brand Marketing / Education). Likely a figure-assembly error; please fix so examples match the described tasks.
- [Appendix F, Table 3] App. F defines overall Score as a sum over tasks rather than a mean; Table 3 reads like an average-scale score (~17–28). Clarify whether reported “Score” is mean×100, sum, or another normalization so weighted metrics are comparable.
- [§5.2, Appendix E.2] Runtime/cost comparisons (§5.2, Fig. 6) mix human wall-clock labor with agent time that includes API latency and scaffold overhead (App. E.2 caveat). Mark LLM time explicitly as reference wall-clock under the scaffold, not pure model compute.
- [§2] Related work is thorough; a compact comparison table (horizon, economic labels, openness, eval target: state vs deliverable) in §2 would help readers place OfficeVal vs OSWorld 2.0, OdysseyBench, and GDPVal quickly.
Circularity Check
No significant circularity: empirical benchmark with independently collected labels and fixed code verifiers, not a self-defining derivation.
full rationale
OmegaUse-OfficeVal is a benchmark/leaderboard paper, not a first-principles derivation. The central claim (Table 3: human score 27.79 vs best LLM 17.91; LLMs cheaper/faster) is an empirical comparison of held-out-style code-verifier scores on agent and human artifacts. Human labor time and task price proxy are annotated separately from model runs (Sec. 4.2, App. C) and are not fitted to produce the quality ranking. Rubrics and verifiers are built from task requirements then calibrated so automated checks align with expert judgment (Sec. 4.3, App. D); that is measurement validation, not defining model scores as inputs by construction. Self-citations (e.g., OmegaUse GUI work) are background and not load-bearing for the headline gap. Methodological asymmetries (best-of human vs single-scaffold LLM; non-GUI scaffold) affect fairness of the gap but are not circular reductions. No self-definitional loop, fitted-input-as-prediction, uniqueness import, or renamed known result appears in the derivation chain.
Axiom & Free-Parameter Ledger
free parameters (5)
- Rubric item weights (+1/+3/+5 and −1/−3/−5) =
three-level positive/negative scale per Appendix D.1
- Human labor time aggregation rule =
avg of two shortest; 1.3× and 0.7× thresholds in Algorithm 1
- Task price proxy aggregation =
explicit price if available else mean after outlier rule
- CNY→USD exchange rate =
1 CNY = 0.14609 USD
- Final task set size after funnel =
100 tasks
axioms (5)
- domain assumption Final deliverable quality for office-suite work can be adequately scored by deterministic code checks derived from fine-grained usability and task-completion rubrics, without human or LLM judges at evaluation time.
- ad hoc to paper A programmatic tool-use scaffold without GUI computer-use is a sufficient setting to compare frontier LLMs on office-suite workflows for the paper’s claims about current agent capability.
- ad hoc to paper Best junior-annotator deliverable per task is an appropriate human reference for the capability gap (as opposed to mean human, first-pass human, or professional freelancers).
- domain assumption Expert-estimated task prices, when practitioner quotes are absent, are valid market-value proxies for value-weighted evaluation.
- domain assumption Privacy-preserving reconstruction of input artifacts with LLM assistance plus human revision preserves task difficulty and intent of original practitioner work.
invented entities (3)
-
OmegaUse-OfficeVal benchmark suite
independent evidence
-
Task price proxy (hybrid explicit/expert)
no independent evidence
-
Two-stage usability gate × weighted completion score
no independent evidence
read the original abstract
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality. The code and dataset are fully open-sourced, and more information is available on our project website: https://omegause-officeval.github.io.
Figures
Reference graph
Works this paper leans on
-
[1]
L´eo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault Le Sellier De Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, and Alexandre Drouin. 2024. WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work Tasks. InAdvances in Neural Information Processing Systems
2024
-
[2]
Rogerio Bonatti, Dan Zhao, Francesco Bonacci, Dillon Dupont, Sara Abdali, Yinheng Li, Yadong Lu, Justin Wagle, Kazuhito Koishida, Arthur Bucker, Lawrence Keunho Jang, and Zheng Hui. 2025. Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale. InProceedings of the 42nd International Conference on Machine Learning. 4874–4910
2025
-
[3]
Sumit Chauhan. 2025. Vibe Working: Introducing Agent Mode and Office Agent in Microsoft 365 Copilot. Microsoft 365 Blog. Official Microsoft 365 blog post, published September 29, 2025. https: //www.microsoft.com/en-us/microsoft-365/blog/2025/09/29/vibe-working-introducing-a gent-mode-and-office-agent-in-microsoft-365-copilot/
2025
-
[4]
Mariya Davydova, Daniel Jeffries, Patrick Barker, Arturo M´arquez Flores, and Sin´ead Ryan. 2025. OSUniverse: Benchmark for Multimodal GUI-navigation AI Agents.arXiv preprint arXiv:2505.03570(2025)
Pith/arXiv arXiv 2025
-
[5]
DeepSeek team. 2026. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence.arXiv preprint arXiv:2606.19348(2026)
arXiv 2026
-
[6]
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su. 2023. Mind2Web: Towards a Generalist Agent for the Web. InAdvances in Neural Information Processing Systems. 28091–28114
2023
-
[7]
Laradji, Manuel Del Verme, Tom Marty, David Vazquez, Nicolas Chapados, and Alexandre Lacoste
Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, Manuel Del Verme, Tom Marty, David Vazquez, Nicolas Chapados, and Alexandre Lacoste. 2024. WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?. InProceedings of the 41st International Conference on Machine Learning. 11642–11662
2024
-
[8]
Nguyen, Muhammad Taqi Raza, Vishal Chowdhary, and Graham Neubig
Apurva Gandhi, Vishwas Suryanarayanan, Raja Hasnain Anwar, Firoz Shaik, Shubhang Desai, Thong Q. Nguyen, Muhammad Taqi Raza, Vishal Chowdhary, and Graham Neubig. 2026. PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks. InProceedings of the 43rd International Conference on Machine Learning
2026
-
[9]
Jiaxin Ge, Zora Zhiruo Wang, Xuhui Zhou, Yi-Hao Peng, Sanjay Subramanian, Qinyue Tan, Maarten Sap, Alane Suhr, Daniel Fried, Graham Neubig, and Trevor Darrell. 2025. AutoPresent: Designing Structured Visuals from Scratch. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2902–2911
2025
-
[10]
GLM-5-Team. 2026. GLM-5: from Vibe Coding to Agentic Engineering.arXiv preprint arXiv:2602.15763 (2026)
Pith/arXiv arXiv 2026
-
[11]
Yiduo Guo, Zekai Zhang, Yaobo Liang, Dongyan Zhao, and Nan Duan. 2024. PPTC Benchmark: Evaluating Large Language Models for PowerPoint Task Completion. InFindings of the Association for Computational Linguistics: ACL 2024. 8682–8701
2024
-
[12]
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024. WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 6864–6890
2024
-
[13]
Jiazhen Jiang, Boxi Cao, Lingyong Yan, Yaojie Lu, Hongyu Lin, Shuaiqiang Wang, Dawei Yin, Xianpei Han, and Le Sun. 2026. DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations.arXiv preprint arXiv:2607.19865(2026). 13 OmegaUse-OfficeVal
Pith/arXiv arXiv 2026
-
[14]
Kimi Team. 2026. Kimi K2.5: Visual Agentic Intelligence.arXiv preprint arXiv:2602.02276(2026)
Pith/arXiv arXiv 2026
-
[15]
Kimi Team. 2026. Kimi K2.6: Advancing Open-Source Coding. Kimi Research Blog. Official model release blog, published April 20, 2026.https://www.kimi.com/blog/kimi-k2-6
2026
-
[16]
Jing Yu Koh, Robert Lo, Lawrence Jang, Vikram Duvvur, Ming Lim, Po-Yu Huang, Graham Neubig, Shuyan Zhou, Russ Salakhutdinov, and Daniel Fried. 2024. VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 881–905
2024
-
[17]
Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Jinkai Hu, Jiayao Li, Rui Gao, Zekun Li, Songquan Zhu, Jingkai Zhou, and Pengyu Zhao
-
[18]
Hongxin Li, Jingran Su, Yuntao Chen, Qing Li, and Zhao-Xiang Zhang. 2023. SheetCopilot: Bringing Software Productivity to the Next Level through Large Language Models. InAdvances in Neural Information Processing Systems. 4952–4984
2023
-
[19]
Jinchao Li, Yunxin Li, Chenrui Zhao, Zhenran Xu, Baotian Hu, and Min Zhang. 2026. WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments. In Findings of the Association for Computational Linguistics: ACL 2026. 15262–15280
2026
-
[20]
Evan Zheran Liu, Kelvin Guu, Panupong Pasupat, Tianlin Shi, and Percy Liang. 2018. Reinforcement Learning on Web Interfaces Using Workflow-Guided Exploration. InInternational Conference on Learning Representations
2018
-
[21]
Tengchao Lv, Dongdong Zhang, Jiayu Ding, Yilin Jia, Yuzhong Zhao, Yupan Huang, Wenshan Wu, Xiangyang Zhou, Shaohan Huang, Nan Yang, Li Dong, Lei Cui, and Furu Wei. 2026. Mind the Gap: Can Frontier LLMs Pass a Standardized Office Proficiency Exam?arXiv preprint arXiv:2606.10956(2026)
Pith/arXiv arXiv 2026
-
[22]
Zeyao Ma, Bohan Zhang, Jing Zhang, Jifan Yu, Xiaokang Zhang, Xiaohan Zhang, Sijia Luo, Xi Wang, and Jie Tang. 2024. SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation. InAdvances in Neural Information Processing Systems. 94871–94908
2024
-
[23]
Mantas Mazeika, Alice Gatti, Cristina Menghini, Udari Madhushani Sehwag, Shivam Singhal, Yury Orlovskiy, Steven Basart, Manasi Sharma, Denis Peskoff, Elaine Lau, Jaehyuk Lim, Lachlan Carroll, Alice Blair, Vinaya Sivakumar, Sumana Basu, Brad Kenstler, Yuntao Ma, Julian Michael, Xiaoke Li, Oliver Ingebretsen, Aditya Mehta, Jean Mottola, John Teichmann, Kevi...
arXiv 2025
-
[24]
Gr´egoire Mialon, Cl´ementine Fourrier, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2024. GAIA: a benchmark for General AI Assistants. InInternational Conference on Learning Representations
2024
-
[25]
MiniMax. 2026. MiniMax M3: Frontier Coding, 1M Context, Native Multimodality—All in One Model. MiniMax Research Blog. Official model release blog, published June 1, 2026. https://www.minimax.io /blog/minimax-m3
2026
-
[26]
Samuel Miserendino, Michele Wang, Tejal Patwardhan, and Johannes Heidecke. 2025. SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?. InProceedings of the 42nd International Conference on Machine Learning. 44412–44450
2025
-
[27]
Kim, Samuel Miserendino, Gildas Chabot, David Li, Patrick Chao, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek
Tejal Patwardhan, Rachel Dias, Elizabeth Proehl, Grace Kim, Michele Wang, Olivia Watkins, Sim ´on Posada Fishman, Marwan Aljubeh, Phoebe Thacker, Laurance Fauconnet, Natalie S. Kim, Samuel Miserendino, Gildas Chabot, David Li, Patrick Chao, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek. 2026. GDPval: Evaluating AI Model Performance on R...
2026
-
[28]
Qwen Team. 2026. Qwen3.7-Plus: Multimodal Agent Intelligence. Qwen Research Blog. Official model release blog, published June 1, 2026.https://qwen.ai/blog?id=qwen3.7-plus
2026
-
[29]
Tianlin Shi, Andrej Karpathy, Linxi Fan, Jonathan Hernandez, and Percy Liang. 2017. World of Bits: An Open-Domain Platform for Web-Based Agents. InProceedings of the 34th International Conference on Machine Learning. 3135–3144. 14 OmegaUse-OfficeVal
2017
-
[30]
Olly Styles, Sam Miller, Patricio Cerda-Mardini, Tanaya Guha, Victor Sanchez, and Bertie Vidgen. 2024. WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting. InConference on Language Modeling
2024
-
[31]
Yiyou Sun et al. 2026. Agents’ Last Exam.arXiv preprint arXiv:2606.05405(2026)
arXiv 2026
-
[32]
Sunstein, Eric Topol, Brendan Foody, and Osvald Nitski
Bertie Vidgen, Abby Fennelly, Evan Pinnix, Julien Benchek, Daniyal Khan, Zach Richards, Austin Bridges, Calix Huang, Kanishka Sahu, Abhishek Kottamasu, Bo Ma, Ben Hunsberger, Isaac Robinson, Akul Datta, Chirag Mahapatra, Dominic Barton, Cass R. Sunstein, Eric Topol, Brendan Foody, and Osvald Nitski. 2025. The AI Productivity Index (APEX).arXiv preprint ar...
arXiv 2025
-
[33]
Bertie Vidgen, Austin Mann, Abby Fennelly, John Wright Stanly, Lucas Rothman, Marco Burstein, Julien Benchek, David Ostrofsky, Anirudh Ravichandran, Debnil Sur, Neel Venugopal, Alannah Hsia, Isaac Robinson, Calix Huang, Olivia Varones, Daniyal Khan, Michael Haines, Austin Bridges, Jesse Boyle, Koby Twist, Zach Richards, Chirag Mahapatra, Brendan Foody, an...
arXiv 2026
-
[34]
Weixuan Wang, Dongge Han, Daniel Madrigal Diaz, Jin Xu, Victor R¨ uhle, and Saravan Rajmohan. 2025. OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows.arXiv preprint arXiv:2508.09124(2025)
Pith/arXiv arXiv 2025
-
[35]
Zilong Wang, Yuedong Cui, Li Zhong, Zimin Zhang, Da Yin, Bill Yuchen Lin, and Jingbo Shang. 2024. OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation.arXiv preprint arXiv:2407.19056(2024)
Pith/arXiv arXiv 2024
-
[36]
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, Yitao Liu, Yiheng Xu, Shuyan Zhou, Silvio Savarese, Caiming Xiong, Victor Zhong, and Tao Yu. 2024. OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments. InAdvances in Neural Informa...
2024
-
[37]
Frank (Fangzheng) Xu, Yufan Song, Boxuan Li, Yuxuan Tang, Kritanjali Jain, Mengxue Bao, Zora Z. Wang, Xuhui Zhou, Zhitong Guo, Murong Cao, Mingyang Yang, Hao Yang Lu, Amaad Martin, Zhe Su, Leander Maben, Raj Mehta, Wayne Chi, Lawrence Jang, Yiqing Xie, Shuyan Zhou, and Graham Neubig. 2025. TheAgentCompany: Benchmarking LLM Agents on Consequential Real Wor...
2025
-
[38]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, ...
Pith/arXiv arXiv 2025
-
[39]
Pei Yang, Hai Ci, and Mike Zheng Shou. 2025. macOSWorld: A Multilingual Interactive Benchmark for GUI Agents. InAdvances in Neural Information Processing Systems
2025
-
[40]
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022. WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents. InAdvances in Neural Information Processing Systems. 20744–20757
2022
-
[41]
2025.𝜏-bench: A Benchmark for Tool- Agent-User Interaction in Real-World Domains
Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan. 2025.𝜏-bench: A Benchmark for Tool- Agent-User Interaction in Real-World Domains. InInternational Conference on Learning Representations
2025
-
[42]
Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong, Weiming Wu, Jiayang Sun, Jiamin Song, Kaiqian Cui, Bowen Wang, Haoyuan Wu, Yitong Li, Dunjie Lu, Haikong Lu, Qi Zhen, Xinyuan Wang, Jiaqi Deng, Yuhao Yang, Cheng Chen, Boyuan Zheng, Alex Su, Xiao Yu, Hao Zou, Saaket Agashe, Xing Han Lu, Manpreet Kaur, Zhengyang Qi, Vincent Sunn Chen, Frederic Sala, Dayiheng Liu, ...
Pith/arXiv arXiv 2026
-
[43]
Z.ai. 2026. GLM-5.2: Built for Long-Horizon Tasks. Z.ai Research Blog. Official model release blog, published June 16, 2026.https://z.ai/blog/glm-5.2
2026
-
[44]
Le Zhang, Yixiong Xiao, Xinjiang Lu, Jingjia Cao, Yusai Zhao, Jingbo Zhou, Lang An, Zikan Feng, Wanxiang Sha, Yu Shi, Congxi Xiao, Jian Xiong, Yankai Zhang, Hua Wu, and Haifeng Wang. 2026. OmegaUse: Building a General-Purpose GUI Agent for Autonomous Task Execution.arXiv preprint arXiv:2601.20380(2026). 15 OmegaUse-OfficeVal
arXiv 2026
-
[45]
Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. 2024. WebArena: A Realistic Web Environment for Building Autonomous Agents. InInternational Conference on Learning Representations
2024
-
[46]
Jian Zhu, Yuzheng Zhang, Zeyao Ma, Bohan Zhang, Armin Schoepf, Daniel Woloch, Peter Yiliu Wang, Guangyu Robert Yang, Samuel Jacob, Siddharth Nagisetty, Abhiram Chundru, Jean Lin, Spencer Mateega, and Jing Zhang. 2026. SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows. arXiv preprint arXiv:2606.29955(2026). A More Related W...
Pith/arXiv arXiv 2026
-
[2026]
MiniMax Sparse Attention.arXiv preprint arXiv:2606.13392(2026)
Pith/arXiv arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.