REVIEW 3 major objections 6 minor 53 references
Frontier AI agents still fail most multi-hour terminal workflows; dense partial-credit grading shows the bottleneck is finishing and verifying long trajectories, not single correct steps.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 15:23 UTC pith:4ADTH7PI
load-bearing objection Solid long-horizon terminal suite with dense subtask rewards; the completion-bottleneck story is useful but still harness- and budget-tied. the 3 major comments →
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Under a shared terminal agent setup on 46 long-horizon tasks with dense subtask rewards, frontier models remain far from saturated: the strongest configuration reaches only about 28% pass rate at a 0.95 reward threshold and about 20% at perfect reward, while the mean across models is only a few percent. Typical runs consume on the order of 240 episodes, nearly 10 million tokens, and about 90 minutes, yet most failures are timeouts after incomplete progress or early exits with weak verification rather than pure local execution collapse.
What carries the argument
Subtask-based dense grading: each task is decomposed into weighted intermediate checks (binary, continuous, or episode-aggregating) whose normalized scores combine into a single reward R in [0,1], so evaluation credits how far an agent progresses through a long workflow, not only whether the final hidden verifier fully passes.
Load-bearing premise
That the authors’ chosen subtask weights, hidden stress tests, packaging, shared agent harness, and fixed time budget measure general long-horizon terminal skill rather than fit to this particular task factory and timeout.
What would settle it
Rerun the same 46 tasks under the same harness and 90-minute budget with a new frontier agent: if several models suddenly clear a large majority of tasks at R≥0.95 with few timeouts or false early finishes, the claim that long-horizon completion remains the central bottleneck for current systems would be undermined.
If this is right
- Binary end-to-end scores alone will keep under-ranking agents that make substantial but incomplete long-horizon progress.
- Agent research should prioritize time budgeting, progress tracking, and calibrated stopping/verification over single-step reasoning alone.
- Cost and episode counts become first-class evaluation axes: higher spend does not automatically raise long-horizon pass rates.
- Near-misses (high but sub-threshold R) become early signals of capability gains before full task solves appear.
Where Pith is reading between the lines
- Training with process-style or subtask rewards that mirror these environment-grounded checks may transfer better than outcome-only fine-tuning for multi-hour terminal work.
- Harness design (context management, replanning, stop rules) may move leaderboard ranks as much as base model size on this distribution.
- A natural extension is human-time-horizon calibration: map each task’s mean human completion time against agent success at fixed budgets.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Long-Horizon-Terminal-Bench (LHTB), a suite of 46 containerized terminal tasks spanning nine domains (experiment reproduction, software engineering, multimodal analysis, interactive games, scientific computing, etc.). Unlike Terminal-Bench-style binary outcome grading, each task is decomposed into weighted subtasks with deterministic, environment-grounded checks (binary, continuous, or episode-aggregating), yielding a dense normalized reward R. Under a shared Harbor/Terminus-2 harness (Codex for GPT-5.3) and a 90-minute timeout, the authors evaluate 17 frontier models and report that runs average ~239 episodes, ~9.8M tokens, and ~88.9 minutes, while even the strongest model (Grok 4.5) reaches only 28.3% pass@1 at R≥0.95 and 19.6% at R=1.0 (mean pass rates 6.4% and 3.2%). Analyses of reward histograms, cost–reward frontiers, timeout vs early-exit composition, and “false finishes” argue that dense grading is necessary to rank models and that long-horizon completion and self-verification—not only local step correctness—are the binding limits. The suite and harness are released.
Significance. If the suite is adopted, LHTB would fill a clear gap between short terminal/coding benchmarks and real multi-hour workflows by combining extended horizons with partial-credit, deterministic subtask grading. The dense-reward design is a concrete methodological contribution: it separates near-misses and partial progress from total failure, exposes weak stopping judgment, and makes cost–efficiency comparisons meaningful. The multi-model evaluation, cost table, unresolved-run decomposition, and public release of containerized tasks with gold solutions and hidden verifiers are substantial engineering assets for the agent community. These strengths hold even if some interpretive claims about bottlenecks need tighter caveats.
major comments (3)
- [§3, §3.4–3.5] §3 (Harness and agents) and §3.4–3.5: The central interpretive claim—that long-horizon completion and self-verification (not local correctness or harness fitness) are the binding limits—rests on a single shared Terminus-2 scaffold for almost all models and a fixed 90-minute budget. Because 79% of unresolved runs are timeouts with mean R only 0.10–0.35, model rankings and the “completion bottleneck” narrative are jointly determined by this harness/time envelope. The paper notes that rankings reflect time efficiency but does not quantify sensitivity (e.g., a second harness on a task subset, or a longer budget on timed-out high-R runs). Either a limited sensitivity study or a substantially softened claim with explicit confounding discussion is needed for the bottleneck conclusion to be load-bearing.
- [§2.3] §2.3 Difficulty calibration: Tasks were calibrated by repeatedly running DeepSeek-V4-Pro under a 1.5-hour budget until “challenging but still solvable,” then evaluated under a 90-minute cutoff. This couples task hardness to one model’s failure modes and to a budget close to the evaluation horizon. Without reporting how many candidates were discarded, how DeepSeek’s partial rewards guided redesign, or whether other models would have induced different difficulty, it is hard to separate general long-horizon competence from fitness to this particular task factory. A clearer calibration protocol and limitation statement are required.
- [Abstract / §3.1] Abstract vs body inconsistency: The abstract (and the arXiv-facing abstract text) reports 15 models, ~9.9M tokens, ~231 episodes, ~85.3 minutes, strongest pass@1 15.2% at R≥0.95 / 10.9% at R=1.0, and mean pass rates 4.3% / 1.7%. The body (§3.1, Table 1, Figure 3) reports 17 models, 9.8M tokens, 239 episodes, 88.9 minutes, Grok 4.5 at 28.3% / 19.6%, and means 6.4% / 3.2%. These are not minor rounding differences; they change the headline results. The abstract must be aligned with the final experimental corpus before acceptance.
minor comments (6)
- [§2.4] §2.4 states that the 46 tasks “span 21 high-level categories,” while the rest of the paper and Figure 2 use a nine-category taxonomy. Reconcile the wording (e.g., 21 fine-grained labels collapsed into nine paper-level categories).
- [Figure 1] Figure 1 caption and surrounding text refer to “237 episodes” and “86.2 minutes,” while Table 1 averages are 239 and 88.9. Align figure callouts with the final aggregate statistics.
- [§2.2] §2.2: The reward formula R = Σ w_k r_k / Σ w_k is clear, but default weight choices and when “higher weight on the final goal” is applied are not specified per task. A short appendix table of K and w_k (or a statement that all weights are equal unless noted) would improve reproducibility.
- [Appendix Table 2] Table 2 difficulty labels (Easy if mean reward ≥0.5) are derived from the same model pool used for ranking; note that these labels will drift as stronger models are added, and consider fixing them to a frozen reference set.
- [Throughout] Typographical issues: “Long-Horizon-T erminal-Bench” spacing artifacts in titles; “sucess” in §2.2; “T esting” / “T asks” in the title line; “abreak-ing” in Table 2 (langchain-version-migration). Clean for camera-ready.
- [§4] Related Work could more explicitly position LHTB against concurrent long-horizon coding suites (FrontierSWE, SWE-Marathon, LongCLI-Bench) on horizon length, grading density, and domain diversity, not only Terminal-Bench and SWE-Bench.
Circularity Check
Empirical benchmark paper with no derivation chain that forces its results by construction; scores come from independent rollouts against deterministic graders.
full rationale
Long-Horizon-Terminal-Bench is an evaluation paper, not a first-principles derivation. Its load-bearing claims are measured outcomes: pass rates, mean rewards, episode counts, token use, and failure-mode breakdowns under a shared Terminus-2 harness and fixed timeout. Task rewards are defined by deterministic, environment-grounded subtask checks (binary, continuous, or episode-aggregating) with gold solutions required to score 1.0 on hidden stress suites; public checks carry low weight. Difficulty calibration with DeepSeek-V4-Pro under a 1.5-hour budget (§2.3) is ordinary benchmark construction and does not make subsequent model scores tautological—the same model is later scored independently and does not saturate the suite. Self-citations in related work are not load-bearing uniqueness theorems or smuggled ansatze that force the leaderboard. Concerns about harness confounding and timeout sensitivity affect external validity of the “completion bottleneck” interpretation, not circularity of a derivation. No step reduces a claimed prediction to its own fitted inputs or definitions by construction.
Axiom & Free-Parameter Ledger
free parameters (4)
- pass threshold τ (default 0.95) =
0.95 (relaxed); also report 1.0
- wall-clock timeout =
~90 minutes evaluation; 1.5h calibration
- subtask weights w_k =
default equal weights
- task difficulty calibration target =
calibrated vs DeepSeek-V4-Pro
axioms (5)
- domain assumption Deterministic environment-grounded subtask checks (files, tests, simulator flags) are a valid proxy for meaningful intermediate progress on professional workflows.
- domain assumption Hidden stress suites (schema aliases, noise, rotated frames, etc.) prevent reward hacking better than public tests alone.
- domain assumption A shared Terminus-2 (or Codex for one model) harness allows fair comparison of base models on long-horizon terminal work.
- standard math Standard container isolation and shell interaction model (Harbor/Terminal-Bench style) is the right evaluation interface.
- ad hoc to paper Difficulty labels Easy/Hard from mean reward ≥0.5 across models are useful summaries of task hardness.
invented entities (2)
-
Long-Horizon-Terminal-Bench (LHTB) task suite
no independent evidence
-
Dense subtask reward R = Σ w_k r_k / Σ w_k with binary/continuous/episode-aggregating checks
no independent evidence
read the original abstract
AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on open-ended workflows. Tasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at a perfect-reward threshold of 1.0, while the mean pass rate across models is 4.3% and 1.7% under the two thresholds, respectively. These results reveal headroom for improvement. We further analyze failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on long-horizon terminal agents.
Reference graph
Works this paper leans on
-
[1]
Seed2.1 officially released: Advancing ai productivity.https://seed.bytedance.com/ en/blog/seed2-1-officially-released-advancing-ai-productivity , June 2026
ByteDance Seed Team. Seed2.1 officially released: Advancing ai productivity.https://seed.bytedance.com/ en/blog/seed2-1-officially-released-advancing-ai-productivity , June 2026. Official model release an- nouncement (Doubao Seed 2.1 / Seed 2.1 Pro)
2026
-
[2]
From language to action: a review of large language models as autonomous agents and tool users.Artificial Intelligence Review, 2026
Sadia Sultana Chowa, Riasad Alvi, Subhey Sadi Rahman, Md Abdur Rahman, Mohaimenul Azam Khan Raiaan, Md Rafiqul Islam, Mukhtar Hussain, and Sami Azam. From language to action: a review of large language models as autonomous agents and tool users.Artificial Intelligence Review, 2026. 11
2026
-
[3]
Frontierswe.Proximal Blog, 2026
Evan Chu, Rajan Agarwal, Abishek Thangamuthu, Brendan Graham, Justus Mattern, Freeman Jiang, Paul Cento, Swarnim Jain, Mersad Abbasi, Mohammad Hossein Rezaei, George Wang, Alex Zhang, Simon Guo, Karina Nguyen, Danna Liu, Arash Bidgoli, Aditya Dalmia, Apoorv Dankar, Ashrut Vaddela, Calvin Chen, Keshav Kumar, Kushagra Vaish, Navid Pour, Rishyanth Kondra, Sa...
2026
-
[4]
Deepseek-v4: Towards highly efficient million-token context intelligence, 2026
DeepSeek-AI. Deepseek-v4: Towards highly efficient million-token context intelligence, 2026. URL https: //arxiv.org/abs/2606.19348
arXiv 2026
-
[5]
Rishi Desai, Jesse Hu, Joan Cabezas, Neel Harsola, Pratyush Shukla, Roey Ben Chaim, Adnan El Assadi, Omkaar Mukund Kamath, Fenil Faldu, Prannay Hebbar, et al. Swe-marathon: Can agents autonomously complete ultra-long-horizon software work?arXiv preprint arXiv:2606.07682, 2026
Pith/arXiv arXiv 2026
-
[6]
A survey on the optimization of large language model-based agents.ACM Computing Surveys, 58(9):1–37, 2026
Shangheng Du, Jiabao Zhao, Jinxin Shi, Zhentao Xie, Xin Jiang, Yanhong Bai, and Liang He. A survey on the optimization of large language model-based agents.ACM Computing Surveys, 58(9):1–37, 2026
2026
-
[7]
Longcli-bench: A preliminary benchmark and study for long-horizon agentic programming in command-line interfaces
Yukang Feng, Jianwen Sun, Zelai Yang, Jiaxin Ai, Chuanhao Li, Zizhen Li, Fanrui Zhang, Kang He, Rui Ma, Jifan Lin, et al. Longcli-bench: A preliminary benchmark and study for long-horizon agentic programming in command-line interfaces. InFindings of the Association for Computational Linguistics: ACL 2026, pages 29952–29963, 2026
2026
-
[8]
Glm-5: from vibe coding to agentic engineering, 2026
GLM-5-Team, :, Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, Chenzheng Zhu, Congfeng Yin, Cunxiang Wang, Gengzheng Pan, Hao Zeng, Haoke Zhang, Haoran Wang, Huilong Chen, Jiajie Zhang, Jian Jiao, Jiaqi Guo, Jingsen Wang, Jingzhao Du, Jinzhu Wu, Kedong Wang, Lei Li, Lin Fan, Lucen Zho...
Pith/arXiv arXiv 2026
-
[9]
Gemini 3.1 pro model card
Google DeepMind. Gemini 3.1 pro model card. https://deepmind.google/models/model-cards/ gemini-3-1-pro/, 2026
2026
-
[10]
Harbor: A framework for evaluating and optimizing agents and models in container environments, 2026
Harbor Framework Team. Harbor: A framework for evaluating and optimizing agents and models in container environments, 2026. URLhttps://doi.org/10.5281/zenodo.20953922
-
[11]
Livecodebench: Holistic and contamination free evaluation of large language models for code
Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. 2024. URLhttps://arxiv.org/abs/2403.07974
Pith/arXiv arXiv 2024
-
[12]
Swe-bench: Can language models resolve real-world github issues? InInternational Conference on Learning Representations, volume 2024, pages 54107–54157, 2024
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? InInternational Conference on Learning Representations, volume 2024, pages 54107–54157, 2024
2024
-
[13]
Process reward models that think.arXiv preprint arXiv:2504.16828, 2025
Muhammad Khalifa, Rishabh Agarwal, Lajanugen Logeswaran, Jaekyeom Kim, Hao Peng, Moontae Lee, Honglak Lee, and Lu Wang. Process reward models that think.arXiv preprint arXiv:2504.16828, 2025. 12
arXiv 2025
-
[14]
Measuring ai ability to complete long tasks, 2025
Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, et al. Measuring ai ability to complete long tasks, 2025. URLhttps://arxiv.org/ abs/2503.14499
Pith/arXiv arXiv 2025
-
[15]
Tue Le, Minh VT Thai, Dung Nguyen Manh, Huy Phan Nhat, and Nghi DQ Bui. Swe-evo: Benchmarking coding agents in long-horizon software evolution scenarios.arXiv preprint arXiv:2512.18470, 2025
Pith/arXiv arXiv 2025
-
[16]
Zongxia Li, Wenhao Yu, Chengsong Huang, Zhenwen Liang, Rui Liu, Fuxiao Liu, Jingxi Che, Dian Yu, Jordan Boyd-Graber, Haitao Mi, et al. Self-rewarding vision-language model via reasoning decomposition.arXiv preprint arXiv:2508.19652, 2025
Pith/arXiv arXiv 2025
-
[17]
Zongxia Li, Dawei Liu, Fuxiao Liu, Yuhang Zhou, Xiyang Wu, Jingxi Chen, Jing Xie, Xiaomin Wu, and Lichao Sun. Comfyclaw: Self-evolving skill harnesses for image generation workflows.arXiv preprint arXiv:2607.01709, 2026
Pith/arXiv arXiv 2026
-
[18]
Let’s verify step by step, 2023
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step, 2023. URLhttps://arxiv.org/abs/2305. 20050
2023
-
[19]
Haojia Lin, Xiaoyu Tan, Yulei Qin, Zihan Xu, Yuchen Shi, Zongyi Li, Gang Li, Shaofei Cai, Siqi Cai, Chaoyou Fu, et al. Cuarewardbench: A benchmark for evaluating reward models on computer-using agent.arXiv preprint arXiv:2510.18596, 2025
arXiv 2025
-
[20]
Minhua Lin, Juncheng Wu, Zijun Wang, Zhan Shi, Yisi Sang, Bing He, Zewen Liu, Tianxin Wei, Zongyu Wu, Zhiwei Zhang, et al. Harness updating is not harness benefit: Disentangling evolution capabilities in self-evolving llm agents.arXiv preprint arXiv:2605.30621, 2026
Pith/arXiv arXiv 2026
-
[21]
Klong: Training llm agent for extremely long-horizon tasks.arXiv preprint arXiv:2602.17547, 2026
Yue Liu, Yingwei Ma, Yibo Miao, Yanhao Li, Yuchong Xie, Xinlong Yang, Zhiyuan Hu, Flood Sung, Jiaheng Zhang, and Bryan Hooi. Klong: Training llm agent for extremely long-horizon tasks.arXiv preprint arXiv:2602.17547, 2026
Pith/arXiv arXiv 2026
-
[22]
Mike A Merrill, Alexander G Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Yeon Shin, Thomas Walshe, E Kelly Buchanan, et al. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces.arXiv preprint arXiv:2601.11868, 2026
Pith/arXiv arXiv 2026
-
[23]
Minimax sparse attention, 2026
MiniMax. Minimax sparse attention, 2026. URLhttps://arxiv.org/abs/2606.13392
Pith/arXiv arXiv 2026
-
[24]
Kimi k2.6: From code to creation, from one to many
Moonshot AI. Kimi k2.6: From code to creation, from one to many. https://www.kimi.com/ai-models/ kimi-k2-6, 2026. Official Moonshot AI model release page (Kimi K2.6)
2026
-
[25]
Kimi k2.7 code: Open-source 1t agentic coding model.https://kimik2ai.com/k2.7/, June 2026
Moonshot AI. Kimi k2.7 code: Open-source 1t agentic coding model.https://kimik2ai.com/k2.7/, June 2026. Kimi K2.7 Code, released June 12, 2026; model IDkimi-k2.7-code
2026
-
[26]
Swt-bench: Testing and validating real-world bug-fixes with code agents.Advances in Neural Information Processing Systems, 37:81857–81887, 2024
Niels Mündler, Mark N Müller, Jingxuan He, and Martin Vechev. Swt-bench: Testing and validating real-world bug-fixes with code agents.Advances in Neural Information Processing Systems, 37:81857–81887, 2024
2024
-
[27]
Codex.https://github.com/openai/codex, 2025
OpenAI. Codex.https://github.com/openai/codex, 2025. OpenAI coding agent / CLI
2025
-
[28]
OpenAI. Gpt-5 system card. https://openai.com/index/gpt-5-system-card/, 2025. System card; also arXiv:2601.03267
Pith/arXiv arXiv 2025
-
[29]
Openclaw, 2026
OpenClaw Contributors. Openclaw, 2026. URLhttps://github.com/openclaw/openclaw. Open-source agent platform
2026
-
[30]
Rubriceval: A rubric-level meta-evaluation benchmark for llm judges in instruction following
Tianjun Pan, Xuan Lin, Wenyan Yang, Qianyu He, Shisong Chen, Licai Qi, Wanqing Xu, Hongwei Feng, Bo Xu, and Yanghua Xiao. Rubriceval: A rubric-level meta-evaluation benchmark for llm judges in instruction following. arXiv preprint arXiv:2603.25133, 2026
arXiv 2026
-
[31]
Qwen3.6.https://qwen.ai/blog?id=qwen3.6, 2026
Qwen Team. Qwen3.6.https://qwen.ai/blog?id=qwen3.6, 2026. Official Qwen model release blog (Qwen3.6 Plus)
2026
-
[32]
Qwen3.7.https://qwen.ai/blog?id=qwen3.7, 2026
Qwen Team. Qwen3.7.https://qwen.ai/blog?id=qwen3.7, 2026. Official Qwen model release blog (Qwen3.7 Max)
2026
-
[33]
Autorubric: Unifying rubric-based llm evaluation.arXiv preprint arXiv:2603.00077, 2026
Delip Rao and Chris Callison-Burch. Autorubric: Unifying rubric-based llm evaluation.arXiv preprint arXiv:2603.00077, 2026. 13
Pith/arXiv arXiv 2026
-
[34]
Hcast: Human-calibrated autonomy software tasks.arXiv preprint arXiv:2503.17354, 2025
David Rein, Joel Becker, Amy Deng, Seraphina Nix, Chris Canal, Daniel O’Connel, Pip Arnott, Ryan Bloom, Thomas Broadley, Katharyn Garcia, et al. Hcast: Human-calibrated autonomy software tasks.arXiv preprint arXiv:2503.17354, 2025
Pith/arXiv arXiv 2025
-
[35]
Akshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab, and Jonas Geiping. The illusion of diminishing returns: Measuring long horizon execution in llms.arXiv preprint arXiv:2509.09677, 2025
arXiv 2025
-
[36]
Dingjie Song, Hanrong Zhang, Dawei Liu, Yixin Liu, Zongxia Li, Zhengqing Yuan, Siqi Zhang, and Lichao Sun. Dr. claw: An ai research workspace from idea to paper, 2026. URLhttps://github.com/OpenLAIR/dr-claw
2026
-
[37]
Tencent hunyuan 3.https://hunyuan.tencent.com/, 2026
Tencent Hunyuan. Tencent hunyuan 3.https://hunyuan.tencent.com/, 2026. Official Tencent Hunyuan model site
2026
-
[38]
Arco: Adaptive rubric with co-evolution for multi-step llm-based agents, 2026
Zihang Tian et al. Arco: Adaptive rubric with co-evolution for multi-step llm-based agents, 2026. URL https://arxiv.org/abs/2606.21262
Pith/arXiv arXiv 2026
-
[39]
Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. Solving math word problems with process-and outcome-based feedback.arXiv preprint arXiv:2211.14275, 2022
Pith/arXiv arXiv 2022
-
[40]
Apex-agents.arXiv preprint arXiv:2601.14242, 2026
Bertie Vidgen, Austin Mann, Abby Fennelly, John Wright Stanly, Lucas Rothman, Marco Burstein, Julien Benchek, David Ostrofsky, Anirudh Ravichandran, Debnil Sur, et al. Apex-agents.arXiv preprint arXiv:2601.14242, 2026
arXiv 2026
-
[41]
Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al
Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al. Openhands: An open platform for ai software developers as generalist agents,
-
[42]
URLhttps://arxiv.org/abs/2407.16741
-
[43]
Xinyu Jessica Wang, Haoyue Bai, Yiyou Sun, Haorui Wang, Shuibai Zhang, Wenjie Hu, Mya Schroder, Bilge Mutlu, Dawn Song, and Robert D Nowak. The long-horizon task mirage? diagnosing where and why agentic systems break.arXiv preprint arXiv:2604.11978, 2026
Pith/arXiv arXiv 2026
-
[44]
Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, Neev Parikh, Thomas Broadley, Lawrence Chan, Michael Chen, Josh Clymer, Jai Dhyani, et al. Re-bench: Evaluating frontier ai r&d capabilities of language model agents against human experts.arXiv preprint arXiv:2411.15114, 2024
Pith/arXiv arXiv 2024
-
[45]
Xiyang Wu, Zongxia Li, Guangyao Shi, Alexander Duffy, Tyler Marques, Matthew Lyle Olson, Tianyi Zhou, and Dinesh Manocha. Co-evolving llm decision and skill bank agents for long-horizon tasks.arXiv preprint arXiv:2604.20987, 2026
Pith/arXiv arXiv 2026
-
[46]
Grok.https://x.ai/, 2026
xAI. Grok.https://x.ai/, 2026. Official xAI site for the Grok model family
2026
-
[47]
Grok 4.5.https://x.ai/, 2026
xAI. Grok 4.5.https://x.ai/, 2026. Official xAI site for Grok 4.5
2026
-
[48]
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024
Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024. URLhttps://arxiv.org/abs/2404.07972
Pith/arXiv arXiv 2024
-
[49]
Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press
John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering, 2024. URLhttps://arxiv.org/ abs/2405.15793
Pith/arXiv arXiv 2024
-
[50]
Yilun Yao, Xinyu Tan, Chao-Hsuan Liu, Yaoming Li, Zhengyang Wang, Wenhan Yu, Zhewen Tan, Yuxuan Tian, Guangxiang Zhao, Lin Sun, et al. Harness-bench: Measuring harness effects across models in realistic agent workflows.arXiv preprint arXiv:2605.27922, 2026
Pith/arXiv arXiv 2026
-
[51]
Self-rewarding language models.arXiv preprint arXiv:2401.10020, 2024
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. Self-rewarding language models.arXiv preprint arXiv:2401.10020, 2024
Pith/arXiv arXiv 2024
-
[52]
Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623, 2023
2023
-
[53]
Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. Webarena: A realistic web environment for building autonomous agents, 2023. URLhttps://arxiv.org/abs/2307.13854. 14 A List of T asks in Long-Horizon-T erminal-Bench Table 2 lists all 46 benchmark tas...
Pith/arXiv arXiv 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.