Pith. sign in

REVIEW 3 major objections 6 minor 53 references

Frontier AI agents still fail most multi-hour terminal workflows; dense partial-credit grading shows the bottleneck is finishing and verifying long trajectories, not single correct steps.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 15:23 UTC pith:4ADTH7PI

load-bearing objection Solid long-horizon terminal suite with dense subtask rewards; the completion-bottleneck story is useful but still harness- and budget-tied. the 3 major comments →

arxiv 2607.08964 v2 pith:4ADTH7PI submitted 2026-07-09 cs.AI

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

classification cs.AI
keywords AI agentslong-horizon tasksterminal benchmarksdense rewardspartial creditagent evaluationself-verificationtool use
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Existing terminal agent tests mostly grade short jobs by final pass or fail, so they hide partial progress and understate how hard long real workflows are. This paper introduces Long-Horizon-Terminal-Bench: 46 containerized terminal tasks across nine domains—research reproduction, software repair, multimodal audits, games, scientific computing, and more—each broken into graded subtasks with objective checks. Agents get a dense reward for how far they advance, not only a binary end score. On shared long-horizon rollouts that average hundreds of steps, millions of tokens, and well over an hour, even the strongest tested model fully or near-fully solves only a minority of tasks, and most models solve almost none. The work argues that long-horizon planning, budgeting time, tracking progress, and self-verification—not merely local command correctness—are the binding limits, and that partial credit is required to see that gap.

Core claim

Under a shared terminal agent setup on 46 long-horizon tasks with dense subtask rewards, frontier models remain far from saturated: the strongest configuration reaches only about 28% pass rate at a 0.95 reward threshold and about 20% at perfect reward, while the mean across models is only a few percent. Typical runs consume on the order of 240 episodes, nearly 10 million tokens, and about 90 minutes, yet most failures are timeouts after incomplete progress or early exits with weak verification rather than pure local execution collapse.

What carries the argument

Subtask-based dense grading: each task is decomposed into weighted intermediate checks (binary, continuous, or episode-aggregating) whose normalized scores combine into a single reward R in [0,1], so evaluation credits how far an agent progresses through a long workflow, not only whether the final hidden verifier fully passes.

Load-bearing premise

That the authors’ chosen subtask weights, hidden stress tests, packaging, shared agent harness, and fixed time budget measure general long-horizon terminal skill rather than fit to this particular task factory and timeout.

What would settle it

Rerun the same 46 tasks under the same harness and 90-minute budget with a new frontier agent: if several models suddenly clear a large majority of tasks at R≥0.95 with few timeouts or false early finishes, the claim that long-horizon completion remains the central bottleneck for current systems would be undermined.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Binary end-to-end scores alone will keep under-ranking agents that make substantial but incomplete long-horizon progress.
  • Agent research should prioritize time budgeting, progress tracking, and calibrated stopping/verification over single-step reasoning alone.
  • Cost and episode counts become first-class evaluation axes: higher spend does not automatically raise long-horizon pass rates.
  • Near-misses (high but sub-threshold R) become early signals of capability gains before full task solves appear.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Training with process-style or subtask rewards that mirror these environment-grounded checks may transfer better than outcome-only fine-tuning for multi-hour terminal work.
  • Harness design (context management, replanning, stop rules) may move leaderboard ranks as much as base model size on this distribution.
  • A natural extension is human-time-horizon calibration: map each task’s mean human completion time against agent success at fixed budgets.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces Long-Horizon-Terminal-Bench (LHTB), a suite of 46 containerized terminal tasks spanning nine domains (experiment reproduction, software engineering, multimodal analysis, interactive games, scientific computing, etc.). Unlike Terminal-Bench-style binary outcome grading, each task is decomposed into weighted subtasks with deterministic, environment-grounded checks (binary, continuous, or episode-aggregating), yielding a dense normalized reward R. Under a shared Harbor/Terminus-2 harness (Codex for GPT-5.3) and a 90-minute timeout, the authors evaluate 17 frontier models and report that runs average ~239 episodes, ~9.8M tokens, and ~88.9 minutes, while even the strongest model (Grok 4.5) reaches only 28.3% pass@1 at R≥0.95 and 19.6% at R=1.0 (mean pass rates 6.4% and 3.2%). Analyses of reward histograms, cost–reward frontiers, timeout vs early-exit composition, and “false finishes” argue that dense grading is necessary to rank models and that long-horizon completion and self-verification—not only local step correctness—are the binding limits. The suite and harness are released.

Significance. If the suite is adopted, LHTB would fill a clear gap between short terminal/coding benchmarks and real multi-hour workflows by combining extended horizons with partial-credit, deterministic subtask grading. The dense-reward design is a concrete methodological contribution: it separates near-misses and partial progress from total failure, exposes weak stopping judgment, and makes cost–efficiency comparisons meaningful. The multi-model evaluation, cost table, unresolved-run decomposition, and public release of containerized tasks with gold solutions and hidden verifiers are substantial engineering assets for the agent community. These strengths hold even if some interpretive claims about bottlenecks need tighter caveats.

major comments (3)
  1. [§3, §3.4–3.5] §3 (Harness and agents) and §3.4–3.5: The central interpretive claim—that long-horizon completion and self-verification (not local correctness or harness fitness) are the binding limits—rests on a single shared Terminus-2 scaffold for almost all models and a fixed 90-minute budget. Because 79% of unresolved runs are timeouts with mean R only 0.10–0.35, model rankings and the “completion bottleneck” narrative are jointly determined by this harness/time envelope. The paper notes that rankings reflect time efficiency but does not quantify sensitivity (e.g., a second harness on a task subset, or a longer budget on timed-out high-R runs). Either a limited sensitivity study or a substantially softened claim with explicit confounding discussion is needed for the bottleneck conclusion to be load-bearing.
  2. [§2.3] §2.3 Difficulty calibration: Tasks were calibrated by repeatedly running DeepSeek-V4-Pro under a 1.5-hour budget until “challenging but still solvable,” then evaluated under a 90-minute cutoff. This couples task hardness to one model’s failure modes and to a budget close to the evaluation horizon. Without reporting how many candidates were discarded, how DeepSeek’s partial rewards guided redesign, or whether other models would have induced different difficulty, it is hard to separate general long-horizon competence from fitness to this particular task factory. A clearer calibration protocol and limitation statement are required.
  3. [Abstract / §3.1] Abstract vs body inconsistency: The abstract (and the arXiv-facing abstract text) reports 15 models, ~9.9M tokens, ~231 episodes, ~85.3 minutes, strongest pass@1 15.2% at R≥0.95 / 10.9% at R=1.0, and mean pass rates 4.3% / 1.7%. The body (§3.1, Table 1, Figure 3) reports 17 models, 9.8M tokens, 239 episodes, 88.9 minutes, Grok 4.5 at 28.3% / 19.6%, and means 6.4% / 3.2%. These are not minor rounding differences; they change the headline results. The abstract must be aligned with the final experimental corpus before acceptance.
minor comments (6)
  1. [§2.4] §2.4 states that the 46 tasks “span 21 high-level categories,” while the rest of the paper and Figure 2 use a nine-category taxonomy. Reconcile the wording (e.g., 21 fine-grained labels collapsed into nine paper-level categories).
  2. [Figure 1] Figure 1 caption and surrounding text refer to “237 episodes” and “86.2 minutes,” while Table 1 averages are 239 and 88.9. Align figure callouts with the final aggregate statistics.
  3. [§2.2] §2.2: The reward formula R = Σ w_k r_k / Σ w_k is clear, but default weight choices and when “higher weight on the final goal” is applied are not specified per task. A short appendix table of K and w_k (or a statement that all weights are equal unless noted) would improve reproducibility.
  4. [Appendix Table 2] Table 2 difficulty labels (Easy if mean reward ≥0.5) are derived from the same model pool used for ranking; note that these labels will drift as stronger models are added, and consider fixing them to a frozen reference set.
  5. [Throughout] Typographical issues: “Long-Horizon-T erminal-Bench” spacing artifacts in titles; “sucess” in §2.2; “T esting” / “T asks” in the title line; “abreak-ing” in Table 2 (langchain-version-migration). Clean for camera-ready.
  6. [§4] Related Work could more explicitly position LHTB against concurrent long-horizon coding suites (FrontierSWE, SWE-Marathon, LongCLI-Bench) on horizon length, grading density, and domain diversity, not only Terminal-Bench and SWE-Bench.

Circularity Check

0 steps flagged

Empirical benchmark paper with no derivation chain that forces its results by construction; scores come from independent rollouts against deterministic graders.

full rationale

Long-Horizon-Terminal-Bench is an evaluation paper, not a first-principles derivation. Its load-bearing claims are measured outcomes: pass rates, mean rewards, episode counts, token use, and failure-mode breakdowns under a shared Terminus-2 harness and fixed timeout. Task rewards are defined by deterministic, environment-grounded subtask checks (binary, continuous, or episode-aggregating) with gold solutions required to score 1.0 on hidden stress suites; public checks carry low weight. Difficulty calibration with DeepSeek-V4-Pro under a 1.5-hour budget (§2.3) is ordinary benchmark construction and does not make subsequent model scores tautological—the same model is later scored independently and does not saturate the suite. Self-citations in related work are not load-bearing uniqueness theorems or smuggled ansatze that force the leaderboard. Concerns about harness confounding and timeout sensitivity affect external validity of the “completion bottleneck” interpretation, not circularity of a derivation. No step reduces a claimed prediction to its own fitted inputs or definitions by construction.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 2 invented entities

The central empirical claim rests on operational choices that define what “long-horizon success” means: containerized Harbor tasks, weighted subtask rewards, hidden stress generation, a 90-minute budget, pass thresholds (especially τ=0.95), and a shared agent harness. These are design parameters and domain assumptions, not free fits to a physical constant. No new physical entities are postulated; the invented objects are the benchmark artifacts themselves.

free parameters (4)
  • pass threshold τ (default 0.95) = 0.95 (relaxed); also report 1.0
    Defines “resolved” for pass@1; paper also reports 0.9 and 1.0. Leaderboard ordering depends on this choice.
  • wall-clock timeout = ~90 minutes evaluation; 1.5h calibration
    Unresolved-run analysis and time-efficiency rankings depend on the fixed budget (~90 minutes; calibration used 1.5 hours).
  • subtask weights w_k = default equal weights
    Overall R is a weighted average of subtask scores; defaults equal with optional higher weight on final goals. Changes ranking of partial progress.
  • task difficulty calibration target = calibrated vs DeepSeek-V4-Pro
    Tasks adjusted by repeated DeepSeek-V4-Pro runs until “challenging but solvable,” selecting 46 of 120 candidates—an author-chosen hardness operating point.
axioms (5)
  • domain assumption Deterministic environment-grounded subtask checks (files, tests, simulator flags) are a valid proxy for meaningful intermediate progress on professional workflows.
    Core of §2.2 grading; without this, dense R does not measure capability.
  • domain assumption Hidden stress suites (schema aliases, noise, rotated frames, etc.) prevent reward hacking better than public tests alone.
    Stated in §2.3 construction recipe as the reason most reward is hidden.
  • domain assumption A shared Terminus-2 (or Codex for one model) harness allows fair comparison of base models on long-horizon terminal work.
    §3 harness setup; related work notes harness effects, but main results treat harness as fixed.
  • standard math Standard container isolation and shell interaction model (Harbor/Terminal-Bench style) is the right evaluation interface.
    Inherited from Terminal-Bench formulation cited throughout §2.
  • ad hoc to paper Difficulty labels Easy/Hard from mean reward ≥0.5 across models are useful summaries of task hardness.
    Appendix Table 2 definition; convenient but circular with the same model pool used for evaluation.
invented entities (2)
  • Long-Horizon-Terminal-Bench (LHTB) task suite no independent evidence
    purpose: Provide long-horizon terminal evaluation with partial credit across 46 tasks and nine categories.
    Primary contribution; independent evidence will be external re-use and third-party leaderboards after release.
  • Dense subtask reward R = Σ w_k r_k / Σ w_k with binary/continuous/episode-aggregating checks no independent evidence
    purpose: Replace sparse pass/fail with graded progress for ranking and failure analysis.
    Defined in §2.2; related to process rewards/rubrics but specialized to terminal state checks.

pith-pipeline@v1.1.0-grok45 · 23367 in / 3660 out tokens · 41660 ms · 2026-07-14T15:23:48.745165+00:00 · methodology

0 comments
read the original abstract

AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluated only by their final outcome. This setup overlooks intermediate progress and partial solutions, yielding sparse reward signals and an incomplete picture of agent capability. We introduce Long-Horizon-Terminal-Bench, a terminal benchmark of 46 long-horizon tasks spanning nine categories, including experiment reproduction, software engineering, multimodal analysis, interactive games, and scientific computing. Each task follows a Terminal-Bench-style setup with a reference solution or simulation engine, but is further decomposed into fine-grained graded subtasks. This design enables dense intermediate rewards and partial credit, allowing evaluation to capture not only whether an agent reaches the final goal, but also how far it progresses on open-ended workflows. Tasks in Long-Horizon-Terminal-Bench typically require hundreds of episodes and minutes to hours of execution, stressing long-horizon planning, long-context management, and iterative debugging rather than one-shot problem solving. We evaluate 15 frontier models and find that agents consume on average 9.9M tokens per task, with roughly 231 episodes and 85.3 minutes of execution time per run, making Long-Horizon-Terminal-Bench more demanding than prior terminal-based benchmarks. Even the strongest tested model achieves 15.2% pass@1 at a partial-reward threshold of 0.95 and 10.9% at a perfect-reward threshold of 1.0, while the mean pass rate across models is 4.3% and 1.7% under the two thresholds, respectively. These results reveal headroom for improvement. We further analyze failure modes and error patterns, and release Long-Horizon-Terminal-Bench to support future progress on long-horizon terminal agents.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

53 extracted references · 25 linked inside Pith

  1. [1]

    Seed2.1 officially released: Advancing ai productivity.https://seed.bytedance.com/ en/blog/seed2-1-officially-released-advancing-ai-productivity , June 2026

    ByteDance Seed Team. Seed2.1 officially released: Advancing ai productivity.https://seed.bytedance.com/ en/blog/seed2-1-officially-released-advancing-ai-productivity , June 2026. Official model release an- nouncement (Doubao Seed 2.1 / Seed 2.1 Pro)

  2. [2]

    From language to action: a review of large language models as autonomous agents and tool users.Artificial Intelligence Review, 2026

    Sadia Sultana Chowa, Riasad Alvi, Subhey Sadi Rahman, Md Abdur Rahman, Mohaimenul Azam Khan Raiaan, Md Rafiqul Islam, Mukhtar Hussain, and Sami Azam. From language to action: a review of large language models as autonomous agents and tool users.Artificial Intelligence Review, 2026. 11

  3. [3]

    Frontierswe.Proximal Blog, 2026

    Evan Chu, Rajan Agarwal, Abishek Thangamuthu, Brendan Graham, Justus Mattern, Freeman Jiang, Paul Cento, Swarnim Jain, Mersad Abbasi, Mohammad Hossein Rezaei, George Wang, Alex Zhang, Simon Guo, Karina Nguyen, Danna Liu, Arash Bidgoli, Aditya Dalmia, Apoorv Dankar, Ashrut Vaddela, Calvin Chen, Keshav Kumar, Kushagra Vaish, Navid Pour, Rishyanth Kondra, Sa...

  4. [4]

    Deepseek-v4: Towards highly efficient million-token context intelligence, 2026

    DeepSeek-AI. Deepseek-v4: Towards highly efficient million-token context intelligence, 2026. URL https: //arxiv.org/abs/2606.19348

  5. [5]

    Swe-marathon: Can agents autonomously complete ultra-long-horizon software work?arXiv preprint arXiv:2606.07682, 2026

    Rishi Desai, Jesse Hu, Joan Cabezas, Neel Harsola, Pratyush Shukla, Roey Ben Chaim, Adnan El Assadi, Omkaar Mukund Kamath, Fenil Faldu, Prannay Hebbar, et al. Swe-marathon: Can agents autonomously complete ultra-long-horizon software work?arXiv preprint arXiv:2606.07682, 2026

  6. [6]

    A survey on the optimization of large language model-based agents.ACM Computing Surveys, 58(9):1–37, 2026

    Shangheng Du, Jiabao Zhao, Jinxin Shi, Zhentao Xie, Xin Jiang, Yanhong Bai, and Liang He. A survey on the optimization of large language model-based agents.ACM Computing Surveys, 58(9):1–37, 2026

  7. [7]

    Longcli-bench: A preliminary benchmark and study for long-horizon agentic programming in command-line interfaces

    Yukang Feng, Jianwen Sun, Zelai Yang, Jiaxin Ai, Chuanhao Li, Zizhen Li, Fanrui Zhang, Kang He, Rui Ma, Jifan Lin, et al. Longcli-bench: A preliminary benchmark and study for long-horizon agentic programming in command-line interfaces. InFindings of the Association for Computational Linguistics: ACL 2026, pages 29952–29963, 2026

  8. [8]

    Glm-5: from vibe coding to agentic engineering, 2026

    GLM-5-Team, :, Aohan Zeng, Xin Lv, Zhenyu Hou, Zhengxiao Du, Qinkai Zheng, Bin Chen, Da Yin, Chendi Ge, Chenghua Huang, Chengxing Xie, Chenzheng Zhu, Congfeng Yin, Cunxiang Wang, Gengzheng Pan, Hao Zeng, Haoke Zhang, Haoran Wang, Huilong Chen, Jiajie Zhang, Jian Jiao, Jiaqi Guo, Jingsen Wang, Jingzhao Du, Jinzhu Wu, Kedong Wang, Lei Li, Lin Fan, Lucen Zho...

  9. [9]

    Gemini 3.1 pro model card

    Google DeepMind. Gemini 3.1 pro model card. https://deepmind.google/models/model-cards/ gemini-3-1-pro/, 2026

  10. [10]

    Harbor: A framework for evaluating and optimizing agents and models in container environments, 2026

    Harbor Framework Team. Harbor: A framework for evaluating and optimizing agents and models in container environments, 2026. URLhttps://doi.org/10.5281/zenodo.20953922

  11. [11]

    Livecodebench: Holistic and contamination free evaluation of large language models for code

    Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. 2024. URLhttps://arxiv.org/abs/2403.07974

  12. [12]

    Swe-bench: Can language models resolve real-world github issues? InInternational Conference on Learning Representations, volume 2024, pages 54107–54157, 2024

    Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? InInternational Conference on Learning Representations, volume 2024, pages 54107–54157, 2024

  13. [13]

    Process reward models that think.arXiv preprint arXiv:2504.16828, 2025

    Muhammad Khalifa, Rishabh Agarwal, Lajanugen Logeswaran, Jaekyeom Kim, Hao Peng, Moontae Lee, Honglak Lee, and Lu Wang. Process reward models that think.arXiv preprint arXiv:2504.16828, 2025. 12

  14. [14]

    Measuring ai ability to complete long tasks, 2025

    Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, et al. Measuring ai ability to complete long tasks, 2025. URLhttps://arxiv.org/ abs/2503.14499

  15. [15]

    Swe-evo: Benchmarking coding agents in long-horizon software evolution scenarios.arXiv preprint arXiv:2512.18470, 2025

    Tue Le, Minh VT Thai, Dung Nguyen Manh, Huy Phan Nhat, and Nghi DQ Bui. Swe-evo: Benchmarking coding agents in long-horizon software evolution scenarios.arXiv preprint arXiv:2512.18470, 2025

  16. [16]

    Self-rewarding vision-language model via reasoning decomposition.arXiv preprint arXiv:2508.19652, 2025

    Zongxia Li, Wenhao Yu, Chengsong Huang, Zhenwen Liang, Rui Liu, Fuxiao Liu, Jingxi Che, Dian Yu, Jordan Boyd-Graber, Haitao Mi, et al. Self-rewarding vision-language model via reasoning decomposition.arXiv preprint arXiv:2508.19652, 2025

  17. [17]

    Comfyclaw: Self-evolving skill harnesses for image generation workflows.arXiv preprint arXiv:2607.01709, 2026

    Zongxia Li, Dawei Liu, Fuxiao Liu, Yuhang Zhou, Xiyang Wu, Jingxi Chen, Jing Xie, Xiaomin Wu, and Lichao Sun. Comfyclaw: Self-evolving skill harnesses for image generation workflows.arXiv preprint arXiv:2607.01709, 2026

  18. [18]

    Let’s verify step by step, 2023

    Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe. Let’s verify step by step, 2023. URLhttps://arxiv.org/abs/2305. 20050

  19. [19]

    Cuarewardbench: A benchmark for evaluating reward models on computer-using agent.arXiv preprint arXiv:2510.18596, 2025

    Haojia Lin, Xiaoyu Tan, Yulei Qin, Zihan Xu, Yuchen Shi, Zongyi Li, Gang Li, Shaofei Cai, Siqi Cai, Chaoyou Fu, et al. Cuarewardbench: A benchmark for evaluating reward models on computer-using agent.arXiv preprint arXiv:2510.18596, 2025

  20. [20]

    Harness updating is not harness benefit: Disentangling evolution capabilities in self-evolving llm agents.arXiv preprint arXiv:2605.30621, 2026

    Minhua Lin, Juncheng Wu, Zijun Wang, Zhan Shi, Yisi Sang, Bing He, Zewen Liu, Tianxin Wei, Zongyu Wu, Zhiwei Zhang, et al. Harness updating is not harness benefit: Disentangling evolution capabilities in self-evolving llm agents.arXiv preprint arXiv:2605.30621, 2026

  21. [21]

    Klong: Training llm agent for extremely long-horizon tasks.arXiv preprint arXiv:2602.17547, 2026

    Yue Liu, Yingwei Ma, Yibo Miao, Yanhao Li, Yuchong Xie, Xinlong Yang, Zhiyuan Hu, Flood Sung, Jiaheng Zhang, and Bryan Hooi. Klong: Training llm agent for extremely long-horizon tasks.arXiv preprint arXiv:2602.17547, 2026

  22. [22]

    Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces.arXiv preprint arXiv:2601.11868, 2026

    Mike A Merrill, Alexander G Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Yeon Shin, Thomas Walshe, E Kelly Buchanan, et al. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces.arXiv preprint arXiv:2601.11868, 2026

  23. [23]

    Minimax sparse attention, 2026

    MiniMax. Minimax sparse attention, 2026. URLhttps://arxiv.org/abs/2606.13392

  24. [24]

    Kimi k2.6: From code to creation, from one to many

    Moonshot AI. Kimi k2.6: From code to creation, from one to many. https://www.kimi.com/ai-models/ kimi-k2-6, 2026. Official Moonshot AI model release page (Kimi K2.6)

  25. [25]

    Kimi k2.7 code: Open-source 1t agentic coding model.https://kimik2ai.com/k2.7/, June 2026

    Moonshot AI. Kimi k2.7 code: Open-source 1t agentic coding model.https://kimik2ai.com/k2.7/, June 2026. Kimi K2.7 Code, released June 12, 2026; model IDkimi-k2.7-code

  26. [26]

    Swt-bench: Testing and validating real-world bug-fixes with code agents.Advances in Neural Information Processing Systems, 37:81857–81887, 2024

    Niels Mündler, Mark N Müller, Jingxuan He, and Martin Vechev. Swt-bench: Testing and validating real-world bug-fixes with code agents.Advances in Neural Information Processing Systems, 37:81857–81887, 2024

  27. [27]

    Codex.https://github.com/openai/codex, 2025

    OpenAI. Codex.https://github.com/openai/codex, 2025. OpenAI coding agent / CLI

  28. [28]

    Gpt-5 system card

    OpenAI. Gpt-5 system card. https://openai.com/index/gpt-5-system-card/, 2025. System card; also arXiv:2601.03267

  29. [29]

    Openclaw, 2026

    OpenClaw Contributors. Openclaw, 2026. URLhttps://github.com/openclaw/openclaw. Open-source agent platform

  30. [30]

    Rubriceval: A rubric-level meta-evaluation benchmark for llm judges in instruction following

    Tianjun Pan, Xuan Lin, Wenyan Yang, Qianyu He, Shisong Chen, Licai Qi, Wanqing Xu, Hongwei Feng, Bo Xu, and Yanghua Xiao. Rubriceval: A rubric-level meta-evaluation benchmark for llm judges in instruction following. arXiv preprint arXiv:2603.25133, 2026

  31. [31]

    Qwen3.6.https://qwen.ai/blog?id=qwen3.6, 2026

    Qwen Team. Qwen3.6.https://qwen.ai/blog?id=qwen3.6, 2026. Official Qwen model release blog (Qwen3.6 Plus)

  32. [32]

    Qwen3.7.https://qwen.ai/blog?id=qwen3.7, 2026

    Qwen Team. Qwen3.7.https://qwen.ai/blog?id=qwen3.7, 2026. Official Qwen model release blog (Qwen3.7 Max)

  33. [33]

    Autorubric: Unifying rubric-based llm evaluation.arXiv preprint arXiv:2603.00077, 2026

    Delip Rao and Chris Callison-Burch. Autorubric: Unifying rubric-based llm evaluation.arXiv preprint arXiv:2603.00077, 2026. 13

  34. [34]

    Hcast: Human-calibrated autonomy software tasks.arXiv preprint arXiv:2503.17354, 2025

    David Rein, Joel Becker, Amy Deng, Seraphina Nix, Chris Canal, Daniel O’Connel, Pip Arnott, Ryan Bloom, Thomas Broadley, Katharyn Garcia, et al. Hcast: Human-calibrated autonomy software tasks.arXiv preprint arXiv:2503.17354, 2025

  35. [35]

    The illusion of diminishing returns: Measuring long horizon execution in llms.arXiv preprint arXiv:2509.09677, 2025

    Akshit Sinha, Arvindh Arun, Shashwat Goel, Steffen Staab, and Jonas Geiping. The illusion of diminishing returns: Measuring long horizon execution in llms.arXiv preprint arXiv:2509.09677, 2025

  36. [36]

    Dingjie Song, Hanrong Zhang, Dawei Liu, Yixin Liu, Zongxia Li, Zhengqing Yuan, Siqi Zhang, and Lichao Sun. Dr. claw: An ai research workspace from idea to paper, 2026. URLhttps://github.com/OpenLAIR/dr-claw

  37. [37]

    Tencent hunyuan 3.https://hunyuan.tencent.com/, 2026

    Tencent Hunyuan. Tencent hunyuan 3.https://hunyuan.tencent.com/, 2026. Official Tencent Hunyuan model site

  38. [38]

    Arco: Adaptive rubric with co-evolution for multi-step llm-based agents, 2026

    Zihang Tian et al. Arco: Adaptive rubric with co-evolution for multi-step llm-based agents, 2026. URL https://arxiv.org/abs/2606.21262

  39. [39]

    Solving math word problems with process-and outcome-based feedback.arXiv preprint arXiv:2211.14275, 2022

    Jonathan Uesato, Nate Kushman, Ramana Kumar, Francis Song, Noah Siegel, Lisa Wang, Antonia Creswell, Geoffrey Irving, and Irina Higgins. Solving math word problems with process-and outcome-based feedback.arXiv preprint arXiv:2211.14275, 2022

  40. [40]

    Apex-agents.arXiv preprint arXiv:2601.14242, 2026

    Bertie Vidgen, Austin Mann, Abby Fennelly, John Wright Stanly, Lucas Rothman, Marco Burstein, Julien Benchek, David Ostrofsky, Anirudh Ravichandran, Debnil Sur, et al. Apex-agents.arXiv preprint arXiv:2601.14242, 2026

  41. [41]

    Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al

    Xingyao Wang, Boxuan Li, Yufan Song, Frank F. Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al. Openhands: An open platform for ai software developers as generalist agents,

  42. [42]

    URLhttps://arxiv.org/abs/2407.16741

  43. [43]

    The long-horizon task mirage? diagnosing where and why agentic systems break.arXiv preprint arXiv:2604.11978, 2026

    Xinyu Jessica Wang, Haoyue Bai, Yiyou Sun, Haorui Wang, Shuibai Zhang, Wenjie Hu, Mya Schroder, Bilge Mutlu, Dawn Song, and Robert D Nowak. The long-horizon task mirage? diagnosing where and why agentic systems break.arXiv preprint arXiv:2604.11978, 2026

  44. [44]

    Re-bench: Evaluating frontier ai r&d capabilities of language model agents against human experts.arXiv preprint arXiv:2411.15114, 2024

    Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, Neev Parikh, Thomas Broadley, Lawrence Chan, Michael Chen, Josh Clymer, Jai Dhyani, et al. Re-bench: Evaluating frontier ai r&d capabilities of language model agents against human experts.arXiv preprint arXiv:2411.15114, 2024

  45. [45]

    Co-evolving llm decision and skill bank agents for long-horizon tasks.arXiv preprint arXiv:2604.20987, 2026

    Xiyang Wu, Zongxia Li, Guangyao Shi, Alexander Duffy, Tyler Marques, Matthew Lyle Olson, Tianyi Zhou, and Dinesh Manocha. Co-evolving llm decision and skill bank agents for long-horizon tasks.arXiv preprint arXiv:2604.20987, 2026

  46. [46]

    Grok.https://x.ai/, 2026

    xAI. Grok.https://x.ai/, 2026. Official xAI site for the Grok model family

  47. [47]

    Grok 4.5.https://x.ai/, 2026

    xAI. Grok 4.5.https://x.ai/, 2026. Official xAI site for Grok 4.5

  48. [48]

    Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024

    Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024. URLhttps://arxiv.org/abs/2404.07972

  49. [49]

    Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press

    John Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. Swe-agent: Agent-computer interfaces enable automated software engineering, 2024. URLhttps://arxiv.org/ abs/2405.15793

  50. [50]

    Harness-bench: Measuring harness effects across models in realistic agent workflows.arXiv preprint arXiv:2605.27922, 2026

    Yilun Yao, Xinyu Tan, Chao-Hsuan Liu, Yaoming Li, Zhengyang Wang, Wenhan Yu, Zhewen Tan, Yuxuan Tian, Guangxiang Zhao, Lin Sun, et al. Harness-bench: Measuring harness effects across models in realistic agent workflows.arXiv preprint arXiv:2605.27922, 2026

  51. [51]

    Self-rewarding language models.arXiv preprint arXiv:2401.10020, 2024

    Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. Self-rewarding language models.arXiv preprint arXiv:2401.10020, 2024

  52. [52]

    Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623, 2023

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623, 2023

  53. [53]

    Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

    Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. Webarena: A realistic web environment for building autonomous agents, 2023. URLhttps://arxiv.org/abs/2307.13854. 14 A List of T asks in Long-Horizon-T erminal-Bench Table 2 lists all 46 benchmark tas...