Pith. sign in

Paper Citation Record · LEDGER

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

As of 24 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 4 inbound Pith citation observations for arXiv:2607.08964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08964 v2

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T15:24:58.589243Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:31:42.307916Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T04:45:37.640056Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0569f4d8-6645-4ae2-b5ab-4736017cec53 · outbound

This paper cites Seed2.1 officially released: Advancing ai productivity.https://seed.bytedance.com/ en/blog/seed2-1-officially-released-advancing-ai-productivity , June 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Seed2.1 officially released: Advancing ai productivity.https://seed.bytedance.com/ en/blog/seed2-1-officially-released-advancing-ai-productivity , June 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:9294bb71ae910b063c245dddf72204b2b6d9d3048a6099601214b16d9e53e7ff

Observation b5c023ec-0e69-4e8d-86bf-e88e100d7df4 · outbound

This paper cites From language to action: a review of large language models as autonomous agents and tool users.Artificial Intelligence Review, 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading From language to action: a review of large language models as autonomous agents and tool users.Artificial Intelligence Review, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:7530ca814cf19720b1439f6215d698d4f84774b37e984dfccdcb0e4c32c85261

Observation 31194ba0-cc09-41eb-9c89-417fd17c9474 · outbound

This paper cites Frontierswe.Proximal Blog, 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Frontierswe.Proximal Blog, 2026

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:31867463e7a0747a5ed09a73af835dd8ee3da971dcdf7cffb7d687ba704739a9

Observation d31106d5-346b-42ea-ac7e-719542066fa7 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence, 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Deepseek-v4: Towards highly efficient million-token context intelligence, 2026

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:ebf9a2b6d8b228eaa623e778c4375e30482efd370c7a09e2fc4ffdc0210ea84d

Observation 927fd05e-9172-4e17-b051-cc1c805660db · outbound

This paper cites SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:af1113637615830f151e126c466a745ca6b836854e77b5ce538f292a2c199a2d

Observation 4c847665-069e-4e53-851d-0784c654d56b · outbound

This paper cites A survey on the optimization of large language model-based agents.ACM Computing Surveys, 58(9):1–37, 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading A survey on the optimization of large language model-based agents.ACM Computing Surveys, 58(9):1–37, 2026

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:ec1dda6ec1e417eadb712f8c35bbb43dd0e81137175b110270014b54dd3a4996

Observation cb65927c-0edc-4c97-a62f-1606e7b108af · outbound

This paper cites Longcli-bench: A preliminary benchmark and study for long-horizon agentic programming in command-line interfaces.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Longcli-bench: A preliminary benchmark and study for long-horizon agentic programming in command-line interfaces

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:584c5a86d0d41a1533944d82e2c2f57600f33573547b19b13fed236a5d549055

Observation 7b25a2df-14a7-466e-be8a-ef426b612a5c · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading GLM-5: from Vibe Coding to Agentic Engineering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:8c5303e0a90e7b261b0c27d77a0dcf125f3ba5a290f4990877c10f6d01010c04

Observation e0971fd6-8ab8-4f0a-a482-038306a8bbba · outbound

This paper cites Gemini 3.1 pro model card.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Gemini 3.1 pro model card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:5a7185535702189d7b3fecff64990443eb6c9c81b6d5d4e8f7e731d7893821fe

Observation 180398ef-19c8-4ff7-9422-72e04cf5c973 · outbound

This paper cites Harbor: A framework for evaluating and optimizing agents and models in container environments, 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Harbor: A framework for evaluating and optimizing agents and models in container environments, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:670b8f33c795bbe6ea1fafbe7ccbc12fbc42446d4b9479f26ad9421441af1913

Observation 30c5921e-b5b4-47d7-a874-be8d0e02f152 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:74a1e536d545bebad14d8a42801896a02e4c34b041d71f6fa2bd3490b43138de

Observation 3872e800-85ec-4435-b412-7e69c6d8c6c0 · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? InInternational Conference on Learning Representations, volume 2024, pages 54107–54157, 2024.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Swe-bench: Can language models resolve real-world github issues? InInternational Conference on Learning Representations, volume 2024, pages 54107–54157, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:bb284838c80e0e71313328a2ade22dd43e412866538a1496e054a343309e8578

Observation 25b2214d-ef77-4038-8578-8bfebd0d98ce · outbound

This paper cites Process reward models that think.arXiv preprint arXiv:2504.16828, 2025.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Process reward models that think.arXiv preprint arXiv:2504.16828, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:5dbabbb8c2f02adefc86a4559833a360a85dc9d3fa17a295cd5ee05570f79bfa

Observation 2001b714-9771-4e11-8bad-590d8ac77e5b · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Measuring AI Ability to Complete Long Software Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:460102305fcb6edb92042d52440dbf91354ba473f83d7bc14d160338461da1fd

Observation 2548c8b6-fd4c-4d28-9ac2-86a0e99e9266 · outbound

This paper cites SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:83f079fbbc664b9bafba4aa7869e8bfbf9b70465c52baf9b6f8aa38bdf4e57cf

Observation 3a164a77-f953-4015-a040-5be20571ce00 · outbound

This paper cites Self-Rewarding Vision-Language Model via Reasoning Decomposition.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:d8280c28de3fad3ccc88d154dccb3d8fe491f76f7f227ed8354e0bd0a28f7a20

Observation 95ec34b4-08e8-4285-8020-0ab013f33d5f · outbound

This paper cites COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:3ceea0430e569dfa961c89591d999cf54939620fc18e42391391fca5464a0474

Observation 7dbc04af-f42b-4da0-9a56-d3636ac96432 · outbound

This paper cites Let’s verify step by step, 2023.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Let’s verify step by step, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:2c8313a9db392adf685a116dae90114ee912743415a99472d292f2d46d287787

Observation 900bd040-4a4e-4c96-a02d-b377ac0e5c77 · outbound

This paper cites Cuarewardbench: A benchmark for evaluating reward models on computer-using agent.arXiv preprint arXiv:2510.18596, 2025.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Cuarewardbench: A benchmark for evaluating reward models on computer-using agent.arXiv preprint arXiv:2510.18596, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:bff6d896a9896bd29dbbb7900afa1ca24fbfbcc59c0ad78502ab53488dc994a8

Observation 0c9ae61e-4b22-4b12-af30-28b97b755bc4 · outbound

This paper cites Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:125497c7d4b0e0632ccaa052850867dde1d03274df8c4fc1ec5358677bfc8b36

Observation 87cb8ece-e5c0-425e-be99-a295c3dd5883 · outbound

This paper cites KLong: Training LLM Agent for Extremely Long-horizon Tasks.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading KLong: Training LLM Agent for Extremely Long-horizon Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:5494692c6e5eb5a3eee2401b9725471bc62da5d55284340154cabbb372ff3e9f

Observation be3d5d25-fb5f-45cc-bd74-e98c3ac0f5e9 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:24ccfe4b597f2f9e49a8b7393b555969046a8ccb2190665422f64459ed16ed66

Observation 163d28d8-f7b0-4600-a851-e2924fbe64e4 · outbound

This paper cites MiniMax Sparse Attention.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading MiniMax Sparse Attention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:bf3ea73bc83d066a50fb11732d4c91823ce03e5ee6a3aecc7730aadfdc921bc5

Observation acb35c20-9cdc-4513-a18c-0ea0468925f0 · outbound

This paper cites Kimi k2.6: From code to creation, from one to many.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Kimi k2.6: From code to creation, from one to many

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:7b6409697e3dede361fe121ee3deb43c41e5422ad365fbcac890fc1b2d5a3952

Observation 4d31dd8b-8611-4bd5-8533-96c993498dbf · outbound

This paper cites Kimi k2.7 code: Open-source 1t agentic coding model.https://kimik2ai.com/k2.7/, June 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Kimi k2.7 code: Open-source 1t agentic coding model.https://kimik2ai.com/k2.7/, June 2026

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:63f097fe05778eb57efb7796ace66bc702fe75b9d76ef2552dcecc0994cf5501

Observation 8f9cfab3-5c5a-4f62-b533-79f43c01a821 · outbound

This paper cites Swt-bench: Testing and validating real-world bug-fixes with code agents.Advances in Neural Information Processing Systems, 37:81857–81887, 2024.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Swt-bench: Testing and validating real-world bug-fixes with code agents.Advances in Neural Information Processing Systems, 37:81857–81887, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:ba6fadcc542ea61be0bfef0ca20b349af7f76d7c0d052b165aca7fd4071ab844

Observation 4ecffd4e-2160-44cb-9433-1984c8ee78ee · outbound

This paper cites Codex.https://github.com/openai/codex, 2025.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Codex.https://github.com/openai/codex, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:3092ec31d4dc9c478c55b2093e6d1c3e0d76a3cc159de3c45d79e847757e4dfe

Observation 5ec6e7fe-9482-4da2-be93-2cafcba960f0 · outbound

This paper cites OpenAI GPT-5 System Card.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading OpenAI GPT-5 System Card

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:2cd78553198a08625f5254deed5cef98686acce3ed325097183485a092593092

Observation c0613d6c-b724-4724-8a69-9299e6a5b7e2 · outbound

This paper cites Openclaw, 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Openclaw, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:380a3478c2b64f0d4b32b18230623888b8828b5484a958ebc191935449149baa

Observation a4702ce7-9f1a-4b50-8dfb-edce1839c82c · outbound

This paper cites Rubriceval: A rubric-level meta-evaluation benchmark for llm judges in instruction following.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Rubriceval: A rubric-level meta-evaluation benchmark for llm judges in instruction following

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:469d31c19bd8884870aad1ede9f06b67b888b043130436a7fa7312c2ffa3b7f6

Observation 88b24471-f688-44c8-93f9-39daa6385223 · outbound

This paper cites Qwen3.6.https://qwen.ai/blog?id=qwen3.6, 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Qwen3.6.https://qwen.ai/blog?id=qwen3.6, 2026

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:f6486cc00688f77e7bf7a05fff247ce3f4f06a1a04656d9f4d6ebcde360ec024

Observation da8b1671-2850-4085-8eab-72a766a0cae4 · outbound

This paper cites Qwen3.7.https://qwen.ai/blog?id=qwen3.7, 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Qwen3.7.https://qwen.ai/blog?id=qwen3.7, 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:111e1bf3ff08ce8d4776767d1b5020a271a7209c0b4d4c30f3ab819da5e5808b

Observation f96bb0c3-74ac-4dc2-9d06-bc11f76acf17 · outbound

This paper cites Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:2e015858d3cde2379b3bc1354b816db82e7085a2b5019cb23658fac18a2fd059

Observation 3d5252ba-4f12-481d-b7f8-31074f118c77 · outbound

This paper cites HCAST: Human-Calibrated Autonomy Software Tasks.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading HCAST: Human-Calibrated Autonomy Software Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:a6159a2cab9defab282bae18313f345a989e2ca53fbd7e89972129e98597e4f1

Observation e2d7f8af-7b38-4604-a28a-1cef99f3bc26 · outbound

This paper cites The illusion of diminishing returns: Measuring long horizon execution in llms.arXiv preprint arXiv:2509.09677, 2025.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading The illusion of diminishing returns: Measuring long horizon execution in llms.arXiv preprint arXiv:2509.09677, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:4fcc34b95c18cd405d75abfad39724885d0dd29660ae740183c329d1a83191a0

Observation 68533192-2278-440d-9c0b-745446854a34 · outbound

This paper cites an unresolved cited work.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:80e06d3b3536a7c495e3ef93aebe00c812f0780b87a265577f13d7942f3dfca6

Observation 7a8772a3-8bd0-4827-92a7-cef97d7c7614 · outbound

This paper cites Tencent hunyuan 3.https://hunyuan.tencent.com/, 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Tencent hunyuan 3.https://hunyuan.tencent.com/, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:841598788b8663a2c71ae28186211a55bece789e28229adecb0249ebe1d994ef

Observation d4472a5d-1bcb-40fb-871f-bc4e1ab7c3ac · outbound

This paper cites ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:eb9a067638d2b34277c13e52ce19184ec43ce8c2fd43d3abb2eaf6aa4a5e0038

Observation 8721f60f-92bd-4315-a963-35e05dee8c87 · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Solving math word problems with process- and outcome-based feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:68f7fbce958002d74935e75a523d2180fedf1ed7aad3315b84c13f583bc348d2

Observation d1756103-768f-4219-ba83-ed2e05c1bf05 · outbound

This paper cites Apex-agents.arXiv preprint arXiv:2601.14242, 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Apex-agents.arXiv preprint arXiv:2601.14242, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:ba60bf3f4b8f765be610ed9af27c9794ca03a877cf30ff08f3f37c61d7291884

Observation cd4a1a6e-b61d-43c7-a38b-b5a543a3243b · outbound

This paper cites Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:f73cc5db16b9f1dd7801b42af12f6623d8f05d0e7968816f49bca571e80f3a83

Observation e871df6f-0e68-4356-90d0-259d43c9e111 · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:a6c7b47f03dbe01d61c43dd499a48d40302b5c37b3c8d28c42e4c82b0e440b78

Observation 39cdc564-835b-49bd-bba2-3e2838575ae9 · outbound

This paper cites The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:129bf98902d4e9042e077d49a67d516030d07475ad02d7a58277ed448e66d0ba

Observation c4e9e2a6-d05e-4c08-b8cf-89f4e4b97842 · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:b95e2523161dd0d5a3987a51cbaf41f9450cffa84765d06f2c3cbc9da15c0d34

Observation a7ec3062-52d4-475b-8b0b-fd0b3c58bc27 · outbound

This paper cites Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:475c1e5a0ff3390e3cf7301244065b6fde8884b9e60fed91aa921cf98df14991

Observation 5e61e23f-8b40-43fa-a326-f7db9dd3d788 · outbound

This paper cites Grok.https://x.ai/, 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Grok.https://x.ai/, 2026

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:2a1b337431e72e5352661baef4a7ed1d2899cccfcd67aa6b55eabb469af57550

Observation 7e932d43-c268-409b-bfe0-ab8638842967 · outbound

This paper cites Grok 4.5.https://x.ai/, 2026.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Grok 4.5.https://x.ai/, 2026

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:c49d1bcf6bf5b92c9caf1931e47f1868bb0367124542b647b84d9ca12aab7f9f

Observation a0273930-b891-47d7-9cce-d9a27083e5ec · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:c4aa01daabad12c01127d688876d39d4eda72af4209e01a985dae4a876f162ea

Observation 8fdd5fb5-5e47-4ca5-af01-524a96a24968 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:969d8b1d5687579a613f020746af9bd9da66000e9230510ad5f4f1d68d9a3244

Observation 1e3afeb1-d4eb-4dbd-85be-d2e879f716e7 · outbound

This paper cites Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:a7ca098715583a8f239429c82bf213b1b489958eb8563d497e355f571926671a

Observation 52d4ddb4-5b6b-4cca-a130-6451205cf9a5 · outbound

This paper cites Self-Rewarding Language Models.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Self-Rewarding Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:e8d87caff7406b40af2df83eb7f4920a88610585dfe2cc4b994d40e021c8e608

Observation 2fb734a2-0497-4a60-b6ea-57d236373dc9 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623, 2023.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623, 2023

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:04c0101da26e2a2e787bf7bacfe1f22a8fb60564edc9cf2982caf807855f08b3

Observation 9825a558-4f43-4a83-b100-3ac4e4893cf0 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:597e1499b142236c1972cec00c4fedbcd4bd13ec235d5bfcc54e51bc0e52a85f

Pith citing papers

Observation 08769bb1-e34f-4a66-b9a0-2e02d1661f6e · inbound

Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance cites this paper.

Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T10:30:30.340852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:30:30.340852Z digest=sha256:0b710e4b0e6ffb198ae114509e8eaadcb3dfae64fe1461ab01df378ac2d13939

Observation f1b89008-24a9-40b4-bcdf-70106faf04c8 · inbound

Recursive Synthesis for Long-Horizon Terminal Tasks cites this paper.

Recursive Synthesis for Long-Horizon Terminal Tasks Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T12:35:01.407591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:35:01.407591Z digest=sha256:1a3ed0c97726b694611a9578412f61c970b21d6b1177d41e6316c9bee056bcb2

Observation f7153a8a-2cb8-43a7-a39e-afe35ad7ade3 · inbound

Recursive Synthesis for Long-Horizon Terminal Tasks cites this paper.

Recursive Synthesis for Long-Horizon Terminal Tasks Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T04:31:42.307916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:31:42.307916Z digest=sha256:01271fa7318a6f12ea58ea2f35703a7cafda1b93f468dccf0f71b8f57f8c60cb

Observation ffc54f7c-ae67-4ca5-b881-0e330d5adf76 · inbound

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks cites this paper.

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:45:37.646242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-07T04:45:35.074893Z digest=sha256:e7eaeca9f62388c3c5ed0ebf0df1c9fb05d039174e4375c43976d8298b6743a3