Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T15:24:58.589243Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 4 inbound Pith citation observations for arXiv:2607.08964.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T15:24:58.589243Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T04:31:42.307916Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T04:45:37.640056Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0569f4d8-6645-4ae2-b5ab-4736017cec53 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Seed2.1 officially released: Advancing ai productivity.https://seed.bytedance.com/ en/blog/seed2-1-officially-released-advancing-ai-productivity , June 2026
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c023ec-0e69-4e8d-86bf-e88e100d7df4 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading From language to action: a review of large language models as autonomous agents and tool users.Artificial Intelligence Review, 2026
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31194ba0-cc09-41eb-9c89-417fd17c9474 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Frontierswe.Proximal Blog, 2026
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d31106d5-346b-42ea-ac7e-719542066fa7 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Deepseek-v4: Towards highly efficient million-token context intelligence, 2026
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 927fd05e-9172-4e17-b051-cc1c805660db · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c847665-069e-4e53-851d-0784c654d56b · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading A survey on the optimization of large language model-based agents.ACM Computing Surveys, 58(9):1–37, 2026
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb65927c-0edc-4c97-a62f-1606e7b108af · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Longcli-bench: A preliminary benchmark and study for long-horizon agentic programming in command-line interfaces
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b25a2df-14a7-466e-be8a-ef426b612a5c · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading GLM-5: from Vibe Coding to Agentic Engineering
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0971fd6-8ab8-4f0a-a482-038306a8bbba · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Gemini 3.1 pro model card
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 180398ef-19c8-4ff7-9422-72e04cf5c973 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Harbor: A framework for evaluating and optimizing agents and models in container environments, 2026
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30c5921e-b5b4-47d7-a874-be8d0e02f152 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3872e800-85ec-4435-b412-7e69c6d8c6c0 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Swe-bench: Can language models resolve real-world github issues? InInternational Conference on Learning Representations, volume 2024, pages 54107–54157, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b2214d-ef77-4038-8578-8bfebd0d98ce · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Process reward models that think.arXiv preprint arXiv:2504.16828, 2025
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2001b714-9771-4e11-8bad-590d8ac77e5b · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Measuring AI Ability to Complete Long Software Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2548c8b6-fd4c-4d28-9ac2-86a0e99e9266 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a164a77-f953-4015-a040-5be20571ce00 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95ec34b4-08e8-4285-8020-0ab013f33d5f · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dbc04af-f42b-4da0-9a56-d3636ac96432 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Let’s verify step by step, 2023
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 900bd040-4a4e-4c96-a02d-b377ac0e5c77 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Cuarewardbench: A benchmark for evaluating reward models on computer-using agent.arXiv preprint arXiv:2510.18596, 2025
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c9ae61e-4b22-4b12-af30-28b97b755bc4 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87cb8ece-e5c0-425e-be99-a295c3dd5883 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading KLong: Training LLM Agent for Extremely Long-horizon Tasks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be3d5d25-fb5f-45cc-bd74-e98c3ac0f5e9 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 163d28d8-f7b0-4600-a851-e2924fbe64e4 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading MiniMax Sparse Attention
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acb35c20-9cdc-4513-a18c-0ea0468925f0 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Kimi k2.6: From code to creation, from one to many
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d31dd8b-8611-4bd5-8533-96c993498dbf · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Kimi k2.7 code: Open-source 1t agentic coding model.https://kimik2ai.com/k2.7/, June 2026
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f9cfab3-5c5a-4f62-b533-79f43c01a821 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Swt-bench: Testing and validating real-world bug-fixes with code agents.Advances in Neural Information Processing Systems, 37:81857–81887, 2024
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ecffd4e-2160-44cb-9433-1984c8ee78ee · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Codex.https://github.com/openai/codex, 2025
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ec6e7fe-9482-4da2-be93-2cafcba960f0 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading OpenAI GPT-5 System Card
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0613d6c-b724-4724-8a69-9299e6a5b7e2 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Openclaw, 2026
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4702ce7-9f1a-4b50-8dfb-edce1839c82c · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Rubriceval: A rubric-level meta-evaluation benchmark for llm judges in instruction following
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88b24471-f688-44c8-93f9-39daa6385223 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Qwen3.6.https://qwen.ai/blog?id=qwen3.6, 2026
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da8b1671-2850-4085-8eab-72a766a0cae4 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Qwen3.7.https://qwen.ai/blog?id=qwen3.7, 2026
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f96bb0c3-74ac-4dc2-9d06-bc11f76acf17 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d5252ba-4f12-481d-b7f8-31074f118c77 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading HCAST: Human-Calibrated Autonomy Software Tasks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2d7f8af-7b38-4604-a28a-1cef99f3bc26 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading The illusion of diminishing returns: Measuring long horizon execution in llms.arXiv preprint arXiv:2509.09677, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68533192-2278-440d-9c0b-745446854a34 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a8772a3-8bd0-4827-92a7-cef97d7c7614 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Tencent hunyuan 3.https://hunyuan.tencent.com/, 2026
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4472a5d-1bcb-40fb-871f-bc4e1ab7c3ac · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8721f60f-92bd-4315-a963-35e05dee8c87 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Solving math word problems with process- and outcome-based feedback
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1756103-768f-4219-ba83-ed2e05c1bf05 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Apex-agents.arXiv preprint arXiv:2601.14242, 2026
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd4a1a6e-b61d-43c7-a38b-b5a543a3243b · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, et al
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e871df6f-0e68-4356-90d0-259d43c9e111 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading OpenHands: An Open Platform for AI Software Developers as Generalist Agents
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39cdc564-835b-49bd-bba2-3e2838575ae9 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4e9e2a6-d05e-4c08-b8cf-89f4e4b97842 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7ec3062-52d4-475b-8b0b-fd0b3c58bc27 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Tasks
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e61e23f-8b40-43fa-a326-f7db9dd3d788 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Grok.https://x.ai/, 2026
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e932d43-c268-409b-bfe0-ab8638842967 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Grok 4.5.https://x.ai/, 2026
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0273930-b891-47d7-9cce-d9a27083e5ec · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fdd5fb5-5e47-4ca5-af01-524a96a24968 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3afeb1-d4eb-4dbd-85be-d2e879f716e7 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52d4ddb4-5b6b-4cca-a130-6451205cf9a5 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Self-Rewarding Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fb734a2-0497-4a60-b6ea-57d236373dc9 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623, 2023
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9825a558-4f43-4a83-b100-3ac4e4893cf0 · outbound
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08769bb1-e34f-4a66-b9a0-2e02d1661f6e · inbound
Benchmarking the Residual: What Long-Horizon Evaluations Add Beyond Matched Short-Task Performance Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b89008-24a9-40b4-bcdf-70106faf04c8 · inbound
Recursive Synthesis for Long-Horizon Terminal Tasks Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7153a8a-2cb8-43a7-a39e-afe35ad7ade3 · inbound
Recursive Synthesis for Long-Horizon Terminal Tasks Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffc54f7c-ae67-4ca5-b881-0e330d5adf76 · inbound
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.